I’ve been building software and contributing to open source since the early days of Linux. Here are some projects I’ve built, contributed to, and experimented with over the years.
⭐️ Featured Projects
Author / Maintainer
Featured Open Source Contributions
⭐️ Data Prep Kit
Open-source tools to cleanse, transform and enrich unstructured data at scale for LLM / RAG workloads.
Repo ↗Open Source Project Contributions
PRs / Issues
Spark Job Server
Submitted multiple patches and pull requests to Spark Job Server.
HBase
Contributed performance and documentation patches to Apache HBase, a distributed NoSQL database, including HBASE-4440 and HBASE-5555.
Earlier Open Source Projects
Dockerized Stacks
I created these Docker-based stacks to make it easier to develop, test, and learn with Big Data and machine learning tools locally.
- Kafka in Docker - Run a lightweight Kafka cluster on a single machine.
- Spark in Docker - Run a mini Spark cluster locally.
- Training Sandbox - Preconfigured environment with Spark, Kafka, TensorFlow, machine learning and deep learning tools, and Anaconda.
- BigDL Docker - Run Intel BigDL in a Dockerized environment.

