<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>My Projects on sujee.dev</title><link>https://sujee.dev/projects/</link><description>Recent content in My Projects on sujee.dev</description><generator>Hugo</generator><language>en-us</language><atom:link href="https://sujee.dev/projects/index.xml" rel="self" type="application/rss+xml"/><item><title>Allycat</title><link>https://sujee.dev/projects/allycat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://sujee.dev/projects/allycat/</guid><description>&lt;img src="allycat-logo1.png" style="float:right; width:300px;"/&gt;
&lt;p&gt;Allycat is an end to end open-source RAG pipeline for website content. It can&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;scape websites&lt;/li&gt;
&lt;li&gt;clean up content&lt;/li&gt;
&lt;li&gt;index content, vectorize and store them in a vector database&lt;/li&gt;
&lt;li&gt;and a UI for queries.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The entire stack is open source. It supports LLMs and embedding models running locally or using an inference service.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/The-AI-Alliance/allycat"&gt;repo&lt;/a&gt;&lt;br&gt;
&lt;img alt="GitHub stars" loading="lazy" src="https://img.shields.io/github/stars/The-AI-Alliance/allycat?style=social"&gt; &lt;img alt="GitHub forks" loading="lazy" src="https://img.shields.io/github/forks/The-AI-Alliance/allycat?style=social"&gt;&lt;/p&gt;
&lt;p&gt;Here is the architecture of Allycat.&lt;/p&gt;
&lt;kbd&gt;
&lt;a href="images/rag-website-1.png"&gt;&lt;img src="images/rag-website-1.png" /&gt;&lt;/a&gt;
&lt;/kbd&gt;
&lt;h2 id="talks--workshops-using-allycat"&gt;Talks / Workshops Using Allycat&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;2025-Nov: Allycat workshop at QConSF&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://qconsf.com/training/nov2025/chat-your-website-using-llm-and-open-stack-allycat"&gt;session details&lt;/a&gt;&lt;/p&gt;</description></item><item><title>Data Prep Kit</title><link>https://sujee.dev/projects/data-prep-kit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://sujee.dev/projects/data-prep-kit/</guid><description>&lt;p&gt;Data Prep Kit accelerates unstructured data preparation for LLM app developers. Developers can use Data Prep Kit to cleanse, transform, and enrich use case-specific unstructured data to pre-train LLMs, fine-tune LLMs, instruct-tune LLMs, or build Retrieval Augmented Generation (RAG) applications for LLMs.&lt;/p&gt;
&lt;p&gt;Data Prep Kit can scale from a single laptop to a cluster scale.&lt;/p&gt;
&lt;p&gt;Repo: &lt;a href="https://github.com/data-prep-kit/data-prep-kit"&gt;data-prep-kit/data-prep-kit&lt;/a&gt;  
&lt;img alt="GitHub stars" loading="lazy" src="https://img.shields.io/github/stars/data-prep-kit/data-prep-kit?style=social"&gt; &lt;img alt="GitHub forks" loading="lazy" src="https://img.shields.io/github/forks/data-prep-kit/data-prep-kit?style=social"&gt;&lt;/p&gt;
&lt;kbd&gt;
&lt;img src="images/data-prep-kit-1.png"/&gt;
&lt;/kbd&gt;
&lt;h2 id="my-contribution-to-data-prep-kit"&gt;My Contribution to Data Prep Kit&lt;/h2&gt;
&lt;p&gt;I worked with the dev team to make Data Prep Kit more accessible to new users and easy to use.&lt;/p&gt;</description></item></channel></rss>