
Cascade
GenAI study companion & hackathon winner
Google GenAI Exchange '24 winner · +30% learning outcomes

AI · Context Engineering · LLM Evals
AI Engineer / Delhi, India
AI Engineer who ships production-grade systems, not demos. I build multi-agent systems, RAG pipelines, and the evals that keep them honest, shipped to production rather than stuck in slideware.
Multi-agent systems that ship
Engineering the right signal
Measuring what actually matters
When I was a kid, I spent more time breaking my bicycle down with wrenches than actually riding it. I just wanted to see how the pieces fit together.
That curiosity is how I ended up in software. I didn't have a grand roadmap. I actually got into AI by joining my college's AI club as the graphic designer. But once I was in the room, I wanted to write the logic, not just draw it. I ended up enrolling in a dummy school just so I could skip the classes I didn't care about and spend my time coding.
Most of what I build starts because something in my daily life is annoying, or because someone I know is struggling with a chore. In my second semester, we had a course on Applied Numerical Methods. Instead of memorizing the formulas from a slide, I wrote a Python visualizer to replicate the math so I could study for the final. When I realized I was logging movies on IMDb while my friends were over on Letterboxd, I wrote a sync engine so we could share what we were watching.
For me, if a tool is clunky or painful to use, it isn't finished. I want the things I make to feel like second nature to use, and I want the code behind them to read like prose that someone actually cared about. I also hate guessing. When I was building RAG pipelines and saw everyone arguing about the best chunking strategies, I built BetterRAG to run the benchmarks and look at the real data.
The context is fresh. Bring me something to take apart.
Systems I designed and shipped, from multi-agent AI platforms to evaluation tooling. Open any one for the brief.
20 / selectedArchitecting a LangGraph NL2MQL multi-agent system with a 5-layer safety framework and full LLM evaluation + observability.
Owned the core backend for 3 AI Companions and the LLM Ops pipeline behind them.
Deployed 10+ automation agents across support, finance, and sales workflows.
Formal grounding: the courses, specializations, and awards behind the work.
10 / credentialsA blind benchmark against four deep-research products - DAR won. Here's the architecture.
Ask four deep research products the same hard question and you get four flavors of the same disappointment. Here's how I built DAR, an agentic research harness that won a blind benchmark against Perplexity, Gemini, Tavily, and Exa.
The jump from useful to real requires three things almost nobody does.
2,000 users in 24 hours. That wasn't the surprising part. Here's how I designed AI personas that feel real, with situation design, trust-gated context, and the blind spot mechanics no one talks about.
2+ hours of daily browsing compressed into 15-minute focused reading sessions.
How I automated my daily learning routine with 4 simple Perplexity Tasks that deliver curated news, AI research, and system design insights, cutting 2+ hours of browsing into 15-minute focused reading sessions.
AI engineer focused on LLM systems, multi-agent architectures, and rigorous evaluation. Always open to ambitious problems, collaborations, and a good technical conversation.