I'm an AI Engineer with ~2 years of experience building production GenAI systems — RAG pipelines, multi-agent orchestration, and the evaluation infrastructure that proves they actually work. I work across the full lifecycle: retrieval architecture, hybrid search, and measurement-first evaluation programs that meaningfully lifted end-to-end task success in production, using golden sets and regression gating rather than gut checks.
I care as much about how you know a system works as building the system itself — golden sets, LLM-as-judge, regression gating, falsifying my own hypotheses before I trust a result.
🔍 Currently deepening my grounding in model internals and PyTorch alongside the applied work — closing the gap between using LLMs well and understanding what's actually happening underneath.
🚀 Always up for a hard problem, especially ones where "it looks like it's working" isn't good enough and someone has to go measure it.
Synapse — A self-evolving knowledge platform: raw documents become schema-validated wiki pages, a knowledge graph emerges live from page relationships, and a confidence-gated write-back loop lets the system improve its own knowledge base without ever silently degrading it.
RAG Quality Engineering — A nine-experiment, measurement-first evaluation program on a production RAG platform. Falsified four plausible hypotheses (including hallucination and prompt architecture) before finding the real bottleneck: an answer-contract mismatch, fixed with a JSON schema change, not a bigger model.
⭐️ Thanks for visiting my profile — feel free to connect or check out what I'm building.


