Retrieval eval harness
A benchmarking framework for RAG retrieval quality across chunking and re-ranking strategies. Cut hallucination rate in a production QA bot by measuring the retrieval step in isolation, not just end-to-end output.
Shivajith Mutteal / GenAI Engineer
I work on GenAI systems — agentic orchestration, MCPs, guardrails, and RAG middleware — and build smaller things on the side because I want to. This site collects both: the professional work and the lab, plus notes along the way. Most of it built to be read, forked, or clicked through, not just described.
A benchmarking framework for RAG retrieval quality across chunking and re-ranking strategies. Cut hallucination rate in a production QA bot by measuring the retrieval step in isolation, not just end-to-end output.
Turns a prompt spec into a structured multi-step LLM pipeline with built-in guardrails and typed intermediate outputs, instead of one long prompt doing five jobs at once.
Distilled a large model's labels into a small real-time classifier for support-ticket triage — 90% lower inference cost with accuracy within 1.5 points of the teacher model.
Run the same prompt across models side by side and compare output, latency, and cost per call in one view.
A small RAG app over a sample document set — built to make the retrieval step visible, not just the final answer.
Smaller things I built because I wanted to, not because a job needed them.
Spaced-repetition drills for the openings I keep blundering, built after losing the same line four times in a row.
Give it a distance and a starting point, get back a runnable loop instead of an out-and-back.
A just-for-fun tool that riffs on whatever's actually in the fridge instead of sending you shopping.