AI & Retrieval-Augmented Generation
Ask TanStack Query
Production RAG with Evals
I built a production RAG assistant that answers questions about TanStack Query with streamed responses, citations to official docs, and a complete evaluation pipeline. The system handles 150+ test questions with automated grading and a scoreboard that shows how each improvement affects quality.
Status: Production
A complete RAG system deployed across multiple cloud providers. Demonstrates the full stack: embeddings ingestion, hybrid search, reranking, streaming generation, automated evaluation, and deployment. The 150+ question eval suite is the real story. It shows how retrieval and answer quality improve over time.
The Problem
Developers learning TanStack Query hit the same walls: docs are comprehensive but scattered across pages, ChatGPT gives answers that sound right but have outdated patterns, and searching for specific API details takes too long.
The real problem runs deeper: I wanted to build and ship a real AI product, measure whether it actually works, and show that measurement in a way hiring managers understand. A website is easy; proving the quality of an AI system is hard.
Most RAG tutorials skip the hard part: evaluation. They show retrieval and generation, but stop before proving the system is grounded in fact. I needed to build the complete loop: retrieval, generation, evaluation, and a way to see how each change affected the score.
How It Works
When someone asks a question, the system runs through seven stages in real time, streaming the answer as it's generated:
The hybrid search stage is key: TanStack Query's API surface has hundreds of hook names. Keyword search (traditional Ctrl+F) catches exact names like `useSuspenseQuery`. Vector search catches *semantic* matches, such as questions about caching that mention context switching but never say "cache". Running both and merging the results works better than either alone.
The reranker is the second critical piece. After hybrid search returns ~20 candidates, Cohere's reranker model reads all of them and re-sorts by fit. This adds no latency in practice (it reads text, not tokens), but the quality jump is dramatic.
What I Built
Three things, built over 18 days:
- The assistant website: Live questions, streaming answers with hover-show citations, a link to each source page
- The eval suite: 150+ real questions with hand-verified answers (sourced from GitHub Discussions where maintainers marked the right answer)
- The scoreboard: One row per quiz run, showing retrieval hit rate, answer correctness, and latency. Proof that each change works
The backend is Python (FastAPI) with Postgres+pgvector for storage. The frontend is Next.js with SSE streaming. Both are deployed: the backend on Render and the frontend on Vercel. The quiz runs on GitHub Actions on every PR.
The whole thing costs about $10–50 to run for a month. Embeddings are pennies. A full quiz run (150 questions) costs under a dollar. That cost transparency is part of the story: I know what this system costs to run, and why.
Key Learnings
This project teaches patterns I use in every RAG system now:
Hybrid search beats semantic alone
Vector similarity finds concepts, keywords catch exact API names. Both together work better than either one.
Reranking is critical
A second model that re-sorts the top 20 results by fit adds dramatic accuracy gains with minimal latency cost.
Grounding requires citations
Every answer claim must map back to a source chunk. This cuts hallucinations and gives users the evidence.
Evals measure what matters
Automated grading with 150+ labeled questions lets you quantify every change. The scorecard is your proof.
Streaming changes UX
Word-by-word response via SSE makes latency invisible. Users see progress instead of waiting for "thinking...".
Data leakage ruins scores
Quiz answers must stay invisible to the system. Leakage makes everything look great but means nothing in production.
Why This Project
It hits the four things AI hiring managers want to see:
- Getting the right information to the AI: Hybrid search + reranking shows I understand retrieval beyond vector similarity
- Measuring quality: A 150-question eval suite with automated grading is rare. Most projects ship without measurement
- Shipping it: Production deployment across three services (Postgres, Render, Vercel) shows end-to-end ownership
- Knowing what it costs: Budget tracking and cost-per-query math show I'm thinking like a builder, not a hobbyist
The scoreboard is the real differentiator. Every other RAG demo stops at "it works." I have numbers proving it works, trending up, with evidence for every change. That's interview material.
Tech Stack
Ready to explore?
The repository has the full source code, the eval suite, and the build tutorial. See how every piece connects.
