Fundamentals · Embeddings & Vector Search
Retrieval Workbench
Embeddings, Pinecone & Grounded Answers
I will build a search system over support policies. It will start with keyword matching, move to embeddings and Pinecone, and finish with a model that answers from the retrieved passages and cites its sources.
Status: Coming Soon
The architecture and build plan are complete, and development has not started. This page describes what I intend to build and how I will measure it. It will be updated with results, code, and a demo as the work happens.
The Problem
Picture a support team whose agents spend their time hunting for the right policy. Finding the policy that answers a question is its own task, and it comes before writing any reply.
Keyword search misses a question like "I was billed twice" when the policy says "duplicate charge". Embedding search can bridge that gap, but it brings its own failure: it always returns something, even when nothing is relevant. The project is about telling those cases apart.
How It Will Work
Ingestion and querying are separate paths, and answering is a separate stage from searching:
The Build Order
I will build it in stages, so each new tool has a job I can explain:
1. Keyword baseline
Start with visible keyword matching so there is something to beat, and a clear failure to explain.
2. Similarity by hand
Calculate cosine similarity on tiny vectors before touching a real embedding model.
3. Real embeddings
Generate text embeddings and save them locally, so the embedding step and the database step stay separate.
4. Pinecone
Move the same vectors into Pinecone and compare its results with the local search.
5. Grounded answers
Only then add a model that answers from the retrieved passages, with sources and the option to abstain.
With only ten policies, a local search is enough on its own. Pinecone is in the project to learn managed indexing and querying, not because the collection needs it.
Planned Design Decisions
Retrieval and answers are different problems
Search will be judged on whether the right policy came back. Answer generation will be judged separately, because rewriting a prompt cannot recover evidence that retrieval never supplied.
Similarity is not confidence
Every query has nearest neighbors, even an unrelated one. A similarity score is not a probability of being right, so the system will need a way to abstain.
Keep the text
A vector alone is not the policy. The original passage text stays available as the evidence the model sees and a reviewer can inspect.
Match models and dimensions
Stored passages and incoming questions have to use the same embedding model and dimensions, or the vectors are not comparable.
Versioned namespaces
A source hash will refuse stale vectors, and a changed source will upload to a new namespace instead of mixing versions.
A filter is not authorization
A category filter narrows results, but a real multi-customer system would derive permitted data from authenticated identity. This project will use fictional policies for a single customer.
How I Will Evaluate It
- •Hit rate at k: how often the expected policy appears in the top results, measured on its own before any answer is generated
- •Separate development and holdout questions, with some questions kept untouched while tuning
- •Unanswerable questions reported separately from ordinary misses, since a search can return something even when nothing is relevant
- •A fresh set of 20 questions: ten straightforward, five paraphrases, and five unanswerable, labeled before evaluation
- •Answers reviewed by hand for source present, claim supported, completeness, and appropriate abstention, using a failure list that separates missing source, wrong ranking, unsupported answer, false refusal, and failure to refuse
- •One change per comparison, such as chunk size or the number of results returned, and results labeled with the dataset size
I will publish results only after they exist, and I will say how small the dataset is.
Scope
This will be a learning project over ten fictional policies for one customer. Ten toy documents do not establish production scale, and the project will not claim they do. The answering stage will not train a model on the policies. It will place retrieved evidence in the model's request.
