Back to Portfolio

Fundamentals · Embeddings & Vector Search

Retrieval Workbench
Embeddings, Pinecone & Grounded Answers

I will build a search system over support policies. It will start with keyword matching, move to embeddings and Pinecone, and finish with a model that answers from the retrieved passages and cites its sources.

EMBEDDINGSPINECONEVECTOR SEARCHGROUNDED ANSWERSEVALUATION
Retrieval Workbench: embeddings, Pinecone, and grounded answers

Status: Coming Soon

The architecture and build plan are complete, and development has not started. This page describes what I intend to build and how I will measure it. It will be updated with results, code, and a demo as the work happens.

The Problem

Picture a support team whose agents spend their time hunting for the right policy. Finding the policy that answers a question is its own task, and it comes before writing any reply.

Keyword search misses a question like "I was billed twice" when the policy says "duplicate charge". Embedding search can bridge that gap, but it brings its own failure: it always returns something, even when nothing is relevant. The project is about telling those cases apart.

How It Will Work

Ingestion and querying are separate paths, and answering is a separate stage from searching:

01Load ten fictional support policies, each with an ID, category, title, and text
02Split each policy into overlapping passages with stable IDs and source metadata
03Generate an embedding for every passage and save the vectors with their source text
04Upload the vectors to a Pinecone cosine index, in a versioned namespace
05Embed an incoming question in the same vector space and search for the nearest passages, with an optional category filter
06Hand the retrieved passages to a separate language model that answers using only that evidence and cites passage IDs
07Abstain when the passages do not support an answer

The Build Order

I will build it in stages, so each new tool has a job I can explain:

1. Keyword baseline

Start with visible keyword matching so there is something to beat, and a clear failure to explain.

2. Similarity by hand

Calculate cosine similarity on tiny vectors before touching a real embedding model.

3. Real embeddings

Generate text embeddings and save them locally, so the embedding step and the database step stay separate.

4. Pinecone

Move the same vectors into Pinecone and compare its results with the local search.

5. Grounded answers

Only then add a model that answers from the retrieved passages, with sources and the option to abstain.

With only ten policies, a local search is enough on its own. Pinecone is in the project to learn managed indexing and querying, not because the collection needs it.

Planned Design Decisions

Retrieval and answers are different problems

Search will be judged on whether the right policy came back. Answer generation will be judged separately, because rewriting a prompt cannot recover evidence that retrieval never supplied.

Similarity is not confidence

Every query has nearest neighbors, even an unrelated one. A similarity score is not a probability of being right, so the system will need a way to abstain.

Keep the text

A vector alone is not the policy. The original passage text stays available as the evidence the model sees and a reviewer can inspect.

Match models and dimensions

Stored passages and incoming questions have to use the same embedding model and dimensions, or the vectors are not comparable.

Versioned namespaces

A source hash will refuse stale vectors, and a changed source will upload to a new namespace instead of mixing versions.

A filter is not authorization

A category filter narrows results, but a real multi-customer system would derive permitted data from authenticated identity. This project will use fictional policies for a single customer.

How I Will Evaluate It

  • •Hit rate at k: how often the expected policy appears in the top results, measured on its own before any answer is generated
  • •Separate development and holdout questions, with some questions kept untouched while tuning
  • •Unanswerable questions reported separately from ordinary misses, since a search can return something even when nothing is relevant
  • •A fresh set of 20 questions: ten straightforward, five paraphrases, and five unanswerable, labeled before evaluation
  • •Answers reviewed by hand for source present, claim supported, completeness, and appropriate abstention, using a failure list that separates missing source, wrong ranking, unsupported answer, false refusal, and failure to refuse
  • •One change per comparison, such as chunk size or the number of results returned, and results labeled with the dataset size

I will publish results only after they exist, and I will say how small the dataset is.

Scope

This will be a learning project over ten fictional policies for one customer. Ten toy documents do not establish production scale, and the project will not claim they do. The answering stage will not train a model on the policies. It will place retrieved evidence in the model's request.

Planned Tech Stack

Python 3.12OpenAI embeddingsPineconeCosine similarityStructured outputsPytest