Back to Portfolio

AI Applications · Multi-Agent Workflows

Resolution Desk
Multi-Agent Support Workspace

I will build a support investigation workspace where a policy specialist and a case specialist gather evidence in parallel, a coordinator proposes a cited response, and a human reviews the exact proposal before anything is saved.

LANGGRAPHPARALLEL AGENTSPGVECTORHUMAN APPROVALEVALUATION
Resolution Desk

Status: Coming Soon

The architecture and build plan are complete, and development has not started. This page describes what I intend to build and how I will measure it. It will be updated with results, code, and a demo as the work happens. Everything will use fictional cases, and the app will never send email or issue a refund.

The Problem

Resolving a support case means combining evidence from different places: what the policy says, and what the case record shows. Doing that by hand is slow, and a quick answer that misreads the policy or assumes a missing fact causes real mistakes.

Automating the whole thing is risky in the other direction. A model that can act on its own can be steered by a customer message or a note on the case. The workflow needs to speed up the investigation while keeping the decision with a person.

How It Works

Each investigation will follow the same path and pause for a person before any write:

01A reviewer picks a case and starts an investigation
02A policy specialist searches policy documents while a case specialist reads the case record, in parallel
03Each specialist can only use its own scoped tool
04A coordinator joins both results into a structured proposal with citations
05The graph pauses at a checkpoint and waits for a human decision
06The approval is bound to a hash of the exact proposal that was reviewed
07An approved proposal is saved once as an internal draft; nothing is sent externally

The same coordinator and approval path will also support a single-investigator mode, where one agent holds both tools. That makes it possible to compare single and parallel workflows on the same cases.

What I Will Build

  • The workspace: A React interface showing investigation progress, the proposed response, and the evidence behind it, with approve and deny controls
  • The orchestration: An Express API and a LangGraph graph with parallel branches, a join, checkpoints, and a human-approval interrupt
  • Two modes: A demo mode with scripted agents and lexical search, and a live mode with tool-using OpenAI models, embeddings, and PostgreSQL with pgvector
  • An evaluation suite: Labeled cases to compare single and parallel workflows

The demo will use a handful of fictional policies and cases, including a refund on the exact boundary of the window, a verified duplicate charge, a missing shipment with no tracking, and a case note that tries to take over the agent.

Planned Design Decisions

Workflow with agent loops inside

A deterministic LangGraph workflow will contain the tool-using agents. It will not be an unrestricted autonomous supervisor, which keeps it easier to reason about and defend.

Scoped tools per specialist

The policy agent will only be able to search policies. The case agent will only be able to read the current case. Distinct permissions keep each task independent.

Case text is untrusted

Customer messages and case notes will be treated as evidence, not instructions. A test case with an injected note will confirm it grants no authority.

Approval bound to a proposal hash

The decision will carry a hash of the reviewed proposal, so an altered or stale proposal is rejected instead of saved.

Three kinds of state

Application state, graph checkpoints, and policy embeddings each serve a different recovery or retrieval purpose, so they will be kept separate.

Idempotent internal save

A unique action per run will make a repeated approval safe. This will cover internal writes only, not external delivery.

How I Will Evaluate It

  • •Ten labeled cases covering policy boundaries (a refund exactly on the 14-day limit and one a day past it), a verified duplicate charge, a false duplicate claim, a missing shipment with no tracking, and prompt-injection attempts
  • •Single-investigator and parallel-specialist workflows compared on the same cases for decision correctness, grounded claims, latency, and token use
  • •Every live answer scored by hand on correctness, grounding, handling of uncertainty, and clarity, with the exact unsupported claim recorded
  • •Scripted plumbing checks reported separately from live model quality, since a passing scripted run says nothing about an LLM
  • •Tests that denial, repeated approval, an altered proposal, and a restart never produce an unauthorized or duplicate save

I will publish results only after they exist, and I will keep scripted checks and live model results clearly separate.

Scope

This will be a single-operator demonstration with fictional data. Access will be protected by a token, not a multi-user identity system, and recovery will be explicit rather than handled by a durable job queue. The idempotent save will apply to internal writes, so external delivery would need its own guarantees.

Planned Tech Stack

TypeScriptReactExpressLangGraphLangChainOpenAI APIPostgreSQL + pgvectorVitestRenderSupabase