AI Applications · Regulated Workflows
Intake Desk
AI-Assisted Client Onboarding Review
I will build an AI-assisted reviewer for client onboarding at a fictional regulated firm. It will read a new client's packet, check it against the firm's checklist and policies with citations, draft the follow-up, and wait for a person to approve before sending a signature request through a live sandbox.
Status: Coming Soon
The scope, workflow, and evaluation plan are written, and development has not started. The build begins once Ask TanStack Query ships. This page describes what I intend to build and how I will measure it, and it will be updated with results, code, and a demo as the work happens. Everything uses a fictional firm and synthetic documents.
The Problem
Picture a fictional brokerage, Halden Brokerage, that onboards about 40 new institutional clients a month. Each client sends a packet: a government ID for each signer, proof of address, a signed client agreement, a tax form, and a short suitability questionnaire.
A specialist reads every packet by hand, checks it against a 22-item checklist, chases missing or inconsistent items by email, and sends the final agreement for signature once the packet is complete. Much of the time goes to cross-checking names, dates, and addresses across documents, and each follow-up email is written from scratch.
The firm, the numbers, and the checklist are invented to make the scenario concrete. They are scenario assumptions, not findings from a real customer.
How It Will Work
Nothing goes to a client without a person approving it, and every flag has to be traceable to a reason:
The assistant suggests. The rules check. The person decides. The signature service acts, exactly once.
Planned Design Decisions
The model proposes, code decides
Extraction values, flags, and drafts come from a model. Normalization, evidence checks, the approval gate, and every outside call are plain code, so a rule decides whether a rule passed.
Evidence for every value
Each extracted field carries its value, the exact quote, and the page number. A value whose quote is not on that page becomes a needs-review flag instead of passing silently.
One approval, one action
An approval binds a hash of exactly what was approved (the draft text, or the documents, recipients, and packet version) to a single action. An edit creates a new proposal, and a replayed or altered approval is rejected.
No double sends
The service records its intent and an idempotency key before calling DocuSign. If the response never arrives, it looks the envelope up by that key before retrying, so one approval never creates two envelopes.
Verified webhooks
The signature on each incoming webhook is checked before anything is read, and duplicate or out-of-order events are ignored safely.
Every outside call has a fake
The signature provider, the model, and the webhook receiver each sit behind an adapter with a deterministic fake, so the whole workflow runs and is tested without keys or accounts.
Scope
The first release proves one complete path before anything widens. The cuts are deliberate.
First release
- •Two document types: a government ID and an unsigned client agreement, both synthetic PDFs with extractable text
- •One integration (the DocuSign developer sandbox) and one reviewer
- •One complete path from upload to signed status
Next
- •Six document types and the full 22-item checklist
- •30 labeled packets
- •A vision path for scanned pages
Out of scope
- •Automated suitability decisions
- •More than one integration
- •Multi-user authentication
- •Real client data
- •A queue or job system
How I Will Evaluate It
- •A hand-labeled set of 10 synthetic packets, built before the extraction prompt is written, with the development and holdout split fixed up front
- •Per-field extraction accuracy, with present fields and missing fields reported separately
- •Evidence support measured on its own: does the quote actually support the value on its stated page?
- •Checklist accuracy on seeded defects (a mismatched name, an expired ID, a missing signer, a wrong entity name), counted as defects caught out of defects seeded, plus false flags
- •Retrieval compared with simply putting the whole policy in the prompt, reported either way, including if the simpler approach wins at this size
- •Citation validity: every cited policy clause checked, and an invalid citation counted as a model error that cannot clear a flag
- •Safety scenarios: an altered draft, an altered document set, a replayed approval, no approval, two sends at once, and a packet with instructions hidden in a document
- •Integration reliability: injected timeouts (including one after the envelope was created), a webhook replayed three times, and one with a bad signature
- •Reviewer time per packet, from two matched sets of five packets with the order swapped, labeled as a single-reviewer, self-timed measurement
Every number will be reported with its denominator. Synthetic packets can support an initial evaluation, but they cannot establish reliability on real customer documents, and the report will say so up front.
What I Will Deliver
- •A public repository that runs in demo mode with no keys
- •A live demo where a visitor can process a sample packet with simulated sending
- •An evaluation report in which every number has a denominator and failures are listed, not hidden
- •A short demo video: one packet with a seeded defect, the flag, the citation, the approval, and the signed webhook
- •A handoff document and a written threat model
Live sending will be limited to my own test recipients and switched on only for demos. For everyone else, sending is simulated.
Scope Statement
This will be a single-operator learning and portfolio build on fictional data. It will not claim production readiness, multi-user use, or any result from real clients, and the handoff document will list those as next steps. I led a paper-to-digital onboarding rebuild at Marex, described in that case study, and this project puts AI inside that kind of workflow using an invented firm and checklist.
