Personal project · Evidence & evaluation / Ongoing
Document Investigation
Change the evidence. Inspect the answer.
An interactive workbench for inspecting retrieval, following source passages, and comparing runs with different evidence.
01Select evidence
Choose from a prepared document collection. The selected document IDs and top-K setting become part of the run record.
02Embed & retrieve
Embed the question and query Pinecone with collection version and document filters. Returned matches are checked against the local corpus to reject excluded or stale evidence.
03Build context
Only the selected passages are supplied to generation. Unlike the portfolio assistant, this workflow adds neither a résumé fallback nor a semantic answer cache.
04Stream & inspect
The UI receives source passages before answer generation. Citation IDs map back to exact source offsets; citation checks validate references, not the factual truth of a claim.
05Compare & replay
Pin a completed run, change the evidence, and compare. Replay reconstructs a captured event log without a new model request. JSON export preserves the event history.
Problem & context
A fluent answer can hide weak evidence. This workbench makes the retrieval inputs, source passages, model output, and failure states available for inspection. It is a portfolio demonstration with a small prepared corpus, including explicitly synthetic webhook examples.
Implementation
Documents are split by headings and paragraphs; long passages use 400-word windows with 60-word overlap. Each passage carries offsets into the original document. A typed event log records stage transitions, sources, answer deltas, and terminal status. Replay rebuilds the UI from those events.
Evaluation approach
Existing deterministic tests cover exact chunk offsets, selection validation, rejection of excluded or stale evidence, citation IDs, cancellation, replay, and fragmented SSE input. These checks evaluate pipeline behavior. They do not establish retrieval precision or answer quality on a representative question set.
Failure analysis
No retrieved evidence produces an explicit insufficient-evidence answer without a generation request. Interrupted runs retain their partial output and a terminal failure state. A citation pointing at a supplied passage can still make an unsupported claim; human source inspection remains necessary.
Results & limitations
Visitors can inspect sources, pin a baseline, alter document selection, replay a captured run, and export its event log. Recorded examples are labelled as recorded. The collection is deliberately small; results do not establish performance on arbitrary uploaded documents or private legal material.
Engineering decision records
01 / Make evidence changes observable
Context. An answer cache or hidden résumé context would make document-exclusion comparisons misleading.
Decision. Use fresh retrieval and generation for each live investigation, with only selected evidence.
Trade-off. Comparisons expose evidence changes, at the cost of another model request per live run.
02 / Replay events rather than regenerate
Context. Model output may differ between runs, and repeated requests incur cost.
Decision. Keep an ordered event log and replay it locally.
Trade-off. A replay is reproducible, but represents a historical run and must be labelled accordingly.