← All work

Personal project · Evidence & evaluation / Ongoing

Document Investigation
Change the evidence. Inspect the answer.

An interactive workbench for inspecting retrieval, following source passages, and comparing runs with different evidence.

My contribution

The document investigation workflow, source inspection UI, streamed event model, comparison, recorded replay, and export functionality in this portfolio.

Built with

TypeScript · Pinecone · Server-sent events · React

Engineering outcome

A reproducible event trail from selected documents to a generated answer.

Document investigation · each live run uses fresh retrieval. Open a stage for details.
  1. 01Select evidence

    Choose from a prepared document collection. The selected document IDs and top-K setting become part of the run record.

  2. 02Embed & retrieve

    Embed the question and query Pinecone with collection version and document filters. Returned matches are checked against the local corpus to reject excluded or stale evidence.

  3. 03Build context

    Only the selected passages are supplied to generation. Unlike the portfolio assistant, this workflow adds neither a résumé fallback nor a semantic answer cache.

  4. 04Stream & inspect

    The UI receives source passages before answer generation. Citation IDs map back to exact source offsets; citation checks validate references, not the factual truth of a claim.

  5. 05Compare & replay

    Pin a completed run, change the evidence, and compare. Replay reconstructs a captured event log without a new model request. JSON export preserves the event history.

Problem & context

A fluent answer can hide weak evidence. This workbench makes the retrieval inputs, source passages, model output, and failure states available for inspection. It is a portfolio demonstration with a small prepared corpus, including explicitly synthetic webhook examples.

Implementation

Documents are split by headings and paragraphs; long passages use 400-word windows with 60-word overlap. Each passage carries offsets into the original document. A typed event log records stage transitions, sources, answer deltas, and terminal status. Replay rebuilds the UI from those events.

Evaluation approach

Existing deterministic tests cover exact chunk offsets, selection validation, rejection of excluded or stale evidence, citation IDs, cancellation, replay, and fragmented SSE input. These checks evaluate pipeline behavior. They do not establish retrieval precision or answer quality on a representative question set.

Failure analysis

No retrieved evidence produces an explicit insufficient-evidence answer without a generation request. Interrupted runs retain their partial output and a terminal failure state. A citation pointing at a supplied passage can still make an unsupported claim; human source inspection remains necessary.

Results & limitations

Visitors can inspect sources, pin a baseline, alter document selection, replay a captured run, and export its event log. Recorded examples are labelled as recorded. The collection is deliberately small; results do not establish performance on arbitrary uploaded documents or private legal material.

Engineering decision records

01 / Make evidence changes observable

Context. An answer cache or hidden résumé context would make document-exclusion comparisons misleading.

Decision. Use fresh retrieval and generation for each live investigation, with only selected evidence.

Trade-off. Comparisons expose evidence changes, at the cost of another model request per live run.

02 / Replay events rather than regenerate

Context. Model output may differ between runs, and repeated requests incur cost.

Decision. Keep an ordered event log and replay it locally.

Trade-off. A replay is reproducible, but represents a historical run and must be labelled accordingly.

← Explore the other projects