Personal project · Retrieval & streaming / Ongoing
Portfolio Assistant
A conversation with the work behind it.
A retrieval-backed portfolio assistant with streamed answers, source context, and an inspectable execution trace.
01Question
The browser submits a question to a same-origin Next.js API route. The server proxy attaches the service token; it is never sent to the browser.
02Embedding
The backend converts the question to a vector using the configured embedding provider. Embedding happens before cache lookup, including on a cache hit.
03Semantic cache
Redis stores question vectors and responses. Cosine similarity is compared with a configurable threshold, defaulting to 0.92. A match reuses the answer, sources, and suggestions.
04Retrieval
On a cache miss, Pinecone returns the closest indexed résumé passages. The prompt includes those passages plus a local résumé context fallback.
05Generation
The configured provider streams answer chunks. Server events report actual stage completion; the browser renders them alongside the answer.
↳ Cache hit → return the saved answer and sources. Cache miss → retrieve → generate → store the response.
Problem & context
A project list is useful for scanning, but visitors often have questions that cut across projects and experience. The assistant adds a conversational route into the portfolio while keeping the written case studies accessible without an AI service.
Implementation
Zustand holds conversation state. A same-origin server proxy forwards requests to the backend. The streaming endpoint emits named events for stage starts, completions, text chunks, and final sources. The client maintains one active request and distinguishes a completed response from an interrupted stream.
Reliability & boundaries
Provider credentials stay on the server. The proxy checks request origin and body size; backend middleware applies authentication and rate limits. Source matches are retrieval context, not proof that every generated sentence is correct. The portfolio assistant does not provide the workbench’s passage-level citation validation.
Cost & performance
A semantic cache hit skips retrieval and generation, but still pays for a query embedding. The cache implementation scans stored vectors, so its lookup cost grows with the cache. There is no published latency or cost benchmark here. Stage durations are observations for individual requests, not a service-level guarantee.
Limitations & next steps
Similarity alone can incorrectly reuse an answer to a related but different question. A useful next step is a versioned answer cache with expiry and a labelled question set for cache false positives. Retrieved résumé data also needs to be refreshed when portfolio content changes.
Engineering decision records
01 / Keep provider access behind the server
Context. The interface needs streamed model output, but provider and service credentials must stay private.
Decision. Use a same-origin Next.js proxy and a separate backend service.
Trade-off. The UI remains provider-independent, with an extra network hop and two services to configure.
02 / Expose execution events
Context. A spinner cannot explain whether a request is retrieving evidence or generating an answer.
Decision. Render the events emitted by the backend, without simulated stage completion.
Trade-off. Visitors can inspect observed progress. A missing or interrupted stream must remain visibly incomplete.