Skip to content
mike.hamata
← work
Agent·2026Prototype

Evidence Desk

A research workspace with inspectable source support, bounded manual and LangGraph agent execution, and reviewed explanations and diagrams.

PythonFastAPILangGraphTypeScriptViteSQLiteOpenAI

Role: Independent builder

AI stack

OpenAI web search · Term-overlap passage selection · Manual / LangGraph stages · Separate support audits

Under the hood ↓

Evidence Desk lets a reader ask a question, inspect the sources behind its answer and explore an explanation. I built both a manual agent loop and a LangGraph path, with durable checkpoints and bounded execution.

Product preview

The hosted workspace opens in Live research mode and shows two retained complete conversations. This publication check did not submit a new research question or make a paid model call. The screenshots below document the separately checked local setup state.

The actual home screen before provider setup, with no research run shown.
The actual home screen before provider setup, with no research run shown.
Research modes and source inspection explained in the interface.
Research modes and source inspection explained in the interface.

Under the hood

Flow: plan → web discovery → safe page reading → passage selection → synthesis → claim review → teaching review → published explanation.

The model names and limits below describe source defaults, not verified hosted settings. The public workspace opens in Live research mode with retained complete conversations; this publication check did not submit a new research run.

Models and source discovery

Research and search default to gpt-4.1-mini through OpenAI Responses. A planner selects up to three focused queries and four subquestions. Quick mode performs one search; Deep mode can perform three, including a gap or counterevidence search. OpenAI's native web-search tool discovers URLs; the application fetches and reads eligible pages itself.

HTML/text/PDF reading records publication metadata, retrieval time and content hashes. Downloads are bounded at 3 MB, and PDF extraction reads at most thirty pages with page labels. Uploaded documents enter provider evidence context; they are not a local-only privacy path. Scanned documents need OCR, which is not implemented.

Passage selection and context

There is no embedding model, vector database, BM25, reciprocal-rank fusion or cross-encoder reranker in the current research path. SQLite stores runs, sources, checkpoints and usage; optional Upstash storage holds scoped notes rather than vectors.

Local query-term overlap ranks 1,200-character windows at 1,000-character strides. The five selected windows form an excerpt bounded at 6,000 characters. A separate sentence selector keeps up to eight numbered passages per source, with an eighty-word passage limit. This inexpensive, inspectable method can miss paraphrases or evidence buried outside the selected windows.

Duplicate content is excluded from evidence context. Metadata and numbered passages share a nominal 10,000-token evidence pool. Complete structured-call payloads have separate bounds; supporting passages are not silently shortened to make an oversized payload fit. Publisher and domain diversity indicators do not prove independent corroboration.

Agent flow, tools and review

A manual saved-stage loop and a LangGraph path share the same research actions. LangChain exposes safe page fetching as a structured tool. Saved SQLite application state, including pending URLs, supports resumption without repeating completed search or review stages; it is separate from LangGraph's in-memory checkpointer.

Synthesis, claim audit, lesson generation and teaching audit are separate stages. Source IDs, passages and bounded quotes are checked before a model reviews support, dates, units and scope. Every lesson block and diagram relationship needs a matching audit decision. Incomplete review withholds teaching content; a hash binds published content to the reviewed version. Diagram geometry is owned by the application rather than executable model output.

Optional illustrations use separate planning and actual-image review. Their source-default review model is gpt-4.1; the image model is gpt-image-2.5-flare. These jobs are not run on every exploration click. Automated review has documented false positives and is not a factual-accuracy guarantee.

Evaluation, safeguards and recorded decisions

The evaluation set contains thirty questions. Its opt-in runner records sources, claim support, status and active time; human accuracy/usefulness and the Google comparison remain uncompleted. Mocked browser checks test interface behavior, not live answer accuracy.

Signed owner-scoped sessions, CSRF checks and guarded public fetches protect records and block private-network destinations. Two research runs can execute concurrently; a run is bounded by twelve calls, 90,000 tokens and 300 active seconds, with conservative spend reservations. No new paid evaluation was run for this case study.

The project records two useful corrections: generic authority terms produced irrelevant sources for household questions, so planning now preserves the user's setting; twenty-word snippets cut mechanism explanations, so internal reading passages were separated from shorter display quotations. Dedicated audits improve inspectability but consume calls and can share the generator's failure modes.

Evidence before presentation

The system retrieves accessible public sources, checks bounded quotations against passages, and reviews support separately from the answer. Explanations, diagrams and optional illustrations reuse retained evidence. Saved output can reopen without another generation call.

Signed browser sessions scope saved records to their owner. Same-origin and CSRF checks protect mutations, private-network fetches are restricted, and provider credentials remain on the server. Limits bound concurrency, calls, sources, time and estimated spend.

What the checks establish

The retained development evidence includes provider-backed manual and LangGraph teaching examples. Tests exercise source and publication boundaries, diagram references, picture integrity and withholding incompletely reviewed output. These are integration and behavior checks, not proof of universal factual accuracy or improved learning.

The remaining boundary

Generated explanations can still oversimplify. Scanned PDFs require OCR, paywalls are not bypassed, and diagrams are conceptual rather than calibrated simulations. Upstash verification and a comparative benchmark remain unfinished. A hosted demonstration is linked above. The screenshots here show the local setup screen; they do not show a completed research run. The hosted interface and retained conversations were checked separately; a new provider-backed run was not tested.