Filings Analyst
A research workspace over SEC filings that answers in numbers it can prove. Every figure is bound to an XBRL fact, a filing cell, or a stated formula, or it is withheld and you are told why. Live on Fly.io, bring your own OpenAI key.
Filings Analyst answers questions over SEC filings: 10-Ks, 10-Qs, and the earnings releases filed as 8-K exhibits. Numeric questions resolve against the filer's own tagged XBRL data. Narrative questions cite the exact passage. When the corpus cannot support an answer, the product says so instead of guessing. It started as the capstone for a course on retrieval-augmented generation and agentic AI, and it turned into a full research workspace, deployed and taking real questions, over ten days of building.

The problem
Public-company filings are the best source of information about a business and the worst thing to read. The tools that exist either summarize without citations, so nothing can be checked, or answer numeric questions by guessing from the prose. Language models are good writers and bad accountants. Ask a general chatbot for a company's operating cash flow and you get a fluent sentence with a number in it, and no way to know which filing, which period, or whether the model made it up.
I spent thirteen years signing off on numbers, so the failure that matters to me is not an awkward sentence. It is a confidently wrong figure under someone's name. The product is built around that one fear.
Numbers do not go through vector search
Embeddings carry meaning, not magnitude. "Revenue was $94.9 billion" and "revenue was $89.5 billion" embed almost identically, so a vector index is the wrong place to look up a number. Numeric questions route instead to the SEC's structured company facts, where every value carries a concept from a controlled taxonomy, a unit, and an exact period.
That data has its own traps, and each one lives in a small pure function with its own tests. The fiscal year and period fields describe the filing, not the fact, so periods resolve from start and end dates only. Every period appears several times across original filings, comparatives, and annual recaps, so one canonical occurrence is chosen. The fourth quarter is never reported as a quarter, so it is derived from the fiscal year minus the three quarters inside it and flagged as derived in the answer's provenance. Freshly filed press-release figures live in a second tier: extracted only when the quoted evidence is an exact substring of the source, stored as provisional, and reconciled against the authoritative XBRL fact once it arrives.

The verifier: cited or withheld
The part I care most about is a deterministic, fail-closed verifier that runs on every draft before a reader sees a word of it. A shared classifier labels every digit string in a sentence as money, percent, per-share, scaled, count, date, year, period, name, identifier, or citation. Only the first five kinds are claims, and every claim must bind to an XBRL fact, a table cell, or a clause of the cited passage that matches on company, subject, period, unit, and value. Computed figures are recomputed from their inputs. Any sentence that recommends, values, or predicts is withheld, so a report can describe a balance sheet but never tell you what to do about it.
Rejection is visible, not silent. Each removed sentence is disclosed to the reader with a plain reason, and the benchmark counts false rejections, meaning sentences the passage did support and the verifier removed anyway. The rules decide by positive definition, what a claim or a name or a period is, rather than by a growing list of exceptions, and every decision every rule makes is recorded as an event so the verifier is maintained from data instead of anecdote. When the design changed, the argument was over what to count as a claim, not over a regex.
An evidence gap is scoped to the question. Having no indexed filing for a company and period is a different outcome from retrieving no matching passage from a filing that is indexed, and both are different from a draft that failed verification. The answer carries a code for which one happened, because "I could not find it" and "the filing does not say it" are not the same statement.
Under the hood
- A LangGraph state machine, not an agent loop. Rewrite folds a follow-up into a standalone question, an analyze step routes it, then one branch runs: structured facts for numeric questions, a statement builder for analysis requests, or hybrid retrieval for narrative ones. Generate is followed by verify, and an unsupported draft gets one bounded retry with the discrepancies named before it becomes a refusal.
- Hybrid retrieval with deterministic fusion. Dense search over pgvector and lexical search over a generated text-search column run concurrently but fuse in a fixed order, so timing never changes ranking. Small child chunks are what gets embedded; larger parent windows with a contextual header are what the model reads. A cross-encoder reranker is optional and measured.
- One origin, one stream. FastAPI serves the JSON API and the React workspace from the same process. A question is a single streamed request: status events narrate the stages, the verified answer arrives as one event, and the web background card streams in underneath it. Draft text never leaves the server.
- Keys travel, keys are never kept. A visitor's OpenAI key rides along as a request header, is used for that request only, and is redacted from every log line and traceback.
- SEC access is a shared, bounded resource. Every caller goes through one token bucket held under the SEC's rate limit, with bounded backoff and an integrity-checked cache. Ingestion is idempotent, and a durable watcher polls EDGAR with locked job claims and resumable progress.
- Postgres is the source of truth. Vectors, keyword indexes, facts, conversation checkpoints, and query telemetry share one database behind an additive, checksummed migration ledger. Conversation identity is stored hashed, and content expires on a thirty-day retention policy.

The worst-day question
This is a public URL with a language model behind it, which is the same cost surface I asked about on Isaac and Praxis: what happens when someone points a script at it? The answer here is structural. Every question runs on the visitor's own key, so the public endpoint cannot run up my bill. Each browser session has a question allowance, the runtime has a bounded concurrency semaphore and queue, and where the operator does fund a call, a spend ceiling reserves the worst-case cost in the database before any provider request is made, so two concurrent callers cannot spend the same remaining dollar.
The same posture runs through the data. Live mode never substitutes synthetic figures after a failure; it reports the failure. Web background is quarantined in its own card, classed by source tier, and forbidden from stating the company's own results, guidance, or valuation, which belong to the filings-grounded answer. Delayed quotes are labeled as separate context. The freshness of every indexed filing is measured from the SEC's acceptance timestamp, and when the corpus misses its own target, the workspace shows the miss rather than hiding it. The footer says what the product is, on every screen: it explains filings, and it is not investment advice.
How I know it works
A grounded system demos well by default, so the evaluation harness was built alongside the pipeline, not after it. Checked-in golden sets cover numeric questions generated from the facts table itself so every expected value is exact, hand-written narrative questions anchored to verbatim filing phrases, unanswerable questions including advice requests and a prompt-injection attempt, statement-analysis reports graded deterministically on every figure binding to a cited fact, and multi-turn conversations. The runner records code, model, configuration, and corpus provenance with every result, because a quality number without its corpus version is a rumor.
Two findings shaped the architecture. A cross-encoder reranker was the one retrieval stage that still moved ranking quality after recall saturated, which matters because the model reads only the top blocks. And against the usual assumption that agents beat pipelines, the fixed pipeline outperformed a tool-calling agent over the same tools on grounding discipline at equal latency, winning on numeric accuracy, abstention, and citation precision; the agent won only on multi-hop questions. The product ships the pipeline. The LLM judge stays labeled uncalibrated until the blind human grades exist, so only the deterministic numbers are ever quoted as claims.
Status
Live at the link above in live mode against the real SEC corpus, on Fly.io with Postgres and pgvector behind it. Bring your own OpenAI key to ask questions; browsing the financials, filings, reader, comparisons, and calendars costs nothing. Coverage is honest and incomplete: financial facts for roughly five hundred companies, indexed filing text for fewer, and every gap labeled by query scope rather than papered over. Earnings-call transcripts are outside the corpus by design and say so when asked.
The business case says it plainly: the verifier is a real, measured edge in the one dimension professional buyers care about, and the product is a strong prototype rather than a company. There are no accounts and no billing. What exists is the hard part, the part that separates writing from verification and proves it on every answer.
What broke
In the first week, the OpenAI account ran out of credit in the middle of a five-year backfill across the S&P 500. Ingestion loads XBRL facts and filing text on separate paths, and only the text path needs the embedding model, so the facts kept landing while every embedding call failed. The failure looked like a rate limit, and the code treated all rate-limit-class errors alike: retry with backoff, move on. So the ticker loop caught the quota error and kept going through every remaining filing, spending each one's retry budget on an account with no money in it. By the time I audited the corpus, there were financial facts for about five hundred companies but indexed text for around three hundred thirty, roughly six thousand failed filings sitting at their attempt limit, and a health check that still reported the corpus as ready, because it was: the facts were there.
The root cause was a classification error, the kind a risk desk trains you to see. An empty wallet and a busy server return the same status code, and only one of them gets better if you wait. The fix gave quota exhaustion its own non-transient error that releases its lock, claims nothing further, stops the batch with an explicit reason, and exits the watcher's poll loop, while an ordinary rate limit keeps its bounded retry. Recovery was done as costed, manifest-bound pilots on single filings, not a mass reset. The lesson is the same one the verifier is built on: decide what a failure is by positive definition, and never let a loop keep consuming a budget it cannot see. A readiness check that reads "ready with warnings" now means exactly that, and the warnings are the point.
Related writing: Book notes: the RAG chapter of AI Engineering, on why hybrid retrieval and chunking at the document's joints carry most of the quality here, and What a risk desk taught me about evaluating AI agents.