Skip to content
mike.hamata
← work
Tool·2026Prototype

Model Observatory

A lab for comparing three model providers, inspecting token usage, and separating prepared examples from measured evidence.

PythonFastAPITypeScriptViteSQLiteOpenAIAnthropicGoogle GenAI

Role: Independent builder

AI stack

Native provider adapters · Structured output · Bounded calculator loop · Human rubric scoring

Under the hood ↓

Model Observatory gives the same prompt to three independent conversations and makes the request record inspectable. I built it as an independent educational project, with prepared examples available before live provider setup.

Product preview

Prepared sample answers from three providers. No API calls or benchmark measurements.
Prepared sample answers from three providers. No API calls or benchmark measurements.
Local tokenizer showing text fragments and token counts.
Local tokenizer showing text fragments and token counts.

Under the hood

Flow: shared request → three independent provider adapters → optional calculator rounds → schema validation → usage ledger and human review.

The model names and limits below are source defaults, rather than a fresh verification of the hosted configuration. The public classroom demo uses prepared answers and makes no provider calls.

Models, embeddings and context

The initial comparison lanes select gpt-4.1-mini, gemini-3.1-flash-lite and claude-haiku-4-5. OpenAI Responses, Google GenAI and Anthropic Messages each use their native SDK. There is no embedding model, document chunker, vector store, retrieval fusion or reranker in this path. Task evidence is supplied in the prompt.

The same new prompt and settings go to all three lanes; each lane keeps its own conversation history. Requests allow up to eight history entries, a 4,000-character prompt and 12,000 UTF-8 bytes of combined prompt/history. Direct, few-example and concise-explanation strategies are explicit settings. The token lab uses a real tokenizer, whose count applies to the selected encoding and input, rather than every provider's full request.

Tools and validation

A hand-written loop handles each provider's tool protocol; there is no LangGraph agent. The only tool is a schema-checked calculator with four arithmetic operations, finite operands and division-by-zero handling. An experiment permits at most three operations across four API rounds. The calculator cannot execute code, access the network or write records.

Native structured-output modes produce an answer, explanation, evidence strings and an insufficient-information flag. Pydantic validates the returned JSON again. Valid structure is recorded separately from answer correctness.

Evaluation and safeguards

Run records retain requested and returned model IDs, raw usage, normalized tokens, dated cost estimates, elapsed time and tool traces. Accounting treats cached input and reasoning/output fields according to each provider's reported usage. SQLite stores experiments and human reviews; JSON/CSV exports preserve the record.

Five fixed tasks cover policy interpretation, arithmetic, extraction, missing information and ticket classification. The comparison runner covers three providers and supports human incorrect/partial/complete scores. The live fifteen-run comparison remains unfinished. Prepared answers have no measured provider latency or usage and establish no provider ranking.

Sample mode is the default. Live calls require explicit opt-in, allowlisted model IDs and atomic budget reservations capped at $2. Provider and experiment timeouts apply; SDK retries are disabled. Uncertain billing retains its reservation. Local records are unencrypted and the provider-backed local app has no multi-user authentication.

Recorded decisions and tradeoffs

The project documents native SDKs as the way to preserve usage and tool-protocol differences. It also distinguishes task quality from cost and latency: a higher price does not establish a better answer. Separate lane histories keep one provider's output out of another's context, but make follow-up comparisons less tightly controlled. Five human-scored tasks can support inspection, not a general benchmark.

The control

Each provider uses its native SDK. Usage is normalized without counting cached or reasoning tokens twice. A local ledger reserves the estimated budget before a request starts; an uncertain billing failure keeps its reservation. Requests, output and tool rounds are bounded, and calculator arguments pass a schema check.

What runs

The public Vercel demo opens in prepared sample mode with no API charges. It is a hosted demonstration of the comparison experience; the screenshots above show the separately checked local build.

The working local interface includes multi-turn sample conversations, a real tokenizer, a conceptual walkthrough of an LLM, dated pricing estimates, and a notebook with review and JSON/CSV export. Sample mode does not call a provider. The diagrams explain concepts; they do not expose proprietary model internals.

Evidence and limits

Offline tests cover provider accounting, concurrent budget reservations, bounded input, sample isolation and sanitized failures. Live verification across all three providers and the five-prompt comparison remain pending. Prepared answers are fixtures, not a benchmark. The provider-backed local application has no multi-user authentication and should stay on loopback. The hosted prepared-sample demo is available at the link above.