"You are a helpful assistant"
Every company now runs on AI answers nobody verifies. I design, build and measure the systems that tell you when your AI is right, when it is wrong, and what it costs. Measured — not promised.
Enterprises deploy models over oceans of unverified output: stale facts, silent failures, meters running. I transform that chaos into governed, measured, decision-grade AI — systems where every answer carries its evidence, every run carries its cost, and every claim survives inspection. Most AI vendors ship a snapshot. I ship a living system: when your data changes, the change is detected, the impact measured, the answers re-certified.
cat 01_the_system.md
Five stages, one rule: the model proposes, the system disposes. Tap a stage to inspect.
Raw answers, documents and data feeds are frozen as they were seen. Nothing is trusted yet, everything can be replayed byte for byte.
Facts are computed deterministically, outside the model. The LLM writes prose — it never gets to invent a number.
Accuracy, freshness, cost per run — against protocols fixed before any call, on test sets never used for tuning. Unknown cost is treated as failure.
Signed human approvals, hard budget ceilings, prompt-injection resistance tested as a requirement. Guardrails live in code — never in prompts.
Every engagement ends in numbers you could publish unchanged — including limitations, confidence, and what could not be verified.
ls ./02_capabilities
Fixed-scope engagements with measurable deliverables. You always know what you'll get, when — and how we'll both know it works.
/ 01
One day, on your real data. I measure where your AI is wrong, stale, or overpriced — against ground truth built for your business. You leave with a numbered report and a straight answer: API or your own model, and exactly what it should cost.
$1,900 — fully deducted from your build.
/ 02
Your AI, deployed and measured. The right engine for each task — frontier APIs where they're needed, your own small model where privacy, cost, or stability wins. Reliability is measured against your ground truth before delivery, not after your customers notice. And it keeps working after handover: data changes are detected, impact is measured, answers are re-certified.
Benchmarks agreed in writing before training — or you don't pay.
$6,000–$12,000 — fixed price agreed upfront.
/ 03
Your AI stays sharp, measurably. Continuous monitoring against your ground truth, model updates and retraining as your data evolves, a numbered report every month, direct support.
$999/month — cancel anytime.
./03_proof --numbers
I apply to my own work exactly what I sell: measure everything, publish the results — including the ones that don't flatter me.
0
stale answers served
A data-change repair system I built: 100% invalidation recall where the index-based baseline caught ~57% and served 11 stale answers out of 40.
1,147
automated tests
On a production-grade AI SaaS built in 20 days — with 19 architecture decision records and a fully reproducible pipeline.
18,448
human judges, 43 countries
Real human data used to calibrate and validate synthetic user research — with held-out sets never used for tuning.
$0.05
measured cost per run
Every pipeline I build reports its real token cost per run. Unknown cost is treated as a failure, not an estimate.
I also publish negative results: benchmarks where my own approach scored at chance level, invalidated scorers, failed hypotheses. That is what "measured" means.
Start by email — tell me what your AI does, you get a first measured read within 48h. I work by email and WhatsApp only: no calls, no video. Not a limitation — a method. Every commitment, every number, every decision stays on record, and you can hold me to it. Every engagement starts the same way: written scope, fixed price, contract and invoice from Nwbrains LLC — before any payment.