AI reliability, measured — not promised
Most companies run AI systems they can't verify. I build the measurement layer: reliability audits, evaluation harnesses, and safe agent architectures for small and mid-size businesses. Every claim I make comes with numbers. Yours will too.
LLM systems serve stale, wrong, or hallucinated answers — and almost nobody measures it. When your source data changes, which of your AI's answers silently become wrong? If you can't answer that question with a number, you're flying blind.
/ 01
Your data changed yesterday. Your RAG pipeline still answers with last month's facts. In one of my benchmarks, a standard cache-invalidation baseline served 11 stale answers out of 40 without detecting them.
/ 02
Agents that email, buy, or publish on your behalf — with guardrails written in prompts. Prompts are suggestions. Guardrails belong in code, with hard budgets and signed approvals.
/ 03
Most teams discover their LLM bill at the end of the month. If you don't know the cost of a single run before it executes, you don't have a system — you have a meter running.
Services
Fixed-scope engagements with measurable deliverables. You always know what you'll get, when, and how we'll both know it works.
/ 01
Two weeks. I measure your AI system's freshness, accuracy, and failure modes against ground truth — including what happens when your data changes. You get a report with numbers, failure cases, and a prioritized fix list.
/ 02
Custom benchmarks for your exact use case: frozen protocols, deterministic ground truth, held-out test sets. Know whether a model or prompt change improves or breaks your system — before your customers find out.
/ 03
Agent systems designed with guardrails in code, not prompts: idempotent actions, hard budget ceilings, human-signed approvals, prompt-injection resistance tested as a requirement. Specified down to the acceptance tests.
Proof
I apply to my own work exactly what I sell: measure everything, publish the results — including the ones that don't flatter me.
0
stale answers served
CoreMesh, my data-change repair system: 100% invalidation recall where the index-based baseline caught ~57% and served 11 stale answers out of 40.
1,147
automated tests
On Plevium, a production-grade AI SaaS built in 20 days — with 19 architecture decision records and a fully reproducible pipeline.
18,448
human judges, 43 countries
Real human data used to calibrate and validate synthetic user research — with held-out sets never used for tuning.
$0.05
measured cost per run
Every pipeline I build reports its real token cost per run. Unknown cost is treated as a failure, not an estimate.
// I also publish negative results: benchmarks where my own approach scored at chance level, invalidated scorers, failed hypotheses. That is what "measured" means.
Method
Three principles. No exceptions, including for myself.
Measure first
Before recommending anything, I build the measurement that will tell us both whether it worked. If a claim can't be measured, it doesn't go in the report.
Verify by design
LLMs are probabilistic. The systems around them shouldn't be: exact computation for facts, deterministic checks for guardrails, replayable captures for audits.
Report honestly
Every engagement ends with a report I could publish unchanged — including limitations, confidence, and what I could not verify. That is the deliverable.
About
"Tautai" means navigator in Polynesian languages.
I'm an engineer-entrepreneur. I help companies navigate AI the way navigators cross oceans: not by hoping the weather holds, but by reading instruments, knowing exactly where you are, and catching drift before it becomes a wreck.
I build and measure AI systems end to end — from evaluation harnesses calibrated against real human judgment, to agent architectures with guardrails enforced in code. When I tell a client a system is reliable, it's because I can show the number that proves it, and the number that says how much it costs.
Contact
A 30-minute discovery call. You leave with at least one thing you can measure this week — whether we work together or not.
[email protected]