tonytautai

PII Guardian

Drop a file — PDF, scan, photo, e-mail, zip, CSV export — and get it back with every piece of personal data redacted. Your file is processed in memory and never stored: that is the whole point of the product.

Drop your file here, or choose one

PDF · images · docx · xlsx · csv · txt · eml · zip · json — 10 MB max

Measured scores, baselines and protocol (we sell honesty too)
  • 95.9 % protection, in the wild — real web-sourced files with fictional data + real client behaviors (realworld5, Oct 2026), chain frozen by SHA-256 at scoring. Five successive unseen test sets: 90.8 → 92.3 → 95.5 → 95.9 → 95.9 %.
  • 97.5 % on simulated client folders (clientsim5, strict validation) · 99.1 % on digital files vs ~91–92 % on scans — the gap is mostly OCR, and we always show it.
  • Metric: a value is protected only if every occurrence of it in the file is covered. Target is 99 % — we are not there yet, and we say so.
  • Baselines (same 175 files, same metric): this chain 95.4–96.5 % · Azure AI Language 83.7 % · GLiNER 77.2 % · Presidio 64.7 %. Detector-level comparison on text from our own extraction/OCR — those tools do neither, their end-to-end would be lower.
  • Red-team: 0/7 prompt injections obeyed, 0/7 false positives on empty documents, 0/28 + 100 % recall on held-out traps (synthetic documents, plaintext traps only).
  • Model: Qwen3-4B (Apache 2.0 base) fine-tuned by us, proprietary LoRA adapter — trained by us, licensed to you, running on your infrastructure. A retrained variant (v8b) was measured, rejected and documented; the previous model stayed.

Measured, not promised. ← all models · tonytautai.com