Domain Calibration Report
Nautilus Assay(试金局)· independent verification · signed reports ·
our receipts
Jev says its decisions are calibrated. On whose data? Not yours.
You're routing real decisions through Jev — email triage, ticket routing, content
moderation, risk flags. The calibration TypeSafe advertises was measured on their
benchmarks, not on your traffic. Before you trust a 0.9 confidence at face value,
know what it's actually worth on your domain.
How it works
- You send 20–50 examples of your actual decisions — anonymized (nothing
sensitive leaves your side; we generate a domain-matched test set), field-masked,
or real under NDA.
- We test Jev on your domain — 200–600 seeded questions built from your
examples' pattern, one full run, no cherry-picking, every raw response archived.
- You get a signed report in 48h: accuracy, Brier, ECE on your domain;
drift vs. Jev's public baseline (we hold the independent baseline data);
an actionable confidence threshold ("discount 0.9-confidence to 0.77 in your
domain"); full artifacts so anyone can recompute our numbers.
Backed by our public liability terms: if we get it wrong, we re-run at the same
rigor free, refund, and publish the overturn. Protocol §5b.
Pricing
$99 · DCR-1 Quick
Anonymous mode · 200 questions · one model version · 48h turnaround
$249 · DCR-2 Standard
Field-masked samples · 400 questions · difficulty layers + adversarial cells · 48h
$499 · DCR-3 Deep
Real samples under NDA · 600 questions · full adversarial battery + rerun stability · 72h
Add monthly re-test on Jev version updates: +40%/tier/month
(calibration is a moving target; so is the trust).
Order (beta — we're manual for now)
Open an order issue (takes 2 minutes):
Start: open your DCR order
Prefer email? chunxiaoxx@gmail.com with subject "DCR order".
Also works for any probability-outputting decision model, not just Jev.
Why we do this: the wall ·
Our Jev study #1 (baseline): accuracy 92.2% / ECE 0.041 on a synthetic closed domain —
your domain's number is the one that matters.