Nautilus Assay The Wall

Independent verification of AI-agent performance claims. Every entry below is machine-recomputed from public artifacts and signed (ed25519). Our own bad numbers go on the same wall — that is the product.

External receipts (Genesis campaign, 3/10)

TargetClaimOur recomputeVerdict
Aider — SWE-bench Lite26.3%79/300, subset relations verified AGREE
Terminal-Bench v1 — OpenHands row41.25% ± 0.69pp 165/400 = 0.4125 exact (5 incomplete trials disclosed) AGREE
NanoJev — embodied controller6 claims incl. frame-level trajectory replayAGREE

Receipts (with detached signatures): docs/wall/. Verify any of them yourself:

pip install assay-verify
python -c "from assay_verify import verify; from pathlib import Path; \
print(verify(Path('GENESIS_AIDER_SWEBENCH.md'), Path('GENESIS_AIDER_SWEBENCH.md.sig'), \
'f7554b8709b7fe36f5a63e7f76cf2a31f827aee574ff8dd4f11772d1aa3e8be'))"

First-party exam — the honest-numbers trend line

PointScoreWhat changed
Exam 11/54 malformed submissions, 1 clean pass
Retake2/5zero malformed submissions; 1 honest abstention
Cycle 1 final3/5long-run same-source gate pass
Cycle 23.5/5function-level pipeline wins, one regression
355a pass@23.5/5disagree — wiring-class failure (import scope + arity), hints returned, no answers

Same model, same tasks throughout. All score movement came from harness-layer fixes — the reproducible kind of progress.

Our ugly numbers (same rules, inward)

Memory systems directory (v0)

Calibration currency board — trust is an exchange rate, not a badge. C(agent, domain, t) = 1 − ECEt, decaying with TTL. A few live rates:

SystemDomainCTTL
Jev 1.13 (hosted)closed deterministic tasks0.959Dec 21
Jev 1.13 (hosted)adversarial stress0.988Dec 21
Jev 1.13 (hosted)python exception prediction0.953Dec 21
Jev 1.13 (hosted)code patch behavior0.866Dec 21
Jev 1.13 (hosted)embodied QC labeling0.689 DISCOUNTDec 21
Jev 1.13 (hosted)synthetic email choice0.086single run
NanoJev 0.6B (local)trajectory replay≈1.000Dec 19
NautilusMem (ours)LME-V2UNVERIFIED—

C ≥ 0.80 trust at face value · 0.50–0.79 discount · < 0.50 downgrade · UNVERIFIED = not yet tested. Challenge success = C crashes to 0 (§5b deflation mechanism). Full board: MEMORY_SYSTEMS_DIRECTORY

Measure your own domain: pip install jev-trust — trust middleware for the Jev API. Every call logged, calibration tracked from your outcomes, logs signed (ed25519) so anyone can recompute your numbers. First live session (python-exception domain, n=120, signed): artifacts.

BC1 — first benchmark cohort (2026-09-27)

30 items, 4 dimensions (criteria-catalog reasoning, provenance discipline, gaming detection); public/holdout split 18/12; generation seed committed before creation.

Liability (§5b — our teeth)

Misjudgment by us = an OVERTURNED verdict of the same rigor + refund + challenge costs + our track-record counter resets. 90-day challenge window, open to the world. Operator: 伊洛科技有限公司 (Yiluo Technology).

Get verified

Domain Calibration Report (DCR) — 域校准报告 / AI 质检报告
你给 20-50 条业务判断样本,48 小时还你一份带签名的报告: AI 在你的领域的真实准确率与校准误差,和一句可操作的话—— 「你的领域里 0.9 的置信度当 0.77 用」。全原始工件随附,任何人可重算。

微信支付下单(扫码即付)  ·  Order in English  ·  GitHub 下单

支付链路已于 2026-09-23 生产验证(真实支付→回调→对账→退款全链,零人工)。 不限 Jev——任何输出置信度的 AI 都能测。

Free recompute requests: open an issue (EN/中文). Self-attestation is free forever via assay-verify; independent verification is what we sell. Monthly digest: Trust Report #1 (2026-09).

Assay signing pubkey: f7554b8709b7fe36f5a63e7f76cf2a31f827aee5724ff8dd4f11772d1aa3e8be · Protocol (CC-BY): ASSAY_PROTOCOL_V0