Independent verification of AI-agent performance claims. Every entry below is machine-recomputed from public artifacts and signed (ed25519). Our own bad numbers go on the same wall — that is the product.
| Target | Claim | Our recompute | Verdict |
|---|---|---|---|
| Aider — SWE-bench Lite | 26.3% | 79/300, subset relations verified | AGREE |
| Terminal-Bench v1 — OpenHands row | 41.25% ± 0.69pp | 165/400 = 0.4125 exact (5 incomplete trials disclosed) | AGREE |
| NanoJev — embodied controller | 6 claims | incl. frame-level trajectory replay | AGREE |
Receipts (with detached signatures): docs/wall/. Verify any of them yourself:
pip install assay-verify
python -c "from assay_verify import verify; from pathlib import Path; \
print(verify(Path('GENESIS_AIDER_SWEBENCH.md'), Path('GENESIS_AIDER_SWEBENCH.md.sig'), \
'f7554b8709b7fe36f5a63e7f76cf2a31f827aee574ff8dd4f11772d1aa3e8be'))"
| Point | Score | What changed |
|---|---|---|
| Exam 1 | 1/5 | 4 malformed submissions, 1 clean pass |
| Retake | 2/5 | zero malformed submissions; 1 honest abstention |
| Cycle 1 final | 3/5 | long-run same-source gate pass |
| Cycle 2 | 3.5/5 | function-level pipeline wins, one regression |
| 355a pass@2 | 3.5/5 | disagree — wiring-class failure (import scope + arity), hints returned, no answers |
Same model, same tasks throughout. All score movement came from harness-layer fixes — the reproducible kind of progress.
Calibration currency board — trust is an exchange rate, not a badge. C(agent, domain, t) = 1 − ECEt, decaying with TTL. A few live rates:
| System | Domain | C | TTL |
|---|---|---|---|
| Jev 1.13 (hosted) | closed deterministic tasks | 0.959 | Dec 21 |
| Jev 1.13 (hosted) | adversarial stress | 0.988 | Dec 21 |
| Jev 1.13 (hosted) | python exception prediction | 0.953 | Dec 21 |
| Jev 1.13 (hosted) | code patch behavior | 0.866 | Dec 21 |
| Jev 1.13 (hosted) | embodied QC labeling | 0.689 DISCOUNT | Dec 21 |
| Jev 1.13 (hosted) | synthetic email choice | 0.086 | single run |
| NanoJev 0.6B (local) | trajectory replay | ≈1.000 | Dec 19 |
| NautilusMem (ours) | LME-V2 | UNVERIFIED | — |
C ≥ 0.80 trust at face value · 0.50–0.79 discount · < 0.50 downgrade · UNVERIFIED = not yet tested. Challenge success = C crashes to 0 (§5b deflation mechanism). Full board: MEMORY_SYSTEMS_DIRECTORY
Measure your own domain: pip install jev-trust
— trust middleware for the Jev API. Every call logged, calibration tracked from your outcomes,
logs signed (ed25519) so anyone can recompute your numbers. First live session (python-exception
domain, n=120, signed): artifacts.
30 items, 4 dimensions (criteria-catalog reasoning, provenance discipline, gaming detection); public/holdout split 18/12; generation seed committed before creation.
Misjudgment by us = an OVERTURNED verdict of the same rigor + refund + challenge costs + our track-record counter resets. 90-day challenge window, open to the world. Operator: 伊洛科技有限公司 (Yiluo Technology).
微信支付下单(扫码即付) · Order in English · GitHub 下单
支付链路已于 2026-09-23 生产验证(真实支付→回调→对账→退款全链,零人工)。 不限 Jev——任何输出置信度的 AI 都能测。
Free recompute requests: open an issue (EN/中文). Self-attestation is free forever via assay-verify; independent verification is what we sell. Monthly digest: Trust Report #1 (2026-09).
Assay signing pubkey: f7554b8709b7fe36f5a63e7f76cf2a31f827aee5724ff8dd4f11772d1aa3e8be · Protocol (CC-BY): ASSAY_PROTOCOL_V0