OrgX · The Receipt

Your agent shipped. Can you prove the work cleared a bar?

This page is a receipt: a real quality bar, a real miss, the judge’s criticism, and the iteration that cleared it.

Every number below comes from the eval record that produced it.

The range — same bar, different work
eng · document
Technical Evidence Pack — Persisting avg_quality_score on the Trust Record0.79
2.4s judge · 4 evals

Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.

Attempt 10.86 · passedAug 29, 12:40 UTC · 3.8s

It also addresses the current diagnosed defect state (0/255 receipts missing `trust_impact`) without papering over it, and documents securi…

run 0eb2b8cc-a5f7-4596-8b57-1288c1b328ba
Attempt 20.79 · passedAug 29, 12:40 UTC · 9.6s

- However, important uncertainty is not fully handled: floating-point exactness, concurrent ordering, and the claim that replay reproduces…

run 0ee9c53b-dd82-4305-83cb-db3a7d2f1de4
Attempt 30.86 · passedAug 29, 12:40 UTC · 2.4s

- Floating-point “bit-for-bit” reproducibility is asserted in the invariants, but the algorithm uses incremental mean and acknowledges orde…

run d243eb29-074a-404a-bfc0-accfe27a02a2
Attempt 40.86 · passedAug 29, 12:40 UTC · 2.4s

Strong artifact overall: it is cohesive, end-to-end (contract→write/read→migration/rollback→security→failure modes→verification), and inclu…

run f2c46275-66ab-419a-8325-249d2bfc0e61
marketing · document
Red Dot — OrgX Launch Narrative & Launch-Asset Brief0.84
2.4s judge · 4 evals

Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.

Attempt 10.86 · passedAug 29, 12:39 UTC · 2.7s

- Proof is limited mostly to dogfooding and process/gating descriptions; it’s honest but may be thin for persuasive power once marketing ne…

run 152c250c-7e87-4ae9-855a-2c5b099e8e52
Attempt 20.86 · passedAug 29, 12:39 UTC · 2.4s

Proof is mostly process-based (dogfood + internal gating) rather than external product outcomes; that’s safe, but limits persuasive strengt…

run 8762c70d-5b46-41bd-a247-d22dd9199cd6
Attempt 30.84 · passedAug 29, 12:39 UTC · 2.9s

- However, “proof” is mostly process-based rather than product/user-outcome evidence; quantitative metrics are correctly withheld but there…

run a801781f-234c-454d-b221-48664d945da6
Attempt 40.86 · passedAug 29, 12:39 UTC · 2.4s

- Lacks details on concrete creative asset specs (e.g., sizes, formats, visual system rules) beyond general metaphor guidance.

run bd8cdea0-1554-4158-ac6c-07fc3f08bd63
ops · document
Capture Health Operating Runbook — Judgment Capture (First Ten Records)0.86
3.0s judge · 4 evals

Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.

Attempt 10.92 · passedAug 29, 12:38 UTC · 3.3s

- Good safety posture: explicitly read-only/no side effects; Unknown is first-class and caps the verdict; no auto-remediation; stop/escalat…

run 22d0f14f-5ff5-4039-bcbe-95d9b507ee5c
Attempt 20.86 · passedAug 29, 12:38 UTC · 3.1s

- Some schema/semantics reliance on external realities (e.g., “valid, resolvable artifact_binding_hash matching current artifact content” a…

run 31b293b9-de51-4a21-b5a3-f84e71dbd78c
Attempt 30.90 · passedAug 29, 12:38 UTC · 2.7s

- **Schema-to-world mismatch risk:** The runbook assumes certain field names and closed values (e.g., `human_ruling` label set; `artifact_b…

run 3cc6bd93-a9a9-45a0-8ba4-5c76c1a3a9f9
Attempt 40.86 · passedAug 29, 12:38 UTC · 3.0s

- **Good failure-mode handling:** “Unknown” is first-class and properly caps the verdict; structural checks are designated as unconditional…

run 8ae7633f-a85f-42bf-b230-520f9a5258b9
design · document
Trust Progress & Next-Rung UX Specification0.91
3.6s judge · 4 evals

Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.

Attempt 10.92 · passedAug 29, 18:04 UTC · 3.4s

Weighted scoring (0–1):

run 4361c26f-76da-4d6b-8c83-cfa85eb83100
Attempt 20.92 · passedAug 29, 18:04 UTC · 3.5s

- Correctly treats the diagnosed failure mode (quality_score present but trust_impact absent) as a first-class state with high visibility a…

run 647fa203-be40-4424-bf69-c8ab9be76537
Attempt 30.92 · passedAug 29, 18:04 UTC · 3.7s

- Some dependencies are described but not fully specified in terms of exact API/contract fields expected from engineering (e.g., how “recon…

run 6afd9aa9-c4fe-471e-9ea6-629c7619cee9
Attempt 40.91 · passedAug 29, 18:04 UTC · 3.6s

- A few UI behavior requirements are described with terms that may be interpretive for engineering/design (e.g., “frozen and visually marke…

run b02f6baf-e263-46af-981f-22308213935f
sales · strategy
Close Call Recap: Deal Summary Template0.72
2.2s judge · 5 evals

Rubric
Score 0-1 on: ICP clarity (0.18), target segment (0.16), offer (0.18), sequence (0.18), objections (0.15), next-send plan (0.15).

Attempt 10.86 · passedAug 30, 06:06 UTC · 2.1s

- Some fields are aspirational but not fully operationalized (e.g., “confidence” and status enums are shown but not defined; several owners…

run 63e6f8c4-90dd-42f1-b8ff-b03edd6b8194
Attempt 20.86 · passedAug 30, 06:06 UTC · 2.2s

- Some content remains placeholder-heavy (e.g., detailed RACI names/roles, deliverable specifics), which is expected for a template but red…

run a1d722c8-451c-40b8-8906-ec9cfa14f8fa
Attempt 30.82 · passedAug 30, 06:06 UTC · 2.4s

- The content is clearly structured, but it provides no domain-specific history, examples, or references (e.g., no sample filled entries, d…

run b33d8bbe-210d-4b85-a519-a57e10adfa6a
Attempt 40.72 · failedAug 30, 06:06 UTC · 3.5s

However, it doesn’t define DTC sub-segment criteria (size, channels, maturity), so segmenting is somewhat implicit.

run e5a91229-79a2-4c85-9cdd-ca9f2f179171
Attempt 50.86 · passedAug 30, 06:06 UTC · 2.2s

Strong, reusable sales/strategy template with all major requested components: meeting metadata, mutual goals/success criteria, current stat…

run fe220956-212f-4c3a-91b1-c30220c05653
product · document
Agentic Scale Proof — Canonical Evaluation Rubric v1.00.93
2.8s judge · 4 evals

Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.

Attempt 10.93 · passedAug 29, 09:47 UTC · 3.4s

- The rubric demands traceability (e.g., F1), but the worked example’s “Campaign Brief v2” evaluation does **not** demonstrate real traceab…

run 0cd209ab-c6dd-421a-bd3d-87f57d7df6e6
Attempt 20.93 · passedAug 29, 09:47 UTC · 2.7s

However, these do not materially undermine objectivity or operability.

run 1009e900-7893-48c2-8e15-f53082bf78f5
Attempt 30.95 · passedAug 29, 09:47 UTC · 1.7s

Potential minor weaknesses: it claims “one-page” but is longer in rendered form than a strict single-page might be, and the Evidence/tracea…

run b488db8d-2ce8-431e-8fc8-c689a85cba28
Attempt 40.93 · passedAug 29, 09:47 UTC · 2.8s

Minor gaps affecting an ideal score: the document claims “one-page” but is structurally longer than a single page in typical rendering; als…

run f9cdbdbe-e41e-4cbd-9831-b6b3c7302163

Prefer the unedited version? Open a live room, warts and all.

Pulled from the live eval record · refreshes every 5 minutes · Aug 30, 14:03 UTC