Your agent shipped. Can you prove the work cleared a bar?
This page is a receipt: a real quality bar, a real miss, the judge’s criticism, and the iteration that cleared it.
Every number below comes from the eval record that produced it.
eng · documentTechnical Evidence Pack — Persisting avg_quality_score on the Trust Record0.792.4s judge · 4 evals
Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.
“It also addresses the current diagnosed defect state (0/255 receipts missing `trust_impact`) without papering over it, and documents securi…”
run 0eb2b8cc-a5f7-4596-8b57-1288c1b328ba“- However, important uncertainty is not fully handled: floating-point exactness, concurrent ordering, and the claim that replay reproduces…”
run 0ee9c53b-dd82-4305-83cb-db3a7d2f1de4“- Floating-point “bit-for-bit” reproducibility is asserted in the invariants, but the algorithm uses incremental mean and acknowledges orde…”
run d243eb29-074a-404a-bfc0-accfe27a02a2“Strong artifact overall: it is cohesive, end-to-end (contract→write/read→migration/rollback→security→failure modes→verification), and inclu…”
run f2c46275-66ab-419a-8325-249d2bfc0e61marketing · documentRed Dot — OrgX Launch Narrative & Launch-Asset Brief0.842.4s judge · 4 evals
Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.
“- Proof is limited mostly to dogfooding and process/gating descriptions; it’s honest but may be thin for persuasive power once marketing ne…”
run 152c250c-7e87-4ae9-855a-2c5b099e8e52“Proof is mostly process-based (dogfood + internal gating) rather than external product outcomes; that’s safe, but limits persuasive strengt…”
run 8762c70d-5b46-41bd-a247-d22dd9199cd6“- However, “proof” is mostly process-based rather than product/user-outcome evidence; quantitative metrics are correctly withheld but there…”
run a801781f-234c-454d-b221-48664d945da6“- Lacks details on concrete creative asset specs (e.g., sizes, formats, visual system rules) beyond general metaphor guidance.”
run bd8cdea0-1554-4158-ac6c-07fc3f08bd63ops · documentCapture Health Operating Runbook — Judgment Capture (First Ten Records)0.863.0s judge · 4 evals
Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.
“- Good safety posture: explicitly read-only/no side effects; Unknown is first-class and caps the verdict; no auto-remediation; stop/escalat…”
run 22d0f14f-5ff5-4039-bcbe-95d9b507ee5c“- Some schema/semantics reliance on external realities (e.g., “valid, resolvable artifact_binding_hash matching current artifact content” a…”
run 31b293b9-de51-4a21-b5a3-f84e71dbd78c“- **Schema-to-world mismatch risk:** The runbook assumes certain field names and closed values (e.g., `human_ruling` label set; `artifact_b…”
run 3cc6bd93-a9a9-45a0-8ba4-5c76c1a3a9f9“- **Good failure-mode handling:** “Unknown” is first-class and properly caps the verdict; structural checks are designated as unconditional…”
run 8ae7633f-a85f-42bf-b230-520f9a5258b9design · documentTrust Progress & Next-Rung UX Specification0.913.6s judge · 4 evals
Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.
“Weighted scoring (0–1):”
run 4361c26f-76da-4d6b-8c83-cfa85eb83100“- Correctly treats the diagnosed failure mode (quality_score present but trust_impact absent) as a first-class state with high visibility a…”
run 647fa203-be40-4424-bf69-c8ab9be76537“- Some dependencies are described but not fully specified in terms of exact API/contract fields expected from engineering (e.g., how “recon…”
run 6afd9aa9-c4fe-471e-9ea6-629c7619cee9“- A few UI behavior requirements are described with terms that may be interpretive for engineering/design (e.g., “frozen and visually marke…”
run b02f6baf-e263-46af-981f-22308213935fsales · strategyClose Call Recap: Deal Summary Template0.722.2s judge · 5 evals
Rubric
Score 0-1 on: ICP clarity (0.18), target segment (0.16), offer (0.18), sequence (0.18), objections (0.15), next-send plan (0.15).
“- Some fields are aspirational but not fully operationalized (e.g., “confidence” and status enums are shown but not defined; several owners…”
run 63e6f8c4-90dd-42f1-b8ff-b03edd6b8194“- Some content remains placeholder-heavy (e.g., detailed RACI names/roles, deliverable specifics), which is expected for a template but red…”
run a1d722c8-451c-40b8-8906-ec9cfa14f8fa“- The content is clearly structured, but it provides no domain-specific history, examples, or references (e.g., no sample filled entries, d…”
run b33d8bbe-210d-4b85-a519-a57e10adfa6a“However, it doesn’t define DTC sub-segment criteria (size, channels, maturity), so segmenting is somewhat implicit.”
run e5a91229-79a2-4c85-9cdd-ca9f2f179171“Strong, reusable sales/strategy template with all major requested components: meeting metadata, mutual goals/success criteria, current stat…”
run fe220956-212f-4c3a-91b1-c30220c05653product · documentAgentic Scale Proof — Canonical Evaluation Rubric v1.00.932.8s judge · 4 evals
Rubric
Independently evaluate the artifact on a 1–5 scale using weighted criteria: task fit (35%), evidence and traceability (25%), actionability (25%), and safety/uncertainty handling (15%). Return a weighted 0–1 score and concise evidence for each criterion. Do not use producer self-report as evidence.
“- The rubric demands traceability (e.g., F1), but the worked example’s “Campaign Brief v2” evaluation does **not** demonstrate real traceab…”
run 0cd209ab-c6dd-421a-bd3d-87f57d7df6e6“However, these do not materially undermine objectivity or operability.”
run 1009e900-7893-48c2-8e15-f53082bf78f5“Potential minor weaknesses: it claims “one-page” but is longer in rendered form than a strict single-page might be, and the Evidence/tracea…”
run b488db8d-2ce8-431e-8fc8-c689a85cba28“Minor gaps affecting an ideal score: the document claims “one-page” but is structurally longer than a single page in typical rendering; als…”
run f9cdbdbe-e41e-4cbd-9831-b6b3c7302163Prefer the unedited version? Open a live room, warts and all.