The substrate matters.
Here’s the data.
Same founder-typed task, run across multiple substrate/model combinations. X axis is cost per run (log). Y axis is the judge’s quality score. Bubble size is latency. Color is model family. Hover for details; click to open the actual artifact. Arrows show the lift from adding OrgX orchestration to the same model.
Live data28 runs · $0.32 spent · 121d ago
How we judgeEvery output is scored 0–10 by Claude Haiku using a task-specific rubric. The judge returns a rationale, strengths, and gaps — all preserved on the run drawer. No human touches the scoring after the fact.
Caveat: “OrgX orchestration” in this first run is a system-prompt wrapper, not the full multi-agent + verifier orchestration. Real orchestration lands in matrix v2.
7 runs·best 8.7/10
Coming soon:Llama 3.3 70BDeepSeek R1Qwen 2.5 72BMistral Largeadd
OPENROUTER_API_KEY to unlockClaude
GPT
Llama
DeepSeek
Qwen
Mistral
Gemma
Grok
- Up-and-right dots are higher quality. Up-and-left dots are higher quality for less money — the efficient frontier.
- Lime-ringed dots are runs with OrgX orchestration on. Compare them to the matching non-ringed dots for the same model to see the lift.
- Arrows trace the shortest OrgX-on counterpart for each model — the “floor raise”.
- Dashed borders encode substrate: solid = raw API, dotted = managed agents, dashed = sandbox, thin solid = agent SDK.