Arena

The substrate matters.
Here’s the data.

Same founder-typed task, run across multiple substrate/model combinations. X axis is cost per run (log). Y axis is the judge’s quality score. Bubble size is latency. Color is model family. Hover for details; click to open the actual artifact. Arrows show the lift from adding OrgX orchestration to the same model.

Live data28 runs · $0.32 spent · 121d ago
How we judgeEvery output is scored 0–10 by Claude Haiku using a task-specific rubric. The judge returns a rationale, strengths, and gaps — all preserved on the run drawer. No human touches the scoring after the fact.
Caveat: “OrgX orchestration” in this first run is a system-prompt wrapper, not the full multi-agent + verifier orchestration. Real orchestration lands in matrix v2.
7 runs·best 8.7/10
4681025¢$1$5Cost per run (log)Quality score
Coming soon:Llama 3.3 70BDeepSeek R1Qwen 2.5 72BMistral Largeadd OPENROUTER_API_KEY to unlock
Claude
GPT
Llama
DeepSeek
Qwen
Mistral
Gemma
Grok
How to read the chart
  • Up-and-right dots are higher quality. Up-and-left dots are higher quality for less money — the efficient frontier.
  • Lime-ringed dots are runs with OrgX orchestration on. Compare them to the matching non-ringed dots for the same model to see the lift.
  • Arrows trace the shortest OrgX-on counterpart for each model — the “floor raise”.
  • Dashed borders encode substrate: solid = raw API, dotted = managed agents, dashed = sandbox, thin solid = agent SDK.