Back to Home
Public benchmark bundles
OrgX Autonomous Initiative Benchmark
Each publish-ready week includes the benchmark scorecard, raw summary, and curated example bundle behind the weekly post, plus the task assumptions and human-baseline provenance needed to verify what the numbers actually mean. If you want to run the same benchmark yourself, the fastest path is to open Benchmark Lab in OrgX and execute the suite there.
local-openai-gpt-5-nano-full-public-judge-202604112026-q115 tasks
232.17x flow multiplier
Domains: cross_functional, design, engineering, marketing, ops, product, sales. Repeat count: 3. Benchmark version: 2026-q1.
local-openai-gpt-5-nano-full-judge-202605302026-q115 tasks
373.32x flow multiplier
Domains: cross_functional, design, engineering, marketing, ops, product, sales. Repeat count: 2. Benchmark version: 2026-q1.