orgx trail · built on
Trail is built on other people’s ideas
Trail reads what your coding agents did, finds the walls they keep relearning, and measures whether a fix worked. Almost every part of that came from someone else first. Here is what each one figured out, what we took, and where it lives in trail. None of them endorse trail; this is our reading of their work.
- Clio: privacy-preserving insights into real-world AI use ↗
Alex Tamkin and the Clio team at Anthropic
2024You can learn how people really use AI without anyone reading their conversations: have a model summarize and cluster conversations, then show a cluster only when enough different people are behind it.
In trailA fix’s public numbers appear only once at least 5 separate people have adopted it, and uploads carry counts, not words (thread titles only if you opt in with --with-titles).
trail share · the /trail/fixes pages · trail sync (metadata only by default) - Error analysis for AI systems (open coding → axial coding → count → judge) ↗
Hamel Husain and Shreya Shankar
2025Look at real traces before building any metric. Write down what went wrong in your own words, group the notes into failure types, count them, and only then automate a judge you have checked against people.
In trailThe walls are failure types counted across real sessions, and trail’s classifier is checked against human labels in a labeling lab. (We wrote our codebook before our notes, which they warn against; the next labeling pass starts from notes.)
trail walls · the labeling lab (lab/serve.mjs) - Sniffly: a dashboard over your local Claude Code logs ↗
Chip Huyen
2025Your agent’s own logs, read locally, can tell you something surprising about how it fails, like how many errors come from looking for files that don’t exist.
In trailLead with one surprising number about your own agents, found locally.
trail card · the Overview - Docent: searching agent transcripts against a rubric, with cited evidence ↗
Transluce
2025Turn anecdotes about agent transcripts into traceable measurements, and prove a finding by fixing it and measuring again.
In trailEvery adopted fix is reported as a measured before and after, not a claim.
trail adopt · the effect line on each wall - Measuring AGENTS.md: what five runs show that one doesn’t (AAIF) ↗
Andrea Griffiths
2026A single before/after comparison of agent runs can point the wrong way and still look convincing; repeat the runs before believing a result.
In trailtrail experiments reports intervals and says “no detectable change” when the interval spans zero. We learned the same lesson the hard way: our first before/after was confounded by a permission-mode switch.
trail experiments · trail bench - Measuring the impact of early-2025 AI on experienced open-source developer productivity ↗
Joel Becker, Nate Rush, Beth Barnes and David Rein (METR)
2025In a randomized trial, developers took 19% longer with AI tools while believing they were faster. How agent work feels is not evidence of how it went.
In trailtrail compares like with like (same client, same permission mode) and marks a result confounded instead of reporting it.
trail adopt effects · trail experiments - AGENTS.md: a simple, open format for guiding coding agents ↗
The AGENTS.md contributors (now stewarded by the Agentic AI Foundation)
2025One plain Markdown file at the root of a repo that every coding agent reads.
In trailFixes are written where agents already look, as a removable block in AGENTS.md or CLAUDE.md.
trail adopt · trail unadopt - Entire: agent checkpoints stored in git, next to the code ↗
Thomas Dohmke and the Entire team
2026The record of what an agent did should travel with the code it changed.
In trailLessons live in the repo, in files a team reviews like any other change.
trail adopt - Agent Trace: an open, vendor-neutral spec for AI code attribution ↗
Cursor
2026A small open spec that any tool can implement beats a format one product owns.
In trailtrail’s upload format (orgx-trail-threads/v1) is a short, versioned contract you can inspect with --dry-run; publishing it as an open spec is next.
trail sync --dry-run - SpecStory CLI: on-disk session formats for a dozen AI coding clients ↗
SpecStory
2026Every AI coding client stores its history differently; write down each format precisely, with the edge cases, so the history can be read and kept.
In trailtrail reads GitHub Copilot (VS Code), Gemini CLI and Factory Droid following SpecStory’s documented, tested formats (Apache-2.0), including Copilot’s snapshot-and-append session logs.
src/adapters-json.mjs - ccusage: token and cost analysis from local agent logs ↗
ryoppippi and the ccusage contributors
2025Answer an anxious question instantly, from files already on your machine, with one npx command and no signup.
In trailnpx, no account, nothing uploaded, a first answer in about 40 seconds.
npx @useorgx/trail - Portable Game Notation (PGN) ↗
Steven J. Edwards
1993A whole game can be written as a short line of moves that people and programs both read.
In trailEach thread is a move string (probe, run, change, check, ship, failed, denied) you can read at a glance and compare.
the braid on every thread - Stigmergy: coordination through traces left in the environment ↗
Pierre-Paul Grassé
1959Termites coordinate without talking to each other: each one responds to what earlier work left behind.
In trailA wall one session hit becomes a trace the next session reads before it starts, instead of every session starting from zero.
trail guard · trail mcp (trail_check)
npx @useorgx/trail credits