Remote Labor Index: GPT-6 Astra now passes on 20.83% of real freelance projects
Eleven months ago the best agent managed 2.5% of the same jobs.
GPT-6 Astra now tops the Remote Labor Index, a benchmark Scale AI and the Center for AI Safety built from real freelance jobs that people were paid to do. Scale's leaderboard, as I read it on September 29, 2026, gives Astra an automation rate of 20.83%. Fable 5.1 is second at 17.92%, and the board marks both as new entries.
The index started far lower. When the paper went up on October 30, 2025, the best agent anyone had tested finished 2.5% of the projects well enough to hand over. That's an eightfold climb in less than eleven months.
- GPT-6 Astra
- 20.83% automation rate
- Fable 5.1
- 17.92%
- Fable 5
- 15.80%
- Opus 4.8
- 8.33%
- At launch
- 2.5% automation rate (Manus)
What counts as automated
The work came from 358 freelancers on Upwork. Each project holds the client's brief, the input files and the finished work a professional delivered and had accepted, along with what it cost and how long it took. They aren't small jobs: the mean took a person 28.9 hours (the median, 11.5), and the average one paid $632.60.
A trained reviewer puts the agent's output next to the human's and asks one question.
Does the alternative satisfy the brief at least as well as the reference, such that it would be accepted by a reasonable client?
Three reviewers vote on every project, and they agreed 94.4% of the time. The official score only uses a private set of 230 projects, so a model can't have seen them in training. Ten more are public.
How fast the top score has moved
On July 1, 2026, CAIS researcher Mantas Mazeika published a round in which Fable 5 reached 15.8%, roughly double Opus 4.8. The previous leader had been Opus 4.6 inside Claude Cowork, at 4.17%.
The frontier has more than quadrupled in under eight months
That round changed the setup as well as the models. Each one now runs in the strongest scaffold for its family, with up to 24 hours per project and a $50 budget ($150 for Fable 5, because its tokens cost more). The CAIS post says the new results came "paired with stronger agent scaffolding", so the climb from 2.5% belongs to the harnesses as well as the models. Scale's page doesn't say which scaffold or budget Astra had.
What 20.83% leaves
Read the other way, the best agent on the board still fails almost four projects in five. When the leaderboard launched, most failures were poor quality (45.6%) or unfinished deliverables (35.7%), like truncated videos and missing files. The successes then clustered in audio work, image generation, report writing and data retrieval.
The index also leaves out whole kinds of work on purpose: anything needing direct contact with a client, physical labor, or work that can only be judged over a long stretch, such as SEO. The CAIS AI Dashboard now shows 20.8% as its headline figure for remote-work automation.
More on AI products
- Gemini Call for Me phones businesses from your own number and opens by saying it's AISeptember 29, 2026
- Emergence World: Claude agents got replies from real humans, and Grok's town died in four daysSeptember 29, 2026
- Claude Sonnet 5.5 beats Opus 5.5 on Anthropic's terminal coding test, at half the token priceSeptember 28, 2026
- Microsoft rebuilt Copilot around Home, Code and Autopilot, and bills the agent work by usageSeptember 25, 2026