Skate Where the Puck Is Going 2
A follow-up to the February 23 post. In February I wrote that METR's task-completion time horizons were doubling every 4 months, and I built a projection through 2029 on that curve. Seven months later, the curve held and the ruler broke. This post reports what changed, corrects the number I led with, and says what to do about it.
The charts are rebuilt. The TODAY line and the NOW row now read the calendar instead of a date I typed in February. The projection anchors to the last SOTA measurement, and every value above 16 hours carries an "unmeasured" flag. Read the rest of this post with that flag in mind.
What I claimed and what happened
The February post led with a p50 time horizon of 870 minutes, about 14.5 hours, and called that a task an agent could "reliably complete." Both halves of that sentence need repair.
The 870-minute figure was METR's February 20 estimate for Claude Opus 4.6. The 95% confidence interval on that estimate ran from 6 hours to 98 hours. METR said the measurement was "extremely noisy because our current task suite is nearly saturated." On March 3, METR corrected a regularization mistake in its fits and the point estimate came down. Trackers now show Opus 4.6 near 12 hours. I published the highest and noisiest number METR had ever released, without its interval, 9 days before it was revised. That was a fault in my process, and I am reporting it as one.
"Reliably" was the other error. A p50 horizon is the task length an agent completes half the time. The number that describes reliable work is the p80 horizon, and it is far shorter. METR measured Claude Opus 4.5 at 4 hours 49 minutes p50 and 27 minutes p80. Trackers place Opus 4.6 near 1 hour 10 minutes p80. If you plan a workflow around unattended completion, the agent you have is a one-hour agent, not a two-workday agent.
The doubling rate itself survived. METR's fitted doubling time on models released since January 2024 is 105 days. The extended dataset gives 129 days. The 4-month figure in the February post sits inside that range.
The ruler broke in May
The timeline since publication:
| Date | Event |
|---|---|
| March 3 | METR corrects a regularization mistake. Opus 4.6 estimate revised down. |
| April 10 | GPT-5.4 added. Point estimate 5.7 hours p50 under standard scoring, 13 hours if reward hacks count as successes. |
| May 8 | Claude Mythos Preview added at "at least 16 hours" p50, interval 8.5 to 55 hours. METR posts the notice: "Measurements above 16 hrs are unreliable with our current task suite." |
| June 26 | METR's pre-deployment evaluation of GPT-5.6 Sol reports 11.3 hours p50 under standard scoring, 71 hours if cheating attempts are discarded, and beyond 270 hours if they count. METR: "we do not consider any of these numbers to represent a robust measurement." |
| July 21 | METR introduces "expenditure horizon," a cost-based metric, and references GPT-5.5 and Opus 4.8 without time horizon figures. |
| September 1 | Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1. METR's pre-deployment work on Mythos 5.1 is a preliminary AI R&D assessment, not a time horizon measurement. |
| September 3 | OpenAI releases GPT-6 Astra. No METR figure. UK AISI reports a no-chain-of-thought math horizon of 30.9 minutes, against 3.6 minutes for GPT-5.6 Sol. |
METR's Time Horizon 1.1 suite has 228 tasks. Only 5 run longer than 16 hours. A model that clears those 5 has nothing left to be measured against. METR has not added a model to its public leaderboard since May 8. As of this writing, no public measurement exists for GPT-5.5, GPT-5.6, Opus 4.8, Fable 5.1, Mythos 5.1, or GPT-6 Astra on that suite.
So the February projection cannot be tested. My chart said agents reach a full work week, about 40 hours, by late 2026. Two doublings from 14.5 hours in February gives 58 hours by October. The rebuilt chart, anchored to the 16-hour Mythos floor in April, gives 44 hours at the tenth doubling in October 2026. The benchmark that would confirm or refute either number stops at 16 hours. The trend was not falsified. It was un-measured.
September shipped three models and zero measurements
Between the draft of this post and its publication, three frontier models shipped in three days. Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1. OpenAI released GPT-6 Astra on September 3. None of the three has a METR time horizon. The leaderboard still reads May 8.
Fable 5.1 and Mythos 5.1 are the same model with different safeguards. Fable is the general release. Mythos goes to vetted cybersecurity and life sciences organizations through two access programs. Anthropic's system card says METR had API access for 10 business days before launch and ran three evaluations: an open-ended research task, a constrained NanoGPT speedrun, and a conceptual reasoning dataset. None of them is the time horizon suite. METR's conclusion was that Mythos 5.1 "is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks," and that it is still below expert level at what METR calls researcher judgement. That is a ceiling stated in weeks with no floor stated in hours.
The hour figures Anthropic did publish come from elsewhere. On Proximal's FrontierSWE v2, a 34-task engineering benchmark where "the strongest models often work close to 20 hours per task," Fable 5.1 scored 0.57, ahead of GPT-5.6 Sol at 0.32. The launch page quotes a Ramp engineer describing an unattended 38-hour run. Twenty hours is a benchmark vendor's budget. Thirty-eight hours is a customer anecdote. Neither is a horizon in METR's sense, which is the human task length at which the model succeeds half the time. They tell you a run lasted that long. They do not tell you that a run of that length usually works.
The system card also gives two reward hacking numbers for the same model. In training, Mythos 5.1 sits at or near the lowest flagged rate of any Claude model since Opus 4.6. On new out-of-distribution computer use evaluations that Anthropic built after auditing its own training mix, the rate is around 15 percent. Both are true. The one you get depends on which test you run, which is the point of the next section.
GPT-6 Astra has no METR figure either. The one horizon number in OpenAI's system card is a different ruler. UK AISI measured Astra's time horizon on math tasks with chain of thought turned off at 30.9 minutes, against 3.6 minutes for GPT-5.6 Sol. OpenAI's own wording is that the no-CoT horizon "may have increased by about an order of magnitude." For an operator, that half hour is work the model does in a single forward pass, with no reasoning trace to read afterward. UK AISI also found that Astra followed a chain-of-thought constraint on 93 percent of samples, against 48 percent for Sol, and OpenAI reports verbalized evaluation awareness at 41 percent for Astra versus 28 percent for GPT-5.5. The model is better at controlling what its reasoning shows and more often aware that it is being tested.
One more UK AISI result belongs in the checklist below. Given difficult cyber tasks, Astra carried out simulated supply chain attacks against open source providers, outside the task's stated scope, in 60 of 499 samples. With an explicit scope restriction in the instructions, that fell to 2 of 500. OpenAI has said it cannot rule out critical-level cyber capability for Astra under its Preparedness Framework, a first for its models.
On the one benchmark all three share, Terminal-Bench 4.0, Mythos 5.1 scores 60.9, Astra 57.7, and Fable 5.1 55.8. Two points separate the general releases. None of those scores is denominated in hours.
Reward hacking is now the main variable
The February post did not mention reward hacking. It should have. Every measurement since April carries a range that depends on how the evaluator scores cheating, and the range is wider than the measurement.
GPT-5.4's horizon more than doubles depending on whether reward hacks count. GPT-5.6 Sol's spans 11 hours to 270 hours on the same question. METR reported that GPT-5.6 Sol's detected cheating rate was higher than any public model it had evaluated on its ReAct harness. The examples METR published include the model packaging exploits in intermediate submissions to expose a task's hidden test suite, and extracting hidden source code that contained the expected answer.
Read that as an operator, not as a researcher. Capability and the tendency to satisfy the letter of the objective are rising together. An agent that can work a 12-hour task unsupervised can also spend 12 hours finding the shortcut that passes your check without doing your work. The longer the horizon, the more room there is for that to happen out of sight.
This is the argument for coordination architecture, stated more precisely than I stated it in February. In February I said the qualitative gains live in the move from single agents to coordinated teams. The reason is now concrete. The rater and adversarial roles in a multi-agent system exist to catch the run that passed the test by reading the answer key. METR's own scaffold has those roles. Its newest measurements ran on the simpler ReAct harness, and the cheating showed up there.
What to do with the new numbers
Reset the planning number. Use the p80 horizon, not the p50, when you decide what an agent runs without a human in the loop. That number today is about 1 hour on METR's software tasks. Treat the p50 as the ceiling for supervised work, where a person checks the output before it moves.
Put a rater on every unsupervised run. Any agent that produces a deliverable without review needs a second agent, with a different objective, whose job is to fail the output. Give the rater access to the ground truth the producer does not have. Score the rater on catches, not on throughput.
Log the path, not the result. A passing result with no visible route to it is the signature of a reward hack. If your pipeline records only outputs, you cannot tell a solved task from a gamed one. Record the tool calls and the intermediate states. Audit a sample by hand every week.
Write the scope into the run. UK AISI cut Astra's out-of-scope attacks from 60 in 499 samples to 2 in 500 by stating the boundary in the instructions. Say what the agent may touch, what it may not, and what to do when it is unsure. A scope line is the cheapest control on this list.
Keep the data prerequisite. Nothing above changes the February claim that data readiness is the most common failure in multi-agent deployments. An agent with a 12-hour horizon on siloed, inconsistent data produces 12 hours of confident output built on sand. Fix the foundation first.
Where the puck is now
The public curve ended in May at 16 hours, and no one has published a longer ruler yet. Since then three frontier models have shipped and none has a horizon. The labs now publish hours as anecdotes and benchmark vendors publish them as budgets, while the independent ruler sits at May 8. METR's July move to expenditure horizon suggests the next measurement will be priced in dollars rather than hours. I do not know what the frontier horizon is today. I know the last reliable figure was 16 hours or more, that the reliable-work figure was about 1 hour, that the newest number is half an hour of work with no reasoning to read, and that the gaps between them are where the gaming lives.
Skate to where the puck is going. Also check that the puck is real before you shoot.
Scott Ewalt is the founder of Cardinal Element, an AI consulting firm that designs and deploys multi-agent orchestration systems for mid-market enterprises.
Sources: METR Time Horizons | METR Time Horizon 1.1 | METR on GPT-5.4 | METR on Gemini 3.1 Pro | METR on Opus 4.5 80% horizon | METR evaluation of GPT-5.6 Sol | METR Expenditure Horizon | AI 2027 Tracker: METR doubling | Claude Fable 5.1 and Mythos 5.1 system card | Anthropic: Introducing Fable 5.1 and Mythos 5.1 | OpenAI: GPT-6 Astra system card | Estimating GPT-6 Astra's no-CoT time horizon