AI Scaling vs Human Scaling
Toby Ord
Do AI models scale as well with more tokens as humans do with more time?
That is, does an AI get the same benefit from 10x the number of available tokens as a human does from 10x the amount of time to do the task? Or might the AI require a larger scale-up of tokens (such as 20x or 100x) to match the performance improvement when scaling up the time available to a human by 10x?
People often think of giving AIs more tokens to complete a task as akin to giving a human worker more time — especially when those tokens are used to extend the AI’s sequential chain of thought or tool use. It’s a useful analogy, but it isn’t clear whether the AI systems are as efficient as the humans are at turning these extra resources into superior performance.
We often lean on this analogy when trying to understand AI’s current capabilities and how they might change. For example, we might think of scaling up AI training as making the AI more knowledgeable and skilled, while scaling up chain-of-thought gives that model more time to complete the task. But if AI systems require a 100x scale-up of tokens to match the improvement a human gets with 10x the time — and a 10,000x scale-up of tokens to match the improvement a human gets with 100x the time — then this analogy has broken down. As we shall see, this would also have consequences for whether AI systems are on track to be able to improve themselves strongly enough to trigger a runaway intelligence explosion.
It’s an important question, but I haven’t seen much good data to test it either way. The data is rare because it also requires measuring human performance on the suite of tasks, including tracking how well the humans do at a wide range of time budgets. Ideally what we’d want is something like this chart from METR which shows how AI performance at research-engineering tasks scales with more inference compute and how human performance scales with more time on the very same chart:
In this case, the humans gain about 4x as much score per doubling of the amount of resources as the AIs do. If this scaling advantage were sustained, it would mean that to get the same improvement a human gets from having 10x as much time, the AI would require about 10,000x the resources.
However, this chart has some serious limitations. A key weakness is that it used very early models, which are known to have short time horizons (only o1-preview is even a reasoning model). The authors made up for this by allowing a length of time (e.g. 16h) to represent either 1 try taking 16 hours, or the best of 2 tries taking 8 hours, or of 4 tries taking 4 hours etc — whichever setup works best for that agent. It is therefore a mixture of chain of thought scaling and best@k scaling. And something similar was also done for the human times. No human was allowed to work for more than 8 hours so the data points for human scaling which appear to show the curve flattening are really showing the effects of parallel labour (even if this had worse results than a single longer attempt would have).
METR did include a chart with more granular data for how human performance improved with more time, which shows it actually bending upwards, even as it approached the 8 hour mark:
So we are left with results that are suggestive of humans scaling much better with time than AI models do with more inference compute, but we’d really like to see it for more recent models which have been trained to allow use of very long chain of thought.
The best data I know of on this doesn’t directly compare humans’ performance on the same benchmark. Instead, it measures AI performance in terms of human performance. It does this using time horizons — measures of AI performance in terms of how long such work would typically take a human.
In July 2026, the UK AISI released an intriguing chart:
It was intended to show that across a range of agents, allowing them more tokens produced a longer 80% time horizon. This horizon is the duration of task — as measured by how long it takes a human — which the AI can successfully complete 80% of the time. For this chart, the tasks were AISI’s internal suite of capture-the-flag hacking tasks.
If you look carefully, this chart also hints at something very interesting about how AIs and humans scale. First, note that both axes are logarithmic — it is a log-log plot. The first thing you should do when you see a log-log plot is to work out the slope of the key diagonal line that goes up one order of magnitude every time it goes across by one order of magnitude. That slope represents a linear relationship between x and y. Any slope lower than that corresponds to scaling with diminishing marginal returns. This line is so important that log-log graphs should include it as a reference line (and ideally also scale their axes evenly so that it is the 45° line). Let’s add it to the chart:
It is now immediately obvious that all of these models are scaling sub-linearly: 10x the number of tokens is giving less than 10x the estimated (human) time horizon. Yet each trend line is roughly straight, which on a log-log plot means the time horizon is (roughly) proportional to the token budget raised to some power. That is:
$H_{80} \propto T^\gamma$
Where $H_{80}$ is the length of the 80% time-horizon, $T$ is the number of tokens, and $\gamma$ is a parameter measuring how quickly the horizon grows as tokens are increased.
If $\gamma = 1$, that would correspond to linear scaling of horizon length, but here these lines all have $\gamma < 1$, which means horizon length is increasing sublinearly. Thus, all these AI systems would appear to be scaling less well than humans.
How much worse are they doing? Let’s list the $\gamma$ for each model, along with how much we’d need to scale up its tokens to achieve a 10x improvement in time-horizon (which is given by $10^{\frac{1}{\gamma}}$):
As expected, the values of $\gamma$ are all less than 1, indicating that all models are scaling worse with more reasoning tokens than humans do with more time. But we can also see that there is a broad upward trend in the value of $\gamma$ as we go down the table, and the models are listed in order of release date. In less than a year, they have gone from requiring 10,000x as many tokens to scale their time horizons by 10x (a similar ratio to the first chart), to needing just 20x as many tokens.
So the AISI chart suggests that the models (up to May 2026) scale worse with more inference compute than humans do with more time, but that the relationship has been improving over time. We can plot how $\gamma$ has been improving with time:
There is clearly some improvement over time here, though there isn’t enough information to distinguish whether the trend is rising throughout this period, or whether we’ve reached a plateau at roughly $\gamma = \frac{2}{3}$. So it is hard to predict whether $\gamma$ is currently about $\frac{2}{3}$, or whether it has reached 1 (and thus parity with how humans scale). Indeed it is entirely possible that improvements to training will produce AI systems with $\gamma > 1$, scaling better than humans. We definitely need more work tracking this.
I want to stress that all of this data and analysis is still just suggestive as there are many limitations with both the data and my quick analysis:
It is just one experiment with one set of tasks
The data points only cover very long chains of thought — much longer than the contexts of the models and thus requiring multiple rounds of compaction. It is possible that the slopes of these curves are lower in this multiple compaction regime than they’d be when using numbers of tokens that fit within their contexts.
There is substantial uncertainty about the slope of each model’s curve and whether they are still increasing for later models. (The error bars in my chart above understate the total uncertainty as they are just showing the error in measuring the slope of the curves.)
It isn’t clear if human-equivalent time-horizons are a good measure for human performance. They are quite a complicated construct and can be misleading. I’d be happier with a chart with actual human data.
But it is suggestive of AI systems scaling differently with more tokens than humans do with more time. In particular, it suggests that AI systems previously scaled much worse than humans but have improved as newer models are released, to a point where it is possible they match or exceed humans’ ability to scale with more time.
Implications
All of this is relevant to how we intuitively understand inference scaling. If the central analogy (of increasing chain of thought as being like increasing the time allotted to a human) is flawed, that is important to know. Or perhaps we should think of the analogy as correct, but that the AI systems are less good at using more time than humans are — a deficiency that it is important to measure, to track over time, and to project into the future. And we shouldn’t forget that they might end up scaling even better than humans.
One particular implication that interests me is the connection to recursive self-improvement (RSI). Researchers studying this use a parameter called $r$ to represent the strength of the feedback loop where a more intelligent AI is even more capable of improving the intelligence of its successor. In their modelling, when $r > 1$ the process explodes, with performance rising towards a vertical asymptote. Researchers define $r$ as equal to $\frac{\lambda}{\beta}$, where $\lambda$ represents how well research scales with parallel labour (more workers) compared to serial labour (more time), and $\beta$ represents the diminishing returns to more research.
These parameters are often calibrated from how the human research community’s outputs have responded to more researchers, or to more time. The general assumption of much of the RSI research is to start modelling RSI at the point where an AI system is at least as good as a human in every respect relevant to doing AI research. So they would, by definition, have the same $\lambda$ value as humans.
I’ve recently written about how we can estimate the current $\lambda$ for swarms of AI agents. i.e. how scaling up the number of agents in a swarm compares to scaling up the number of tokens they get. From what I can tell, the current AI $\lambda$ is pretty similar to the human $\lambda$. But the considerations here suggest that adjusting for this may not get you all the way from how AI scales with more compute to how humans scale with more sequential time.
As a simplified example, suppose $\lambda = 0.5$ and $\gamma = 0.5$. Then to get the same improvement in performance as scaling up human time by 10x, we’d need a single agent with a chain of thought 100x as long (via $\gamma$), and achieving this performance by instead scaling up the number of agents working in parallel would require 10,000x as many agents (via $\lambda$).
More generally, if we let $N$ be the number of agents in a swarm:
$T \propto N^\lambda$
$H_{80} \propto T^\gamma$
So:
$H_{80} \propto (N^\lambda)^\gamma = N^{\lambda\gamma}$
If we were to adjust the definition of $r$ to account for this (so that it can convert a scale-up of number of AI agents into the equivalent scale up of serial time for human researchers), it would be $r = \frac{\lambda\gamma}{\beta}$. When $\gamma < 1$, this makes an intelligence explosion harder than we’d have thought, though if in the future $\gamma > 1$, it would be easier. This makes measuring and tracking $\gamma$ look even more important.
AI Use
I used Claude Opus 5.5 to estimate the slopes of the lines from the AISI chart and to plot how these vary by release date. I also used it to fact-check the final piece.
9 October 2026