AI task length keeps doubling: what METR's research means for business data
METR, an independent AI evaluation nonprofit, measures an agent's time horizon: the length of task, timed by how long skilled people take, that the agent completes about half the time. Its 2025 research reported that this horizon has been doubling at a steady pace for several years. Longer autonomous work needs training records of complete, multi-step tasks.
The short answer: the tasks AI agents can finish keep getting longer
METR, a nonprofit that evaluates what frontier AI systems can do on their own, published research in 2025 titled Measuring AI Ability to Complete Long Tasks. It introduced a simple yardstick, the time horizon: how long a task takes a skilled person, measured at the point where an AI agent succeeds about half the time. Plotted across model releases, that horizon has grown exponentially, doubling at a fairly steady interval measured in months rather than years.
This page does not reprint METR's specific figures, because METR updates its measurements as new models appear and publishes confidence intervals that matter. Read METR's own publications for the current estimate. What follows explains what the measure means, what it does not show, and why it matters to companies that hold records of real work.
How METR measures a time horizon
The method is easier to follow than the headlines suggest.
- Assemble a suite of tasks of very different lengths, from quick lookups to multi-hour projects, mostly in software engineering and research.
- Time skilled people completing each task to set its human duration.
- Run AI agents, equipped with tools, on the same tasks and record which they complete.
- Fit a curve linking the agent's chance of success to each task's human duration.
- Read off the duration at which success falls to about one in two: that is the model's time horizon.
- Repeat for models released over several years and plot horizons by release date to see the trend.
The output is one number per model that can be compared across generations, which is why the doubling framing spread so quickly.
What the trend shows and what it does not
| Common reading | What the research supports | What it does not show |
|---|---|---|
| Agents can now do long tasks | The task length agents complete at about even odds has grown steadily | Reliability: horizons measured at higher success rates are much shorter |
| This applies to all work | Results on software engineering and research tasks | Messy business work with vague goals, several people and missing information |
| The trend will continue | A consistent pattern across several years of releases | Any certainty; trends can slow, stall or speed up |
| Agents will replace jobs | Nothing directly about employment | Cost, oversight, liability or adoption inside companies |
| Longer horizons need more data | Not measured by METR | This is an inference, made on this page, about training needs |
The last row matters for honesty. METR measures capability, not data demand. The link from longer tasks to demand for records is reasoning, set out below, not a finding of the research.
Why longer tasks change what training data is worth
A short task fits in one exchange: a question and an answer. A long task is a chain: gather information, make a decision, hand off to a colleague, hit an exception, recover, finish. Agents that work for hours need to learn from examples of whole chains, and the public web mostly holds finished outputs such as published articles, final reports and polished code, not the path that produced them.
Companies hold the path. The table maps task length to the records that capture a complete chain.
| Task length for a person | Illustrative business task | Records that show the whole chain |
|---|---|---|
| Minutes | Route an inbound customer email to the right team | The email, its routing label and the reply |
| Hours | Reconcile a vendor statement against the ledger | Statement, ERP entries, adjustment notes and approvals |
| Days | Resolve an escalated support case | Ticket thread, chat messages, linked engineering issue and closure note |
| Weeks | Onboard a new enterprise customer | Project plan, task history, emails, training records and sign-off |
| Months | Win a complex B2B contract | CRM stage history, proposals, redlines, internal approvals and the outcome |
The further down the table, the scarcer the material outside companies, and the more a long, connected archive is worth. The explainer on why AI agents fail at real business tasks shows what goes wrong when agents lack this kind of example, and why AI data demand keeps growing puts the trend in market context.
What advisors can tell clients without overselling
The useful message is a measured trend from an independent evaluator, not a prediction about the client's industry. A short version for an owner or CFO:
For a wider annual view of AI capability, investment and policy, Stanford HAI's 2026 AI Index is a common reference; quote any figure from the report itself rather than from summaries.
Before raising the topic, check whether the client's records reflect long tasks:
- Threads link across systems, such as a ticket that references an engineering issue or a CRM deal tied to its email trail.
- Timestamps show the order of steps, not just the final state.
- Outcomes are recorded: resolved, won, approved, shipped, reopened.
- History runs back several years, including systems that were archived.
- The company created the records itself rather than holding them for its clients.
Skip the conversation for now if the archives were deleted when old tools were cancelled, if the records mostly belong to the client's own customers, or if a previous AI-training license already covers the same data. A long task history only helps when someone at the company can still export it and has the authority to license it.
Limits and open questions
Treat the doubling trend as a useful signal with clear limits.
- Task mix: the measured tasks lean toward software and research, where success is easy to check. Business workflows may behave differently.
- Success threshold: a horizon at about even odds says little about the dependability a finance or operations team needs.
- Forecasting: extrapolating any exponential curve is risky; the trend could bend either way.
- Data link: whether longer horizons translate into more demand for licensed records depends on how developers choose to train, which outsiders cannot observe directly.
Rights and provenance still decide whether any record can be licensed at all; developers' provenance checks are described in what ethically sourced AI training data means.
Where referral partners come in
A referral partner's job is to notice companies whose archives show long, connected work, then make an introduction. SourceX looks for US companies with 50+ full-time employees at peak (contractors excluded), a multi-year operating history, records spread across many systems, clear rights to license them and an authorized sponsor such as the owner, CEO or CFO. The company fit checker offers a fast first screen with no commitment, and how it works lays out each stage after the introduction.
Partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company; the reward becomes payable only after the buyer pays and SourceX receives its fee, and no reward is guaranteed.
Next step
For a client whose systems tell the full story of long projects, register as a partner and make the introduction. Owners who prefer to start on their own can apply at sourcex.si/apply through your referral link, which keeps the introduction credited to you.
- Step 1Share your linkSend your personal link to a company you know.
- Step 2Company appliesThe company applies itself at /apply.
- Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
- Step 4You get your rewardYour share of SourceX fees becomes payable.
Common questions
What is METR?
METR, short for Model Evaluation and Threat Research, is a nonprofit research organization that evaluates what frontier AI systems can do autonomously, with a focus on multi-step capabilities and the risks they could pose. Its time-horizon work is one of several public efforts to measure agent progress in a way that can be compared across model generations.
What does a time horizon at 50 percent success mean?
It is the length of task, measured by how long a skilled person takes, at which an AI agent succeeds about half the time. Shorter tasks are completed more often and longer ones less often. The measure shows where an agent's ability crosses even odds, not how dependable it is, so horizons measured at higher success rates are considerably shorter.
Does the doubling trend mean agents can already do a full week of work?
Not on the evidence of the trend alone. Measured horizons apply to the specific task suites tested, at roughly even odds of success, under conditions set by researchers. Business work involves unclear goals, several people and missing information. Check METR's latest published measurements rather than extending a curve, and expect real deployments to lag benchmark results.
Why does task length matter to a company with old records?
Training agents for longer tasks requires examples of long tasks done from start to finish, with each step and decision visible. Those examples rarely exist on the public web. A company that kept years of connected records, such as tickets linked to engineering work or deals linked to email threads, holds exactly that kind of material, subject to rights and redaction.
Is the trend a reason to rush a licensing decision?
No. A company should decide on its own timeline, with its counsel, after reviewing scope, price and terms, and nothing is binding until it signs. The research is useful context for why buyers value complete workflow records, not a deadline. A sensible first step is a preliminary fit check and a list of which systems and years of history still exist.
Related pages
Free resources
- Client data licensing eligibility checker — A transparent preliminary screen for one company.
- Enterprise value calculator — Enterprise value from equity value, debt and cash.
- Earnout scenario calculator — Probability-weighted earnout value and its present value.
- All free tools · MCP resource center
By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09
Know a US company with valuable proprietary data?
Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.
Refer a company →I own a business
Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.
Start an assessment