What is GDPval, and why do real occupational tasks matter for AI?

GDPval is a public AI evaluation, released in September 2025, that measures how well models produce real professional work products, such as reports, plans and spreadsheets, for occupations drawn from the largest sectors of the US economy. Experienced practitioners wrote the tasks and blind-graded the results, making authentic workplace deliverables the yardstick for AI progress.

GDPval in plain language

GDPval is an evaluation that asks AI models to do the kind of work professionals are paid for, then has experienced practitioners judge the results. A task hands the model a request plus the reference files a person in that job would receive, such as last year's figures in a spreadsheet or a set of meeting notes, and asks for a finished deliverable: a memo, a slide deck, a schedule, a forecast.

The name points to the design. Occupations were drawn from the sectors that contribute most to US gross domestic product, with a focus on roles whose work happens mainly on a computer. GDPval was released publicly in September 2025 with a research paper and an openly published sample of tasks. This page describes the design in general terms and gives no task counts or scores; read those, and confirm the design details, in the primary release, because they change as new models are tested.

For anyone who works with established companies, one point stands out: the yardstick is real work product, not exam questions.

How does GDPval work?

Each task follows a request, deliverable and review pattern that anyone who has run a professional practice will recognize.

ComponentWhat it involvesWhy it matters
OccupationsKnowledge-work roles from the largest sectors of the US economyTies AI progress to work with real economic value
Task authorsPractitioners with long experience in each occupationTasks reflect how the job is actually done, not how a textbook describes it
InputsA written request plus reference files such as documents, spreadsheets or imagesThe model has to use context, not just general knowledge
DeliverablesFiles a manager could use: reports, presentations, spreadsheets, diagramsOutput is judged as finished work
GradingExperts from the same occupation compare the model's deliverable with a human expert's without knowing which is whichQuality is defined by practitioners, not by an answer key
Open sampleA subset of tasks is published for others to studyAnyone can inspect what is being measured

How is GDPval different from earlier AI benchmarks?

Earlier benchmarks mostly asked models to choose or produce a correct answer. GDPval asks whether the output would be accepted by someone who does the job for a living.

Benchmark styleWhat it testsWhat it misses
Exam and quiz benchmarksKnowledge and reasoning on questions with one right answerWhether the model can turn messy inputs into usable work
Coding benchmarksWhether a code change passes a repository's testsNon-software work and judgment calls with no automatic check
Occupational task evaluations such as GDPvalQuality of a finished deliverable, judged by practitionersWork that unfolds over days with colleagues and clients
Agent workflow evaluationsMulti-step tool use toward a goalOften run in simulated systems rather than real company ones

The coding row is covered in how AI coding agents are trained, and the multi-day gap in why agents need records of multi-week projects.

Why do real occupational tasks matter for AI?

Building a task like this takes knowledge an AI developer cannot easily produce alone. Someone has to know what a good coverage summary, structural calculation note or month-end variance commentary looks like, and what separates it from a weak one. That knowledge lives inside the firms that produce those deliverables every week.

Evaluations built on genuine work products therefore pull attention toward company records. An established professional-services firm, engineering practice or operations team holds thousands of examples of the pattern GDPval uses: a request, the files used to answer it, the deliverable, and a reviewer's verdict in the form of redlines, approvals or a client's reply. Those records also show what happened next, which a single task cannot.

Simulated tasks can approximate this, with trade-offs set out in synthetic environments versus real business logs. Recency counts as well, because methods, rules and tools change; data freshness is part of the same picture.

What GDPval means for a referral partner

GDPval is a public research benchmark. SourceX has no role in it, and nothing here suggests that its authors license data from any company. For a partner, its value is as a clear public illustration of the evidence AI developers use to judge their models.

Use it as a lens on the companies you know. The request, deliverable, review test:

  • Request: does the company keep the original asks, such as client briefs, statements of work, tickets or internal requests, and not just final files?
  • Deliverable: are finished work products stored with dates and versions in shared drives, document systems or project tools?
  • Review: can you see what reviewers said, through redlines, approval steps, QA scorecards or client sign-off?
  • Outcome: is there a record of what followed, such as a renewal, a change order, a claim paid or a project closed?
  • Baseline: is it a US company with 50+ full-time employees at peak (contractors excluded), several years of documented operations, the right to license what it created, and an owner, CEO or CFO who can approve?

A company that ticks most boxes is worth running through the company fit checker, a preliminary, non-binding screen. If it looks promising, you make the introduction and the company works with SourceX on qualification, its data inventory, pricing and buyer review, as described in how it works. You never export, upload or describe the company's confidential files.

If a deal closes, partners earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals, capped at $100,000 per referred company, and the reward becomes payable only after the buyer pays and SourceX receives its fee. No reward is guaranteed.

What GDPval does not tell you

Treat GDPval as one well-designed measurement, not as a verdict on jobs or on what any company's records are worth.

  • Each task is a single request and a single deliverable. Real work involves clarifying questions, drafts, feedback and handoffs across days, which a one-shot task cannot capture.
  • Expert preference is a judgment. Two reviewers in the same occupation can disagree, especially on style.
  • It covers computer-based work. Physical, on-site and relationship-heavy parts of a role sit outside its scope.
  • Doing well on tasks drawn from a role is not the same as automating that role.
  • Benchmarks age, and once tasks are public they can leak into training data. Annual reviews such as Stanford HAI's AI Index, which covers technical performance alongside investment, adoption and policy, help track how benchmark results are read over time.

Next step

If you know a firm whose archive is full of requests, deliverables and reviews, register as a partner and introduce it, or send the owner to apply directly at sourcex.si/apply using your referral link.

  1. Step 1Share your linkSend your personal link to a company you know.
  2. Step 2Company appliesThe company applies itself at /apply.
  3. Step 3Buyer selects and paysThe buyer selects and pays for the data and SourceX receives its fee.
  4. Step 4You get your rewardYour share of SourceX fees becomes payable.

Common questions

Is GDPval a forecast of which jobs AI will replace?

No. It measures how model deliverables compare with expert deliverables on specific, well-defined tasks. A job is made of many tasks plus coordination, relationships and accountability that a one-shot evaluation does not test. Read GDPval as evidence of progress on certain kinds of knowledge work, not as an automation forecast for any role, firm or industry.

Is SourceX involved in GDPval or connected to the organization that released it?

No. GDPval is an independent public research release, and SourceX had no part in designing, building or scoring it. It is discussed here only because it shows, in public, the kind of real work product AI developers use to judge their models. Nothing on this page implies that any AI developer licenses data through SourceX.

Why do professional-services firms come up so often in GDPval discussions?

Because their output is already packaged as deliverables. Consulting reports, engineering calculations, accounting workpapers and legal memos each start with a request and end with a reviewed product, which mirrors how GDPval tasks are built. Firms that keep those requests, drafts and reviews for years hold a record of expert work that is hard to recreate from public sources.

Could company records like these be used for evaluation rather than training?

Yes, depending on the license. AI developers use records both to train models and to test them on held-out examples, and the permitted uses are set in the agreement the company negotiates and signs. The company approves the scope, the de-identification and redaction rules and the price before anything is delivered, so evaluation-only use is a question for those terms.

Where can I see the actual GDPval tasks and results?

The release included a research paper and an openly published sample of tasks, which are the best places to see exactly what is measured and how deliverables were graded. Go to the primary release rather than summaries, because task counts, occupation lists and model scores are updated as new models are evaluated, and secondary write-ups often lag behind.

Free resources

By SourceX Partnerships Team · Published 2026-10-09 · Updated 2026-10-09

Know a US company with valuable proprietary data?

Become a referral partner from anywhere we support, get your link and introduce an owner or authorized decision-maker.

Refer a company →

I own a business

Explore licensing your company's data to AI developers worldwide. Start a short assessment; no uploads needed.

Start an assessment