Operations and workflow data for AI training

Operations and workflow data records how work actually gets done: task histories with decisions and outcomes, field-service jobs, quality inspections, and recordings of people doing real work. AI labs license it to train agents and robots on real-world processes that public data barely covers.

Last updated October 3, 2026

Listings

DataTypical sourcesModalityAvailability
Task histories with context, actions and outcomes
Work items from request to completion, with each action taken and the result.
ServiceNow, Jira Service Management, SalesforceStructured data and textOn request
Field-service work orders with technician notes and photos
Jobs from dispatch to completion: diagnosis notes, parts used, photos and outcomes.
ServiceTitan, Salesforce Field Service, IFSStructured data, text and imagesOn request
Quality inspection records with photos and defect logs
Inspection checklists, defect findings, photos and corrective actions.
SafetyCulture, Procore, ETQ RelianceImages and structured dataOn request
Screen recordings of real software workflows
Recordings and event logs of people completing real tasks in business software.
Task-mining and session-recording tools a company already usesVideo and event logsOn request
Recordings of hands-on work
New video of skilled physical work, coordinated with partner businesses.
New recordings arranged for each projectVideoOn request

What’s included

  • Work items with context, actions, decisions and outcomes
  • Field-service jobs with technician notes, parts and photos
  • Inspection checklists, defect photos and corrective actions
  • Screen recordings and event logs of real software tasks
  • New video recordings of hands-on work, coordinated with partner businesses

Public datasets and licensed data

Public agent benchmarks are built on crowdworker demonstrations or simulated companies. They’re useful tests, but small: hundreds to a few thousand tasks. Licensed workflow data adds years of real task histories from operating businesses, with outcomes attached.

Public datasetReleasedWhat it containsLimits for enterprise useLicense
Mind2WebOhio State University, 20232,350 tasks on 137 real websites across 31 domains, with step-by-step demonstrations by crowdworkers.Consumer websites and crowdworker demonstrations, not company workflows.CC BY 4.0 (research use)
WebArenaCarnegie Mellon University, 2023812 multi-step tasks on self-hosted shopping, forum, GitLab and content-management sites.A benchmark environment with simulated sites and data.Apache 2.0
OSWorldHKU, Salesforce Research, CMU and Waterloo, 2024369 computer tasks on Ubuntu desktop and web applications, plus 43 Windows tasks.A test set of hundreds of tasks, not a record of real work.Apache 2.0
TheAgentCompanyCarnegie Mellon University, 2024175 tasks inside a simulated software company with its own code host, files, task tracker and chat.The company and its data are simulated.MIT

Preparation and rights

Preparation requirements are agreed with each supplier before any work begins. For this kind of data they typically include:

  • Pseudonymizing customer and employee identities and redacting free-text notes
  • Getting consent from people who appear in recordings, under the company’s policies and local law
  • Blurring or removing faces, addresses and documents visible in photos and video

Every dataset has a documented owner and confirmed licensing rights, and is delivered only after an executed agreement and the supplier’s authorization. See data governance on sourcex.si.

Use cases

More on sourcex.si

Questions

What is workflow data for AI agents?

Records of how real work moves from request to completion: who did what, in which system, with what information, and how it turned out. Agents learn multi-step processes from it, and outcomes give a way to score them.

Can SourceX arrange new recordings?

New recordings of hands-on work can be coordinated with partner businesses. Describe the tasks and settings you need in your request.

Need operations and workflow data?

Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.