Company documents and knowledge for AI training
Company documents and knowledge bases hold how an organization actually works: SOPs, runbooks, wikis, and documents with their revision history. AI labs license them to train agents that read, follow and update real internal documentation.
Listings
| Data | Typical sources | Modality | Availability |
|---|---|---|---|
| SOPs, runbooks and playbooks | Confluence, Notion, SharePoint | Documents | On request |
| Internal wikis and knowledge bases | Confluence, Notion, SharePoint | Documents | On request |
| Documents, spreadsheets and decks with revision history | Google Drive, SharePoint, OneDrive | Documents and spreadsheets | On request |
What’s included
- SOPs, runbooks and playbooks with versions and owners
- Internal wikis and knowledge bases with links and edit history
- Documents, spreadsheets and decks with revisions and comments
Public datasets and licensed data
There’s no widely used public collection of real internal SOPs or wikis. Public sets cover vendor technical notes, government service pages or scanned documents from decades-old litigation archives. Licensed archives add living internal documentation with authorship and revision history. Company archives commonly span 3–15+ years of files.
| Public dataset | Released | What it contains | Limits for enterprise use | License |
|---|---|---|---|---|
| TechQA | IBM, 2020 | 1,400 real forum questions answered from 801,998 public IBM technical notes. | Public vendor documentation, not internal procedures. | See source |
| doc2dial | IBM, 2020 | 4,500+ dialogues grounded in 450+ public US government service documents. | Public government pages, not company documentation. | CC BY 3.0 |
| RVL-CDIP | Ryerson University, 2015 | 400,000 scanned business-document images in 16 classes, from the Legacy Tobacco Document Library. | Old scanned images labeled by document type only. | See source |
Preparation and rights
Preparation requirements are agreed with each supplier before any work begins. For this kind of data they typically include:
- Scoping files by folder, space or type before any review starts
- Excluding personal, HR, legal and security-sensitive material
- Excluding third-party documents, such as vendor manuals and purchased reports, unless they can be licensed
Every dataset has a documented owner and confirmed licensing rights, and is delivered only after an executed agreement and the supplier’s authorization. See data governance on sourcex.si.
Use cases
More on sourcex.si
Questions
Why license real SOPs instead of generating synthetic ones?
Real documentation carries the contradictions, outdated pages and shorthand that make enterprise work hard. Synthetic corpora are useful for testing, but they don’t reflect how procedures drift over years.
Which systems do documents come from?
Confluence, Notion, SharePoint, Google Drive, OneDrive, Dropbox and Box are the most common.
Need company documents and knowledge?
Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.