Internal wikis and knowledge bases
Internal wikis and knowledge bases, such as Confluence spaces, Notion workspaces and internal help centers, with page hierarchy, links, comments and edit history. Labs use them to train and test retrieval and question-answering systems on the messy, interlinked knowledge real companies keep.
What a record contains
One page: content, its place in the space hierarchy, inbound and outbound links, comments and edit history.
| Field | What it holds |
|---|---|
page_id, space, parent_id | Where the page sits |
title, body | Content |
links_out[] | Pages it links to |
comments[] | Discussion on the page |
edits[] | Author role and date for each edit |
{
"page_id": "WK-3381",
"space": "Payments",
"title": "How settlement retries work",
"links_out": ["WK-3379", "WK-2210"],
"edits": [{"by": "eng_role", "at": "2022-04-11"}, {"by": "eng_role", "at": "2024-11-05"}],
"comments": [{"at": "2024-06-02", "body": "Is this still true after the queue migration?"}]
}How AI labs use it
- Enterprise retrieval and RAG evaluation
- Real questions over interlinked, partly outdated pages.
- Question answering
- Pair pages with the support or chat questions they answered.
- Knowledge maintenance
- Spot stale pages using edit history.
Typical preparation requirements
Agreed with the supplier before any work begins. Typical requirements include:
- HR, legal and security spaces excluded or reviewed
- Personal pages excluded
- Customer and partner names pseudonymized
Every dataset has a documented owner and confirmed licensing rights. See data governance on sourcex.si.
What makes a strong package
- Linked to tickets or chat questions
- Long edit histories
- Decision records
Compared with public datasets
Public sets such as TechQA and doc2dial are useful references, but limited as enterprise training data. The documents and knowledge category page compares them with licensed data.
Who typically holds it
- Software companies
- Consulting firms
- Any company with a mature internal wiki
Know a company like this?
Introduce the company to SourceX. If its data deal closes, you can earn up to $100,000 in referral fees, paid after the buyer accepts the data and SourceX receives payment.
Refer a companyQuestions
Synthetic enterprise corpora exist. Why license real wikis?
Synthetic corpora are useful for testing, but real wikis carry the contradictions, stale pages and shorthand that make enterprise retrieval hard.
Need this data for a model?
Describe what you need: domain, volume, history, format and licensing terms. SourceX looks for companies that hold it and can license it.