The concept
RLHF — reinforcement learning from human feedback — improves models by showing them human judgments: which of two answers is better, whether an output meets a standard, how an expert would respond.
Why it is expensive
- It requires skilled human judgment, not just raw text.
- It must be produced deliberately — it cannot be scraped.
- Quality control is intensive.
Industry reporting documents AI developers spending heavily on exactly this kind of human judgment data (Reuters).
The company-data connection
Inside established companies, judgment is recorded constantly: which response resolved the ticket, which approach won the deal, which decision passed review. That recorded judgment — properly anonymized — is valuable preference data.
What this means for partners
Companies with 50+ full-time employees at peak generate recorded judgment daily. Introduce one to SourceX; if a deal closes and SourceX collects its fee, you earn 25% of that fee, up to $100,000 per company.
Refer a company · How rewards work
Common questions
What is RLHF data?
RLHF data involves human judgments about which AI model outputs are better, whether they meet standards, or how an expert would respond. This feedback helps improve AI models through reinforcement learning.
Why is RLHF data expensive?
RLHF data is expensive because it requires skilled human judgment, cannot be scraped, and needs intensive quality control. AI developers are known to spend heavily on this kind of human judgment data.
What kind of companies generate valuable preference data?
Established companies generate valuable preference data through their daily operations. This includes recorded judgments like which customer response resolved a ticket or which approach won a deal.
How do partners earn rewards from SourceX?
Partners can earn 25% of SourceX's collected fee, up to $100,000 per company, by introducing companies with 50+ full-time employees at peak. Rewards are paid if a deal closes and SourceX collects its fee.