Fractional CTOs, consultants, and strategic advisors are uniquely positioned to identify valuable data assets within companies. SourceX simplifies the process of licensing this data for AI applications. To help you prepare US companies with relevant data for an introduction, we've developed a metadata-only GitHub inventory template.
The Need for a Metadata-Only Approach
When introducing a company to SourceX for data licensing, the initial focus is on understanding what kind of data they possess, how much of it, and from what periods. It's critical that this preliminary assessment happens without any actual data leaving the company's control or passing through a partner's hands. This template facilitates that process, allowing you to highlight potential value without any risk of data exposure.
Partner rewards are a share of SourceX's fee and are never deducted from what the company receives (`F07`). The partner reward becomes payable only after the buyer pays and SourceX receives its fee (`F03`).
How the Template Works
This GitHub inventory template is designed to be populated with metadata only. This means listing categories of records, date ranges, and potential volume indicators, but never actual data points, personally identifiable information, or confidential company specifics. It's a structured way to catalog what data could be licensed, not the data itself.
Partner Role and Data Handling
Crucially, as a SourceX referral partner, you never handle the company's data. Your role is to identify and introduce US companies that meet the criteria (`F06`), facilitating their connection to SourceX (`F04`). SourceX then manages the entire data licensing transaction, from sourcing and rights review to delivery and payment (`F09`). We do not train AI models (`F10`).
Who Qualifies for SourceX
For a successful introduction, target US companies that meet the following baseline criteria (`F06`):
- Size: 50+ full-time employees at peak (contractors excluded).
- History: Years of operating records.
- Rights: Rights to license the data.
- Sponsor: An authorized sponsor within the company.
GitHub Metadata Inventory Template
Here’s a basic structure you can use within a GitHub repository (e.g., as a Markdown file `DATA_INVENTORY.md`). Companies can adapt this to their specific internal structure.
```markdown
SourceX Data Licensing Metadata Inventory
This document outlines metadata for potential data licensing opportunities with SourceX. It contains NO actual data, only descriptions of data types and their characteristics.
Overview
- Company Name: [Company Name]
- Primary Business: [Brief description of primary business operations]
- Contact for Licensing Discussions: [Name/Role - This would be for internal company use until ready for SourceX]
Data Record Categories
Category 1: [e.g., Customer Interaction Records]
- Description: [Briefly describe the type of data, e.g., anonymized customer service interactions, support tickets, chat logs]
- Data Types Included: [e.g., Text (conversations), Timestamps, Interaction IDs]
- Date Range Available: [e.g., January 2018 - Present]
- Estimated Volume/Frequency: [e.g., ~1 million records/month, TBs per year, daily feeds]
- Potential Exclusions/Sensitivities: [e.g., PII removed, sensitive topics filtered]
- Current Storage/Format: [e.g., Cloud database, CSVs, JSON, internal API]
- Potential Use Cases: [e.g., Sentiment analysis, trend identification, AI model training]
Category 2: [e.g., Operational Sensor Data]
- Description: [e.g., Telemetry from IoT devices, equipment performance logs]
- Data Types Included: [e.g., Numeric (readings), Timestamps, Device IDs, Geolocation (if non-sensitive)]
- Date Range Available: [e.g., Q3 2020 - Present]
- Estimated Volume/Frequency: [e.g., GBS per day, millions of data points hourly]
- Potential Exclusions/Sensitivities: [e.g., Proprietary algorithms removed]
- Current Storage/Format: [e.g., Time-series database, streaming data]
- Potential Use Cases: [e.g., Predictive maintenance, anomaly detection]
Category 3: [Add more categories as needed]
- ...
Data Rights & Compliance Considerations
- Source of Data: [e.g., Directly collected, aggregated from partners]
- Licensing Rights: [Confirm internal assessment of rights to license this data for commercial purposes]
- Anonymization/Pseudonymization Status: [Describe current state or plans]
- Relevant Regulations: [e.g., GDPR, CCPA, HIPAA - Note: SourceX handles specific compliance aspects but initial awareness is helpful]
Next Steps (Internal)
- Review with legal counsel for licensing readiness.
- Identify authorized sponsor for SourceX introduction.
- Prepare for initial metadata discussion with SourceX.
```
Partner Rewards
As a partner, you earn 25% of the eligible platform fees SourceX actually collects from the referred company's licensing deals (`F01`). Rewards are capped at $100,000 cumulative per referred company (`F02`). Anyone can join, from any supported country (`F05`). Licensed professionals should always check their own rules on referral fees.
This metadata-only approach allows for efficient, secure initial qualification, focusing on the potential of the data without compromising confidentiality. Once the company is ready, you can make the introduction to SourceX.