Vendor review · crowd annotation
Nexdata
One of the largest Chinese AI-data vendors — a decade-plus-old catalog machine with 1,000+ off-the-shelf datasets and collection and annotation services at industrial scale.
- Massive off-the-shelf catalog - 1,000+ datasets, with company-cited holdings of ~3M speech hours and 800TB of image/video
- Language breadth - major-language speech corpora at the 100,000-hour scale plus many smaller languages and dialects
- Mature operations - over a decade of collection and annotation across automotive, home, retail, and GenAI use cases
- Compliance asserted, not shown - GDPR/CCPA/PIPL claims without public per-dataset consent or provenance records
- China-based supply chain - extra diligence for sovereignty-, export-, or biometric-sensitive buyers
- Catalog opacity - collection context and speaker consent terms are summarized in listings rather than documented
What Nexdata is
Nexdata is an AI training-data company founded in 2011 — public profiles describe it as the international brand of Beijing-based Datatang — selling 1,000+ ready-made datasets plus custom collection, annotation, and curation for automotive, smart-home, retail, conversational-AI, and GenAI customers.
Modalities and languages they cover
Speech, image, video, text, and point-cloud data. The company cites petabyte-scale LLM/GenAI data, around 3 million hours of speech, 800TB of image/video, and unsupervised speech pools it puts at 100,000+ hours per language across several major languages (English, French, Japanese, Korean, Arabic, German, Spanish).
Sourcing model
In-house and crowd collection at scale plus long-running annotation operations — a catalog-and-services model built up since 2011, not owner-licensed marketplace supply.
Rights & provenance posture
Nexdata states its datasets comply with GDPR, CCPA, and China’s PIPL and that data is collected with authorization. What is public is a company-level compliance claim rather than per-dataset consent documentation; provenance records, performer releases, and the Datatang corporate structure are the items Western legal teams typically verify during procurement.
Pricing
Mostly quote-based. Third-party marketplace listings on Datarade indicate roughly $10,000–$20,000 per off-the-shelf dataset purchase; Nexdata itself does not publish a price list.
Where Nexdata falls behind
Documentation depth and jurisdiction. Compliance is asserted at the company level rather than evidenced at the file level, and a China-domiciled supply chain adds review overhead for buyers with sovereignty, export, or biometric-consent obligations. This is volume sourcing, not consent-first licensing with paperwork that travels with the asset.
Sources
Alternatives
Frequently asked questions
Is Nexdata the same company as Datatang?
Public profiles describe Nexdata as the international brand of Beijing-based Datatang, with shared history back to 2011. Confirm the contracting entity and governing law when procuring.
How much do Nexdata datasets cost?
There is no public price list. Datarade listings indicate roughly $10,000–$20,000 per dataset purchase, with custom collection and annotation quoted per project.
Does Nexdata data comply with GDPR?
The company states its datasets comply with GDPR, CCPA, and PIPL. That is a vendor claim — ask for dataset-specific consent and authorization documentation to verify it for your use case.
Want data you can actually defend in diligence?
fiund licenses real-world audio and video at the source, with the rights cleared before anything moves.
Send a brief