Vendor review · crowd annotation
Pangeanic
A Spanish language-technology veteran selling multilingual datasets, bespoke collection, and alignment services — with a sovereign-AI, keep-it-in-your-infrastructure bent.
- Deep multilingual expertise - 20+ years in language services with strong low-resource and regional coverage
- Sovereign-AI options - on-premises and air-gapped deployment for buyers with strict data-control requirements
- Full-stack services - collection, annotation, evaluation, RLHF, and anonymization under one roof
- Text-first - speech exists, but image/video and media licensing are not the core business
- Mixed provenance - commissioned, legacy, and cleaned-open-data assets carry different rights profiles per dataset
- No published pricing or catalog-level licence terms to compare in advance
What Pangeanic is
Pangeanic S.L. is a Spanish language-services and NLP company supplying AI training datasets (text, speech, image, video, documents, instruction and alignment data), custom collection, annotation via its PECAT platform, and its ECO language-AI platform; listed use cases include the Barcelona Supercomputing Center and Spain’s tax agency (AEAT).
Modalities and languages they cover
Strongest in multilingual text — parallel and monolingual corpora drawn from a repository the company puts at 10 billion aligned segments — plus speech/ASR/TTS data and regional programs (Arabic, Chinese, European, UK, African, Southeast Asian languages). Image, video, and document data are newer, thinner lines.
Sourcing model
A services mix rather than a marketplace: an in-house corpus built over two decades of translation work, native speakers recruited through ECO to write and record to spec, curated non-crawlable and cleaned open data, and bespoke collection projects.
Rights & provenance posture
Pangeanic markets datasets with full ownership and copyright and an ethical-AI framing, and its anonymization specialty helps on the privacy side. But sourcing spans commissioned writing, legacy translation corpora, and reworked open data — public materials do not document consent or licence chains per dataset, so provenance diligence lands in the contract.
Pricing
Not published. Off-the-shelf datasets and custom collection are both quoted per project.
Where Pangeanic falls behind
Text is the center of gravity. Per-speaker consent records, likeness handling, and audio/video licensing depth are not publicly documented the way the multilingual text offering is — and assets derived from open data and made IP-free deserve extra scrutiny before high-stakes training use.
Sources
Alternatives
Frequently asked questions
What data is Pangeanic best known for?
Multilingual text: parallel corpora for machine translation, monolingual LLM corpora, and instruction/alignment data, plus speech datasets for ASR and TTS — including regional programs for Arabic, Chinese, African, and Southeast Asian languages.
Does Pangeanic own the rights to its datasets?
The company describes its datasets as sold with full ownership and copyright and ethically sourced. Because sources range from commissioned writing to cleaned open data, confirm provenance and explicit AI-training terms per dataset.
Does Pangeanic do custom collection?
Yes — bespoke collection to language, dialect, domain, and demographic specs, with annotation, evaluation, and QA workflows layered on top.
Want data you can actually defend in diligence?
fiund licenses real-world audio and video at the source, with the rights cleared before anything moves.
Send a brief