Vendor review · crowd annotation

Pangeanic

A Spanish language-technology veteran selling multilingual datasets, bespoke collection, and alignment services — with a sovereign-AI, keep-it-in-your-infrastructure bent.

The verdict. A solid partner for multilingual text and speech data, low-resource language coverage, and RLHF/evaluation work — especially for European buyers with sovereignty requirements; media rights depth should be confirmed per dataset.
  • Deep multilingual expertise - 20+ years in language services with strong low-resource and regional coverage
  • Sovereign-AI options - on-premises and air-gapped deployment for buyers with strict data-control requirements
  • Full-stack services - collection, annotation, evaluation, RLHF, and anonymization under one roof
  • Text-first - speech exists, but image/video and media licensing are not the core business
  • Mixed provenance - commissioned, legacy, and cleaned-open-data assets carry different rights profiles per dataset
  • No published pricing or catalog-level licence terms to compare in advance

What Pangeanic is

Pangeanic S.L. is a Spanish language-services and NLP company supplying AI training datasets (text, speech, image, video, documents, instruction and alignment data), custom collection, annotation via its PECAT platform, and its ECO language-AI platform; listed use cases include the Barcelona Supercomputing Center and Spain’s tax agency (AEAT).

Modalities and languages they cover

Strongest in multilingual text — parallel and monolingual corpora drawn from a repository the company puts at 10 billion aligned segments — plus speech/ASR/TTS data and regional programs (Arabic, Chinese, European, UK, African, Southeast Asian languages). Image, video, and document data are newer, thinner lines.

Sourcing model

A services mix rather than a marketplace: an in-house corpus built over two decades of translation work, native speakers recruited through ECO to write and record to spec, curated non-crawlable and cleaned open data, and bespoke collection projects.

Rights & provenance posture

Pangeanic markets datasets with full ownership and copyright and an ethical-AI framing, and its anonymization specialty helps on the privacy side. But sourcing spans commissioned writing, legacy translation corpora, and reworked open data — public materials do not document consent or licence chains per dataset, so provenance diligence lands in the contract.

Pricing

Not published. Off-the-shelf datasets and custom collection are both quoted per project.

Where Pangeanic falls behind

Text is the center of gravity. Per-speaker consent records, likeness handling, and audio/video licensing depth are not publicly documented the way the multilingual text offering is — and assets derived from open data and made IP-free deserve extra scrutiny before high-stakes training use.

Sources

Alternatives

Frequently asked questions

What data is Pangeanic best known for?

Multilingual text: parallel corpora for machine translation, monolingual LLM corpora, and instruction/alignment data, plus speech datasets for ASR and TTS — including regional programs for Arabic, Chinese, African, and Southeast Asian languages.

Does Pangeanic own the rights to its datasets?

The company describes its datasets as sold with full ownership and copyright and ethically sourced. Because sources range from commissioned writing to cleaned open data, confirm provenance and explicit AI-training terms per dataset.

Does Pangeanic do custom collection?

Yes — bespoke collection to language, dialect, domain, and demographic specs, with annotation, evaluation, and QA workflows layered on top.

Want data you can actually defend in diligence?

fiund licenses real-world audio and video at the source, with the rights cleared before anything moves.

Send a brief