An ethically-crawled corpus for Europe’s languages.
We believe AI models don’t need to be built on stolen data and don't have to cover mostly English.
We crawl the European web the right way — respecting robots.txt and AI opt-out protocols, skipping IP farms, identifying our crawler openly, and running on clean electricity.
Get in touch →Built on four commitments
Every dataset we ship is judged against the same bar, regardless of language or source.
Ethical by design
Every crawl respects robots.txt, skips IP farms, and keeps request rates polite.
Multilingual, European
Coverage across the continent’s languages — EU and beyond — including the minority and regional languages LLMs usually miss.
Clean electricity
Our servers run in regions powered by low-carbon grids.
Research-grade
Documented, compliant, and legally sound — built for serious NLP work.
What “ethical” means, in practice
Not a slogan — a checklist we hold every crawl to.
- ✓Respects robots.txt and other protocols on every domain, every time
- ✓No anonymous user agents — we identify our crawler, always
- ✓ No IP farms, no scraping shortcuts
- ✓ Real politeness delays between requests
- ✓ Infrastructure hosted in clean-electricity regions
- ✓ Documented provenance for every source
- ✓ European sovereignty — hosted and run on EU infrastructure
Get in touch
We’re at an early stage of our mission. If you are interested in our data or need a specific language or domain crawled, drop us an email and we’ll be in touch.
Tell us a little about your training data needs — the languages, domains or volumes you’re after — and we’ll get back to you.