Ethical, multilingual AI training data

An ethically-crawled corpus for Europe’s languages.

We believe AI models don’t need to be built on stolen data and don't have to cover mostly English.

We crawl the European web the right way — respecting robots.txt and AI opt-out protocols, skipping IP farms, identifying our crawler openly, and running on clean electricity.

Get in touch →

Built on four commitments

Every dataset we ship is judged against the same bar, regardless of language or source.

Ethical by design

Every crawl respects robots.txt, skips IP farms, and keeps request rates polite.

Multilingual, European

Coverage across the continent’s languages — EU and beyond — including the minority and regional languages LLMs usually miss.

Clean electricity

Our servers run in regions powered by low-carbon grids.

Research-grade

Documented, compliant, and legally sound — built for serious NLP work.

What “ethical” means, in practice

Not a slogan — a checklist we hold every crawl to.

  • Respects robots.txt and other protocols on every domain, every time
  • No anonymous user agents — we identify our crawler, always
  • No IP farms, no scraping shortcuts
  • Real politeness delays between requests
  • Infrastructure hosted in clean-electricity regions
  • Documented provenance for every source
  • European sovereignty — hosted and run on EU infrastructure

Get in touch

We’re at an early stage of our mission. If you are interested in our data or need a specific language or domain crawled, drop us an email and we’ll be in touch.

Email us
contact@ethicrawl.ai

Tell us a little about your training data needs — the languages, domains or volumes you’re after — and we’ll get back to you.