Abstract
We describe the progress of the High Performance Language Technologies (HPLT) project, a 3-year EU-funded project that started in September 2022. We focus on the up-to-date results on the release of free text datasets derived from web crawls, one of the central objectives of the project. The second release used a revised processing pipeline, and an enlarged set of input crawls. From 4.5 petabytes of web crawls we extracted 7.6T tokens of monolingual text in 193 languages, plus 380 million parallel sentences in 51 language pairs. We also release MultiHPLT, a cross-combination of the parallel data, which produces 1,275 pairs, as well as releasing the containing documents for all parallel sentences in order to enable research in document-level MT. We report changes in the pipeline, analysis and evaluation results for the second parallel data release based on machine translation systems. All datasets are released under a permissive CC0 licence.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of Machine Translation Summit XX: Volume 2 |
| Editors | Pierrette Bouillon, Johanna Gerlach, Sabrina Girletti, Lise Volkart, Raphael Rubino, Rico Sennrich, Samuel Läubli, Martin Volk, Miquel Esplà-Gomis, Vincent Vandeghinste, Helena Moniz, Sara Szoc |
| Place of Publication | Geneva, Switzerland |
| Publisher | European Association for Machine Translation (EAMT) |
| Pages | 101-102 |
| Number of pages | 2 |
| ISBN (Print) | 9782970189718 |
| Publication status | Published - 01 Jun 2025 |
| Externally published | Yes |
| Event | Machine Translation Summit XX - Geneva, Switzerland Duration: 23 Jun 2025 → 27 Jun 2025 |
Conference
| Conference | Machine Translation Summit XX |
|---|---|
| Country/Territory | Switzerland |
| City | Geneva |
| Period | 23/06/2025 → 27/06/2025 |
Fingerprint
Dive into the research topics of 'HPLT's second data release'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver