Skip to main navigation Skip to search Skip to main content

Findings of the WMT25 multilingual instruction shared task: persistent hurdles in reasoning, generation, and evaluation

  • Tom Kocmi
  • , Ekaterina Artemova
  • , Eleftherios Avramidis
  • , Eleftheria Briakou
  • , Pinzhen Chen
  • , Marzieh Fadaee
  • , Markus Freitag
  • , Roman Grundkiewicz
  • , Yupeng Hou
  • , Philipp Koehn
  • , Julia Kreutzer
  • , Saab Mansour
  • , Stefano Perrella
  • , Lorenzo Proietti
  • , Parker Riley
  • , Eduardo Sánchez
  • , Patrícia Schmidtová
  • , Mariya Shmatova
  • , Vilém Zouhar

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

The WMT25 Multilingual Instruction Shared Task (MIST) introduces a benchmark to evaluate large language models (LLMs) across 30 languages. The benchmark covers five types of problems: machine translation, linguistic reasoning, open-ended generation, cross-lingual summarization, and LLM-as-a-judge.We provide automatic evaluation and collect human annotations, which highlight the limitations of automatic evaluation and allow further research into metric meta-evaluation. We run on our benchmark a diverse set of open- and closed-weight LLMs, providing a broad assessment of the multilingual capabilities of current LLMs. Results highlight substantial variation across sub-tasks and languages, revealing persistent challenges in reasoning, cross-lingual generation, and evaluation reliability. This work establishes a standardized framework for measuring future progress in multilingual LLM development.
Original languageEnglish
Title of host publicationProceedings of the Tenth Conference on Machine Translation
EditorsBarry Haddow, Tom Kocmi, Philipp Koehn, Christof Monz
Place of PublicationSuzhou, China
PublisherAssociation for Computational Linguistics
Pages414-435
Number of pages22
ISBN (Print) 9798891763418
Publication statusPublished - 01 Nov 2025
Externally publishedYes

Fingerprint

Dive into the research topics of 'Findings of the WMT25 multilingual instruction shared task: persistent hurdles in reasoning, generation, and evaluation'. Together they form a unique fingerprint.

Cite this