XTREME-S: Evaluating Cross-lingual Speech Representations

Alexis Conneau,Ankur Bapna,Yu Zhang,Min Ma,Patrick von Platen,Anton Lozhkov,Colin Cherry,Ye Jia,Clara Rivera,Mihir Kale,Daan Van Esch,Vera Axelrod,Simran Khanuja,Jonathan H. Clark,Orhan Firat,Sebastian Ruder,Jason Riesa,Melvin Johnson

Conference of the International Speech Communication Association (INTERSPEECH)（2022）

引用 6|浏览118

暂无评分

摘要

We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval. Covering 102 languages from 10+ language families, 3 different domains and 4 task families, XTREME-S aims to simplify multilingual speech representation evaluation, as well as catalyze research in "universal" speech representation learning. This paper describes the new benchmark and establishes the first speech-only and speech-text baselines using XLS-R and mSLAM on all downstream tasks. We motivate the design choices and detail how to use the benchmark. Datasets and fine-tuning scripts are made easily accessible.

查看译文

关键词

representations,speech,cross-lingual

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要