TANDO: A Corpus for Document-level Machine Translation.

Harritxu Gete,Thierry Etchegoyhen,David Ponce,Gorka Labaka,Nora Aranberri,Ander Corral,Xabier Saralegi,Igor Ellakuria,Maite Martín

International Conference on Language Resources and Evaluation (LREC)（2022）

引用 0|浏览26

暂无评分

摘要

Document-level Neural Machine Translation aims to increase the quality of neural translation models by taking into account contextual information. Properly modelling information beyond the sentence level can result in improved machine translation output in terms of coherence, cohesion and consistency. Suitable corpora for context-level modelling are necessary to both train and evaluate context-aware systems, but are still relatively scarce. In this work we describe TANDO, a document-level corpus for the under-resourced Basque-Spanish language pair, which we share with the scientific community. The corpus is composed of parallel data from three different domains and has been prepared with context-level information. Additionally, the corpus includes contrastive test sets for fine-grained evaluations of gender and register contextual phenomena on both source and target language sides. To establish the usefulness of the corpus, we trained and evaluated baseline Transformer models and context-aware variants based on context concatenation. Our results indicate that the corpus is suitable for fine-grained evaluation of document-level machine translation systems.

查看译文

关键词

Document-level Machine Translation, Parallel Corpus, Contrastive tests, Basque, Spanish

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要