Building an Italian-Chinese Parallel Corpus for Machine Translation from the Web

Rita Tse, Silvia Mirri, Su Kit Tang, Giovanni Pau, Paola Salomoni

研究成果: Conference contribution同行評審

15 引文 斯高帕斯(Scopus)

摘要

In an increasingly globalized world, being able to understand texts in different languages (even more so in different alphabets and charsets) has become a necessity. This can be strategic even while moving and travelling across different countries, characterized by different languages. With this in mind, bilingual corpora become critical resources since they are the basis of every state-of-the-art automatic translation system; moreover, building a parallel corpus is usually a complex and very expensive operation. This paper describes an innovative approach we have defined and adopted to automatically build an Italian-Chinese parallel corpus, with the aim of using it for training an Italian-Chinese Neural Machine Translation. Our main idea is to scrape parallel texts from the Web: we defined a general pipeline, describing each specific step from the selection of the appropriate data sources to the sentence alignment method. A final evaluation was conducted to evaluate the goodness of our approach and its results show that 90% of the sentences were correctly aligned. The corpus we have obtained consists of more than 6,000 sentence pairs (Italian and Chinese), which are the basis for building a Machine Translation system.

原文English
主出版物標題Proceedings of the 6th EAI International Conference on Smart Objects and Technologies for Social Good, GOODTECHS 2020
發行者Association for Computing Machinery
頁面265-268
頁數4
ISBN(電子)9781450375597
DOIs
出版狀態Published - 14 9月 2020
事件6th EAI International Conference on Smart Objects and Technologies for Social Good, GOODTECHS 2020 - Virtual, Online, Mexico
持續時間: 14 9月 202016 9月 2020

出版系列

名字ACM International Conference Proceeding Series

Conference

Conference6th EAI International Conference on Smart Objects and Technologies for Social Good, GOODTECHS 2020
國家/地區Mexico
城市Virtual, Online
期間14/09/2016/09/20

指紋

深入研究「Building an Italian-Chinese Parallel Corpus for Machine Translation from the Web」主題。共同形成了獨特的指紋。

引用此