arXiv:2508.02555cs.CL2025-08被引 1

构建多语言主题对齐语料库,提升无平行文本场景下的跨语言理解能力。

Building and Aligning Comparable Corpora

  • 基于维基百科与欧新社数据,构建英法阿三语可比语料库。
  • 跨语言LSI相似度方法在文档对齐上优于词典匹配法。
  • 可实现事件级与主题级对齐,适合新闻与多语言信息挖掘。

可比语料库是由多语言中主题对齐但非翻译关系的文档构成,适用于缺乏平行语料的领域或语言。本文利用英文、法文和阿拉伯文的维基百科及欧新社网站数据构建可比语料库,并采用两种跨语言相似度度量方法进行自动对齐:基于双语词典的方法与基于潜在语义索引(LSI)的方法。实验表明,跨语言LSI(CL-LSI)优于词典方法。进一步,从英国广播公司(BBC)和半岛电视台(JSC)分别收集英阿新闻文本,使用CL-LSI进行对齐。评估显示,该方法不仅能实现跨语言文档在主题层面的对齐,还能在事件层面完成精准匹配。

原文摘要 · Abstract (English)

Comparable corpus is a set of topic aligned documents in multiple languages, which are not necessarily translations of each other. These documents are useful for multilingual natural language processing when there is no parallel text available in some domains or languages. In addition, comparable documents are informative because they can tell what is being said about a topic in different languages. In this paper, we present a method to build comparable corpora from Wikipedia encyclopedia and EURONEWS website in English, French and Arabic languages. We further experiment a method to automatically align comparable documents using cross-lingual similarity measures. We investigate two cross-lingual similarity measures to align comparable documents. The first measure is based on bilingual dictionary, and the second measure is based on Latent Semantic Indexing (LSI). Experiments on several corpora show that the Cross-Lingual LSI (CL-LSI) measure outperforms the dictionary based measure. Finally, we collect English and Arabic news documents from the British Broadcast Corporation (BBC) and from ALJAZEERA (JSC) news website respectively. Then we use the CL-LSI similarity measure to automatically align comparable documents of BBC and JSC. The evaluation of the alignment shows that CL-LSI is not only able to align cross-lingual documents at the topic level, but also it is able to do this at the event level.

可比语料跨语言对齐新闻分析多语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。