arXiv:2512.20204cs.CLcs.AI2025-12

构建跨语言会议语料库,自动检测对话中的误解。

Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings

  • 创建5小时跨语言对话数据集,含原语语音与英译文。
  • 用大模型检测误解,召回率达77%,准确率47%。
  • 适合研究多语言会议、自动翻译与误判识别的学者。

语音处理与翻译技术有望帮助无共同语言的人员开展会议交流。为评估此类系统的性能,亟需一个多样化且真实的评估语料库。为此,我们构建并发布了跨语言对话语料库,涵盖无共同语言者在自动同声翻译支持下的真实会议对话。语料库包含5小时原始语言语音录音,对应12种语言的自动语音识别(ASR)结果与人工转录文本,以及自动翻译和人工校正后的英文译文。为支持跨语言摘要研究,语料库还包含会议纪要(minutes)。此外,本文提出自动误解检测任务,通过人工标注量化跨语言会议中的误解情况,并测试当前大型语言模型的自动检测能力。结果显示,Gemini模型能以77%的召回率和47%的准确率识别出存在误解的文本片段。

原文摘要 · Abstract (English)

Speech processing and translation technology have the potential to facilitate meetings of individuals who do not share any common language. To evaluate automatic systems for such a task, a versatile and realistic evaluation corpus is needed. Therefore, we create and present a corpus of cross-lingual dialogues between individuals without a common language who were facilitated by automatic simultaneous speech translation. The corpus consists of 5 hours of speech recordings with ASR and gold transcripts in 12 original languages and automatic and corrected translations into English. For the purposes of research into cross-lingual summarization, our corpus also includes written summaries (minutes) of the meetings. Moreover, we propose automatic detection of misunderstandings. For an overview of this task and its complexity, we attempt to quantify misunderstandings in cross-lingual meetings. We annotate misunderstandings manually and also test the ability of current large language models to detect them automatically. The results show that the Gemini model is able to identify text spans with misunderstandings with recall of 77% and precision of 47%.

跨语言会议纪要误解检测语音翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。