测试大模型在危机翻译中保持紧急程度的能力,发现其表现不稳定且低资源语言更差。
LLM-Powered Automatic Translation and Urgency in Crisis Scenarios
- 用多语言危机数据和新标注的紧急度数据集评估翻译性能
- 低资源语言翻译质量下降明显,紧急程度判断差异可达从不紧急到关键级
- 人类评估者对紧急度判断一致,而大模型结果随语言变化剧烈
大型语言模型(LLMs)在危机准备与响应中的多语言沟通应用日益增多,但其在高风险场景下的适用性仍缺乏充分评估。本研究考察了先进机器翻译系统与大模型在危机领域翻译中的表现,重点关注紧急程度的保留——这是有效危机沟通与分级的关键属性。基于多语言危机数据集(TICO-19,30种语言)及新构建的100个场景、29种语言的紧急度标注数据集,我们发现专用翻译模型与大模型均出现显著质量下降,尤其在低资源语言中更为严重。除翻译质量外,通过人工标注研究发现显著不对称:人类评估者对相同场景的紧急度判断不受提示语言影响,保持一致;而基于大模型的紧急度分类结果在不同语言间差异巨大,有时覆盖从‘不紧急’到‘关键’的全范围。这些结果揭示了将通用语言技术用于危机分级的重大风险,并强调需要建立多语言、以人为本的评估框架。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. However, their suitability for high-stakes crisis contexts remains insufficiently evaluated. This work examines the performance of state-of-the-art LLMs and machine translation systems in crisis-domain translation, with a focus on preserving urgency, a critical property for effective crisis communication and triage. Using multilingual crisis data (TICO-19, 30 languages) and a newly introduced urgency-annotated dataset of 100 scenarios translated into 29 languages, we show that dedicated translation models and LLMs exhibit substantial quality degradation, particularly for low-resource languages. Beyond translation quality, we conduct a human annotation study revealing a striking asymmetry: human assessors maintain consistent urgency judgments regardless of prompt language, while LLM-based urgency classifications vary widely across languages for identical scenarios, at times spanning the full range from Not Urgent to Critical. These findings highlight significant risks in deploying general-purpose language technologies for crisis triage and underscore the need for multilingual, human-centered evaluation frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。