用小语料+领域适配,让大模型输出更易懂的灾难应急英文。
Translating Under Pressure: Domain-Aware LLMs for Crisis Communication

- 从通用语料中筛选过滤,扩充小规模灾难题材双语数据
- 微调小模型并优化输出,使英文达到CEFR A2级可读性
- 适合资源有限但需快速多语种应急沟通的场景
灾荒与人为灾害期间,及时可靠的多语言沟通至关重要,但高质量平行语料稀缺限制了有效解决方案的发展。本文提出一种领域自适应流程:通过从通用语料库中检索并筛选,扩展小规模参考语料库;利用所得数据微调小型语言模型进行灾难题域翻译,并采用偏好优化使输出偏向CEFR A2级英语。自动与人工评估均表明该方法在保持较强语义准确性的前提下,显著提升可读性。结果表明,简化英语结合领域适配,可在无法实现全多语覆盖时,作为实际可用的应急沟通通用语。
原文摘要 · Abstract (English)
Timely and reliable multilingual communication is critical during natural and human-induced disasters, but developing effective solutions for crisis communication is limited by the scarcity of curated parallel data. We propose a domain-adaptive pipeline that expands a small reference corpus, by retrieving and filtering data from general corpora. We use the resulting dataset to fine-tune a small language model for crisis-domain translation and then apply preference optimization to bias outputs toward CEFR A2-level English. Automatic and human evaluation shows that this approach improves readability, while maintaining strong adequacy. Our results indicate that simplified English, combined with domain adaptation, can function as a practical lingua franca for emergency communication when full multilingual coverage is not feasible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。