arXiv:2511.01360cs.CL2025-11

构建首个印地语系癌症误解多语言评测集,测试大模型在跨语言下的推理鲁棒性。

Safer in Translation? Presupposition Robustness in Indic Languages

  • 将英文癌症误解数据集翻译成5种南亚语言,保持隐含预设一致性
  • 评估多个主流大模型在含虚假前提的医疗问题上的回答准确性
  • 针对低资源语言场景,揭示大模型跨语言推理的潜在风险

越来越多用户依赖大语言模型获取医疗建议,因此评估其响应的有效性与准确性至关重要。现有医学评测基准大多为英文,导致多语言LLM评估存在显著空白。本文提出Cancer-Myth-Indic,一个通过将500项英文癌症误解样本(均匀覆盖原类别)翻译至五种南亚地区常用但资源匮乏的语言(每语言500项,共2500项)构建的多语言评测集。母语译者遵循风格指南以保留翻译中的隐含预设,题目包含与癌症相关的错误预设。我们在此预设压力下评估多个流行LLM的表现,旨在填补多语言医疗LLM评估的空白。

原文摘要 · Abstract (English)

Increasingly, more and more people are turning to large language models (LLMs) for healthcare advice and consultation, making it important to gauge the efficacy and accuracy of the responses of LLMs to such queries. While there are pre-existing medical benchmarks literature which seeks to accomplish this very task, these benchmarks are almost universally in English, which has led to a notable gap in existing literature pertaining to multilingual LLM evaluation. Within this work, we seek to aid in addressing this gap with Cancer-Myth-Indic, an Indic language benchmark built by translating a 500-item subset of Cancer-Myth, sampled evenly across its original categories, into five under-served but widely used languages from the subcontinent (500 per language; 2,500 translated items total). Native-speaker translators followed a style guide for preserving implicit presuppositions in translation; items feature false presuppositions relating to cancer. We evaluate several popular LLMs under this presupposition stress.

多语言评估医疗LLM跨语言鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。