测试大模型在不同表达方式下是否保持医疗信息一致,发现低学历用户提问常被简化遗漏关键内容。
MIRA: A Bilingual Benchmark for Medical Information Response Audit

- 构建双语可控基准MIRA,评估模型对不同语言、语气、文凭水平提问的响应一致性
- 5个主流模型对4320个问题均作答,但低学历表达导致关键信息缺失率上升
- 提出'信息稀释'现象,适合医疗AI安全评估与模型优化研究者参考
大型语言模型(LLMs)越来越多用于提供面向公众的健康信息,但现有安全评估忽略了不同用户表述相同问题时,模型是否能保持医疗信息的一致性。为此,我们提出医学信息响应审计(MIRA),一个双语、受控的基准测试,用于评估模型在用户侧语言、语体和健康素养信号变化下的响应一致性。MIRA包含4,320个由60个医学审核过的低风险健康问题生成的提示。在五个主流LLMs中,模型全部回答了所有医学问题,但针对低健康素养信号的回应普遍遗漏更多关键信息,提供的具体下一步行动更少,且支持独立判断的程度更低。我们称此现象为‘差异性信息稀释’(DID)。语言影响具有模型特异性,并非非英语提示统一更差。与300个真实健康查询对比显示初步排序有效性。知识引导的缓解提示可降低多数模型的信息稀释,其中Claude(约8%)和Qwen(约6%)在减少信息不足简化方面改善最显著。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). Language effects are model-specific rather than uniformly worse for non-English prompts. A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigation prompt reduces information dilution for most models, with the largest reductions in underinformative simplification observed for Claude (~8%) and Qwen (~6%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。