AI健康建议准确性受语言、话题和信息源影响,英语表现好不代表全局可靠。
How much does context affect the accuracy of AI health advice?
- 用7个主流大模型测试跨语言、跨话题的健康声明验证能力
- 英语及欧洲语言准确率最高,非欧语言显著下降,科学文献类最差
- 新冠和政府发布的内容更易判断,适合多语种、细分领域评估
大型语言模型(LLMs)在提供健康建议方面日益普及,但其在不同语言、主题和信息来源下的准确性证据仍有限。我们评估了七个广泛使用的LLMs在两个数据集上的表现:(i) 1,975条英国与欧盟监管机构认证的营养与健康声明,翻译成21种语言;(ii) 9,088条记者审核过的公共卫生声明,涵盖新冠疫情、堕胎、政治与一般健康,来自政府通告、科研摘要与媒体。模型通过多次运行的多数投票对每条声明进行支持或不支持分类。按语言、主题、来源和模型分析准确率。认证声明在英语及相近欧洲语言中准确率最高,随与英语句法距离增加而下降。真实世界公共卫生声明的准确率明显更低,且系统性地随主题和来源变化:对新冠疫情和政府归因声明表现最好,对一般健康和科研摘要最差。英语下高准确率掩盖了明显的上下文依赖差距。训练数据暴露、编辑框架与主题特化调优可能是差异成因,其幅度与跨语言差异相当。大模型在健康声明验证中的准确性强烈依赖于语言、主题与信息源。英语表现无法可靠泛化到其他情境,部署前需开展多语言、领域特定评估。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to provide health advice, yet evidence on how their accuracy varies across languages, topics and information sources remains limited. We assess how linguistic and contextual factors affect the accuracy of AI-based health-claim verification. We evaluated seven widely used LLMs on two datasets: (i) 1,975 legally authorised nutrition and health claims from UK and EU regulatory registers translated into 21 languages; and (ii) 9,088 journalist-vetted public-health claims from the PUBHEALTH corpus spanning COVID-19, abortion, politics and general health, drawn from government advisories, scientific abstracts and media sources. Models classified each claim as supported or unsupported using majority voting across repeated runs. Accuracy was analysed by language, topic, source and model. Accuracy on authorised claims was highest in English and closely related European languages and declined in several widely spoken non-European languages, decreasing with syntactic distance from English. On real-world public-health claims, accuracy was substantially lower and varied systematically by topic and source. Models performed best on COVID-19 and government-attributed claims and worst on general health and scientific abstracts. High performance on English, canonical health claims masks substantial context-dependent gaps. Differences in training data exposure, editorial framing and topic-specific tuning likely contribute to these disparities, which are comparable in magnitude to cross-language differences. LLM accuracy in health-claim verification depends strongly on language, topic and information source. English-language performance does not reliably generalise across contexts, underscoring the need for multilingual, domain-specific evaluation before deployment in public-health communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。