评估24个大模型在多语种医疗聊天场景下的表现,发现本土模型不总靠谱。
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
- 用统一检索增强生成框架测试24个模型在真实医患对话中的表现
- 印地语等本地语言回答的准确性普遍低于英语,平均低18%
- 混合语和文化相关问题让模型更难应对,适合医疗AI评估研究者参考
大型语言模型(LLMs)的能力评估日益受到关注,但真实场景下的多模型比较仍较少。现有跨语言评估常依赖翻译基准,难以体现源语言的语言与文化特征。本研究基于印度患者与医疗聊天机器人的真实交互数据,在印度英语及四种印地语系语言中对24个LLM进行了全面评估。采用统一的检索增强生成(RAG)框架生成回复,并通过自动化指标与人工评估在四个应用相关维度上进行评测。结果表明模型性能差异显著;指令微调的印地语系模型在印地语查询上表现并不总是优于其他模型。进一步实证显示,针对印地语查询的回复事实正确性普遍低于英语查询。定性分析还发现,数据集中混合语表达与文化相关问题给模型带来显著挑战。
原文摘要 · Abstract (English)
Assessing the capabilities and limitations of large language models (LLMs) has garnered significant interest, yet the evaluation of multiple models in real-world scenarios remains rare. Multilingual evaluation often relies on translated benchmarks, which typically do not capture linguistic and cultural nuances present in the source language. This study provides an extensive assessment of 24 LLMs on real world data collected from Indian patients interacting with a medical chatbot in Indian English and 4 other Indic languages. We employ a uniform Retrieval Augmented Generation framework to generate responses, which are evaluated using both automated techniques and human evaluators on four specific metrics relevant to our application. We find that models vary significantly in their performance and that instruction tuned Indic models do not always perform well on Indic language queries. Further, we empirically show that factual correctness is generally lower for responses to Indic queries compared to English queries. Finally, our qualitative work shows that code-mixed and culturally relevant queries in our dataset pose challenges to evaluated models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。