首个波斯语医疗问答基准,评估大模型在真实患者问题上的表现
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language
- 构建波斯语医疗问答数据集,涵盖68,138条清洗后的问答对
- 使用LLM评分框架验证,发现多语言医疗问答仍存在显著误差
- 适合研究多语言医疗AI、低资源语言模型的开发者参考
医疗消费者问答(CQA)对赋能患者获取个性化、可靠的健康信息至关重要。尽管大语言模型(LLMs)在医学问答领域取得进展,面向消费者的多语言资源,特别是低资源语言如波斯语,仍然稀缺。为此,我们提出PerMedCQA,首个针对波斯语的真实世界消费者医疗问答评估基准。该数据集源自大型医学问答论坛,包含68,138条经清洗的问答对,原始数据为87,780条。我们评估了多个先进的多语言及指令微调大模型,并采用MedJudge——一种由LLM驱动的新型基于评分标准的评估框架,其结果经专家人工标注验证。实验揭示了多语言医疗问答中的关键挑战,为构建更准确、更具上下文感知能力的医疗辅助系统提供了重要洞察。数据已公开于https://huggingface.co/datasets/NaghmehAI/PerMedCQA。
原文摘要 · Abstract (English)
Medical consumer question answering (CQA) is crucial for empowering patients by providing personalized and reliable health information. Despite recent advances in large language models (LLMs) for medical QA, consumer-oriented and multilingual resources, particularly in low-resource languages like Persian, remain sparse. To bridge this gap, we present PerMedCQA, the first Persian-language benchmark for evaluating LLMs on real-world, consumer-generated medical questions. Curated from a large medical QA forum, PerMedCQA contains 68,138 question-answer pairs, refined through careful data cleaning from an initial set of 87,780 raw entries. We evaluate several state-of-the-art multilingual and instruction-tuned LLMs, utilizing MedJudge, a novel rubric-based evaluation framework driven by an LLM grader, validated against expert human annotators. Our results highlight key challenges in multilingual medical QA and provide valuable insights for developing more accurate and context-aware medical assistance systems. The data is publicly available on https://huggingface.co/datasets/NaghmehAI/PerMedCQA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。