arXiv:2510.18468cs.CL2025-10被引 1

构建首个意大利语医疗问答基准,助力医学AI理解真实医患对话。

IMB: An Italian Medical Benchmark for Question Answering

  • 用大模型提升意大利语医患对话的清晰度与一致性。
  • 在77类医学话题上构建超78万条对话数据集。
  • 适合研究多语言医疗问答、领域适配模型的研究者。

在线医疗论坛长期作为患者获取专业医疗建议的重要平台,积累了大量宝贵知识。然而,论坛交流的非正式性与语言复杂性给自动化问答系统带来挑战,尤其在非英语语境下。我们构建了两个全面的意大利语医学基准: extbf{IMB-QA},包含来自77个医学领域的782,644条患者-医生对话; extbf{IMB-MCQA},包含25,862道医学专科考试的多选题。实验表明,大型语言模型(LLMs)可有效提升论坛数据的清晰度与一致性,同时保留原意与对话风格。在开放问答与多选题任务中,对比多种LLM架构,发现基于检索增强生成(RAG)与领域微调的策略,优于更大规模通用模型。结果表明,医疗AI系统更依赖领域专长与高效信息检索,而非单纯扩大模型规模。我们已将两个数据集与评估框架开源至GitHub,以支持多语言医学问答研究:https://github.com/PRAISELab-PicusLab/IMB。

原文摘要 · Abstract (English)

Online medical forums have long served as vital platforms where patients seek professional healthcare advice, generating vast amounts of valuable knowledge. However, the informal nature and linguistic complexity of forum interactions pose significant challenges for automated question answering systems, especially when dealing with non-English languages. We present two comprehensive Italian medical benchmarks: \textbf{IMB-QA}, containing 782,644 patient-doctor conversations from 77 medical categories, and \textbf{IMB-MCQA}, comprising 25,862 multiple-choice questions from medical specialty examinations. We demonstrate how Large Language Models (LLMs) can be leveraged to improve the clarity and consistency of medical forum data while retaining their original meaning and conversational style, and compare a variety of LLM architectures on both open and multiple-choice question answering tasks. Our experiments with Retrieval Augmented Generation (RAG) and domain-specific fine-tuning reveal that specialized adaptation strategies can outperform larger, general-purpose models in medical question answering tasks. These findings suggest that effective medical AI systems may benefit more from domain expertise and efficient information retrieval than from increased model scale. We release both datasets and evaluation frameworks in our GitHub repository to support further research on multilingual medical question answering: https://github.com/PRAISELab-PicusLab/IMB.

医疗问答多语言数据集领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。