arXiv:2605.00969cs.SDcs.AI2026-05中稿 · ICML

构建医疗音频问答基准,测试模型在真实临床场景下的多轮推理能力

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

论文配图:MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
图 1 · 摘自论文原文
  • 涵盖多种医疗音源,含生理声、合成语音与长短对话
  • 46,701个问答对,覆盖多选、多轮与开放题,检验复杂推理能力
  • 顶尖模型如Gemini-2.5-pro准确率仅68.1%,凸显医疗推理挑战

由于隐私法规和领域专业知识带来的高标注成本,医学音频数据难以获取,现有基准常无法反映复杂的临床场景。为此,我们提出MedMosaic,一个面向语言与音频推理模型的医学音频问答数据集,旨在评估模型在真实临床约束下的表现。该数据集包含多种医学音频类型:与病情相关的生理声音、模拟语音失真的人工合成语音,以及长短不一的真实临床对话,以建模不同上下文长度。共包含46,701个问答对,涵盖多项选择、多轮序列和开放式问答,支持对多跳推理与答案生成能力的系统性评估。对13个音频与多模态推理模型的基准测试显示,所有系统在推理任务上仍面临挑战,且不同题型间性能差异显著。即使最先进的模型Gemini-2.5-pro,准确率也仅达68.1%。这些结果揭示了当前医疗推理能力的局限性,强调亟需更稳健、专为领域设计的多模态推理模型。

原文摘要 · Abstract (English)

Medical audio data is difficult to collect due to privacy regulations and high annotation costs arising from domain expertise. Thus, existing benchmarks tend to underrepresent complex medical audio scenarios. To address this challenge, we present MedMosaic, a medical audio question-answering dataset designed to benchmark language and audio reasoning models under realistic clinical constraints. MedMosaic features a diverse range of medical audio types, including condition-related physiological sounds, carefully constructed synthetic voices to mimic speech with artifacts as well as real short and long length clinical conversations to model varying context lengths. The dataset also features a total of 46,701 question-answer pairs, spanning categories such as multiple-choice, sequential multi-turn, and open-ended question-answers, enabling systematic evaluation of multi-hop reasoning and answer generation capabilities. Benchmarking 13 audio and multimodal reasoning models reveals that reasoning remains challenging for all evaluated systems, with substantial performance variation across question types. In particular, even state-of-the-art model like Gemini-2.5-pro can only achieve 68.1% accuracy approximately. These findings underscore persistent limitations in medical reasoning and highlight the need for more robust, domain-specific multimodal reasoning models. A sample of benchmark data is available here: https://shorturl.at/Lyp33

医疗音频多模态推理问答系统基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。