构建呼吸音频问答基准,评估模型在真实场景下的表现
RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity
- 统一数据生成与评估流程,整合900万多样化问答对
- 发现现有模型在设备、语音差异下性能显著下降
- 适合医疗AI、多模态研究者用于测试真实场景鲁棒性
随着对话式多模态AI工具在患者数据处理中日益普及,亟需在真实条件下衡量进展并揭示失效模式。尽管呼吸音频对移动健康筛查至关重要,但现有研究对呼吸音频问答的探索仍有限,且评估范围狭窄,缺乏跨模态、设备和问题类型的真实异质性。为此,我们提出 extbf{呼吸音频问答(RA-QA)基准},包含标准化数据生成流程、全面的多模态问答集合及统一评估协议。该基准将公开的呼吸音频数据集整合为包含900万格式多样问答对的集合,覆盖诊断与上下文属性。我们对通用音频-语言模型及领域专用架构进行基准测试,建立可复现参考点,并揭示当前方法在异质条件下的失效问题。
原文摘要 · Abstract (English)
As conversational multimodal AI tools are increasingly adopted to process patient data for health assessment, robust benchmarks are needed to measure progress and expose failure modes under realistic conditions. Despite the importance of respiratory audio for mobile health screening, respiratory audio question answering remains underexplored, with existing studies evaluated narrowly and lacking real-world heterogeneity across modalities, devices, and question types. We hence introduce the \textbf{Respiratory-Audio Question-Answering (RA-QA) benchmark}, including a standardized data generation pipeline, a comprehensive multimodal QA collection, and a unified evaluation protocol. RA-QA harmonizes public RA datasets into a collection of 9 million format-diverse QA pairs covering diagnostic and contextual attributes. We benchmark general audio-language models as well as domain-specific architectures, establishing reproducible reference points and showing how current approaches fail under heterogeneity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。