arXiv:2603.06542cs.SDcs.AI2026-03

针对呼吸音频问答的多源异构问题,提出分层专业化模型提升鲁棒性。

RAMoEA-QA: Hierarchical Specialization for Robust Respiratory Audio Question Answering

  • 构建两级分层框架,根据输入特性动态选择处理路径。
  • 在域内测试中准确率达0.72,较单路径模型提升显著。
  • 适合临床与自录数据混合场景,尤其适用于跨设备迁移任务。

对话式生成AI在医疗领域日益受到关注,要求模型能整合异构患者信号并支持多样交互风格,同时输出临床有意义的结果。在呼吸护理中,通过传感设备获取的非侵入式音频记录为筛查和长期监测提供了可扩展途径,但其异质性尤为突出:录音受设备、环境及采集协议影响,查询意图、答案格式与预测目标也各不相同。现有的呼吸评估生物医学音频-语言问答系统虽已出现,但通常采用单一路径模型,所有输入均经相同声学与语言路径处理,无法适应不同录音条件与查询类型。且多数在有限场景下评估,未充分检验其在真实分布偏移下的鲁棒性,包括采集域、模态与临床任务的变化。为此,我们提出RAMoEA-QA,首个支持输入依赖性专业化的统一分层两阶段呼吸音频问答模型。我们在涵盖临床与自录、多设备采集设置、多种问题格式以及离散与连续目标的统一评估框架中验证该设计。在域内及受控偏移评估中,RAMoEA-QA优于匹配的单路径基线与路由控制,在判别任务上达到0.72的域内测试准确率(单路径基线为0.61和0.67),同时在回归任务表现最佳,并在数据集、模态与任务偏移下展现出更强的平均迁移能力,如在慢阻肺模态偏移设置中准确率提升最高达23个百分点。

原文摘要 · Abstract (English)

Conversational generative AI is increasingly explored in healthcare, where models must integrate heterogeneous patient signals and support diverse interaction styles while producing clinically meaningful outputs. In respiratory care, non-invasive audio recordings captured with sensing devices offer a scalable route to screening and longitudinal monitoring, but heterogeneity is particularly acute: recordings vary across devices, environments, and acquisition protocols, and queries may vary in intent, answer format, and prediction objective. Existing biomedical audio-language question answering systems for respiratory assessment are starting to emerge, but they are typically built as single-path models, processing all inputs through the same acoustic and language pathway despite variation in recording conditions and query types. They are also usually evaluated in relatively limited settings, leaving open their robustness under realistic distribution shifts, including changes in acquisition domains, modality, and clinical task. To address this gap, we introduce RAMoEA-QA, the first RA QA model designed to support input-dependent specialization across heterogeneous recordings and query types within a unified hierarchical two-stage framework. We study this design in a unified RA QA setting spanning clinical and self-recorded, multi-device acquisition settings, question formats, and both discrete and continuous targets. Across in-domain and controlled-shift evaluations, RAMoEA-QA improves over matched monolithic baselines and routing controls, reaching 0.72 in in-domain test accuracy (vs. 0.61 and 0.67 for single-path baselines) on discriminative tasks, while also achieving the best regression performance and stronger average transfer under dataset, modality, and task shifts, including gains of up to 23 percentage points in accuracy on the COPD modality-shift setting.

语音问答呼吸健康多模态鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。