用大模型分析语音诊断嗓音疾病,兼顾准确与安全。
VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware Evaluation
- 基于微调的音频大模型,结合医院真实患者数据
- 在嗓音疾病分类上优于现有方法,准确率显著提升
- 强调诊断公平性与跨语言能力,适合医疗健康领域
嗓音健康对人们沟通与社交至关重要,但全球范围内嗓音障碍患者仍普遍缺乏便捷的诊断与治疗途径。本文提出VocalAgent,一种基于音频大语言模型(LLM)的嗓音健康诊断系统。通过在三个来自医院患者的实地采集数据集上微调Qwen-Audio-Chat,构建了多维度评估框架,包括安全性评估以缓解诊断偏见、跨语言性能分析及模态消融实验。VocalAgent在嗓音障碍分类任务中表现优于当前最优基线模型。其基于大模型的方法具备可扩展性,有助于推动健康诊断的普及,同时强调伦理与技术验证的重要性。
原文摘要 · Abstract (English)
Vocal health plays a crucial role in peoples' lives, significantly impacting their communicative abilities and interactions. However, despite the global prevalence of voice disorders, many lack access to convenient diagnosis and treatment. This paper introduces VocalAgent, an audio large language model (LLM) to address these challenges through vocal health diagnosis. We leverage Qwen-Audio-Chat fine-tuned on three datasets collected in-situ from hospital patients, and present a multifaceted evaluation framework encompassing a safety assessment to mitigate diagnostic biases, cross-lingual performance analysis, and modality ablation studies. VocalAgent demonstrates superior accuracy on voice disorder classification compared to state-of-the-art baselines. Its LLM-based method offers a scalable solution for broader adoption of health diagnostics, while underscoring the importance of ethical and technical validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。