arXiv:2606.17339cs.AIcs.CL2026-06

构建跨疾病临床语音AI评估基准,推动通用语音表征发展

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

论文配图:SpeechDx: A Multi-Task Benchmark for Clinical Speech AI
图 1 · 摘自论文原文
  • 按语音生成阶段划分27项任务,统一评估不同健康状况下的语音表现
  • 大模型在多数任务中表现最佳,但无模型能跨条件稳定泛化
  • 适合研究通用临床语音表示、跨疾病诊断的AI开发者使用

语音通过同时激活神经、运动、呼吸和发声系统,为健康状态提供独特洞察。当前临床语音人工智能研究多局限于单一疾病,导致结果难以比较且泛化能力难评估。我们提出SpeechDx,一个涵盖12个数据集、27项任务的大规模临床语音AI基准,覆盖多种健康状况。为实现对共享临床机制的评估,SpeechDx按语音生成阶段(概念化、构词、发音)组织任务。基准通过少量标注数据任务及跨数据集评估,区分真实临床模式与数据偏差。我们系统评估了12个先进音频编码器在所有任务及零样本跨条件迁移下的表现。结果显示:大规模语音模型构成最强基线,领域特定模型仅在任务匹配时提升性能,而现有表征无法可靠跨临床场景泛化。SpeechDx建立了一个共享评估框架,助力通用临床语音表示的发展。

原文摘要 · Abstract (English)

Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems. Current clinical speech AI methods have largely progressed through isolated condition-specific studies, making results difficult to compare and generalization difficult to assess. We introduce SpeechDx, a large-scale benchmark for clinical speech AI spanning 12 datasets and 27 tasks across diverse health conditions. To enable evaluation across shared clinical mechanisms, SpeechDx structures tasks by the stage of speech production they disrupt: conceptualization, formulation, and articulation. The benchmark tests generalization by including tasks with limited labeled data and evaluating the same health condition across multiple datasets, distinguishing clinically meaningful patterns from dataset artefacts. We systematically evaluate 12 state-of-the-art audio encoders across all tasks and under zero-shot cross-condition transfer. Results show that large-scale speech models represent the strongest overall baselines, domain-specific models improve performance only on closely matched tasks, and no current representation generalizes reliably across the clinical speech landscape. SpeechDx establishes a shared evaluation framework for tracking progress toward general-purpose clinical speech representations

语音诊断多任务学习临床AI基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。