arXiv:2509.16765cs.CLcs.AI2025-09EMNLP被引 2

首个语音病理语言模型综合评测,揭示性能差异与优化路径。

The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology

  • 构建五大临床场景基准,每项含千条人工标注数据。
  • 微调后模型性能提升超10%,但对男性说话人更优。
  • 发现思维链提示会降低复杂分类任务准确率,适合特定人群。

美国国立卫生研究院数据显示,超过340万儿童患有需临床干预的言语障碍,而言语语言病理学家(SLPs)数量仅为患儿的约1/20,凸显照护缺口与技术支援的迫切需求。当前多模态语言模型(MLMs)虽具潜力,但在高风险临床环境中的表现仍缺乏深入理解。为此,我们联合领域专家构建了真实应用场景分类体系,并推出首个涵盖五类核心用例的综合性评测基准,每类含1,000条人工标注数据。该基准包含噪声、性别、口音等条件下的鲁棒性与敏感性测试。对15个前沿MLMs的评估显示:无模型在所有任务中持续领先;存在系统性偏差,模型对男性说话人表现更佳;且思维链提示在标签空间大、决策边界窄的任务中会降低分类性能。进一步在领域数据上微调可使模型性能提升超10%。结果揭示了当前MLMs在言语病理应用中的潜力与局限,强调需针对性研究与开发。

原文摘要 · Abstract (English)

According to the U.S. National Institutes of Health, more than 3.4 million children experience speech disorders that require clinical intervention. The number of speech-language pathologists (SLPs) is roughly 20 times fewer than the number of affected children, highlighting a significant gap in children's care and a pressing need for technological support that improves the productivity of SLPs. State-of-the-art multimodal language models (MLMs) show promise for supporting SLPs, but their use remains underexplored largely due to a limited understanding of their performance in high-stakes clinical settings. To address this gap, we collaborate with domain experts to develop a taxonomy of real-world use cases of MLMs in speech-language pathologies. Building on this taxonomy, we introduce the first comprehensive benchmark for evaluating MLM across five core use cases, each containing 1,000 manually annotated data points. This benchmark includes robustness and sensitivity tests under various settings, including background noise, speaker gender, and accent. Our evaluation of 15 state-of-the-art MLMs reveals that no single model consistently outperforms others across all tasks. Notably, we find systematic disparities, with models performing better on male speakers, and observe that chain-of-thought prompting can degrade performance on classification tasks with large label spaces and narrow decision boundaries. Furthermore, we study fine-tuning MLMs on domain-specific data, achieving improvements of over 10\% compared to base models. These findings highlight both the potential and limitations of current MLMs for speech-language pathology applications, underscoring the need for further research and targeted development.

语音病理多模态模型评测基准微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。