arXiv:2606.31310cs.CLcs.MM2026-06

通过隐空间序数原型对齐,提升语音语言评估的准确性与可解释性。

LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment

论文配图:LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment
图 1 · 摘自论文原文
  • 在隐空间引入序数原型正则化,强制模型学习语言能力的等级结构。
  • 在SLA任务上实现0.361的RMSE,媲美千亿参数大模型但无需微调。
  • 适合关注可解释性与高效建模的语音评估研究者使用。

随着模型规模扩大和多模态输入的发展,多模态大语言模型(MLLMs)已成为语音语言评估(SLA)的有前景范式。然而,现有方法常忽视语言习得的内在序数结构。本文提出隐空间序数原型对齐(LOPA),一种基于原型的正则化方法,直接在隐空间施加序数几何先验。结合语义锚定层路由(SALR),自冻结的Whisper编码器动态提取多层级表示,框架在测试中达到0.361的RMSE,性能媲美百亿参数系统,且无需基于LLM的微调。进一步分析表明,SALR与LOPA的协同作用实现了可解释、符合评估标准的偏好预测,为当前以扩展为中心的模型提供了一种高效且具有序数感知能力的替代方案。

原文摘要 · Abstract (English)

Fueled by increasing model scale and multimodal inputs, Multimodal Large Language Models (MLLMs) have emerged as a promising paradigm for Spoken Language Assessment (SLA). While effective, this paradigm often overlooks the intrinsic ordinal structure of language acquisition. This paper works around the necessity of large-scale MLLMs by introducing Latent Ordinal Prototype Alignment (LOPA) for SLA, a prototype-based regularizer that enforces an ordinal geometric prior directly on the latent space. Coupled with Semantic-Anchored Layer Routing (SALR), which adaptively harvests multi-depth representations from a frozen Whisper encoder, our framework achieves an RMSE of 0.361. This performance rivals billion-parameter systems without the need for LLM-based fine-tuning. Further analysis reveals that SALR's synergy with LOPA offers interpretable, criterion-aligned preferences, thereby supporting an efficient and ordinal-aware modeling alternative to current scaling-centric models for SLA.

语音评估序数建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。