用大模型同时评语速、发音、逻辑并给出自然语言解释
A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales
- 用评分标准引导的语音大模型,联合预测句子和音素级准确度
- 在762个样本上达到与单粒度模型相当的评分精度
- 生成的解释在句级合理,但词音级参考不足且对齐差
自动二语语音评估可打分,但缺乏可解释性。我们提出一种基于评分标准引导的语音大模型,采用监督微调与有界直接偏好优化相结合的混合目标进行训练,能联合预测句子级(准确度、流利度、语调)的序数标签、词/音素级准确度,并在同一响应中生成自然语言理由。在SpeechOcean762数据集上,该方法表现匹配或优于单粒度模型,且保持与现有方法竞争力。我们从自一致性与真实标签对齐两个维度分析理由可靠性:通过情感一致性(合理性)和提及一致性(忠实性)衡量。结果显示,句级理由合理,但词/音素级理由参考稀疏且与标记对齐较弱。
原文摘要 · Abstract (English)
Automated L2 speech assessment can assign proficiency labels, but often lacks interpretability. We propose a rubric-guided SpeechLLM for multi-aspect, multi-granular assessment, trained with a hybrid objective combining supervised fine-tuning and Bounded Direct Preference Optimization. The model jointly predicts ordinal labels at the sentence-level (accuracy, fluency, prosody), word/phoneme-level accuracy, and generates a natural-language rationale in the same response. On SpeechOcean762, our approach matches or outperforms single-granularity models while remaining competitive with prior approaches. We analyze rationale reliability along two axes: self-consistency with model predictions and alignment with ground-truth labels, using sentiment consistency (plausibility) and mention-based agreement (faithfulness). Rationales are plausible at the sentence level, but faithfulness degrades at the word/phoneme level: references are sparse and weakly aligned with token-level labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。