arXiv:2606.15325cs.CL2026-06

大模型语音诊断常被刻板印象误导,依赖外部证据才可靠。

Prior over Evidence: Stereotype-Driven Diagnosis in LLM-Based L2 Pronunciation Feedback

论文配图:Prior over Evidence: Stereotype-Driven Diagnosis in LLM-Based L2 Pronunciation Feedback
图 1 · 摘自论文原文
  • 用多种证据输入测试大模型,发现其推理常与真实发音脱节。
  • 39.6%的错误判断有内部一致的错误理由,仅15.8%正确判断理由可信。
  • 只有直接对应目标维度的声学特征才能提升诊断准确性,适合做语音分析辅助。

大型语言模型在第二语言英语发音反馈中日益普及,普遍假设其诊断基于实际语音证据而非预训练先验。本研究在1,800个来自六种母语背景的L2-Arctic语音样本上,测试了三种具备音频能力的LLM在四种发音维度下的表现,覆盖从纯文本到数值声学特征及原始音频的五种证据条件。每个(语音×模型×条件×维度)组合通过三项指标评估:评分准确率(RA)、证据一致性(EC)和有据正确性(GC)。结果表明:第一,评分准确率与有据推理解耦——39.6%的判断存在内部一致但错误的推理,仅15.8%正确判断有合理依据;第二,音素级反馈趋同于一组固定的难发音音素,跨越所有母语背景和证据条件;第三,只有直接探测目标维度的声学特征能提升评分——将音高变化的有据性从(0.18–0.19)提升至(0.45–0.62),而重音与音素正确性因需目标-现实对齐,仍无法建立有效关联。不带文本化音高的原始波形无法复现该提升。结果表明,当前通用大模型更适合作为外部计算证据的口语化输出工具,而非独立诊断引擎。

原文摘要 · Abstract (English)

Large language models are increasingly deployed for written pronunciation feedback in second-language (L2) English learning, under the assumption that their diagnoses are grounded in the supplied speech evidence rather than in priors from pretraining. This assumption is tested on 1,800 L2-Arctic utterances spanning six L1 backgrounds, three audio-capable LLMs, four pronunciation dimensions, and five evidence conditions ranging from a text-only baseline to numeric acoustic features and raw audio. Each (utterance x model x condition x dimension) cell is scored on three metrics: Rating Accuracy (RA) against gold labels, Evidence Coherence (EC) assessing internal consistency without ground truth, and Grounded Correctness (GC) evaluated against gold evidence. Results show three findings across models. First, rating accuracy and grounded reasoning decouple: 39.6% of judged cells contain internally coherent reasoning that supports a wrong rating, against only 15.8% where the reasoning supports a correct rating. Second, phoneme-level feedback converges to a fixed inventory of L2-English difficulty phones that recurs across all six L1 backgrounds and all evidence conditions. Third, acoustic evidence improves the rating only when the supplied feature directly probes the target dimension: textualised F0 range raises pitch-variation grounding from (0.18-0.19) to (0.45-0.62) across all three models, while stress and phoneme correctness, which require target-to-realisation alignment, remain ungrounded. The same audio waveform without textualised F0 values does not reproduce this improvement. These findings indicate that current general-purpose LLMs are more reliable as verbalisers of externally computed pronunciation evidence than as standalone diagnostic engines.

语音反馈大模型诊断声学证据语言学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。