通过音唇同步分析提升二语发音评估的可解释性。
When Speech Meets Lips: Interpretable Audio-Visual Synchronization for L2 Pronunciation Assessment

- 用跨模态注意力融合语音与口型,显式建模时间对齐。
- 提出滞后轨迹和稳定指数,量化发音与口型同步程度。
- 适合需要精准反馈的二语教学与发音训练场景。
自动发音评估系统虽已借助基于Transformer的模型和自监督语音表示取得良好效果,但多数方法仅依赖声学信号,忽视了语音与发音动作间的时间同步问题,限制了对二语发音训练中时序错配的诊断反馈。本文提出一种可解释的音唇同步框架,通过特征编码、跨注意力融合、滞后估计、稳定性量化与可视化,显式建模语音-口型时间对齐。框架引入帧级滞后轨迹与滞后稳定性指数(LSI),用于量化同步鲁棒性。我们对30名参与者(10名教师、20名不同母语背景学生)进行了访谈,验证其有效性。该框架将隐式的对齐关系转化为可解释表示,连接自动评分与可操作的计算机辅助发音训练反馈。数据集与补充材料见https://www.robots.ox.ac.uk/~vgg/data/lip_reading/。
原文摘要 · Abstract (English)
Automatic Pronunciation Assessment (APA) systems have achieved strong performance with transformer-based models and self-supervised speech representations. However, most methods rely only on acoustic signals and overlook temporal synchronization between speech and articulatory movements, limiting diagnostic feedback on timing mismatches important for L2 pronunciation training. We propose an interpretable audio-visual synchronization framework that explicitly models speech-lip temporal alignment through feature encoding, cross-attention fusion, lag estimation, stability quantification, and visualization. The framework introduces frame-level lag trajectories and a Lag Stability Index (LSI) to quantify synchronization robustness. We also interviewed 30 participants, including 10 instructors and 20 students with diverse first-language backgrounds, to assess its effectiveness. By transforming implicit alignment into interpretable representations, the framework connects automatic scoring with actionable Computer-Aided Pronunciation Training feedback. Datasets and supplemental materials are available at https://www.robots.ox.ac.uk/~vgg/data/lip_reading/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。