用拼音做中间表示,提升中文唇语识别准确率
VALLR-Pin: Uncertainty-Factorized Visual Speech Recognition for Mandarin with Pinyin Guidance
- 双解码器联合预测汉字与拼音,增强视觉语言表征
- 用大模型对拼音和候选字进行纠错,解决同音字歧义
- 在多说话人场景下显著提升精度,适合语音识别场景
唇语识别(VSR)旨在从无声唇动视频中转录语音内容,中文因发音相似词多、同音词普遍而尤为困难。本文提出VALLR-Pin,一种两阶段中文唇语识别框架,在VALLR基础上引入拼音作为中间表示。第一阶段通过共享视觉编码器驱动双解码器,联合预测汉字及对应拼音序列,增强视觉-语言表征鲁棒性;第二阶段采用基于大模型的精炼模块,结合预测拼音与字符候选列表(N-best),解决同音字带来的歧义。为适配视觉识别错误,使用模型生成的拼音-文本对构建合成指令数据微调大模型,实现错误感知修正。在公开中文唇语识别基准测试中,VALLR-Pin在多说话人条件下持续提升转录准确率,验证了拼音引导与轻量级大模型精炼的有效性。
原文摘要 · Abstract (English)
Visual speech recognition (VSR) aims to transcribe spoken content from silent lip-motion videos and is particularly challenging in Mandarin due to severe viseme ambiguity and pervasive homophones. We propose VALLR-Pin, a two-stage Mandarin VSR framework that extends the VALLR architecture by explicitly incorporating Pinyin as an intermediate representation. In the first stage, a shared visual encoder feeds dual decoders that jointly predict Mandarin characters and their corresponding Pinyin sequences, encouraging more robust visual-linguistic representations. In the second stage, an LLM-based refinement module takes the predicted Pinyin sequence together with an N-best list of character hypotheses to resolve homophone-induced ambiguities. To further adapt the LLM to visual recognition errors, we fine-tune it on synthetic instruction data constructed from model-generated Pinyin-text pairs, enabling error-aware correction. Experiments on public Mandarin VSR benchmarks demonstrate that VALLR-Pin consistently improves transcription accuracy under multi-speaker conditions, highlighting the effectiveness of combining phonetic guidance with lightweight LLM refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。