首个罗马尼亚语唇读数据集,支持低资源场景下监督质量与多模态鲁棒性研究。
VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

- 构建200小时罗马尼亚语唇读数据集,含自动伪标签与人工标注子集。
- 伪标签可随数据量扩展持续提升性能,优于固定规模的人工标注。
- 验证多模态融合在噪声环境下的鲁棒性,且模型表征可有效迁移至孤立词识别。
我们提出VSRo-200,首个面向罗马尼亚语视觉语音识别(唇读)的大规模数据集,包含200小时真实世界播客视频。所有样本均使用微调的罗马尼亚语语音识别模型生成伪标签,其中100小时额外由人工转录,可在统一框架下控制分析监督质量。基于该数据集,我们建立了低资源环境下视觉语音识别的基准。系统研究了监督质量的影响:尽管人工标注在固定数据量下表现更优,但伪标签可通过数据扩展实现持续提升。进一步通过精心设计的分布外(OOD)测试集评估域偏移下的鲁棒性,并在噪声条件下分析音视频联合识别(AVSR),结果显示多模态融合显著优于纯音频模型。最后,证明在VSRo-200上学习的表征能有效迁移至孤立词识别基准LRRo,显著超越此前报告结果。总体而言,VSRo-200为低资源视觉语音识别中的监督、域泛化与多模态融合研究提供了新平台。
原文摘要 · Abstract (English)
We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast videos. All samples are annotated with pseudo-labels generated by a fine-tuned Romanian ASR model, while a subset of 100 hours is additionally transcribed by humans, enabling controlled analysis of supervision quality under a unified framework. Building on this dataset, we establish a benchmark for visual speech recognition in low-resource settings. We systematically study the impact of supervision quality, showing that while human annotations provide better performance at fixed data scales, pseudo-labels enable continued improvements through scalability. We further evaluate robustness under domain shift using curated out-of-distribution (OOD) test sets, and analyze audio-visual speech recognition (AVSR) under noisy conditions, where multimodal fusion significantly improves robustness compared to audio-only models. Finally, we demonstrate that representations learned on VSRo-200 transfer effectively to the LRRo benchmark for isolated word recognition, substantially outperforming previously reported results. Overall, VSRo-200 provides a new testbed for studying supervision, domain generalization, and multimodal fusion in low-resource visual speech recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。