构建首个大规模波斯语唇读数据集,助力听障辅助技术发展
LRW-Persian: Lip-reading in the Wild Dataset for Persian Language
- 自动化全流程清洗,确保41.4万视频样本高质量
- 覆盖743个词、67个电视节目,含头姿/年龄/性别等元数据
- 适配低资源语言研究,支持跨语言迁移与多模态分析
唇读作为提升语音识别系统鲁棒性及为听力障碍者开发辅助技术的重要方向,仍面临非英语资源匮乏的问题。本文提出LRW-Persian,目前最大规模的波斯语自然场景词级唇读数据集,包含743个目标词汇和超过414,000个视频样本,源自67个电视节目的1,900小时以上影像资料。该数据集具备去身份训练与测试划分、广泛地区与方言覆盖,以及每段视频的头部姿态、年龄、性别等丰富元数据。为保障大规模数据质量,建立了基于自动语音识别(ASR)的端到端自动化清洗流程,涵盖说话人定位、质量筛选与姿态/掩码检测。我们还在该数据集上微调两种主流唇读模型,建立基准性能,揭示波斯语视觉语音识别的挑战。该数据集填补了低资源语言的关键空白,支持严格基准测试、跨语言迁移研究,并为非主流语言的多模态语音研究提供基础。数据集已公开:https://lrw-persian.vercel.app。
原文摘要 · Abstract (English)
Lipreading has emerged as an increasingly important research area for developing robust speech recognition systems and assistive technologies for the hearing-impaired. However, non-English resources for visual speech recognition remain limited. We introduce LRW-Persian, the largest in-the-wild Persian word-level lipreading dataset, comprising $743$ target words and over $414{,}000$ video samples extracted from more than $1{,}900$ hours of footage across $67$ television programs. Designed as a benchmark-ready resource, LRW-Persian provides speaker-disjoint training and test splits, wide regional and dialectal coverage, and rich per-clip metadata including head pose, age, and gender. To ensure large-scale data quality, we establish a fully automated end-to-end curation pipeline encompassing transcription based on Automatic Speech Recognition(ASR), active-speaker localization, quality filtering, and pose/mask screening. We further fine-tune two widely used lipreading architectures on LRW-Persian, establishing reference performance and demonstrating the difficulty of Persian visual speech recognition. By filling a critical gap in low-resource languages, LRW-Persian enables rigorous benchmarking, supports cross-lingual transfer, and provides a foundation for advancing multimodal speech research in underrepresented linguistic contexts. The dataset is publicly available at: https://lrw-persian.vercel.app.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。