VAANI构建了覆盖165个地区的多模态语音数据集,助力印度语种的普惠语音技术。
VAANI: Capturing the language landscape for an inclusive digital India
- 用图像提示采集自发语音,结合多主题图像采集
- 含31255小时语音、2043小时转录音频,覆盖105种语言
- 首次大规模覆盖众多印度小语种,适合语音研究者使用
语音技术有望弥合数字鸿沟,但现有数据集未能反映印地语系语言的地域与语言多样性。我们推出项目VAANI,一个大规模多模态数据集,覆盖印度165个地区。通过图像提示诱导自然语音回应,同时独立采集涵盖各地多样主题的图像。数据集经过自动化与人工结合的多阶段质量控制,确保音频质量与转写准确率。共发布约28.9万张图像、31,255小时语音及2,043小时转录音频,涵盖28个邦和3个中央直辖区的105种语言。许多语言首次以如此规模呈现,使VAANI成为包容性语音技术的基础资源。该数据集支持多语言、多模态模型研发,推动低资源语言在语音识别、语言理解与跨模态学习中的研究。
原文摘要 · Abstract (English)
Voice based technologies have the potential to bridge digital accessibility gaps; however, existing datasets fail to capture the linguistic and regional diversity of Indic languages. We present Project VAANI, a large scale multimodal dataset designed to represent India's linguistic landscape across 165 districts. Speech data is collected using image based prompts to elicit spontaneous responses, while images are curated through a separate pipeline covering diverse themes across regions. The dataset undergoes a rigorous multi stage quality control process, combining automated and manual evaluation to ensure high audio quality and transcription accuracy. We release approximately 289K images, 31,255 hours of speech, and 2,043 hours of transcribed audio spanning 105 languages from 28 states and 3 union territories. Many of these languages are represented at this scale for the first time, making VAANI a foundational resource for inclusive speech technology. The dataset enables the development of robust, multilingual, and multimodal models, and supports research in speech recognition, language understanding, and cross-modal learning for underrepresented languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。