构建覆盖104个地区的多模态印地语语音识别基准数据集
Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi

- 基于图像提示采集真实场景下的自发语音,覆盖多元人口
- 每段音频有三份独立转录,支持多参考评估
- 适合研究包容性语音识别与跨区域语言建模的团队
语音识别系统需要系统的评测机制。尽管已有多个开源印地语语音识别数据集,但现有基准在地理多样性、人口代表性及转录鲁棒性方面仍显不足。本文引入一个涵盖印度104个地区的包容性多模态印地语语音识别基准数据集。数据通过图像提示激发自发语音,在真实声学环境下采集自不同人口群体。每段音频配有三份独立转录,支持考虑可接受拼写和词汇差异的多参考评估,提升评测的稳健性、包容性与真实性。我们对多个开源与专有语音识别模型进行了基准测试,并报告其在该数据集上的性能表现。
原文摘要 · Abstract (English)
Benchmarking is critical for the systematic evaluation and comparison of automatic speech recognition (ASR) systems. While several open-source datasets are available for Hindi ASR, existing benchmarks remain limited in geographic diversity, demographic representation, and transcription robustness. We introduce an inclusive, multimodal Hindi ASR benchmark collected from 104 districts across India. The dataset consists of spontaneous speech elicited using image prompts and recorded in real-world acoustic conditions across diverse demographic groups. Each audio segment is annotated with three independent transcriptions, enabling multi-reference evaluation that accounts for permissible orthographic and lexical variations. This design supports more robust, inclusive, and realistic ASR evaluation. We benchmark multiple open-source and proprietary ASR models and report their comparative performance on the benchmark dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。