arXiv:2607.23808cs.CLcs.AI2026-07

首个覆盖印度22种官方语言的语音对话分割与识别基准数据集

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

论文配图:Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
图 1 · 摘自论文原文
  • 构建108小时多说话人自然语料,涵盖近场、远场及真实场景音频
  • 包含人工校对的时间对齐转录与说话人标注,支持代码混杂等本土语音特征
  • 为印度语种语音技术研究提供开源基准,适合多语言语音系统开发者

本文提出Indic DiarBench,一个覆盖印度全部22种官方语言的语音说话人分离与自动语音识别(ASR)基准数据集。该数据集包含约108小时的自然多说话人音频,来源包括近场会议、远场录音及真实环境采集。所有标注均经人工校对,并配有时间对齐的说话人归属转录。数据集捕捉了印度口语中常见的英语混用、方言差异和频繁说话人重叠等特征。为建立联合ASR与说话人分离的基线性能,我们评估了主流系统,包括商用语音API和多模态大语言模型。Indic DiarBench作为开放资源发布,旨在推动面向印度语言的包容性、多语言语音技术研究。

原文摘要 · Abstract (English)

In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.

语音识别说话人分离多语言印度语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。