梳理50个印度语种语音数据集的跨任务可用性,发现大量未被利用的元数据潜力。
Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
- 按9类下游任务评估50个印度语言语音数据集的适用性
- 发现多数数据集含可支持多任务的未充分利用元数据
- 为低资源语言研究提供数据优先级和增强建议
随着包容性语音技术需求上升,多语言自然语言处理(NLP)研究亟需多语种数据集。然而,对低资源语言中现有任务特定资源的认知不足阻碍了研究进展,尤其在语言多样性突出的印度更为显著。通过跨任务分析现有印度语音数据集,可缓解数据稀缺问题。这需要考察数据集在多个下游任务中的适用性,而非仅关注单一任务。以往调查多聚焦单一任务,缺乏全面的跨任务分析。因此,我们提出Task-Lens,一项针对50个覆盖26种语言的印度语音数据集的跨任务调研,评估其对九类下游语音任务的准备度。首先,分析哪些数据集具备特定任务所需的元数据与属性;其次,提出任务对齐的增强方法以释放数据集的全部下游潜力;最后,识别当前资源严重不足的任务和语言。结果表明,许多印度语音数据集包含可支持多个下游任务的未利用元数据。通过揭示跨任务关联与缺口,Task-Lens使研究人员能够探索现有数据集的更广泛应用,并优先建设关键任务与语言的数据集。
原文摘要 · Abstract (English)
The rising demand for inclusive speech technologies amplifies the need for multilingual datasets for Natural Language Processing (NLP) research. However, limited awareness of existing task-specific resources in low-resource languages hinders research. This challenge is especially acute in linguistically diverse countries, such as India. Cross-task profiling of existing Indian speech datasets can alleviate the data scarcity challenge. This involves investigating the utility of datasets across multiple downstream tasks rather than focusing on a single task. Prior surveys typically catalogue datasets for a single task, leaving comprehensive cross-task profiling as an open opportunity. Therefore, we propose Task-Lens, a cross-task survey that assesses the readiness of 50 Indian speech datasets spanning 26 languages for nine downstream speech tasks. First, we analyze which datasets contain metadata and properties suitable for specific tasks. Next, we propose task-aligned enhancements to unlock datasets to their full downstream potential. Finally, we identify tasks and Indian languages that are critically underserved by current resources. Our findings reveal that many Indian speech datasets contain untapped metadata that can support multiple downstream tasks. By uncovering cross-task linkages and gaps, Task-Lens enables researchers to explore the broader applicability of existing datasets and to prioritize dataset creation for underserved tasks and languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。