分析顶会论文数据来源,揭示数据溯源困境并呼吁重视数据根基。
A quest through interconnected datasets: lessons from highly-cited ICASSP papers
- 深度追踪顶级论文数据集源头,发现大量信息不透明或路径纠缠。
- 多数论文未明确标注数据真实来源,部分需跨多层文档追溯。
- 适合关注模型可信性与数据伦理的研究者与教育工作者阅读。
随着语音机器学习成果被应用于具有社会影响力的场景,理解所用数据的质量与来源变得至关重要。然而,在应用机器学习领域,明确数据来源并未获得学术出版的直接奖励,也未纳入常规课程体系。本文研究了国际声学、语音与信号处理会议(ICASSP)近五年被引次数最高的五篇论文中的数据使用情况。通过深入的深度优先分析,我们追溯了这些论文所用数据集的真实来源,常需超越论文中公开的信息,最终发现许多数据来源模糊甚至相互缠绕。在当前大模型及生成式AI快速发展的背景下,对数据可追溯性的关注日益增加。因此,我们呼吁学术界不仅应追求更大模型的工程突破,更应为明确模型基础、强调数据溯源的行为创造空间并给予认可。
原文摘要 · Abstract (English)
As audio machine learning outcomes are deployed in societally impactful applications, it is important to have a sense of the quality and origins of the data used. Noticing that being explicit about this sense is not trivially rewarded in academic publishing in applied machine learning domains, and neither is included in typical applied machine learning curricula, we present a study into dataset usage connected to the top-5 cited papers at the International Conference on Acoustics, Speech, and Signal Processing (ICASSP). In this, we conduct thorough depth-first analyses towards origins of used datasets, often leading to searches that had to go beyond what was reported in official papers, and ending into unclear or entangled origins. Especially in the current pull towards larger, and possibly generative AI models, awareness of the need for accountability on data provenance is increasing. With this, we call on the community to not only focus on engineering larger models, but create more room and reward for explicitizing the foundations on which such models should be built.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。