提出143项指标,系统评估多模态数据集的可信与伦理属性。
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
- 设计143个可验证指标,从文档中提取数据集的可信与伦理特征。
- 仅少数数据集明确记录了知情同意、隐私保护等伦理信息。
- 众包和直接采集的数据集更可能包含伦理说明,爬取数据则忽视伦理。
数据透明是负责任AI的关键,但影响可信与伦理的多模态数据属性洞察仍匮乏且难以跨数据集比较。为此,我们提出可信与伦理数据集指标(TEDI),支持对数据集文档进行系统性、实证分析。TEDI包含143个细粒度指标,刻画多模态数据集及其收集过程的可信与伦理属性,均基于可验证信息设计。我们手动标注并分析了超过100个包含人类语音的多模态数据集,补充了数据来源、规模和模态信息,以揭示塑造可信与伦理维度的关键因素。结果发现,仅有少数数据集记录了知情同意、隐私保护及有害内容等关键伦理信息。伦理指标的披露程度取决于数据收集方式:众包和直接采集的数据集更可能提及这些内容,而以爬取为主的收集方式虽规模大,却普遍忽略伦理细节。本研究为提升数据集在可信与伦理维度的透明度提供方法与实证支持,并为未来自动化提取文档信息奠定基础。
原文摘要 · Abstract (English)
Dataset transparency is a key enabler of responsible AI, but insights into multimodal dataset attributes that impact trustworthy and ethical aspects of AI applications remain scarce and are difficult to compare across datasets. To address this challenge, we introduce Trustworthy and Ethical Dataset Indicators (TEDI) that facilitate the systematic, empirical analysis of dataset documentation. TEDI encompasses 143 fine-grained indicators that characterize trustworthy and ethical attributes of multimodal datasets and their collection processes. The indicators are framed to extract verifiable information from dataset documentation. Using TEDI, we manually annotated and analyzed over 100 multimodal datasets that include human voices. We further annotated data sourcing, size, and modality details to gain insights into the factors that shape trustworthy and ethical dimensions across datasets. We find that only a select few datasets have documented attributes and practices pertaining to consent, privacy, and harmful content indicators. The extent to which these and other ethical indicators are addressed varies based on the data collection method, with documentation of datasets collected via crowdsourced and direct collection approaches being more likely to mention them. Scraping dominates scale at the cost of ethical indicators, but is not the only viable collection method. Our approach and empirical insights contribute to increasing dataset transparency along trustworthy and ethical dimensions and pave the way for automating the tedious task of extracting information from dataset documentation in future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。