arXiv:2607.09668cs.LG2026-07

提出真实数据是人为构建的,非客观存在

Position: Every Ground Truth is a Human Construction, not an Objective Truth

论文配图:Position: Every Ground Truth is a Human Construction, not an Objective Truth
图 1 · 摘自论文原文
  • 强调真实数据由人与技术共同构造,非天然存在
  • 指出数据集具有情境依赖性,使用时需明确边界
  • 倡导提升模型可靠性,推动透明与跨学科合作

真实数据集在机器学习模型的训练与评估中起基础作用。本文主张,真实数据并非中立的客观测量,而是由人类与技术共同构建的产物。机器学习领域应正视这些常被隐藏或忽略的选择,承认参考数据集具有情境依赖性而非普适性。关注真实数据的生成背景,有助于更清醒地判断数据集及所塑造模型的适用范围、时间与场景,从而提升模型的‘情境可靠性’——即清晰阐述模型的局限与优势。重视真实数据的建构过程,可增强系统透明度、问责性,并促进跨学科协作。

原文摘要 · Abstract (English)

Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models. This position paper argues that ground truths are not neutral objective measurements that are naturally given, but instead that they are constructed by arrangements of humans and technologies. We argue that the ML community will benefit from articulating and discussing these often invisible or unreported choices and acknowledging that reference data sets are contingent, not universal. Focusing on the situated and context-dependent nature of ground truths can improve reliability by enabling a better informed perspective on where, when, and how the datasets, and the models they have shaped, can best be used. We argue for increasing `situated reliability' which includes articulating the limits and strengths of models and their truth claims. Finally, paying more attention to the construction of ground truths can support transparency, accountability, and interdisciplinary work.

数据构建可靠性伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。