arXiv:2411.13653cs.AIstat.ML2024-11NeurIPS被引 1

现有数据收集方式无法验证复杂社会系统中模型的有效性

No Free Delivery Service: Epistemic limits of passive data collection in complex social systems

  • 证明在复杂社会系统中,训练-测试范式对风险估计器无效
  • 在MovieLens数据集上验证了因果与反事实估计器的不可靠性
  • 适用于推荐系统和大模型推理,警示盲目扩展无效

机器学习的快速验证依赖于训练-测试范式,但现代AI系统常违背其基本假设。本文表明,在复杂社会系统的典型推断场景中,该范式不仅缺乏理论依据,且对包括反事实和因果估计器在内的所有风险估计器均无效,概率极高。这些形式化不可能性结果揭示了一个根本性的认识论问题:在当前数据收集方式下,我们无法判断模型是否有效。这涵盖推荐系统和大语言模型推理等关键任务,单纯扩大规模或使用有限基准无法解决此问题。通过广泛使用的MovieLens基准验证结论,并讨论参与式数据治理与开放科学等潜在应对策略。

原文摘要 · Abstract (English)

Rapid model validation via the train-test paradigm has been a key driver for the breathtaking progress in machine learning and AI. However, modern AI systems often depend on a combination of tasks and data collection practices that violate all assumptions ensuring test validity. Yet, without rigorous model validation we cannot ensure the intended outcomes of deployed AI systems, including positive social impact, nor continue to advance AI research in a scientifically sound way. In this paper, I will show that for widely considered inference settings in complex social systems the train-test paradigm does not only lack a justification but is indeed invalid for any risk estimator, including counterfactual and causal estimators, with high probability. These formal impossibility results highlight a fundamental epistemic issue, i.e., that for key tasks in modern AI we cannot know whether models are valid under current data collection practices. Importantly, this includes variants of both recommender systems and reasoning via large language models, and neither naïve scaling nor limited benchmarks are suited to address this issue. I am illustrating these results via the widely used MovieLens benchmark and conclude by discussing the implications of these results for AI in social systems, including possible remedies such as participatory data curation and open science.

模型验证社会系统因果推断数据伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。