arXiv:2603.12288cs.LGcs.AI2026-03

高维脏数据也能让模型可靠,关键在数据结构与模型能力的配合。

From Garbage to Gold: A Data-Architectural Theory of Predictive Robustness

  • 将噪声分为预测误差和结构不确定性,证明高维错误数据可克服两类噪声。
  • 信息性共线性提升模型收敛效率,维度越高越利于有限样本下的推断。
  • 适合做工业级数据建模、追求真实场景鲁棒性的研究者和工程师。

表格机器学习存在悖论:现代模型虽使用高维、共线、含错的数据,仍能达到顶尖性能,违背‘垃圾进,垃圾出’的常识。本文融合信息论、潜在因子模型与心理测量学,阐明预测鲁棒性不仅依赖数据清洁度,更源于数据架构与模型容量的协同作用。将预测空间的噪声拆分为‘预测误差’和‘结构不确定性’(由随机生成映射导致的信息缺失),证明利用高维的错误预测变量可渐近克服两类噪声;而清洗低维数据则因结构不确定性存在根本性局限。揭示‘信息性共线性’(源于共同潜在原因的依赖关系)能增强可靠性与收敛效率,并说明更高维度降低潜在推断负担,使有限样本下可行。针对实际约束,提出‘主动数据中心人工智能’以高效识别支持鲁棒性的预测变量。推导系统性误差区间的边界,解释为何吸收‘异常’依赖关系的模型可缓解假设偏离问题。将潜在架构与良性过拟合关联,迈出统一应对结果误差与预测空间噪声鲁棒性的第一步,同时明确传统数据清洁方法在特定情形下仍有效。重新定义数据质量为组合层面的架构而非个体精度,为从实时未清理的企业‘数据沼泽’中学习提供理论依据,推动从‘模型迁移’到‘方法迁移’的范式转变,突破静态泛化瓶颈。

原文摘要 · Abstract (English)

Tabular machine learning presents a paradox: modern models achieve state-of-the-art performance using high-dimensional (high-D), collinear, error-prone data, defying the "Garbage In, Garbage Out" mantra. To help resolve this, we synthesize principles from Information Theory, Latent Factor Models, and Psychometrics, clarifying that predictive robustness arises not solely from data cleanliness, but from the synergy between data architecture and model capacity. Partitioning predictor-space "noise" into "Predictor Error" and "Structural Uncertainty" (informational deficits from stochastic generative mappings), we prove that leveraging high-D sets of error-prone predictors asymptotically overcomes both types of noise, whereas cleaning a low-D set is fundamentally bounded by Structural Uncertainty. We demonstrate why "Informative Collinearity" (dependencies from shared latent causes) enhances reliability and convergence efficiency, and explain why increased dimensionality reduces the latent inference burden, enabling feasibility with finite samples. To address practical constraints, we propose "Proactive Data-Centric AI" to identify predictors that enable robustness efficiently. We also derive boundaries for Systematic Error Regimes and show why models that absorb "rogue" dependencies can mitigate assumption violations. Linking latent architecture to Benign Overfitting, we offer a first step towards a unified view of robustness to Outcome Error and predictor-space noise, while also delineating when traditional DCAI's focus on label cleaning remains powerful. By redefining data quality from item-level perfection to portfolio-level architecture, we provide a theoretical rationale for "Local Factories" -- learning from live, uncurated enterprise "data swamps" -- supporting a deployment paradigm shift from "Model Transfer" to "Methodology Transfer'' to overcome static generalizability limitations.

数据架构鲁棒性高维数据模型能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。