arXiv:2602.23446cs.LGcs.AI2026-02

人类监督本质是信息瓶颈,导致模型无法消除的错误底限。

Human Supervision as an Information Bottleneck: A Unified Theory of Error Floors in Human-Guided Learning

  • 将人类反馈视为信息受限通道,构建统一理论解释误差下限
  • 六种分析框架均证明:非充分监督必然带来正的误差下限
  • 引入外部信号(如检索、工具)可突破瓶颈,降低或消除误差

大型语言模型主要依赖人工生成数据和反馈训练,但仍存在由标注噪声、主观偏好及自然语言表达能力有限引发的持续性错误。我们提出,这些限制源于监督通道的结构性缺陷,而非模型规模或优化问题。本文建立统一理论,证明当人类监督通道不足以表征潜在评估目标时,该通道会成为信息压缩环节,导致任何受其主导的学习者都存在严格为正的超额风险下限。通过算子理论、PAC-Bayes、信息论、因果推断、范畴论及强化学习中的人类反馈博弈分析等六种互补框架,我们发现该下限由标注噪声、偏好扭曲和语义压缩三部分结构分解而成。理论表明,仅靠扩大模型规模无法消除人类对齐中的持久误差。实验在真实偏好数据、合成已知目标任务及可外部验证基准上验证了预测的结构特征:仅人类监督存在持续误差底限,而足够信息的辅助通道(如检索、程序执行)能有效提升监督容量,消除或显著降低超额误差。

原文摘要 · Abstract (English)

Large language models are trained primarily on human-generated data and feedback, yet they exhibit persistent errors arising from annotation noise, subjective preferences, and the limited expressive bandwidth of natural language. We argue that these limitations reflect structural properties of the supervision channel rather than model scale or optimization. We develop a unified theory showing that whenever the human supervision channel is not sufficient for a latent evaluation target, it acts as an information-reducing channel that induces a strictly positive excess-risk floor for any learner dominated by it. We formalize this Human-Bounded Intelligence limit and show that across six complementary frameworks (operator theory, PAC-Bayes, information theory, causal inference, category theory, and game-theoretic analyses of reinforcement learning from human feedback), non-sufficiency yields strictly positive lower bounds arising from the same structural decomposition into annotation noise, preference distortion, and semantic compression. The theory explains why scaling alone cannot eliminate persistent human-aligned errors and characterizes conditions under which auxiliary non-human signals (e.g., retrieval, program execution, tools) increase effective supervision capacity and collapse the floor by restoring information about the latent target. Experiments on real preference data, synthetic known-target tasks, and externally verifiable benchmarks confirm the predicted structural signatures: human-only supervision exhibits a persistent floor, while sufficiently informative auxiliary channels strictly reduce or eliminate excess error.

人类反馈误差下限信息瓶颈监督机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。