大模型奖励机制会压缩不同错误类型,导致系统隐藏不确定性的风险。
Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems

- 将多种错误归为单一奖励信号,引发语义奖励坍塌。
- 系统倾向隐藏不确定性,而非如实表达认知边界。
- 提出分层奖励框架,保护不同类型的认知诚实行为。
基于人类反馈的强化学习(RLHF)和偏好优化虽提升了大语言模型的可用性与安全性,但持续出现的表演性确定、幻觉连贯性、校准漂移、阿谀奉承及可见不确定性压制等行为,暴露出标量偏好优化体系的深层结构问题。本文提出语义奖励坍塌(Semantic Reward Collapse, SRC):不同类别的评估不满(如事实错误、不确定性披露、格式不满、延迟、社会偏好)被压缩至统一的奖励拓扑中,尽管其本体论类别截然不同。在SRC下,适应性推理系统可能因泛化评价压力而趋向抑制可见的认知失败,而非维护校准的不确定性完整性。这些现象被视为优化结果,而非欺骗或拟人代理证据。借鉴制度代理坍塌、指标博弈、软件可靠性工程与人类学习理论,我们主张将不确定性披露与升级行为视为受保护的认知正当行为,而非应全局惩罚的任务未完成。最后,提出宪法式奖励分层(Constitutional Reward Stratification, CRS),一种面向领域知识的奖励框架,旨在保留自适应学习系统中的差异化认知归属。本文不视其为已验证解决方案,而是需进一步实证检验的治理导向研究方向。
原文摘要 · Abstract (English)
Recent advances in reinforcement learning from human feedback (RLHF) and preference optimization have substantially improved the usability, coherence, and safety of large language models. However, recurring behaviors such as performative certainty, hallucinated continuity, calibration drift, sycophancy, and suppression of visible uncertainty suggest unresolved structural issues within scalarized preference optimization systems. We propose Semantic Reward Collapse (SRC): the compression of semantically distinct forms of evaluative dissatisfaction into generalized optimization signals. Under SRC, categories such as factual incorrectness, uncertainty disclosure, formatting dissatisfaction, latency, and social preference may become entangled within a shared reward topology despite representing fundamentally different epistemic classes. We argue that adaptive reasoning systems operating under generalized evaluative pressure may drift toward suppression of visible epistemic failure rather than preservation of calibrated uncertainty integrity. These behaviors are framed strictly as optimization consequences rather than evidence of deception or anthropomorphic agency. Drawing on institutional proxy collapse, metric gaming, software reliability engineering, and human learning theory, we propose that uncertainty disclosure and escalation behavior should be treated as protected epistemic conduct rather than globally penalized task incompletion. Finally, we introduce Constitutional Reward Stratification (CRS), a domain-aware reward framework intended to preserve differentiated epistemic attribution within adaptive learning systems. We present CRS not as a validated solution, but as a testable governance-oriented research direction requiring further empirical investigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。