arXiv:2601.12310cs.AI2026-01

用环境存活机制替代奖励,实现无需人工标注的稳定自训练。

Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection

  • 以环境存活为唯一筛选标准,行为通过真实资源约束下的持续影响被选择。
  • 模型自发形成可复现策略,实现长期改进,避免语义漂移与奖励作弊。
  • 适合研究自主系统、具身智能和无监督学习的开发者与研究员。

自训练系统常因缺乏外部数据质量评判标准而退化,导致奖励欺骗与语义漂移。本文提出一种在稀疏外部反馈和有限记忆条件下仍稳定的自训练架构,并实证分析其学习动态与失效模式。该架构仅通过环境生存性来中介学习:候选行为在真实资源约束下执行,只有那些既持久存在又能维持未来交互可能性的行为才被保留。环境不提供语义反馈、密集奖励或任务特定监督;选择仅通过行为作为世界改变事件的差异化存活实现,使代理优化不可行,奖励作弊在演化上不稳定。语义动态分析表明,改进主要源于有效重复策略在巩固与修剪机制下的持续留存,我们称之为负空间学习(NSL);模型还自发发展出元学习策略(如故意失败以获取信息性错误反馈),且无需显式指导。本工作证明,基于环境的选择能实现可持续的开放式自我提升,为构建无需人类标注数据或复杂奖励设计的鲁棒通用自主系统提供了可行路径。

原文摘要 · Abstract (English)

Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic drift. This paper provides a proof-of-concept system architecture for stable self-training under sparse external feedback and bounded memory, and empirically characterises its learning dynamics and failure modes. We introduce a self-training architecture in which learning is mediated exclusively by environmental viability, rather than by reward, objective functions, or externally defined fitness criteria. Candidate behaviours are executed under real resource constraints, and only those whose environmental effects both persist and preserve the possibility of future interaction are propagated. The environment does not provide semantic feedback, dense rewards, or task-specific supervision; selection operates solely through differential survival of behaviours as world-altering events, making proxy optimisation impossible and rendering reward-hacking evolutionarily unstable. Analysis of semantic dynamics shows that improvement arises primarily through the persistence of effective and repeatable strategies under a regime of consolidation and pruning, a paradigm we refer to as negative-space learning (NSL), and that models develop meta-learning strategies (such as deliberate experimental failure in order to elicit informative error messages) without explicit instruction. This work establishes that environment-grounded selection enables sustainable open-ended self-improvement, offering a viable path toward more robust and generalisable autonomous systems without reliance on human-curated data or complex reward shaping.

自训练具身智能元学习无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。