arXiv:2512.22200cs.LG2025-12

用类情绪信号让智能体自适应,提升真实环境下的学习效率。

Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents

  • 引入类情绪的内生反馈机制,替代传统静态奖励函数。
  • 在非平稳环境中,样本效率提升且收敛更稳定。
  • 适合需要自主探索与持续适应的复杂场景应用。

当前从深度强化学习到大语言模型的主流方法依赖外部定义的静态奖励函数,虽在封闭静态环境中表现超人,但在开放真实环境中却脆弱易崩。标准智能体缺乏内在自主性:无法在稀疏反馈下有效探索,难以应对分布漂移,且需大量人工调参。本文提出,鲁棒自主性的缺失在于缺乏类生物情绪的高阶稳态调控机制。我们提出情感激励学习信号(EILS),一种统一框架,将零散的优化启发式整合为生物启发的内生反馈引擎。不同于将情绪视为语义标签,EILS将情绪建模为连续的稳态评估信号,如好奇心、压力和信心。这些信号作为交互历史衍生的向量状态,实时动态调节智能体的优化景观:好奇心调控熵以防止模式崩溃,压力调节可塑性以克服惰性,信心自适应信任区域以稳定收敛。我们假设该闭环稳态调节机制能显著提升智能体在样本效率与非平稳适应方面的性能。

原文摘要 · Abstract (English)

The ruling method in modern Artificial Intelligence spanning from Deep Reinforcement Learning (DRL) to Large Language Models (LLMs) relies on a surge of static, externally defined reward functions. While this "extrinsic maximization" approach has rendered superhuman performance in closed, stationary fields, it produces agents that are fragile in open-ended, real-world environments. Standard agents lack internal autonomy: they struggle to explore without dense feedback, fail to adapt to distribution shifts (non-stationarity), and require extensive manual tuning of static hyperparameters. This paper proposes that the unaddressed factor in robust autonomy is a functional analog to biological emotion, serving as a high-level homeostatic control mechanism. We introduce Emotion-Inspired Learning Signals (EILS), a unified framework that replaces scattered optimization heuristics with a coherent, bio-inspired internal feedback engine. Unlike traditional methods that treat emotions as semantic labels, EILS models them as continuous, homeostatic appraisal signals such as Curiosity, Stress, and Confidence. We formalize these signals as vector-valued internal states derived from interaction history. These states dynamically modulate the agent's optimization landscape in real time: curiosity regulates entropy to prevent mode collapse, stress modulates plasticity to overcome inactivity, and confidence adapts trust regions to stabilize convergence. We hypothesize that this closed-loop homeostatic regulation can enable EILS agents to outperform standard baselines in terms of sample efficiency and non-stationary adaptation.

强化学习自主智能体情绪模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。