arXiv:2604.02348cs.LG2026-04中稿 · AAMAS 2025被引 1

让强化学习理解环境与自身状态的动态互动,提升真实世界泛化能力。

Contextual Intelligence The Next Leap for Reinforcement Learning

  • 区分环境施加与智能体驱动的上下文,构建分层认知框架。
  • 提出多时间尺度建模与抽象上下文融合,支持跨场景零样本迁移。
  • 适合研究通用智能体、安全部署系统或复杂环境控制的学者。

强化学习在游戏、机器人和连续控制中取得了显著成果,但其策略常难以超越训练分布泛化,限制了实际应用。近期情境强化学习(cRL)表明,让智能体感知环境特征(即上下文)可提升零样本迁移能力。然而,现有方法将上下文视为单一静态观测,制约了泛化性能。本文提出新的上下文分类法,区分外源性(环境施加)与自源性(智能体驱动)因素,并指出三个关键研究方向:(1) 学习异构上下文,显式利用分层结构以实现对世界影响的双向推理;(2) 多时间尺度建模,识别外源变量缓慢变化或静止,而自源变量可在单个回合内快速变化,需采用不同学习机制;(3) 集成抽象高阶上下文,如角色、资源与规制制度、不确定性等非物理描述符,这些因素对行为具有决定性影响。我们主张将上下文作为第一性建模原语,使智能体能够推理自身身份、环境约束及其随时间演化,推动新一代情境感知智能体的发展,实现安全高效的真实世界部署。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has produced spectacular results in games, robotics, and continuous control. Yet, despite these successes, learned policies often fail to generalize beyond their training distribution, limiting real-world impact. Recent work on contextual RL (cRL) shows that exposing agents to environment characteristics -- contexts -- can improve zero-shot transfer. So far, the community has treated context as a monolithic, static observable, an approach that constrains the generalization capabilities of RL agents. To achieve contextual intelligence we first propose a novel taxonomy of contexts that separates allogenic (environment-imposed) from autogenic (agent-driven) factors. We identify three fundamental research directions that must be addressed to promote truly contextual intelligence: (1) Learning with heterogeneous contexts to explicitly exploit the taxonomy levels so agents can reason about their influence on the world and vice versa; (2) Multi-time-scale modeling to recognize that allogenic variables evolve slowly or remain static, whereas autogenic variables may change within an episode, potentially requiring different learning mechanisms; (3) Integration of abstract, high-level contexts to incorporate roles, resource & regulatory regimes, uncertainties, and other non-physical descriptors that crucially influence behavior. We envision context as a first-class modeling primitive, empowering agents to reason about who they are, what the world permits, and how both evolve over time. By doing so, we aim to catalyze a new generation of context-aware agents that can be deployed safely and efficiently in the real world.

强化学习上下文感知智能体泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。