arXiv:2603.23964cs.AI2026-03

分析2000+篇论文,揭示强化学习环境从物理模拟到语言驱动的演进规律

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments

  • 构建多维分类体系,量化分析2000+核心文献的演化路径
  • 发现领域分裂为大模型主导的语义先验与特定任务泛化两大生态
  • 识别跨任务协同与零样本泛化的认知特征,助力下一代智能体设计

强化学习的显著进展与其训练评估环境密切相关。本文通过程序化处理海量学术文献,系统梳理超过2000篇核心论文,提出一种定量方法,描绘从孤立物理仿真向通用、语言驱动的基础智能体演进的图景。采用新型多维分类体系,对基准测试在不同应用领域及所需认知能力上的表现进行系统分析。自动化语义与统计分析揭示出一个数据验证的深刻范式转变:领域分化为以大语言模型(LLMs)为主导的‘语义先验’生态系统和‘特定领域泛化’生态系统。此外,我们刻画了两类领域的‘认知指纹’,揭示跨任务协同、多领域干扰及零样本泛化的内在机制。本研究为设计下一代具身语义模拟器提供严谨的量化路线图,弥合连续物理控制与高层逻辑推理之间的鸿沟。

原文摘要 · Abstract (English)

The remarkable progress of reinforcement learning (RL) is intrinsically tied to the environments used to train and evaluate artificial agents. Moving beyond traditional qualitative reviews, this work presents a large-scale, data-driven empirical investigation into the evolution of RL environments. By programmatically processing a massive corpus of academic literature and rigorously distilling over 2,000 core publications, we propose a quantitative methodology to map the transition from isolated physical simulations to generalist, language-driven foundation agents. Implementing a novel, multi-dimensional taxonomy, we systematically analyze benchmarks against diverse application domains and requisite cognitive capabilities. Our automated semantic and statistical analysis reveals a profound, data-verified paradigm shift: the bifurcation of the field into a "Semantic Prior" ecosystem dominated by Large Language Models (LLMs) and a "Domain-Specific Generalization" ecosystem. Furthermore, we characterize the "cognitive fingerprints" of these distinct domains to uncover the underlying mechanisms of cross-task synergy, multi-domain interference, and zero-shot generalization. Ultimately, this study offers a rigorous, quantitative roadmap for designing the next generation of Embodied Semantic Simulators, bridging the gap between continuous physical control and high-level logical reasoning.

强化学习智能体环境演化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。