用生成模型重构决策框架,提升复杂行为建模能力。
Generative Models in Decision Making: A Survey
- 按功能划分控制器、模型、优化器、评估器四类角色
- 支持多模态行为建模,解决传统RL表达力不足问题
- 适合研究通用物理智能与高风险场景应用
生成模型从根本上改变了决策范式,将问题从单一奖励最大化转变为高保真轨迹生成与分布匹配。这一转变克服了经典强化学习中单峰策略分布表达力有限的缺陷,尤其在捕捉多样化数据集中的复杂多模态行为时更具优势。然而现有研究常将这些模型视为孤立算法改进,缺乏统一框架整合。本文提出基于控制即推断的概率框架,通过变分分解轨迹后验,定义四种功能角色:用于策略推理的控制器、用于动态先验的模型、用于迭代轨迹优化的优化器、以及用于轨迹引导与价值评估的评估器。该功能中心框架使我们能跨维度批判性分析代表性生成方法。同时考察其在具身智能、自动驾驶和科学人工智能等高风险领域的部署,揭示世界模型中的动态幻觉与代理利用等系统性风险。最后,展望通向通用物理智能的路径,指出推断效率、可信度及物理基础模型涌现的关键挑战。
原文摘要 · Abstract (English)
Generative models have fundamentally reshaped the landscape of decision-making, reframing the problem from pure scalar reward maximization to high-fidelity trajectory generation and distribution matching. This paradigm shift addresses intrinsic limitations in classical Reinforcement Learning (RL), particularly the limited expressivity of standard unimodal policy distributions in capturing complex, multi-modal behaviors embedded in diverse datasets. However, current literature often treats these models as isolated algorithmic improvements, rarely synthesizing them into a single comprehensive framework. This survey proposes a principled taxonomy grounding generative decision-making within the probabilistic framework of Control as Inference. By performing a variational factorization of the trajectory posterior, we conceptualize four distinct functional roles: Controllers for amortized policy inference, Modelers for dynamics priors, Optimizers for iterative trajectory refinement, and Evaluators for trajectory guidance and value assessment. Unlike existing architecture-centric reviews, this function-centric framework allows us to critically analyze representative generative families across distinct dimensions. Furthermore, we examine deployment in high-stakes domains, specifically Embodied AI, Autonomous Driving, and AI for Science, highlighting systemic risks such as dynamics hallucination in world models and proxy exploitation. Finally, we chart the path toward Generalist Physical Intelligence, identifying pivotal challenges in inference efficiency, trustworthiness, and the emergence of Physical Foundation Models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。