arXiv:2603.25415cs.AIcs.RO2026-03

改进智能体导航策略,提升语义场景图生成效率与完整性

Modernising Reinforcement Learning-Based Navigation for Embodied Semantic Scene Graph Generation

  • 采用细粒度动作分解与多头策略,优化探索决策
  • 相比基线,场景图完整度提升21%,碰撞率显著降低
  • 适合研究具身智能、场景理解与强化学习应用者

语义世界模型使具身智能体能超越几何表征,对物体、关系和空间上下文进行推理。在有机计算中,这类模型是目标驱动自适应的核心,尤其在不确定性与资源受限环境下。核心挑战是在有限行动预算内获取最大化模型质量与下游效用的观测。语义场景图(SSGs)为此提供结构化且紧凑的表示。然而,在有限行动周期内构建高质量的SSG,需权衡信息增益与导航成本,并判断何时额外行动收益递减。本文提出一种模块化具身语义场景图生成导航组件,通过替换策略优化算法并重新审视离散动作设计,研究了紧凑与更细粒度的动作集,比较单头策略与分因子多头策略在原子动作上的表现。评估课程学习与可选深度碰撞监督,考察了SSG完整度、执行安全性及导航行为。结果表明,仅更换优化算法即在相同奖励设计下使SSG完整度相对提升21%。深度信息主要影响执行安全(无碰撞运动),而完整度基本不变。结合现代优化与细粒度分因子动作表示,实现最佳完整度-效率平衡。

原文摘要 · Abstract (English)

Semantic world models enable embodied agents to reason about objects, relations, and spatial context beyond purely geometric representations. In Organic Computing, such models are a key enabler for objective-driven self-adaptation under uncertainty and resource constraints. The core challenge is to acquire observations maximising model quality and downstream usefulness within a limited action budget. Semantic scene graphs (SSGs) provide a structured and compact representation for this purpose. However, constructing them within a finite action horizon requires exploration strategies that trade off information gain against navigation cost and decide when additional actions yield diminishing returns. This work presents a modular navigation component for Embodied Semantic Scene Graph Generation and modernises its decision-making by replacing the policy-optimisation method and revisiting the discrete action formulation. We study compact and finer-grained, larger discrete motion sets and compare a single-head policy over atomic actions with a factorised multi-head policy over action components. We evaluate curriculum learning and optional depth-based collision supervision, and assess SSG completeness, execution safety, and navigation behaviour. Results show that replacing the optimisation algorithm alone improves SSG completeness by 21\% relative to the baseline under identical reward shaping. Depth mainly affects execution safety (collision-free motion), while completeness remains largely unchanged. Combining modern optimisation with a finer-grained, factorised action representation yields the strongest overall completeness--efficiency trade-off.

具身智能场景图生成强化学习导航策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。