arXiv:2601.23285cs.ROcs.AI2026-01被引 1

端到端优化用户意图与辅助策略,提升机器人协作的适应性与效率。

End-to-end Optimization of Belief and Policy Learning in Shared Autonomy Paradigms

  • 基于贝叶斯推理与上下文编码,联合优化意图推断与辅助决策。
  • 在复杂场景中实现6.3%更高成功率、41%更优路径效率。
  • 适合需要动态辅助的机器人任务,如人机协作操控与目标模糊环境。

共享自主系统需有效推断用户意图并确定恰当的辅助水平。以往方法依赖固定混合比例或分离意图推断与辅助决策,导致非结构化环境中性能不佳。本文提出BRACE(贝叶斯强化辅助与上下文编码),通过支持端到端梯度传播的架构,联合优化贝叶斯意图推断与上下文自适应辅助。该框架将协同控制策略基于环境上下文和完整的目标概率分布。分析表明:(1)最优辅助水平随目标不确定性增加而降低,随环境约束强度增加而升高;(2)将信念信息融入策略学习可带来相对于串行方法的二次预期遗憾优势。我们在三阶段评估中对比了现有先进方法(IDA、DQN):(1)2D人机回路光标任务中的核心人机交互动态;(2)机械臂的非线性动力学;(3)目标模糊与环境约束下的综合操作。实验结果表明,本方法在成功率上优于现有方法6.3%,路径效率提升41%;相比无辅助控制,成功率提高36.3%,路径效率提升87%。结果证实,集成优化在复杂、目标模糊场景中效益最大,且可在多种需目标导向辅助的机器人领域泛化,推动了自适应共享自主的前沿水平。

原文摘要 · Abstract (English)

Shared autonomy systems require principled methods for inferring user intent and determining appropriate assistance levels. This is a central challenge in human-robot interaction, where systems must be successful while being mindful of user agency. Previous approaches relied on static blending ratios or separated goal inference from assistance arbitration, leading to suboptimal performance in unstructured environments. We introduce BRACE (Bayesian Reinforcement Assistance with Context Encoding), a novel framework that fine-tunes Bayesian intent inference and context-adaptive assistance through an architecture enabling end-to-end gradient flow between intent inference and assistance arbitration. Our pipeline conditions collaborative control policies on environmental context and complete goal probability distributions. We provide analysis showing (1) optimal assistance levels should decrease with goal uncertainty and increase with environmental constraint severity, and (2) integrating belief information into policy learning yields a quadratic expected regret advantage over sequential approaches. We validated our algorithm against SOTA methods (IDA, DQN) using a three-part evaluation progressively isolating distinct challenges of end-effector control: (1) core human-interaction dynamics in a 2D human-in-the-loop cursor task, (2) non-linear dynamics of a robotic arm, and (3) integrated manipulation under goal ambiguity and environmental constraints. We demonstrate improvements over SOTA, achieving 6.3% higher success rates and 41% increased path efficiency, and 36.3% success rate and 87% path efficiency improvement over unassisted control. Our results confirmed that integrated optimization is most beneficial in complex, goal-ambiguous scenarios, and is generalizable across robotic domains requiring goal-directed assistance, advancing the SOTA for adaptive shared autonomy.

共享自主人机交互贝叶斯推理机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。