让机器人理解上下文中的用户偏好,实现安全高效的自适应导航。
Interpreting Context-Aware Human Preferences for Multi-Objective Robot Navigation
- 用视觉语言模型和大语言模型解析环境与自然语言反馈,生成可更新的行为规则。
- 将上下文与规则转化为偏好向量,实时调整多目标强化学习策略。
- 在真实场景中验证,提升人机协作的透明性与可控性,适合交互式服务机器人。
在人类共存环境中运行的机器人不仅需达成安全、高效等任务目标,还需根据人类偏好调整行为。然而,人类偏好通常以自然语言表达且依赖环境上下文,难以直接融入底层控制策略。本文提出一种端到端框架,结合基础模型与多目标强化学习(MORL)导航策略,实现高层语义推理与底层运动控制的融合。视觉语言模型(VLM)从车载视觉观测中提取结构化环境上下文,大语言模型(LLM)将用户自然语言反馈转换为可解释的、上下文相关的规则,并存储于持久可更新的规则记忆中。偏好翻译模块将上下文信息与规则映射为数值偏好向量,用于参数化预训练的MORL策略,实现实时导航适应。通过组件级定量评估、用户研究及多种室内环境的真实机器人部署,结果表明该系统能可靠捕捉用户意图,生成一致的偏好向量,并在不同场景下实现可控行为调整。整体框架提升了机器人在共享环境中的适应性、透明性和可用性,同时保持安全、实时的控制响应。
原文摘要 · Abstract (English)
Robots operating in human-shared environments must not only achieve task-level navigation objectives such as safety and efficiency, but also adapt their behavior to human preferences. However, as human preferences are typically expressed in natural language and depend on environmental context, it is difficult to directly integrate them into low-level robot control policies. In this work, we present a pipeline that enables robots to understand and apply context-dependent navigation preferences by combining foundational models with a Multi-Objective Reinforcement Learning (MORL) navigation policy. Thus, our approach integrates high-level semantic reasoning with low-level motion control. A Vision-Language Model (VLM) extracts structured environmental context from onboard visual observations, while Large Language Models (LLM) convert natural language user feedback into interpretable, context-dependent behavioral rules stored in a persistent but updatable rule memory. A preference translation module then maps contextual information and stored rules into numerical preference vectors that parameterize a pretrained MORL policy for real-time navigation adaptation. We evaluate the proposed framework through quantitative component-level evaluations, a user study, and real-world robot deployments in various indoor environments. Our results demonstrate that the system reliably captures user intent, generates consistent preference vectors, and enables controllable behavior adaptation across diverse contexts. Overall, the proposed pipeline improves the adaptability, transparency, and usability of robots operating in shared human environments, while maintaining safe and responsive real-time control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。