arXiv:2602.01880cs.RO2026-02中稿 · January 2026被引 1

让机器人通过视觉理解家庭情境,做出符合清洁、舒适与安全价值的实时决策。

Multimodal Large Language Models for Real-Time Situated Reasoning

  • 用GPT-4o结合机器人视觉输入,判断何时该开始清洁。
  • 在真实家居环境中,仅凭有限视觉信息推断上下文与用户偏好。
  • 适合关注智能机器人自主决策与人机价值观对齐的研究者。

本文研究多模态大语言模型如何支持实时情境与价值感知决策。我们结合GPT-4o语言模型与TurtleBot 4平台,模拟智能家居清洁机器人。模型通过视觉输入评估环境,判断是否适宜启动清洁。系统展现出对家庭活动、社会规范及用户偏好的推理能力,并能做出符合清洁、舒适、安全等价值的细致决策。我们在真实家居环境中演示了该系统,证明其可从有限视觉输入中推断上下文与价值。结果表明,多模态大语言模型在提升机器人自主性与情境感知方面前景广阔,但也面临一致性、偏见与实时性能挑战。

原文摘要 · Abstract (English)

In this work, we explore how multimodal large language models can support real-time context- and value-aware decision-making. To do so, we combine the GPT-4o language model with a TurtleBot 4 platform simulating a smart vacuum cleaning robot in a home. The model evaluates the environment through vision input and determines whether it is appropriate to initiate cleaning. The system highlights the ability of these models to reason about domestic activities, social norms, and user preferences and take nuanced decisions aligned with the values of the people involved, such as cleanliness, comfort, and safety. We demonstrate the system in a realistic home environment, showing its ability to infer context and values from limited visual input. Our results highlight the promise of multimodal large language models in enhancing robotic autonomy and situational awareness, while also underscoring challenges related to consistency, bias, and real-time performance.

多模态机器人决策大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。