arXiv:2512.15957cs.CVcs.AI2025-12中稿 · IEEE/CVF Winter Co…

用视觉语言模型预测多人行为,提升机器人环境理解能力

Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models

  • 融合视觉上下文与场景图空间信息,增强人-环境交互预测
  • 在合成与真实数据上提升66.9%预测准确率,优于现有基线
  • 适合需要多主体行为理解的机器人导航与交互任务

准确预测人类行为对移动机器人在人群环境中运行至关重要。以往研究多聚焦于从第一人称视角预测单人动作,但许多机器人应用需从第三人称视角理解多人行为。为此,我们提出CAMP-VLM(上下文感知的多人行为预测):一种基于视觉语言模型(VLM)的框架,结合视觉输入中的上下文特征与场景图提供的空间意识,以增强对人-场景交互的预测。由于缺乏适合第三人称视角多人行为预测的数据集,我们使用逼真模拟器生成的合成数据对CAMP-VLM进行微调,并在合成与真实序列上评估模型泛化能力。通过监督微调(SFT)和直接偏好优化(DPO),CAMP-VLM在预测准确率上相比最佳基线最高提升66.9%。

原文摘要 · Abstract (English)

Accurately predicting human behaviors is crucial for mobile robots operating in human-populated environments. While prior research primarily focuses on predicting actions in single-human scenarios from an egocentric view, several robotic applications require understanding multiple human behaviors from a third-person perspective. To this end, we present CAMP-VLM (Context-Aware Multi-human behavior Prediction): a Vision Language Model (VLM)-based framework that incorporates contextual features from visual input and spatial awareness from scene graphs to enhance prediction of humans-scene interactions. Due to the lack of suitable datasets for multi-human behavior prediction from an observer view, we perform fine-tuning of CAMP-VLM with synthetic human behavior data generated by a photorealistic simulator, and evaluate the resulting models on both synthetic and real-world sequences to assess their generalization capabilities. Leveraging Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), CAMP-VLM outperforms the best-performing baseline by up to 66.9% in prediction accuracy.

行为预测视觉语言模型多主体交互机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。