arXiv:2504.00839cs.ROcs.AI2025-04中稿 · IEEE International…被引 5

用多模态大模型预测人类行为,提升机器人交互的上下文理解能力。

Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights

  • 构建模块化框架,融合多模态输入与上下文学习提升预测能力。
  • 在目标帧上达到92.8%语义相似度和66.1%精确标签准确率。
  • 适用于需理解复杂场景的人机协作任务,如服务机器人。

在共享环境中预测人类行为对实现安全高效的机器人交互至关重要。传统数据驱动方法依赖特定领域数据集、活动类型和预测时长进行预训练。相比之下,大语言模型(LLMs)的突破为描述各类人类活动并实现跨场景泛化预测提供了可能。特别是多模态大语言模型(MLLMs)能整合多种信息源,增强上下文感知与场景理解能力。然而,直接应用通用MLLM进行预测面临输入序列处理能力有限、提示设计敏感及微调成本高等挑战。本文系统分析了预训练MLLM在上下文感知人类行为预测中的应用,提出一种模块化多模态人类活动预测框架,可评估不同MLLM、输入变体、上下文学习(ICL)及自回归技术。评估显示,最佳配置在目标帧上达到92.8%语义相似度和66.1%精确标签准确率。

原文摘要 · Abstract (English)

Predicting human behavior in shared environments is crucial for safe and efficient human-robot interaction. Traditional data-driven methods to that end are pre-trained on domain-specific datasets, activity types, and prediction horizons. In contrast, the recent breakthroughs in Large Language Models (LLMs) promise open-ended cross-domain generalization to describe various human activities and make predictions in any context. In particular, Multimodal LLMs (MLLMs) are able to integrate information from various sources, achieving more contextual awareness and improved scene understanding. The difficulty in applying general-purpose MLLMs directly for prediction stems from their limited capacity for processing large input sequences, sensitivity to prompt design, and expensive fine-tuning. In this paper, we present a systematic analysis of applying pre-trained MLLMs for context-aware human behavior prediction. To this end, we introduce a modular multimodal human activity prediction framework that allows us to benchmark various MLLMs, input variations, In-Context Learning (ICL), and autoregressive techniques. Our evaluation indicates that the best-performing framework configuration is able to reach 92.8% semantic similarity and 66.1% exact label accuracy in predicting human behaviors in the target frame.

行为预测多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。