arXiv:2502.10046cs.GRcs.CV2025-02

用大模型让虚拟角色头动得更自然,像真人一样看环境、做判断。

ViRAC: A Vision-Reasoning Agent Head Movement Control Framework in Arbitrary Virtual Environments

  • 利用视觉语言模型的常识推理能力,自动生成符合情境的头部动作。
  • 在多个场景中表现优于现有方法,与真人头部运动数据更接近。
  • 无需人工设计规则,适合需要高真实感的虚拟人项目。

创造能与环境自然互动的逼真虚拟角色是计算机图形学的长期目标。本文针对虚拟角色头部旋转这一关键行为,提出基于视觉-推理的大模型驱动框架(ViRAC),通过利用大规模视觉语言模型(VLM)和大语言模型(LLM)中内化的常识知识与推理能力,实现对环境线索的动态响应,而无需显式建模每一种认知机制。相比依赖数据或显著性分析的旧方法,该框架在多样环境中表现出更强的上下文感知能力,避免了行为僵硬或遗漏重要场景元素的问题。实验表明,ViRAC生成的头部运动更自然、更具认知合理性;定量评估显示其与真实人类头部运动数据的对齐度更高,用户研究也证实其在真实感与行为可信度上显著提升。

原文摘要 · Abstract (English)

Creating lifelike virtual agents capable of interacting with their environments is a longstanding goal in computer graphics. This paper addresses the challenge of generating natural head rotations, a critical aspect of believable agent behavior for visual information gathering and dynamic responses to environmental cues. Although earlier methods have made significant strides, many rely on data-driven or saliency-based approaches, which often underperform in diverse settings and fail to capture deeper cognitive factors such as risk assessment, information seeking, and contextual prioritization. Consequently, generated behaviors can appear rigid or overlook critical scene elements, thereby diminishing the sense of realism. In this paper, we propose \textbf{ViRAC}, a \textbf{Vi}sion-\textbf{R}easoning \textbf{A}gent Head Movement \textbf{C}ontrol framework, which exploits the common-sense knowledge and reasoning capabilities of large-scale models, including Vision-Language Models (VLMs) and Large-Language Models (LLMs). Rather than explicitly modeling every cognitive mechanism, ViRAC leverages the biases and patterns internalized by these models from extensive training, thus emulating human-like perceptual processes without hand-tuned heuristics. Experimental results in multiple scenarios reveal that ViRAC produces more natural and context-aware head rotations than recent state-of-the-art techniques. Quantitative evaluations show a closer alignment with real human head-movement data, while user studies confirm improved realism and cognitive plausibility.

虚拟角色视觉推理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。