arXiv:2508.15047cs.AIcs.GR2025-08被引 2

用大模型驱动对话与导航,让虚拟人群自然形成社交行为。

Emergent Crowds Dynamics from Language-Driven Multi-Agent Interactions

  • 用大语言模型生成角色对话,结合性格和情绪决策行动
  • 群体自发分组与解散,实现复杂社会互动
  • 适合做虚拟现实、游戏或城市仿真中的智能人群

基于代理的群集动画与仿真已较为成熟,但多数方法仅依赖转向和固定目标规划,忽视了语言与社交互动对人类行为的影响。本文提出一种新方法,利用大语言模型(LLMs)控制代理的运动。该方法包含对话系统与语言驱动导航两部分:定期根据角色性格、身份、欲望及关系,调用以代理为中心的LLM生成交互对话;再结合对话内容、性格、情绪、视觉与身体状态,动态决定每个代理的导航与转向。实验在两个复杂场景中验证,观察到代理自动分组与拆散,且对话成为信息传递机制。结果表明,该框架可在任意环境设置下生成更真实的群集模拟,使社交行为自然涌现。

原文摘要 · Abstract (English)

Animating and simulating crowds using an agent-based approach is a well-established area where every agent in the crowd is individually controlled such that global human-like behaviour emerges. We observe that human navigation and movement in crowds are often influenced by complex social and environmental interactions, driven mainly by language and dialogue. However, most existing work does not consider these dimensions and leads to animations where agent-agent and agent-environment interactions are largely limited to steering and fixed higher-level goal extrapolation. We propose a novel method that exploits large language models (LLMs) to control agents' movement. Our method has two main components: a dialogue system and language-driven navigation. We periodically query agent-centric LLMs conditioned on character personalities, roles, desires, and relationships to control the generation of inter-agent dialogue when necessitated by the spatial and social relationships with neighbouring agents. We then use the conversation and each agent's personality, emotional state, vision, and physical state to control the navigation and steering of each agent. Our model thus enables agents to make motion decisions based on both their perceptual inputs and the ongoing dialogue. We validate our method in two complex scenarios that exemplify the interplay between social interactions, steering, and crowding. In these scenarios, we observe that grouping and ungrouping of agents automatically occur. Additionally, our experiments show that our method serves as an information-passing mechanism within the crowd. As a result, our framework produces more realistic crowd simulations, with emergent group behaviours arising naturally from any environmental setting.

多智能体语言驱动群集仿真大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。