arXiv:2608.04622cs.CV2026-08

用双智能体协作生成人体,视角大变时仍保持细节清晰

DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation

论文配图:DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation
图 1 · 摘自论文原文
  • 设计双智能体:一个推理隐含特征,一个感知视角错位
  • 在深衣和市场1501数据集上,视角剧变时纹理对齐率提升显著
  • 适合需要高保真人体生成的场景,如虚拟试衣、动画制作

AI智能体已成为生成图像的新范式,使系统能进行复杂语义推理而非被动像素映射。在姿态引导的人体生成中,传统方法在剧烈视角变化下不可避免产生严重视觉伪影,根源在于缺乏逻辑推断未见区域和建模复杂空间形变的认知能力。为此,我们提出DAC-Pose,一种新型代理驱动的多模态框架,将单视图人体生成重构为协同双智能体系统。该框架集成两个互补组件:先验语义推理(PSR)智能体与差异感知视觉编码(DAVE)智能体。PSR作为认知引擎,通过协作推理推断未见区域的细粒度属性;同时,DAVE作为专用视觉感知智能体,量化并编码视角引起的空间错位,持续将鲁棒的空间约束反馈至生成过程。这一语义推断与视觉感知间的自主反馈回路,确保了高保真细节合成。在DeepFashion和Market-1501基准上的大量实验验证了该代理驱动范式的优越性。特别地,DAC-Pose在剧烈视角变化下显著提升了纹理对齐与身份一致性。代码已公开于https://github.com/AIVRC/DAC-Pose。

原文摘要 · Abstract (English)

AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than passive pixel-level mapping. In pose-guided human generation, conventional methods inevitably produce severe visual artifacts under drastic viewpoint shifts, fundamentally because they lack the cognitive capacity to logically deduce unseen regions and model complex spatial deformations. To bridge this gap, we propose DAC-Pose, a novel agent-driven multimodal framework that reformulates single-view human generation as a collaborative dual-agent system. DAC-Pose integrates two complementary components, namely, the Prior Semantic Reasoning (PSR) agent and the Discrepancy-Aware Visual Encoding (DAVE) agent. Functioning as a cognitive engine, PSR utilizes collaborative reasoning to deduce the fine-grained attributes of unseen regions. Concurrently, acting as a specialized visual perception agent, DAVE quantifies and encodes viewpoint-induced spatial misalignments, continuously feeding robust spatial constraints back into the generative process. This autonomous feedback loop between semantic deduction and visual perception ensures high-fidelity detail synthesis. Extensive experiments on the DeepFashion and Market-1501 benchmarks validate the superiority of our agent-driven paradigm. Notably, DAC-Pose excels in preserving texture alignment and identity consistency under drastic viewpoint changes. The code is available at https://github.com/AIVRC/DAC-Pose.

人体生成智能体协作姿态引导视觉对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。