arXiv:2606.10902cs.CVcs.AI2026-06

让图像生成模型精准控制定制对象的姿势,避免姿态不准或外观不一致。

Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject Customization

论文配图:Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject Customization
图 1 · 摘自论文原文
  • 用多组图像-姿态配对实现无需微调的3D感知上下文学习
  • 在真实物体和3D资产上,姿态准确率和身份一致性显著提升
  • 适合需要精确姿势控制的个性化图像生成场景

主体定制是现代图像生成的基础任务。通过提供少量参考图像和文本提示,用户可生成特定对象在任意场景中的图像。然而,现有方法在定制主体的姿态控制方面仍存在不足,常出现姿态不准或跨姿态外观不一致的问题。这表明2D主干网络对物体的体感理解仍具挑战性。为此,我们提出Pose-ICL,一种无需微调的框架,利用3D感知的上下文学习(ICL)通过多组图像-姿态配对直接适应新主体。其核心机制——表面锚定位置嵌入(SAPE),通过将图像标记锚定在体积分箱的表面坐标,赋予模型显式的3D感知能力。专门设计的优化使其与现有DiT模型无缝兼容。在3D资产和真实主体上的大量评估表明,Pose-ICL在姿态准确性和身份一致性上均显著优于当前方法。

原文摘要 · Abstract (English)

Subject Customization is a foundational task in modern image generation. By providing a few reference images and a text prompt, users can generate images of a specific object in any desired scene. However, existing methods still struggle to achieve effective pose control for customized subjects. In practice, they often exhibit inaccurate poses or inconsistent cross-pose appearances. These limitations suggest that understanding objects in a volumetric manner remains a significant challenge for 2D-native backbones. To address this challenge, we propose Pose-ICL, a tuning-free framework that leverages 3D-aware In-Context Learning (ICL) to directly adapt to new subjects through multiple paired image-pose references. Its core mechanism,Surface-Anchored Position Embedding (SAPE), equips the model with explicit 3D awareness by anchoring image tokens to the surface coordinates of a volumetric bounding box. Dedicated refinements ensure its seamless compatibility with existing DiT models. Extensive evaluations on both 3D assets and real-world subjects demonstrate that Pose-ICL significantly outperforms current methods in both pose accuracy and identity consistency.

姿态控制主体定制3D感知上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。