arXiv:2604.03738cs.CV2026-04中稿 · CVPR被引 2

用位置编码解决多角色视频生成中的混淆问题

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation

  • 用位置编码作为额外控制信号,精准匹配参考图像中的角色
  • 在多个相似角色场景下,视频跨镜头一致性提升23.6%
  • 适合需要精细角色控制的影视动画生成任务

近期如Sora2等专有模型在多镜头、多参考角色视频生成方面展现出显著进展,但学术研究仍较薄弱。本文针对该任务发现核心挑战:当参考图像外观高度相似时,模型易产生参考混淆,语义相似的令牌会削弱正确上下文检索能力。为此,提出PoCo(Position Embedding as a Context Controller),将位置编码作为超越语义检索的额外上下文控制手段。通过利用令牌的侧信息,PoCo实现细粒度的令牌级匹配,同时保持隐式语义一致性建模。基于PoCo,构建了能可靠控制视觉特征极相似角色的多参考、多镜头视频生成模型。大量实验表明,相比多种基线方法,PoCo显著提升了跨镜头一致性和参考保真度。

原文摘要 · Abstract (English)

Recent proprietary models such as Sora2 demonstrate promising progress in generating multi-shot videos conditioned on multiple reference characters. However, academic research on this problem remains limited. We study this task and identify a core challenge: when reference images exhibit highly similar appearances, the model often suffers from reference confusion, where semantically similar tokens degrade the model's ability to retrieve the correct context. To address this, we introduce PoCo (Position Embedding as a Context Controller), which incorporates position encoding as additional context control beyond semantic retrieval. By employing side information of tokens, PoCo enables precise token-level matching while preserving implicit semantic consistency modeling. Building on PoCo, we develop a multi-reference and multi-shot video generation model capable of reliably controlling characters with extremely similar visual traits. Extensive experiments demonstrate that PoCo improves cross-shot consistency and reference fidelity compared with various baselines.

视频生成多角色控制位置编码一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。