让视频中人与物的互动更自然,支持抽象对象如标志的精准生成。
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

- 用多模态大模型增强语义理解,统一处理人-物关系与参考图像输入
- 在自注意力中引入全局多模态引导,提升特征对齐精度
- 适合需要高保真人物和复杂交互生成的视频创作场景
人-物中心视频个性化(HOCVP)是主体驱动视频生成的核心任务。现有方法存在两大瓶颈:一是跨主体个性化中难以兼顾主体保真度与人-物交互准确性,尤其在物体为抽象概念(如品牌标志)时;二是虽有内部参考(如OCR图、多视角输入)可提升保真度,但缺乏理解其潜在对应关系的机制。为此,我们提出HOMIE框架,统一处理跨主体与内主体输入。相比先前方法,HOMIE采用更优的多模态大模型(MLLM)融合策略,在不牺牲文本编码器可控性或避免昂贵重对齐的前提下,提取参考级关系知识。具体地,我们在自注意力中引入全局多模态引导,更好对齐MLLM生成的语义特征与VAE令牌;同时提出模态-参考嵌入,区分来自MLLM特征与VAE令牌的表示,并关联内主体参考图像令牌。大量实验验证,本方法在各类HOCVP任务上均达到领先性能。
原文摘要 · Abstract (English)
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。