arXiv:2512.08294cs.CV2025-12被引 3

用视频数据提升图像生成中主体身份一致性,尤其在复杂场景下表现更好。

OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation

论文配图:OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
图 1 · 摘自论文原文
  • 基于视频跨帧一致性构建250万样本数据集,强化主体身份先验
  • 通过分割引导补画与框引导修复,保留主体身份并支持灵活编辑
  • 适用于需要高保真主体还原的图像生成与多主体场景应用

尽管主体驱动图像生成取得进展,现有模型仍易偏离参考身份,且在多主体复杂场景中表现不佳。为此,我们提出OpenSubject,一个基于视频构建的大规模语料库,包含250万样本和435万张图像,用于主体驱动生成与操作。该数据集通过四阶段流程构建:(i) 视频筛选,采用分辨率与美学过滤获取高质量片段;(ii) 跨帧主体挖掘与配对,结合视觉语言模型进行类别一致、局部定位与多样性感知配对;(iii) 身份保持参考图像合成,引入分割图引导补画生成生成输入,框引导修复生成操作输入,并结合几何增强与不规则边界侵蚀;(iv) 验证与标注,使用视觉语言模型验证合成样本,失败样本按阶段(iii)重合成,并构建短/长文本描述。此外,我们建立基准评测体系,评估身份保真度、提示遵循度、操作一致性与背景一致性。大量实验表明,使用OpenSubject训练可显著提升复杂场景下的生成与操作性能。

原文摘要 · Abstract (English)

Despite the promising progress in subject-driven image generation, current models often deviate from the reference identities and struggle in complex scenes with multiple subjects. To address this challenge, we introduce OpenSubject, a video-derived large-scale corpus with 2.5M samples and 4.35M images for subject-driven generation and manipulation. The dataset is built with a four-stage pipeline that exploits cross-frame identity priors. (i) Video Curation. We apply resolution and aesthetic filtering to obtain high-quality clips. (ii) Cross-Frame Subject Mining and Pairing. We utilize vision-language model (VLM)-based category consensus, local grounding, and diversity-aware pairing to select image pairs. (iii) Identity-Preserving Reference Image Synthesis. We introduce segmentation map-guided outpainting to synthesize the input images for subject-driven generation and box-guided inpainting to generate input images for subject-driven manipulation, together with geometry-aware augmentations and irregular boundary erosion. (iv) Verification and Captioning. We utilize a VLM to validate synthesized samples, re-synthesize failed samples based on stage (iii), and then construct short and long captions. In addition, we introduce a benchmark covering subject-driven generation and manipulation, and then evaluate identity fidelity, prompt adherence, manipulation consistency, and background consistency with a VLM judge. Extensive experiments show that training with OpenSubject improves generation and manipulation performance, particularly in complex scenes.

图像生成主体一致视频数据多主体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。