arXiv:2512.02554cs.CV2025-12被引 1

生成保持身份一致的行人图像,提升重识别数据质量

OmniPerson: Unified Identity-Preserving Pedestrian Generation

  • 统一生成模型支持多模态、多参考图与文本控制
  • 多参考融合器实现任意数量参考图的身份保真生成
  • 首个大规模可控行人生成数据集,适合数据增强研究者

行人重识别(ReID)因数据隐私和标注成本高,缺乏大规模高质量训练数据。现有生成方法难以保证身份一致性且控制能力弱,限制了数据增强效果。为此,我们提出OmniPerson,首个面向可见光/红外图像与视频ReID任务的统一身份保真行人生成框架。贡献包括:1)提出OmniPerson统一生成模型,支持RGB/IR模态、任意数量参考图、两种人体姿态及文本控制,并具备RGB到IR转换与图像超分辨率能力;2)设计多参考融合器(Multi-Refer Fuser),从多视角参考图中提炼统一身份,确保生成行人高保真;3)构建首个大规模多参考可控行人生成数据集PersonSyn,提出自动化标注流程,将公开的仅含身份信息的ReID基准转化为富含密集多模态监督的数据资源。实验表明,OmniPerson在行人生成上达到最新水平,视觉质量和身份一致性均领先。用其生成数据增强现有数据集,可持续提升ReID模型性能。代码、预训练模型与数据集将开源。

原文摘要 · Abstract (English)

Person re-identification (ReID) suffers from a lack of large-scale high-quality training data due to challenges in data privacy and annotation costs. While previous approaches have explored pedestrian generation for data augmentation, they often fail to ensure identity consistency and suffer from insufficient controllability, thereby limiting their effectiveness in dataset augmentation. To address this, We introduce OmniPerson, the first unified identity-preserving pedestrian generation pipeline for visible/infrared image/video ReID tasks. Our contributions are threefold: 1) We proposed OmniPerson, a unified generation model, offering holistic and fine-grained control over all key pedestrian attributes. Supporting RGB/IR modality image/video generation with any number of reference images, two kinds of person poses, and text. Also including RGB-to-IR transfer and image super-resolution abilities.2) We designed Multi-Refer Fuser for robust identity preservation with any number of reference images as input, making OmniPerson could distill a unified identity from a set of multi-view reference images, ensuring our generated pedestrians achieve high-fidelity pedestrian generation.3) We introduce PersonSyn, the first large-scale dataset for multi-reference, controllable pedestrian generation, and present its automated curation pipeline which transforms public, ID-only ReID benchmarks into a richly annotated resource with the dense, multi-modal supervision required for this task. Experimental results demonstrate that OmniPerson achieves SoTA in pedestrian generation, excelling in both visual fidelity and identity consistency. Furthermore, augmenting existing datasets with our generated data consistently improves the performance of ReID models. We will open-source the full codebase, pretrained model, and the PersonSyn dataset.

行人生成身份一致数据增强多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。