打造可任意搭配服饰的虚拟试穿模型,支持多品类、多场景真实生成。
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

- 构建专用数据引擎与三阶段训练流程,实现多参考图像的风格一致生成
- 在单件与多件试穿任务中均达顶尖水平,真实感与身份一致性出色
- 支持自由组合服饰并遵循编辑指令,适合电商与设计场景使用
我们提出Oxygen-TryOn,一种统一的任意物品虚拟试穿基础模型。不同于通用图像编辑器,该模型专为试穿设计,通过专用数据引擎和针对性训练实现。给定一个或多个参考物品(干净产品图或真实穿着图)及一张目标主体图像,即可合成主体穿戴各类时尚单品的逼真图像,覆盖几乎所有服装类别。此前系统仅限特定品类且需实验室环境,近期多参考方法仍以单品为中心;而Oxygen-TryOn支持多样物品与场景,包括全身/半身视图、可变数量参考图、自由多件组合,并精准保留主体身份与物品外观。我们摒弃基于掩码的修复方式,将试穿重构为多参考、理解驱动的生成任务。构建了大规模高质量试穿数据收集、制造、标注与过滤的数据引擎,设计了持续预训练(CPT)、监督微调(SFT)与强化学习(RL)三阶段训练方案。其中,强化学习阶段采用混合奖励机制,结合自研试穿奖励模型与基于评分标准的通用模型,共同优化细节一致性与指令级质量。模型还能在同一流程中执行姿态变化等通用编辑指令。在公开基准与自建Oxygen-TryOn Bench上,其单件试穿表现达到当前最优,多件试穿领先,超越主流闭源系统(Nano Banana Pro, GPT-Image-2, Seedream5 Lite)与开源模型(FLUX.2)。
原文摘要 · Abstract (English)
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。