arXiv:2609.07137cs.CVcs.AI2026-09

用多教师强化学习提升3D生成模型的几何质量

Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer

论文配图:Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer
图 1 · 摘自论文原文
  • 通过多教师蒸馏与在线策略蒸馏,融合多个专家模型能力
  • 在所有几何质量指标上均优于教师模型,平均性能更优
  • 适合需要高精度3D生成的研究者和工业应用

基于流匹配扩散Transformer(DiT)的图像到3D生成模型可生成高保真网格,但其后训练策略仍不明确。强化学习中存在两大瓶颈:难以定义全面的3D几何质量奖励函数,以及联合优化异构目标时的梯度干扰。受大语言模型和图像生成中在线策略蒸馏(OPD)可行性的启发,我们提出Flow3D-OPD,一种两阶段后训练框架,将多教师蒸馏引入3D几何生成。第一阶段,利用半策略增强预训练模型基础能力,并设计代理验证器评估3D几何质量;基于该验证器,通过直接偏好优化(DPO)培养领域专用教师模型。第二阶段,通过硬任务路由采样与梯度累积的在线策略蒸馏,将异构专长整合至统一学生模型,缓解联合优化中的梯度干扰。无需复杂修改,该方法在所有几何质量维度上均实现一致提升,且平均性能超越所有教师模型。大量实验表明,该方法为3D生成中的强化学习提供有效范式。

原文摘要 · Abstract (English)

Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fidelity meshes, yet their post-training strategy remains largely unexplored. There exist several critical bottlenecks in reinforcement learning: the inherent difficulty of defining comprehensive rewards for 3D geometric quality, and the gradient interference that arises when jointly optimizing heterogeneous objectives. Inspired by the practicability of on-policy distillation (OPD) in large language models and image generation, we propose \textbf{Flow3D-OPD}, a two-stage post-training framework that introduces multi-teacher distillation into 3D geometry generation. In the first stage, we utilize the semi-policy to enhance the foundational capability of the pretrained model and then design an agentic verifier for 3D geometric quality evaluation. Based on the verifier, we could cultivate domain-specialized teacher models via direct preference optimization (DPO). In the second stage, we consolidate heterogeneous expertise into a unified student model through on-policy distillation with hard task-routing sampling and gradient accumulation, which could mitigate the gradient interference in joint optimization. Without relying on elaborate modifications, our straightforward yet effective design achieves consistent improvements across all geometric quality dimensions and surpasses all teacher models in the average metric. Extensive experiments demonstrate that our approach provides an effective paradigm for reinforcement learning in 3D generation.

3D生成扩散模型强化学习蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。