arXiv:2603.00526cs.CV2026-03中稿 · CVPR被引 1

提出异步强化学习框架,提升3D网格生成效率与质量

Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation

  • 设计首个异步在线强化学习框架,训练速度提升3.75倍
  • 提出优势引导排序偏好优化算法,兼顾效率与泛化能力
  • 适用于艺术风格和密集网格生成,适合3D内容创作场景

强化学习在文本和图像生成中已取得显著成果,但在3D生成领域仍鲜有探索。现有方法多采用离线直接偏好优化(DPO),存在训练效率低、泛化能力差的问题。本文旨在提升3D网格生成中强化学习的训练效率与生成质量。首先,设计首个专为3D网格生成优化的异步在线强化学习框架,相较同步强化学习提速3.75倍。其次,提出优势引导排序偏好优化(ARPO)算法,在训练效率与泛化能力之间实现更优平衡,优于当前用于3D网格生成的DPO和组相对策略优化(GRPO)。最后,基于异步ARPO,构建Mesh-Pro模型,引入新型对角感知混合三角-四边形标记化表示与基于射线的几何完整性奖励。实验表明,Mesh-Pro在艺术风格及密集网格生成任务上达到当前最优性能。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has demonstrated remarkable success in text and image generation, yet its potential in 3D generation remains largely unexplored. Existing attempts typically rely on offline direct preference optimization (DPO) method, which suffers from low training efficiency and limited generalization. In this work, we aim to enhance both the training efficiency and generation quality of RL in 3D mesh generation. Specifically, (1) we design the first asynchronous online RL framework tailored for 3D mesh generation post-training efficiency improvement, which is 3.75$\times$ faster than synchronous RL. (2) We propose Advantage-guided Ranking Preference Optimization (ARPO), a novel RL algorithm that achieves a better trade-off between training efficiency and generalization than current RL algorithms designed for 3D mesh generation, such as DPO and group relative policy optimization (GRPO). (3) Based on asynchronous ARPO, we propose Mesh-Pro, which additionally introduces a novel diagonal-aware mixed triangular-quadrilateral tokenization for mesh representation and a ray-based reward for geometric integrity. Mesh-Pro achieves state-of-the-art performance on artistic and dense meshes.

3D生成强化学习网格生成艺术风格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。