arXiv:2605.07971cs.CVcs.LG2026-05被引 1

用离散扩散模型生成和编辑3D体素,提升可解释性和编辑效率。

DVD: Discrete Voxel Diffusion for 3D Generation and Editing

论文配图:DVD: Discrete Voxel Diffusion for 3D Generation and Editing
图 1 · 摘自论文原文
  • 将体素占用视为离散变量,避免连续转离散的阈值问题。
  • 通过预测熵识别模糊区域,支持数据筛选与质量评估。
  • 轻量微调实现单轮采样内体素修复与编辑,无需额外计算。

我们提出离散体素扩散(DVD),用于基于SLat的3D生成与编辑流水线中的稀疏体素生成、评估与编辑。尽管离散扩散在图像类生成中尚未取代连续扩散,但本工作证明其可作为稀疏体素骨架的有效第一阶段先验。通过将体素占用视为原生离散变量,DVD避免了连续到离散的阈值转换,提供简洁的体素生成、不确定性估计与编辑框架。除质量提升外,显式类别建模使生成过程更具可解释性。进一步地,我们利用预测熵作为鲁棒的不确定性度量,识别模糊体素区域与复杂样本,支持数据过滤与质量评估。最后,提出一种基于块结构扰动模式的轻量级微调策略,使模型在单次采样中即可完成体素补全与编辑,仅需极小额外计算且无需额外模型评估。代码已开源:https://github.com/TeCai/DVD。

原文摘要 · Abstract (English)

We introduce Discrete Voxel Diffusion (DVD), a discrete diffusion framework to generate, assess, and edit sparse voxels for SLat (Structured LATent) based 3D generative pipelines. Although discrete diffusion has not generally displaced continuous diffusion in image-like generation, we show that it can be an effective first-stage prior for sparse voxel scaffolds. By treating voxel occupancy as a native discrete variable, DVD avoids continuous-to-discrete thresholding and provides a simple framework for voxel generation, uncertainty estimation, and editing. Beyond quality gains, DVD provides more interpretable generation dynamics through explicit categorical modeling. Furthermore, we leverage the predictive entropy as a robust uncertainty metric to identify ambiguous voxel regions and complicated samples, facilitating tasks such as data filtering and quality assessment. Finally, we propose a lightweight fine-tuning strategy using block-structured perturbation patterns. This approach empowers the model to inpaint and edit voxels within a single sampling round, requiring negligible auxiliary computation and no additional model evaluations. Code is available at https://github.com/TeCai/DVD.

3D生成扩散模型体素编辑不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。