arXiv:2606.27305cs.CV2026-06

用人类偏好直接优化3D人脸生成,不依赖网格或文本

Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aware Face GAN

论文配图:Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aware Face GAN
图 1 · 摘自论文原文
  • 直接基于密度场奖励微调3D人脸生成器
  • 单个标注者偏好即可让74.4%生成结果更受欢迎
  • 无需网格、文本或多视角图,保留2D外观

现有3D生成的强化学习人类反馈(RLHF)多针对显式表面表示,通常需将辐射场转为网格并依赖大量表面监督数据。本文直接对预训练的3D感知生成模型进行微调,利用从神经辐射场(NeRF)密度值(σ)学习到的奖励信号,无需外部网格或形状先验。该奖励模型无需预训练,仅用少量偏好样本即可快速训练,并显著提升3D几何质量。在无条件3D感知人脸生成模型EG3D上,奖励直接读取连续3D密度场,仅提供几何优化信号,无需文本条件、网格提取或多视角渲染。通过密度一致性约束,在保持2D外观基本不变的同时重塑几何结构,2D质量略有下降(FID-50k从4.09升至6.66)。作为概念验证,由单个标注者偏好训练的生成器,在成对比较中获得74.4%的人类偏好。

原文摘要 · Abstract (English)

Reinforcement learning from human feedback (RLHF) for 3D generation is now established across a number of works, but most existing pipelines optimise explicit surface representations, often by converting radiance fields into meshes and training heavily on surface-supervised data. We instead fine-tune a pretrained 3D-aware generative model directly from a learned reward over radiance-field density ($σ$) values, with no externally supplied mesh or shape prior. The reward model requires no pretraining, trains easily on a small set of preference samples, and yields robust improvement in 3D geometry. Working on an unconditional 3D-aware face GAN (EG3D), our reward reads the continuous 3D density field of the neural radiance field (NeRF) directly and supplies a geometry-only learning signal, requiring neither text conditioning, mesh extraction, nor multi-view rendering. A density-consistency constraint keeps the 2D appearance qualitatively similar while the geometry is reshaped, at a measurable but bounded distributional cost (FID-50k rises from 4.09 to 6.66): the fine-tuned generator, trained from the preferences of a single annotator as a proof of concept, produces face geometries preferred by users in 74.4% of pairwise comparisons.

3D生成人脸生成强化学习人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。