arXiv:2510.08096cs.CV2025-10中稿 · VCIP 2025

用3D高斯点云自动修正极端角度人脸分割标签,提升模型鲁棒性。

Efficient Label Refinement for Face Parsing Under Extreme Poses Using 3D Gaussian Splatting

  • 通过双3D高斯模型联合拟合图像与初始分割图,实现多视角一致性约束。
  • 仅需少量初始图像和简单后处理,即可生成高质量标注数据。
  • 无需3D真值标注,适合真实场景中复杂姿态的人脸分割优化。

极端视角下的人脸分割因标注数据稀缺而面临挑战,手动标注成本高且难以规模化。本文提出一种新型标签精炼流程,利用3D高斯点云(3DGS)从噪声多视角预测中生成准确的分割掩码。通过联合拟合两个3DGS模型——一个用于RGB图像,一个用于初始分割图——我们的方法利用共享几何结构强制多视角一致性,从而合成具有多样化姿态的训练数据,仅需少量后期处理。在该精炼数据集上微调人脸分割模型后,对挑战性头姿的准确性显著提升,同时保持标准视图下的优异性能。大量实验(包括人工评估)表明,尽管无需真实3D标注且仅使用少量初始图像,本方法仍优于现有最先进方法。该方案为提升真实场景中人脸分割的鲁棒性提供了可扩展、高效的新路径。

原文摘要 · Abstract (English)

Accurate face parsing under extreme viewing angles remains a significant challenge due to limited labeled data in such poses. Manual annotation is costly and often impractical at scale. We propose a novel label refinement pipeline that leverages 3D Gaussian Splatting (3DGS) to generate accurate segmentation masks from noisy multiview predictions. By jointly fitting two 3DGS models, one to RGB images and one to their initial segmentation maps, our method enforces multiview consistency through shared geometry, enabling the synthesis of pose-diverse training data with only minimal post-processing. Fine-tuning a face parsing model on this refined dataset significantly improves accuracy on challenging head poses, while maintaining strong performance on standard views. Extensive experiments, including human evaluations, demonstrate that our approach achieves superior results compared to state-of-the-art methods, despite requiring no ground-truth 3D annotations and using only a small set of initial images. Our method offers a scalable and effective solution for improving face parsing robustness in real-world settings.

人脸分割3D高斯标签精炼多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。