arXiv:2511.11231cs.CVcs.AI2025-11

用扩散模型提升眼神重定向精度,生成更自然的训练数据。

3D Gaussian and Diffusion-Based Gaze Redirection

  • 结合扩散Transformer与弱监督中间视角,实现平滑眼神变化。
  • 在真实数据上将眼神误差降低4.1%,达到6.353度新低。
  • 适合需要高质量合成数据的视觉模型训练者使用。

高保真眼神重定向对生成增强数据以提升眼神估计算法泛化能力至关重要。当前基于3D高斯点阵(3DGS)的模型如GazeGaussian已达领先水平,但在渲染细微连续眼神偏移时仍存在困难。本文提出DiT-Gaze框架,融合扩散Transformer(DiT)、跨眼神角度的弱监督策略以及正交性约束损失。DiT提升图像生成质量,弱监督通过合成中间眼神角度构建平滑的眼神方向流形;正交性约束损失则数学上强制分离眼神、头部姿态与表情的内部表征。大量实验表明,DiT-Gaze在感知质量与重定向精度上均达新SOTA,将现有眼神误差降低4.1%,降至6.353度,并为合成训练数据提供更优方案。代码与模型将向研究社区开放以供基准测试。

原文摘要 · Abstract (English)

High-fidelity gaze redirection is critical for generating augmented data to improve the generalization of gaze estimators. 3D Gaussian Splatting (3DGS) models like GazeGaussian represent the state-of-the-art but can struggle with rendering subtle, continuous gaze shifts. In this paper, we propose DiT-Gaze, a framework that enhances 3D gaze redirection models using a novel combination of Diffusion Transformer (DiT), weak supervision across gaze angles, and an orthogonality constraint loss. DiT allows higher-fidelity image synthesis, while our weak supervision strategy using synthetically generated intermediate gaze angles provides a smooth manifold of gaze directions during training. The orthogonality constraint loss mathematically enforces the disentanglement of internal representations for gaze, head pose, and expression. Comprehensive experiments show that DiT-Gaze sets a new state-of-the-art in both perceptual quality and redirection accuracy, reducing the state-of-the-art gaze error by 4.1% to 6.353 degrees, providing a superior method for creating synthetic training data. Our code and models will be made available for the research community to benchmark against.

眼神重定向扩散模型3D高斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。