只用几秒关键视频片段,就能生成高质量3D会说话人脸。
ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation
- 自动筛选音频、口型、视角最丰富的短片段作为参考
- 处理和训练时间减少5倍以上,质量几乎不变
- 适合需要快速生成个性化3D人脸的实用场景
基于神经辐射场(NeRF)和3D高斯泼溅(3DGS)的说话人脸生成方法在个性化头像合成方面取得了显著进展。然而,现有方法通常需要数分钟的参考视频进行精细预处理和拟合,导致准备时间长达数小时,限制了实际应用。本文重新审视一个基础但未被充分探索的问题:高质量个性化说话人脸生成是否真需数分钟长的参考视频?我们的探索性研究发现,仅需精心挑选的几秒参考片段,即可达到与完整视频相当的效果。这表明参考数据的信息量比时长更重要。受此启发,我们提出ISExplore(信息片段探索)策略,通过音频特征多样性、唇部运动幅度和视角多样性三个维度,自动识别最具信息量的短参考片段。大量实验表明,该方法使基于NeRF和3DGS的方法数据处理与训练时间减少超5倍,同时保持高保真生成质量。本方法为个性化说话人脸生成提供了高效实用的解决方案,并揭示了3D说话人脸生成中的数据效率新思路。
原文摘要 · Abstract (English)
Talking Face Generation (TFG) methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have recently achieved impressive progress in personalized talking head synthesis. However, existing methods typically require several minutes of reference video for meticulous preprocessing and fitting, resulting in hours of preparation time and limiting their practical applicability. In this paper, we revisit a fundamental yet underexplored question: do high-quality personalized TFG models truly require minutes-long reference videos? Our exploratory study reveals that a carefully selected reference segment of only a few seconds can often achieve performance comparable to that of using the full reference video. This finding suggests that the informativeness of reference data is more critical than its duration. Motivated by this observation, we propose ISExplore (Informative Segment Explore), a simple yet effective segment selection strategy that automatically identifies the most informative short reference segment based on three key data quality dimensions: audio feature diversity, lip movement amplitude, and viewpoint diversity. Extensive experiments demonstrate that ISExplore reduces data processing and training time by over 5x for both NeRF- and 3DGS-based methods, while preserving high-fidelity generation quality. Our method provides a practical and efficient solution for personalized TFG and offers new insights into data efficiency in 3D talking face generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。