arXiv:2512.01340cs.CV2025-12被引 3

首个多主体说话人生成质量评估框架,提升多角色口型同步与身份一致性。

EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans

  • 构建首个大规模多主体说话人图像质量评估数据集,含5492个样本。
  • 识别12类常见失真,发现多角色生成中存在显著画质退化问题。
  • 融合多模态同步感知,实现对整体质量、人物特征和身份一致性的精准评估。

语音驱动的说话人生成(Talker)当前在多主体驱动能力上存在局限。将该范式扩展为可同时驱动多个主体的“多说话人”(Multi-Talker),能增强音视频交互的丰富性与沉浸感。然而现有多说话人模型仍因技术限制导致明显质量下降,影响用户体验。为此,我们构建了首个大规模多说话人生成质量评估数据集THQA-MT,包含15个代表性多说话人模型生成的5492个多主体说话人图像(MTHs),使用400张在线收集的真实人像作为驱动源。通过主观实验分析不同多说话人之间的感知差异,识别出12种常见失真类型。进一步提出EvalTalker框架,具备全局质量、人类特征与身份一致性感知能力,并集成Qwen-Sync实现多模态同步感知。实验表明,EvalTalker与主观评分具有更强相关性,为高质量多说话人生成与评估研究提供了坚实基础。

原文摘要 · Abstract (English)

Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker," capable of animating multiple subjects simultaneously, introduces richer interactivity and stronger immersion in audiovisual communication. However, current Multi-Talkers still exhibit noticeable quality degradation caused by technical limitations, resulting in suboptimal user experiences. To address this challenge, we construct THQA-MT, the first large-scale Multi-Talker-generated Talking Human Quality Assessment dataset, consisting of 5,492 Multi-Talker-generated THs (MTHs) from 15 representative Multi-Talkers using 400 real portraits collected online. Through subjective experiments, we analyze perceptual discrepancies among different Multi-Talkers and identify 12 common types of distortion. Furthermore, we introduce EvalTalker, a novel TH quality assessment framework. This framework possesses the ability to perceive global quality, human characteristics, and identity consistency, while integrating Qwen-Sync to perceive multimodal synchrony. Experimental results demonstrate that EvalTalker achieves superior correlation with subjective scores, providing a robust foundation for future research on high-quality Multi-Talker generation and evaluation.

说话人生成多主体驱动质量评估多模态同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。