用多模态融合预测人对AI配音的感知,效果接近人工评分。
Can Hierarchical Cross-Modal Fusion Predict Human Perception of AI Dubbed Content?
- 分层融合音视频文本特征,捕捉同步、语义、情感等细节。
- 在1.2万条双语配音数据上训练,与人工评分相关性超0.75。
- 轻量适配器实现高效微调,适合大规模自动评估场景。
评估AI生成的配音内容本质上是多维度的,受同步性、可理解性、说话人一致性、情感契合度和语义上下文影响。人类平均意见得分(MOS)仍是黄金标准,但成本高且难以规模化。本文提出一种分层多模态架构,整合音频、视频与文本的互补线索。模型从音频中提取说话人身份、语调等细粒度特征,从视频中捕获面部表情与场景线索,从文本中获取语义上下文,并通过模内与模间层级逐步融合。轻量级LoRA适配器实现跨模态参数高效微调。为克服主观标签稀缺问题,通过主动学习优化权重,聚合客观指标生成代理MOS。该架构在1.2万条印地语-英语双向配音片段上训练,并用人工MOS进行微调,实现强感知一致性(PCC > 0.75),提供了一种可扩展的AI配音自动评估方案。
原文摘要 · Abstract (English)
Evaluating AI generated dubbed content is inherently multi-dimensional, shaped by synchronization, intelligibility, speaker consistency, emotional alignment, and semantic context. Human Mean Opinion Scores (MOS) remain the gold standard but are costly and impractical at scale. We present a hierarchical multimodal architecture for perceptually meaningful dubbing evaluation, integrating complementary cues from audio, video, and text. The model captures fine-grained features such as speaker identity, prosody, and content from audio, facial expressions and scene-level cues from video and semantic context from text, which are progressively fused through intra and inter-modal layers. Lightweight LoRA adapters enable parameter-efficient fine-tuning across modalities. To overcome limited subjective labels, we derive proxy MOS by aggregating objective metrics with weights optimized via active learning. The proposed architecture was trained on 12k Hindi-English bidirectional dubbed clips, followed by fine-tuning with human MOS. Our approach achieves strong perceptual alignment (PCC > 0.75), providing a scalable solution for automatic evaluation of AI-dubbed content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。