用阿拉伯语训练的模型,能准确识别英德语人脸与声音匹配。
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
- 基于ImageBind+LoRA的多模态基础模型,实现跨语言关联。
- 仅用阿拉伯语音训练,却在英德语测试集上达24.73%错误率。
- 适合做多语言语音-人脸对齐的轻量级部署系统。
本文介绍提交至FAME2026挑战赛的UZH-CL系统。挑战聚焦于多语言环境下跨模态验证,特别是未见语言和未听语言场景。我们探索两种架构:从零训练的双编码器基线系统,采用对比损失和正交投影损失;以及基于ImageBind的多模态基础模型结合LoRA微调。为应对数据稀缺与语言限制,我们从VoxBlink收集了外部阿拉伯语数据集。最佳系统ImageBind-LoRA展现出卓越的跨语言泛化能力:尽管仅在阿拉伯语音上微调,仍取得评估集(英语和德语)24.73%的等错误率(EER),在竞赛中获得第二名。
原文摘要 · Abstract (English)
This paper describes the UZH-CL system submitted to the FAME2026 Challenge. The challenge focuses on cross-modal verification under unique multilingual conditions, specifically unseen and unheard languages. Our approach investigates two distinct architectures, consisting of a baseline dual-encoder system trained from scratch using contrastive and orthogonal projection losses, and a foundation model approach leveraging ImageBind with LoRA. To address the data scarcity and language constraints of the challenge, we curated an external Arabic dataset from VoxBlink. Our best-performing system, ImageBind-LoRA, demonstrates remarkable cross-lingual generalization: despite being fine-tuned exclusively on Arabic audio, it achieved an EER of 24.73% on the evaluation set (English and German), securing 2nd place in the competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。