改进跨模态融合与正交投影,提升多语言人脸-语音关联性能
RFOP: Rethinking Fusion and Orthogonal Projection for Face-Voice Association
- 重构双模态融合与正交投影机制,聚焦语义相关特征
- 在英德数据集上达33.1%的EER,排名FAME 2026第三
- 适用于多语言场景下的跨模态身份匹配任务
多语言环境下的人脸-语音关联挑战(FAME 2026)旨在研究多语言场景中的人脸与语音关联任务。该挑战引入英德语种的人脸-语音配对用于评估。为此,我们重新审视了人脸与语音关联中的融合与正交投影方法,通过有效聚焦两个模态内的相关语义信息来提升性能。所提方法在英德数据集上表现优异,最终在FAME 2026挑战中以33.1%的EER位列第三。
原文摘要 · Abstract (English)
Face-voice association in multilingual environment challenge 2026 aims to investigate the face-voice association task in multilingual scenario. The challenge introduces English-German face-voice pairs to be utilized in the evaluation phase. To this end, we revisit the fusion and orthogonal projection for face-voice association by effectively focusing on the relevant semantic information within the two modalities. Our method performs favorably on the English-German data split and ranked 3rd in the FAME 2026 challenge by achieving the EER of 33.1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。