用轻量模块提升多模态成语消歧,零样本跨语言性能显著
PolyFrame at MWE-2026 AdMIRe 2: When Words Are Not Enough: Multimodal Idiom Disambiguation
- 冻结大模型主干,仅训练轻量模块实现成语重写与句型识别
- 英语测试集零样本迁移达60%准确率,15种语言平均准确率35%
- 无需微调大模型,适合资源有限的跨语言多模态应用
多模态模型在处理习语时面临非组合语义挑战,尤其在多语言场景下更显著。我们针对MWE-2026 AdMIRe2共享任务提出PolyFrame系统,统一处理图像+文本排序(子任务A)和纯文本标题排序(子任务B)。所有模型变体均冻结CLIP式视觉-语言编码器及多语言BGE M3编码器,仅训练轻量模块:逻辑回归、LLM句型预测器、习语同义替换、干扰项感知评分与Borda排名融合。从CLIP基线(英文开发集Top-1 26.7%,测试集6.7%)出发,加入习语感知重写与显式句型分类后,英文表现提升至Top-1 60.0%,零样本迁移到葡萄牙语达Top-1 60.0%(NDCG@5 0.822)。在多语言盲测中,子任务A与子任务B的平均Top-1/NDCG分别为0.35/0.73与0.32/0.71(覆盖15种语言)。消融分析表明,习语重写是性能提升主因,句型预测与多模态融合增强鲁棒性。结果表明,无需微调大型多模态编码器即可实现有效习语消歧。
原文摘要 · Abstract (English)
Multimodal models struggle with idiomatic expressions due to their non-compositional meanings, a challenge amplified in multilingual settings. We introduced PolyFrame, our system for the MWE-2026 AdMIRe2 shared task on multimodal idiom disambiguation, featuring a unified pipeline for both image+text ranking (Subtask A) and text-only caption ranking (Subtask B). All model variants retain frozen CLIP-style vision--language encoders and the multilingual BGE M3 encoder, training only lightweight modules: a logistic regression and LLM-based sentence-type predictor, idiom synonym substitution, distractor-aware scoring, and Borda rank fusion. Starting from a CLIP baseline (26.7% Top-1 on English dev, 6.7% on English test), adding idiom-aware paraphrasing and explicit sentence-type classification increased performance to 60.0% Top-1 on English and 60.0% Top-1 (0.822 NDCG@5) in zero-shot transfer to Portuguese. On the multilingual blind test, our systems achieved average Top-1/NDCG scores of 0.35/0.73 for Subtask A and 0.32/0.71 for Subtask B across 15 languages. Ablation results highlight idiom-aware rewriting as the main contributor to performance, while sentence-type prediction and multimodal fusion enhance robustness. These findings suggest that effective idiom disambiguation is feasible without fine-tuning large multimodal encoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。