arXiv:2606.01671cs.CL2026-06ACL

用混合专家模型提升低资源语言习语理解能力

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models

论文配图:When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models
图 1 · 摘自论文原文
  • 设计混合专家框架,融合选中与未选专家输出以缓解稀疏性
  • 在多模态模型上实现5%-6%性能提升,更好保留习语文化含义
  • 适合研究多语言习语、跨文化理解的AI学者

在多语言教育背景下,习语学习是理解创意、文化价值观、历史背景及多元视角的重要途径。本文聚焦于印地语、孟加拉语和泰语等低资源东南亚语言中的习语建模难题,这些语言因深层隐喻复杂性而难以进行计算建模与跨语言迁移。为此,我们构建了包含3,533条多语言习语的多模态习语语料库Varnika,涵盖七种习语语气,并融合文本与视觉表征。提出HybridMoE混合专家框架,在集成多个习语专家意见的同时,通过受控混合机制整合选中与未选专家输出,缓解专家稀疏问题,并引入掩码多模态嵌入的习语属性信号增强推理。为全面评估,设计IDIO-TONE与习语验证得分,采用三阶段评价体系:(i)字面翻译保真度,(ii)视觉-语义对齐,(iii)习语意义保留。实证结果表明,HybridMoE在先进视觉语言模型上实现5–6%性能提升,显著改善多语言多模态场景下比喻语言与文化嵌入意义的表征能力。

原文摘要 · Abstract (English)

In the contemporary epoch of multilingual education, learning idioms provides a fascinating gateway towards creativity, cultural values, historical context, and diverse perspectives inherent to various linguistic traditions. This paper showcases the navigation of retaining figurative and cultural semantics in low-resource Southeast Asian languages such as Hindi, Bengali, and Thai, where culturally rich idioms pose significant obstacles for computational modeling and cross-linguistic transfer due to their deep metaphorical complexity. To tackle such complexity, we present Varnika, a reconstructed multimodal idiom corpus comprising 3,533 multilingual idioms, enriched with seven idiomatic tones aligned with both textual and visual representations. Additionally, to infer informative idiomatic understanding, we introduce a Hybrid Mixture-of-Experts (HybridMoE) framework that embeds multiple idiomatic expert opinions while mitigating expert sparsity by integrating outputs from both selected and unselected experts through controlled hybridization, further augmented with Idiomatic Property Signals via masked multimodal embeddings. To analyze the performance across multiple dimensions, we propose the IDIO-TONE and Idiomatic Validation Score, a three-stage evaluation pipeline measuring (i) literal translation fidelity, (ii) visual-semantic alignment, and (iii) idiomatic meaning retention. Empirical evaluations highlight that HybridMoE achieves 5--6\% performance gains across advanced vision language models, demonstrating improved representation of figurative language and culturally embedded meaning in multilingual multimodal settings

习语理解多模态混合专家低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。