arXiv:2604.17632cs.IR2026-04ACL被引 1

研究混合语言查询检索,发现现有系统在跨语言切换时表现大幅下降。

Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers

论文配图:Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers
图 1 · 摘自论文原文
  • 构建人工标注的混合语言检索数据集CSR-L,真实还原语言混用场景。
  • 多模型评估显示,跨语言切换使检索性能最高下降27%。
  • 现有多语言技术无法完全解决此问题,揭示当前系统脆弱性。

代码切换是全球交流中普遍存在的语言现象,但现代信息检索系统仍主要针对单语环境设计与评估。为弥合这一关键差距,我们开展了一项全面的代码切换信息检索研究。提出CSR-L(代码切换检索基准-精简版),通过人工标注构建数据集,以捕捉混合语言查询的真实自然性。在统计、密集和晚期交互范式下的评估表明,代码切换构成根本性性能瓶颈,即使强大的多语言模型也显著退化。我们证明,这种失败源于纯语言与混合语言文本在嵌入空间中的显著差异。进一步扩展研究,提出CS-MTEB,涵盖11种多样化任务,观察到性能最高下降27%。最后,我们表明标准多语言技术如词汇扩展无法完全弥补这些缺陷。这些发现突显了当前系统的脆弱性,并确立代码切换为未来信息检索优化的关键前沿。

原文摘要 · Abstract (English)

Code-switching is a pervasive linguistic phenomenon in global communication, yet modern information retrieval systems remain predominantly designed for, and evaluated within, monolingual contexts. To bridge this critical disconnect, we present a holistic study dedicated to code-switching IR. We introduce CSR-L (Code-Switching Retrieval benchmark-Lite), constructing a dataset via human annotation to capture the authentic naturalness of mixed-language queries. Our evaluation across statistical, dense, and late-interaction paradigms reveals that code-switching acts as a fundamental performance bottleneck, degrading the effectiveness of even robust multilingual models. We demonstrate that this failure stems from substantial divergence in the embedding space between pure and code-switched text. Scaling this investigation, we propose CS-MTEB, a comprehensive benchmark covering 11 diverse tasks, where we observe performance declines of up to 27%. Finally, we show that standard multilingual techniques like vocabulary expansion are insufficient to resolve these deficits completely. These findings underscore the fragility of current systems and establish code-switching as a crucial frontier for future IR optimization.

信息检索代码切换多语言评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。