通过语义压缩提升多语言生成式检索的准确率与效率
Multilingual Generative Retrieval via Cross-lingual Semantic Compression
- 将多语言关键词合并为共享语义单元,统一跨语言标识符
- 在mMarco100k和mNQ320k上分别提升6.83%和4.77%准确率
- 减少文档标识符长度74.51%~78.2%,适合多语言信息检索场景
生成式信息检索在单语场景中表现优异,但在多语言场景仍面临跨语言标识符错位和标识符膨胀两大挑战。为此,我们提出基于跨语言语义压缩的多语言生成式检索框架MGR-CSC,将语义等价的多语言关键词统一为共享原子以对齐语义并压缩标识符空间,并引入动态多步约束解码策略。MGR-CSC通过一致的标识符分配增强跨语言对齐,通过减少冗余提升解码效率。实验表明,该方法在mMarco100k和mNQ320k数据集上分别提升检索准确率6.83%和4.77%,同时文档标识符长度分别缩减74.51%和78.2%。
原文摘要 · Abstract (English)
Generative Information Retrieval is an emerging retrieval paradigm that exhibits remarkable performance in monolingual scenarios.However, applying these methods to multilingual retrieval still encounters two primary challenges, cross-lingual identifier misalignment and identifier inflation. To address these limitations, we propose Multilingual Generative Retrieval via Cross-lingual Semantic Compression (MGR-CSC), a novel framework that unifies semantically equivalent multilingual keywords into shared atoms to align semantics and compresses the identifier space, and we propose a dynamic multi-step constrained decoding strategy during retrieval. MGR-CSC improves cross-lingual alignment by assigning consistent identifiers and enhances decoding efficiency by reducing redundancy. Experiments demonstrate that MGR-CSC achieves outstanding retrieval accuracy, improving by 6.83% on mMarco100k and 4.77% on mNQ320k, while reducing document identifiers length by 74.51% and 78.2%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。