调整训练数据语言组合,提升跨语言检索性能。
Improving Korean-English Cross-Lingual Retrieval: A Data-Centric Study of Language Composition and Model Merging
- 构建韩英平行语料,测试不同语言组合对检索影响。
- 特定语言配对提升跨语言检索,但会降低单语检索性能。
- 模型融合可平衡两者,适合多语言应用开发人员。
随着多语言文本信息的广泛应用,跨语言信息检索(CLIR)成为关键研究方向。然而,训练数据的语言构成对CLIR和单语信息检索(IR)性能的影响仍缺乏系统研究。为此,我们构建了语言对齐的韩英平行数据集,并采用不同语言组合训练检索模型。实验表明,训练数据的语言构成显著影响检索性能,存在重要跨语言关联:特定语言对能提升CLIR效果,但会损害单语IR表现。我们的研究证明,模型融合能有效缓解这一权衡,在保持单语检索能力的同时实现优异的跨语言检索结果。研究揭示了训练数据语言配置对两类任务的深远影响,并提出模型融合是一种优化多任务性能的有效策略。
原文摘要 · Abstract (English)
With the increasing utilization of multilingual text information, Cross-Lingual Information Retrieval (CLIR) has become a crucial research area. However, the impact of training data composition on both CLIR and Mono-Lingual Information Retrieval (IR) performance remains under-explored. To systematically investigate this data-centric aspect, we construct linguistically parallel Korean-English datasets and train retrieval models with various language combinations. Our experiments reveal that the language composition of training data significantly influences IR performance, exhibiting important inter-lingual correlations: CLIR performance improves with specific language pairs, while Mono-Lingual IR performance declines. Our work demonstrates that Model Merging can effectively mitigate this trade-off, achieving strong CLIR results while preserving Mono-Lingual IR capabilities. Our findings underscore the effects of linguistic configuration of training data on both CLIR and Mono-Lingual IR, and present Model Merging as a viable strategy to optimize performance across these tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。