提出EqualizeIR框架,缓解检索模型对语言复杂度的偏差问题。
EqualizeIR: Mitigating Linguistic Biases in Retrieval Models
- 用语言偏差弱学习器捕捉数据中的语言偏见
- 通过正则化与预测优化提升模型整体性能
- 适合关注公平性与鲁棒性的检索系统研究者
本研究发现,现有信息检索(IR)模型在语言结构简单或复杂的查询上表现不均,对复杂或简单语言的查询存在显著性能偏差。为此,我们提出EqualizeIR框架,通过构建语言偏差弱学习器来识别数据中的语言偏见,并利用该学习器对强模型进行正则化与预测修正,防止其过度拟合特定语言模式。我们提出了四种构建语言偏差模型的方法。在多个数据集上的大量实验表明,该方法有效降低了语言简单与复杂查询间的性能差距,同时提升了整体检索效果。
原文摘要 · Abstract (English)
This study finds that existing information retrieval (IR) models show significant biases based on the linguistic complexity of input queries, performing well on linguistically simpler (or more complex) queries while underperforming on linguistically more complex (or simpler) queries. To address this issue, we propose EqualizeIR, a framework to mitigate linguistic biases in IR models. EqualizeIR uses a linguistically biased weak learner to capture linguistic biases in IR datasets and then trains a robust model by regularizing and refining its predictions using the biased weak learner. This approach effectively prevents the robust model from overfitting to specific linguistic patterns in data. We propose four approaches for developing linguistically-biased models. Extensive experiments on several datasets show that our method reduces performance disparities across linguistically simple and complex queries, while improving overall retrieval performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。