发现稠密检索的鲁棒性与有效性存在可量化的缩放规律,优化策略可打破二者权衡。
On the Scaling of Robustness and Effectiveness in Dense Retrieval
- 通过实验揭示鲁棒性与有效性均遵循幂律缩放规律。
- 不同优化策略下,鲁棒性与有效性可共同提升,实现帕累托最优。
- 调整优化权重可实现高效联合训练,无需额外资源即可媲美多倍资源扩展。
鲁棒性与有效性是构建真实应用场景中稠密检索模型的关键因素。已知二者存在权衡关系。近期研究揭示了有效性随模型规模和数据规模增长的幂律关系。鲁棒性是否也遵循缩放规律?若然,能否同时提升二者而不陷入权衡?为此,我们开展全面实验研究。发现:(i) 鲁棒性(包括分布外与对抗鲁棒性)同样遵循缩放规律。(ii) 鲁棒性与有效性的缩放模式不同,联合提升需付出显著资源成本。因此,我们转向第三种影响因素——优化策略。发现:(i) 采用不同优化策略时,鲁棒性与有效性的联合性能呈现帕累托前沿;(ii) 当优化策略偏离帕累托效率,联合性能以次优方向缩放;(iii) 通过调节优化权重实现帕累托效率,可达成帕累托训练,使联合性能缩放最为高效。即使不增加资源,该方法性能也相当于在过度偏重某一方的策略下多次资源扩展的结果。最后,验证了这些发现对实际部署中高效、平衡的稠密检索模型具有指导意义。
原文摘要 · Abstract (English)
Robustness and Effectiveness are critical aspects of developing dense retrieval models for real-world applications. It is known that there is a trade-off between the two. Recent work has addressed scaling laws of effectiveness in dense retrieval, revealing a power-law relationship between effectiveness and the size of models and data. Does robustness follow scaling laws too? If so, can scaling improve both robustness and effectiveness together, or do they remain locked in a trade-off? To answer these questions, we conduct a comprehensive experimental study. We find that:(i) Robustness, including out-of-distribution and adversarial robustness, also follows a scaling law.(ii) Robustness and effectiveness exhibit different scaling patterns, leading to significant resource costs when jointly improving both. Given these findings, we shift to the third factor that affects model performance, namely the optimization strategy, beyond the model size and data size. We find that: (i) By fitting different optimization strategies, the joint performance of robustness and effectiveness traces out a Pareto frontier. (ii) When the optimization strategy strays from Pareto efficiency, the joint performance scales in a sub-optimal direction. (iii) By adjusting the optimization weights to fit the Pareto efficiency, we can achieve Pareto training, where the scaling of joint performance becomes most efficient. Even without requiring additional resources, Pareto training is comparable to the performance of scaling resources several times under optimization strategies that overly prioritize either robustness or effectiveness. Finally, we demonstrate that our findings can help deploy dense retrieval models in real-world applications that scale efficiently and are balanced for robustness and effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。