arXiv:2602.03344cs.LGcs.AI2026-02被引 2

模型性能越高,鲁棒性自然涌现,无需额外设计。

Robustness as an Emergent Property of Task Performance

  • 性能提升过程中,鲁棒性随任务掌握程度自发产生。
  • 跨数据集与配置实验中,性能与鲁棒性呈强正相关。
  • 适合关注模型实际能力而非刻意优化鲁棒性的研究者。

鲁棒性常被视为现实应用中的关键挑战,但本文发现,当模型在任务上接近高表现时,鲁棒性会自然实现。通过在多个模型、多种数据集及配置(如同义改写、不同温度)下的实证分析,我们观察到性能与鲁棒性之间存在强正相关。进一步研究表明,鲁棒性主要由任务特定能力驱动,而非模型固有属性,这挑战了当前将鲁棒性视为独立能力的做法。因此,从宏观视角看,随着新任务趋于饱和,其鲁棒性也将随之涌现。对研究者而言,这意味着可减少对鲁棒性的显式干预;对实践者而言,表明文献中多数任务仍不可靠,但在已掌握的简单任务上,模型已足够稳健,可部署于真实场景。

原文摘要 · Abstract (English)

Robustness is often regarded as a critical future challenge for real-world applications, where stability is essential. However, as models often learn tasks in a similar order, we hypothesize that easier tasks will be easier regardless of how they are presented to the model. Indeed, in this paper, we show that as models approach high performance on a task, robustness is effectively achieved. Through an empirical analysis of multiple models across diverse datasets and configurations (e.g., paraphrases, different temperatures), we find a strong positive correlation. Moreover, we find that robustness is primarily driven by task-specific competence rather than inherent model-level properties, challenging current approaches that treat robustness as an independent capability. Thus, from a high-level perspective, we may expect that as new tasks saturate, model robustness on these tasks will emerge accordingly. For researchers, this implies that explicit efforts to measure and improve robustness may warrant reduced emphasis, as such robustness is likely to develop alongside performance gains. For practitioners, it acts as a sign that indeed the tasks that the literature deals with are unreliable, but on easier past tasks, the models are reliable and ready for real-world deployment.

鲁棒性模型性能任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。