arXiv:2509.19557cs.CLcs.LG2025-09

改进RoBERTa在实体匹配中的置信度,提升模型判断可靠性。

Confidence Calibration in Large Language Model-Based Entity Matching

  • 用温度缩放等方法校准大模型置信度
  • 校准后误差降低最高达23.83%
  • 适合关注模型可信度的开发者使用

本研究探讨大语言模型在实体匹配任务中置信度校准的问题。通过实证比较基准RoBERTa模型在实体匹配任务中的原始置信度与采用温度缩放、蒙特卡洛丢弃和集成方法校准后的置信度。实验基于Abt-Buy、DBLP-ACM、iTunes-Amazon和Company数据集。结果表明,原始模型存在轻微过自信现象,预期校准误差(Expected Calibration Error)在0.0043至0.0552之间。通过温度缩放可有效缓解该问题,使误差最高降低23.83%。

原文摘要 · Abstract (English)

This research aims to explore the intersection of Large Language Models and confidence calibration in Entity Matching. To this end, we perform an empirical study to compare baseline RoBERTa confidences for an Entity Matching task against confidences that are calibrated using Temperature Scaling, Monte Carlo Dropout and Ensembles. We use the Abt-Buy, DBLP-ACM, iTunes-Amazon and Company datasets. The findings indicate that the proposed modified RoBERTa model exhibits a slight overconfidence, with Expected Calibration Error scores ranging from 0.0043 to 0.0552 across datasets. We find that this overconfidence can be mitigated using Temperature Scaling, reducing Expected Calibration Error scores by up to 23.83%.

置信度校准实体匹配大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。