arXiv:2509.07309cs.CLcs.LG2025-09被引 1

为自由生成任务提供更精细的性能评估区间,提升结果可信度。

PIE: Performance Interval Estimation for Free-Form Generation Tasks

  • 用回归方法结合置信度特征预测多个连续指标的评分
  • 在11个数据集上点估计误差更低,不确定性区间更校准
  • 适合需要可靠质量评估的生成模型研究与应用

置信度估计旨在判断模型输出是否正确。对于答案精确的任务,二元正确性判断合理;但自由生成任务通常更复杂,输出质量具有细粒度和多维度特点。为此,我们提出性能区间估计(PIE),既能预测任意连续评价指标的点估计值,又能给出校准后的不确定性区间。我们比较了两种方法:大模型作为评判者(LLM-as-judge)与经典回归结合置信度特征。在涵盖摘要、翻译、代码生成、函数调用和问答等11个数据集上的评估显示,回归方法在两点上表现更优:一是指标评分的点估计误差更低;二是不确定性区间校准效果更好。为支持复现与后续研究,我们公开了数据与代码。

原文摘要 · Abstract (English)

Confidence estimation infers a probability for whether each model output is correct or not. While predicting such binary correctness is sensible for tasks with exact answers, free-form generation tasks are often more nuanced, with output quality being both fine-grained and multi-faceted. We thus propose Performance Interval Estimation (PIE) to predict both: 1) point estimates for any arbitrary set of continuous-valued evaluation metrics; and 2) calibrated uncertainty intervals around these point estimates. We then compare two approaches: LLM-as-judge vs. classic regression with confidence estimation features. Evaluation over 11 datasets spans summarization, translation, code generation, function-calling, and question answering. Regression is seen to achieve both: i) lower error point estimates of metric scores; and ii) well-calibrated uncertainty intervals. To support reproduction and follow-on work, we share our data and code.

生成评估不确定性回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。