arXiv:2509.15356cs.LG2025-09EMNLP被引 4

研究大模型零样本预测的可靠性,帮用户判断何时能放心用。

Predicting Language Models' Success at Zero-Shot Probabilistic Prediction

  • 通过大规模实验分析大模型在不同表格任务中的零样本表现
  • 发现模型在基础任务上表现好时,其预测概率更可信
  • 提出无需标注数据的指标,可预判新任务是否适合用大模型

近期研究探索了大语言模型(LLMs)作为零样本模型生成个体特征的能力(例如作为风险模型或补充调查数据集)。然而,用户应在何种情况下相信大模型能提供高质量预测?为回答此问题,我们对大语言模型在广泛表格预测任务上的零样本预测能力进行了大规模实证研究。结果表明,大模型的表现具有高度可变性,不仅在同一数据集内的任务间差异大,跨数据集也如此。然而,当大模型在基础预测任务上表现良好时,其预测概率成为个体准确性的更强信号。随后,我们构建了任务级预测指标,旨在区分大模型可能表现良好与不适宜的任务。发现其中一些指标仅需无标签数据即可评估,对新任务的大模型表现具有强预测能力。

原文摘要 · Abstract (English)

Recent work has investigated the capabilities of large language models (LLMs) as zero-shot models for generating individual-level characteristics (e.g., to serve as risk models or augment survey datasets). However, when should a user have confidence that an LLM will provide high-quality predictions for their particular task? To address this question, we conduct a large-scale empirical study of LLMs' zero-shot predictive capabilities across a wide range of tabular prediction tasks. We find that LLMs' performance is highly variable, both on tasks within the same dataset and across different datasets. However, when the LLM performs well on the base prediction task, its predicted probabilities become a stronger signal for individual-level accuracy. Then, we construct metrics to predict LLMs' performance at the task level, aiming to distinguish between tasks where LLMs may perform well and where they are likely unsuitable. We find that some of these metrics, each of which are assessed without labeled data, yield strong signals of LLMs' predictive performance on new tasks.

大模型零样本预测评估可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。