arXiv:2602.00977cs.CLcs.LG2026-02中稿 · The ACM Web Confer…被引 2

用单次推理捕捉模型内部结构信号,提升大模型输出可信度判断

Trust in One Round: Confidence Estimation for Large Language Models via Structural Signals

  • 基于模型最后一层隐状态的多尺度结构信号,构建无需采样的置信度估计框架
  • 在四个跨领域基准上,置信度预测的AUROC和AUPR均优于现有方法
  • 适合资源受限场景下对高风险应用(如医疗、科学)的可靠结果验证

大语言模型在高社会、科学或安全成本领域部署日益广泛,但标准置信度估计方法(如词元似然、语义相似性、多样本一致性)在分布偏移、领域专有文本和计算限制下表现脆弱。本文提出结构置信(Structural Confidence),一种单次前向传播、模型无关的置信度估计框架,通过分析模型最终层隐状态轨迹的谱特征、局部变化和全局形状等多尺度结构信号,捕捉概率与句向量遗漏的内部稳定性模式。我们在四个异构基准上进行广泛评估:FEVER(事实验证)、SciFact(科学主张)、WikiBio-hallucination(传记一致性)和TruthfulQA(真实性导向问答)。结果表明,该框架在AUROC和AUPR指标上显著优于主流基线。更重要的是,相比需多次随机生成和额外模型的采样一致性方法,本方法仅需一次确定性前向传播,为资源受限、高社会影响的LLM应用提供了高效、鲁棒的后验置信度估计基础。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in domains where errors carry high social, scientific, or safety costs. Yet standard confidence estimators, such as token likelihood, semantic similarity and multi-sample consistency, remain brittle under distribution shift, domain-specialised text, and compute limits. In this work, we present Structural Confidence, a single-pass, model-agnostic framework that enhances output correctness prediction based on multi-scale structural signals derived from a model's final-layer hidden-state trajectory. By combining spectral, local-variation, and global shape descriptors, our method captures internal stability patterns that are missed by probabilities and sentence embeddings. We conduct extensive, cross-domain evaluation across four heterogeneous benchmarks-FEVER (fact verification), SciFact (scientific claims), WikiBio-hallucination (biographical consistency), and TruthfulQA (truthfulness-oriented QA). Our Structural Confidence framework demonstrates strong performance compared with established baselines in terms of AUROC and AUPR. More importantly, unlike sampling-based consistency methods which require multiple stochastic generations and an auxiliary model, our approach uses a single deterministic forward pass, offering a practical basis for efficient, robust post-hoc confidence estimation in socially impactful, resource-constrained LLM applications.

置信度估计大模型单次推理结构信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。