arXiv:2601.21410stat.MLcs.LG2026-01

让模型学会判断何时信任大模型的语义先验,避免错误引导。

Learning When to Trust LLM Priors: A Validated Framework for Semantic Prior Integration

  • 基于交叉验证动态校准大模型先验的影响强度
  • 在多种任务中提升可靠先验性能,自动抑制无效或错误先验
  • 适合需要高可靠性、避免幻觉的统计学习场景

大型语言模型(LLMs)蕴含丰富的语义知识,可用于监督学习,但其输出作为统计先验不可靠:可能噪声大、设定错误或产生幻觉。现有方法或直接信任这些信号,使预测易受误导;或仅限于单一模型类别。本文提出Statsformer,一种可验证的框架,用于在监督统计学习中决定何时信任大模型衍生的语义先验。该框架将大模型生成的特征评分映射到线性与非线性预测器组成的异构库中的多类先验注入机制,并通过留出集验证自适应校准每种先验引导的学习器影响,使有用信息提升预测性能,同时削弱弱、错设或对抗性先验的影响。最终系统具备类似“预言家”的保证:在统计误差范围内,最终预测器表现不低于其库内候选者(包括无先验学习器)的最优凸组合。在多样预测任务中,有信息量的先验提升性能,不可靠先验则被自动降权。这使Statsformer成为面向可靠性的大模型驱动统计学习方法:不直接信任大模型知识,而是在数据验证后才允许其影响最终预测。

原文摘要 · Abstract (English)

Large language models (LLMs) encode rich semantic knowledge that can be useful for supervised learning, but their outputs are unreliable as statistical priors: they may be noisy, misspecified, or hallucinated. Existing LLM-informed learning methods either trust such signals directly, leaving predictions vulnerable to unreliable LLM guidance, or restrict semantic integration to a single model class. We introduce Statsformer, a validated framework for learning when to trust LLM-derived semantic priors in supervised statistical learning. Statsformer maps LLM-derived feature scores into a family of learner-specific prior-injection mechanisms across a heterogeneous library of linear and nonlinear predictors. It then uses out-of-fold validation to adaptively calibrate the influence of each prior-informed learner, allowing useful semantic information to improve prediction while attenuating weak, misspecified, or adversarial priors. This yields a guardrailed statistical learning system with an oracle-style guarantee: up to statistical error, the final predictor performs no worse than the best convex combination of its in-library candidates, including prior-free learners. Across diverse prediction tasks, informative LLM priors improve performance, while unreliable priors are automatically downweighted. These results position Statsformer as a reliability-oriented approach to LLM-informed statistical learning: rather than trusting LLM knowledge directly, it validates semantic priors against data before allowing them to influence the final predictor.

大模型先验统计学习可信融合模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。