arXiv:2606.07624cs.LG2026-06被引 2

用统计过程控制提升大模型部署的可信度

Sequential statistical inference for Large Language Models: Representation, validity, and monitoring

  • 将大模型交互视为依赖性随机过程,而非孤立问答对
  • 在重复使用和模型更新下仍保持可靠的不确定性保障
  • 通过序列警报检测幻觉率、公平性等关键指标变化

本文认为,序列统计推断可自然提升大模型的可信度。实际部署中,大模型系统会根据不断演化的上下文反复被查询,并整合用户或工具反馈,模型更新或分布变化可能导致行为漂移。讨论围绕三个任务展开:表示,将大模型交互建模为依赖性随机过程而非孤立提示-响应对;有效性,在依赖、重复使用和自适应条件下仍保持有意义的不确定性保证;监控,利用序列告警和变点检测识别校准、幻觉率、拒答行为、公平性等任务相关属性的变化。该视角补充了近期综述,将可信的大模型部署视为统计过程控制问题。

原文摘要 · Abstract (English)

This discussion argues that sequential statistical inference can naturally contribute to LLM trustworthiness. In deployment, LLM systems are queried repeatedly, conditioned on evolving contexts, and incorporate user or tool feedback, and may exhibit behavioral shifts after model updates or distribution changes. The discussion is organized around three tasks: representation, modeling LLM interactions as dependent stochastic processes rather than isolated prompt--response pairs; validity, developing uncertainty guarantees that remain meaningful under dependence, repeated use, and adaptation; and monitoring, using sequential alarms and change-point detection to identify shifts in calibration, hallucination rates, refusal behavior, fairness, or other task-relevant properties. This perspective complements recent surveys by viewing trustworthy LLM deployment as a problem of statistical process control.

大模型可信度统计推断过程监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。