arXiv:2604.27082cs.AIcs.LG2026-04被引 1

用统计方法精准评估模型替换效果,确保上线后不翻车。

When Your LLM Reaches End-of-Life: A Framework for Confident Model Migration in Production Systems

  • 用贝叶斯方法校准自动评测与人工判断的一致性
  • 在仅有限人工标注下准确比较模型表现
  • 适合企业大规模部署模型时的迁移决策

我们提出一种框架,用于在生产级大语言模型(LLM)达到生命周期终点或需替换时进行迁移。核心贡献是基于贝叶斯统计的方法,将自动化评估指标与人工判断对齐,即使在人工标注数据有限的情况下也能自信地进行模型对比。我们在一个商业问答系统上验证该框架,该系统每月服务530万次交互,覆盖六个全球区域;评估了正确性、拒绝行为和风格一致性,成功识别出合适的替代模型。该框架适用于任何部署LLM产品的企事业单位,提供一种有原则、可复现的模型迁移方法,兼顾质量保障与评估效率。随着LLM生态快速演进,组织需管理多个模型、区域和应用场景的AI服务组合,此能力日益关键。

原文摘要 · Abstract (English)

We present a framework for migrating production Large Language Model (LLM) based systems when the underlying model reaches end-of-life or requires replacement. The key contribution is a Bayesian statistical approach that calibrates automated evaluation metrics against human judgments, enabling confident model comparison even with limited manual evaluation data. We demonstrate this framework on a commercial question-answering system serving 5.3M monthly interactions across six global regions; evaluating correctness, refusal behavior, and stylistic adherence to successfully identify suitable replacement models. The framework is broadly applicable to any enterprise deploying LLM-based products, providing a principled, reproducible methodology for model migration that balances quality assurance with evaluation efficiency. This is a capability increasingly essential as the LLM ecosystem continues to evolve rapidly and organizations manage portfolios of AI-powered services across multiple models, regions, and use cases.

模型迁移评估框架生产落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。