arXiv:2608.00419cs.LGcs.AI2026-08中稿 · version of an arti…

解决大模型实时部署中的幻觉与过时问题,实现企业级可靠运行。

Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

  • 构建统一运维架构,融合数据流、持续学习与人工反馈闭环。
  • 在医疗金融场景中降低延迟成本,准确率提升且支持审计回滚。
  • 提出自适应检索与稀疏路由学习机制,提升模型长期可靠性。

在实时、受监管的环境中部署大语言模型面临知识过时、灾难性遗忘、幻觉及反馈弱等问题。本文提出一种统一的模式驱动型LLMOps架构,集成实时数据摄入、持续学习、检索增强生成(RAG)与人机协同反馈,形成一体化运营流程。四个核心贡献对应成熟软件设计模式:自适应摄入模式编排器(AIPO),在FreshStreamBench上验证;STAR+FAR持续学习方法,采用稀疏时间适配器路由与新鲜度感知重放;SAGE——一种面向服务等级目标(SLO)的自适应检索策略,可预测每查询的段落预算以满足尾部延迟要求;以及基于强化学习的人工反馈驱动收敛阶段。该方案有效缓解延迟-成本-准确率权衡,支持高风险领域如医疗与金融的可审计性和回滚能力。

原文摘要 · Abstract (English)

Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback loops. We present a unified, pattern-driven LLMOps architecture integrating real-time data ingestion, continual learning, retrieval-augmented generation (RAG), and human-in-the-loop feedback into a single operational pipeline. Four contributions map to established software design patterns: an adaptive ingestion pattern orchestrator (AIPO) evaluated with FreshStreamBench; STAR+FAR continual learning with sparse temporal adapter routing and freshness-aware replay; SAGE, an SLO-aware adaptive retrieval policy predicting a per-query passage budget to meet tail-latency targets; and an automated feedback-driven convergence stage with RLHF triggers. The result reduces latency-cost-accuracy trade-offs while supporting auditability and rollback for high-risk sectors such as health care and finance.

大模型部署LLMOps持续学习RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。