arXiv:2511.19933cs.AI2025-11被引 12

梳理了15种LLM系统在实际应用中的隐性故障模式,助力建设更可靠的AI系统。

Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications

  • 从系统层面归纳15类真实场景下的LLM故障模式,如推理漂移、上下文边界退化等。
  • 发现现有评测无法覆盖稳定性、可复现性等问题,存在明显评估缺口。
  • 适合关注AI系统可靠性、部署运维与成本控制的研究者和工程师阅读。

大型语言模型(LLMs)正快速集成至决策支持工具、自动化流程和智能软件系统中。然而,其在生产环境中的行为仍不清晰,故障模式与传统机器学习模型有本质差异。本文提出一个系统级分类框架,涵盖15种真实应用中出现的隐性故障模式,包括多步推理漂移、潜在不一致、上下文边界退化、错误工具调用、版本漂移及成本驱动的性能崩溃。基于此分类,我们分析了评估与监控实践间的日益扩大的差距:现有基准测试主要衡量知识或推理能力,但对稳定性、可复现性、漂移或工作流集成缺乏洞察。同时,我们探讨了部署LLM面临的关键挑战——可观测性限制、成本约束及更新引发的回归问题,并提出构建可靠、可维护、成本敏感的LLM系统的设计原则。通过将LLM可靠性视为系统工程问题而非单纯模型问题,本研究为未来评估方法、AI系统鲁棒性及可信部署提供了分析基础。

原文摘要 · Abstract (English)

Large language models (LLMs) are being rapidly integrated into decision-support tools, automation workflows, and AI-enabled software systems. However, their behavior in production environments remains poorly understood, and their failure patterns differ fundamentally from those of traditional machine learning models. This paper presents a system-level taxonomy of fifteen hidden failure modes that arise in real-world LLM applications, including multi-step reasoning drift, latent inconsistency, context-boundary degradation, incorrect tool invocation, version drift, and cost-driven performance collapse. Using this taxonomy, we analyze the growing gap in evaluation and monitoring practices: existing benchmarks measure knowledge or reasoning but provide little insight into stability, reproducibility, drift, or workflow integration. We further examine the production challenges associated with deploying LLMs - including observability limitations, cost constraints, and update-induced regressions - and outline high-level design principles for building reliable, maintainable, and cost-aware LLM systems. Finally, we outline high-level design principles for building reliable, maintainable, and cost-aware LLM-based systems. By framing LLM reliability as a system-engineering problem rather than a purely model-centric one, this work provides an analytical foundation for future research on evaluation methodology, AI system robustness, and dependable LLM deployment.

LLM可靠性系统设计故障分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。