为大模型调试提供系统化方法,提升问题定位与修复效率
A Systematic Approach for Large Language Models Debugging

- 将大模型视为可观察系统,统一评估、解释与错误分析流程
- 支持在无标准基准下迭代诊断弱点并优化提示与参数
- 适合需要可复现、透明部署的AI应用开发者
大语言模型已成为现代AI工作流的核心,支撑从开放式文本生成到复杂代理推理的应用。然而,由于其黑箱性和概率特性,以及在多样任务与场景中诊断错误的困难,调试仍是一大挑战。本文提出一种系统化的LLM调试方法,将模型视为可观测系统,提供从问题检测到模型优化的结构化、模型无关路径。通过整合评估、可解释性与错误分析实践,该方法使从业者能迭代诊断模型缺陷,优化提示、参数,并调整数据用于微调或评估,即使在缺乏标准化基准与评估标准的场景下依然有效。我们认为,这种结构化方法不仅加速故障排查,更促进部署过程中的可复现性、透明性与可扩展性。
原文摘要 · Abstract (English)
Large language models (LLMs) have become central to modern AI workflows, powering applications from open-ended text generation to complex agent-based reasoning. However, debugging these models remains a persistent challenge due to their opaque and probabilistic nature and the difficulty of diagnosing errors across diverse tasks and settings. This paper introduces a systematic approach for LLM debugging that treats models as observable systems, providing structured, model-agnostic methods from issue detection to model refinement. By unifying evaluation, interpretability, and error-analysis practices, our approach enables practitioners to iteratively diagnose model weaknesses, refine prompts and model parameters, and adapt data for fine-tuning or assessment, while remaining effective in contexts where standardized benchmarks and evaluation criteria are lacking. We argue that such a structured methodology not only accelerates troubleshooting but also fosters reproducibility, transparency, and scalability in the deployment of LLM-based systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。