arXiv:2509.06822cs.AIcs.CL2025-09Conference of the …被引 18

提出可迭代推理的故障定位框架,自动发现复杂大模型系统的错误环节。

RAFFLES: Reasoning-based Attribution of Faults for LLM Systems

  • 构建多组件迭代管道,由裁判和专用评估器协同诊断故障
  • 在多个数据集上准确率超80%,显著优于现有方法
  • 适合需要自动化故障检测的复杂AI系统研发与测试人员

随着复杂、互联的长时序大模型系统日益普及,识别其故障位置与时间变得极为困难。当前评估方法多依赖简单指标、端到端结果,且高度依赖人工视角。为匹配系统复杂性,评估框架需具备推理、探查、迭代与理解系统内部逻辑的能力。本文提出RAFFLES,一种离线评估架构,采用迭代式推理机制。该系统以中央裁判为核心,通过多阶段迭代,系统性识别故障,并由专门评估器验证候选故障及其解释合理性。我们在多个基准上进行了评估:使用Who&When数据集检测多智能体系统中的步骤级故障,以及使用ReasonEval数据集诊断数学推理中的步骤级错误。RAFFLES在Who&When手写构造和算法生成数据集上的准确率分别超过20%和50%,在ReasonEval数据集上超过80%。这些结果标志着向实现自主系统自动化故障检测迈出了关键一步,减少对人工密集审查的依赖。

原文摘要 · Abstract (English)

The advent of complex, interconnected long-horizon LLM systems has made it incredibly tricky to identify where and when these systems break down. Evaluation capabilities that currently exist today are limited in that they often focus on simple metrics, end-to-end outcomes, and are dependent on the perspectives of humans. In order to match the increasing complexity of these many component systems, evaluation frameworks must also be able to reason, probe, iterate, and understand the nuanced logic passing through these systems. In this paper, we present RAFFLES, an offline evaluation architecture that incorporates iterative reasoning. Specifically, RAFFLES operates as an iterative, multi-component pipeline, using a central Judge to systematically identify faults and a set of specialized Evaluators to assess the quality of the candidate faults as well as rationales of the Judge. We evaluated RAFFLES with several benchmarks - the Who&When dataset to identify step-level faults in multi-agent systems and the ReasonEval datasets to diagnose step-level mathematical reasoning errors. RAFFLES outperforms strong baselines, achieving an accuracy of over 20% and 50% on the Who&When Hand-Crafted and Algorithmically-Generated datasets, and over 80% on the ReasonEval datasets. These results demonstrate a key step towards introducing automated fault detection for autonomous systems over labor-intensive manual review.

大模型评估故障定位自动化测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。