arXiv:2602.11304cs.IRcs.AI2026-02KDD被引 1

评测大模型在复杂金融数据中多工具协作的分析缺陷,揭示高风险错误。

CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis

  • 构建覆盖198个真实加密问题的评测基准,涵盖11类任务
  • 发现主流大模型在动态数据整合中存在7类高阶错误,影响决策可靠性
  • 提供可扩展的评估框架,适合研究智能分析师系统的开发者

现代分析师型代理需处理复杂、高文本量输入,包括数十份检索文档、工具输出及时效性数据。现有工作虽有工具调用评测和知识增强系统的真实性检查,但极少研究多工具输出在动态、结构化与非结构化数据融合中的表现。本文以加密货币为高密度数据代表领域,提出:(1)CryptoAnalystBench,一个包含198个生产级加密与DeFi查询的分析师对齐基准,覆盖11个类别;(2)配备相关工具的智能体测试环境,用于生成多个前沿大模型的回答;(3)基于引用验证和大模型作为裁判的评估流程,涵盖用户定义的四个成功维度:相关性、时间相关性、深度与数据一致性。通过人工标注,提炼出七类高阶错误类型,这些错误无法被事实性检查或大模型评分可靠捕捉。结果显示,即使最先进的系统仍持续存在此类错误,可能危及高风险决策。基于该分类,优化了裁判评分标准。尽管裁判评分与人类标注者在具体分值上不完全一致,但能稳定识别关键失败模式,支持开发者与研究者进行可扩展反馈。我们公开发布CryptoAnalystBench,包含标注查询、评估流程、裁判标准与错误分类,并提出缓解策略与评估长篇多工具系统的关键挑战。

原文摘要 · Abstract (English)

Modern analyst agents must reason over complex, high token inputs, including dozens of retrieved documents, tool outputs, and time sensitive data. While prior work has produced tool calling benchmarks and examined factuality in knowledge augmented systems, relatively little work studies their intersection: settings where LLMs must integrate large volumes of dynamic, structured and unstructured multi tool outputs. We investigate LLM failure modes in this regime using crypto as a representative high data density domain. We introduce (1) CryptoAnalystBench, an analyst aligned benchmark of 198 production crypto and DeFi queries spanning 11 categories; (2) an agentic harness equipped with relevant crypto and DeFi tools to generate responses across multiple frontier LLMs; and (3) an evaluation pipeline with citation verification and an LLM as a judge rubric spanning four user defined success dimensions: relevance, temporal relevance, depth, and data consistency. Using human annotation, we develop a taxonomy of seven higher order error types that are not reliably captured by factuality checks or LLM based quality scoring. We find that these failures persist even in state of the art systems and can compromise high stakes decisions. Based on this taxonomy, we refine the judge rubric to better capture these errors. While the judge does not align with human annotators on precise scoring across rubric iterations, it reliably identifies critical failure modes, enabling scalable feedback for developers and researchers studying analyst style agents. We release CryptoAnalystBench with annotated queries, the evaluation pipeline, judge rubrics, and the error taxonomy, and outline mitigation strategies and open challenges in evaluating long form, multi tool augmented systems.

大模型评测智能代理加密金融多工具协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。