arXiv:2604.10015cs.AIcs.CE2026-04被引 1

评测大模型长时金融任务中的工具调用能力,发现用对工具不等于能有效推理。

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

论文配图:FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks
图 1 · 摘自论文原文
  • 构建800条专家标注轨迹,从动作、效率、过程到输出全面评估模型表现
  • 13个模型在信息利用和最终答案质量上普遍表现差,暴露推理短板
  • 首个金融工具调用轨迹偏好数据集,微调后中间推理提升但最终答案仍受限

近期研究显示,工具调用能力使大语言模型(LLMs)能与外部环境交互以完成长周期金融任务。现有基准多集中于有限场景,依赖调用级指标,无法捕捉轨迹级推理质量。为此,我们提出FinTrace,一个包含800条专家标注轨迹的基准,覆盖34类真实金融任务,涵盖多个难度等级。FinTrace采用基于评分标准的评估协议,包含九项指标,分属行动正确性、执行效率、过程质量和输出质量四大维度,实现对模型工具调用行为的细粒度评估。对13个LLMs的评估表明,前沿模型虽能准确选择工具,但所有模型在信息利用和最终答案质量上均表现不佳,暴露出调用正确工具与有效推理其输出之间的显著差距。为超越诊断,我们构建了FinTrace-Training,首个面向金融工具调用的轨迹级偏好数据集,包含8,196条经筛选的轨迹及工具增强上下文与偏好对。使用监督微调结合直接偏好优化(DPO)对Qwen-3-8B/32B进行微调,结果显示训练后中间推理指标持续提升,且DPO更有效抑制失败模式。然而,端到端的答案质量仍是瓶颈,表明轨迹级改进尚未完全传导至最终输出质量。

原文摘要 · Abstract (English)

Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks. While existing benchmarks have begun evaluating financial tool calling, they focus on limited scenarios and rely on call-level metrics that fail to capture trajectory-level reasoning quality. To address this gap, we introduce FinTrace, a benchmark comprising 800 expert-annotated trajectories spanning 34 real-world financial task categories across multiple difficulty levels. FinTrace employs a rubric-based evaluation protocol with nine metrics organized along four axes -- action correctness, execution efficiency, process quality, and output quality -- enabling fine-grained assessment of LLM tool-calling behavior. Our evaluation of 13 LLMs reveals that while frontier models achieve strong tool selection, all models struggle with information utilization and final answer quality, exposing a critical gap between invoking the right tools and reasoning effectively over their outputs. To move beyond diagnosis, we construct FinTrace-Training, the first trajectory-level preference dataset for financial tool-calling, containing 8,196 curated trajectories with tool-augmented contexts and preference pairs. We fine-tune Qwen-3-8B/32B using supervised fine-tuning followed by direct preference optimization (DPO) and show that training on FinTrace-Training consistently improves intermediate reasoning metrics, with DPO more effectively suppressing failure modes. However, end-to-end answer quality remains a bottleneck, indicating that trajectory-level improvements do not yet fully propagate to final output quality.

大模型评估金融AI工具调用轨迹评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。