评测金融AI决策的可复现性,发现结果相同但执行路径差异大。
DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

- 通过可复现性协议衡量决策与工具路径的一致性
- 95%决策一致但工具路径仅67%一致,轨迹一致性更低至45%
- 适合关注AI决策透明性与合规性的研发和审计人员
金融AI代理在改变工具、顺序或记录参数与结果的情况下仍可能重复同一决策。仅评估结果会忽略这种变化,而这对回放与变更控制至关重要。DFAH-Bench 实现了确定性-忠实性保障测试(DFAH),其中‘忠实性’指可观察执行过程在回放时的保真度,而非答案正确性。该协议筛选出可比较且充分观测的回放样本,测量决策一致率(DAR)与工具路径一致率(TAR)。我们分析了来自719个合成合规与金融DataOps组的4,157个回溯案例,以及包含570个合格案例的190个组的论据感知前瞻性扩展。在扩展中,决策一致率达94.2-95.1%,精确工具名路径一致率为66.9-69.4%,差距达25.8-27.3个百分点;参数与结果轨迹一致率下降至45.0-51.5%。即使在决策完全一致的组中,任务加权下路径仍存在66.7-68.9%的差异。DFAH-Bench使稳定决策背后的执行过程对回放、调查与变更审查可见。
原文摘要 · Abstract (English)
A financial AI agent can repeat a decision while changing the tools, order, or recorded arguments and results used to reach it. Outcome-only evaluation misses this variation, even when it matters for replay and change control. DFAH-Bench operationalizes the Determinism-Faithfulness Assurance Harness (DFAH), where faithfulness means fidelity of observable execution under replay, not answer correctness. The protocol qualifies comparable, sufficiently observed replays and measures decision agreement (DAR) and tool-path agreement (TAR) over the same eligible groups. We analyze 4,157 retrospective episodes from configurations with observed tool use across 719 synthetic compliance and financial DataOps groups, together with an argument-aware prospective extension comprising 570 eligible episodes across 190 groups. In that extension, decisions agree 94.2-95.1% while exact tool-name paths agree 66.9-69.4%, producing 25.8-27.3 percentage-point gaps; argument-and-result trajectory agreement falls to 45.0-51.5%. Even among unanimous-decision groups, paths vary in 66.7-68.9% under task weighting. DFAH-Bench makes the execution behind a stable decision visible for replay, investigation, and change review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。