测试大模型在理财工作流中的表现,发现流程可靠性比算力更重要。
Benchmarking LLM Agents for Wealth-Management Workflows
- 构建12组理财任务对,含检索、分析与沟通,带明确标准和自动评分
- 高自主性任务中模型表现显著下降,证明端到端可靠性是瓶颈
- 适合研究金融AI助手评估的学者或产品经理参考
现代工作依赖多种数字协作工具,但常规流程仍常因人为错误和延迟而受阻。本文将TheAgentCompany扩展为金融专用环境,探究通用大语言模型能否在准确性和经济性上完成典型理财任务。研究引入合成领域数据,增强同事模拟,并原型化自动任务生成管道。目标是创建可有效衡量代理在助理级理财工作中适配度的评估集。构建了12组涵盖检索、分析与合成/沟通的理财任务对,具有明确接受标准和确定性评分机制。引入新的金融专用数据集,并为每项任务设置高自主性与低自主性两种变体。研究结论指出,模型限制更多源于端到端工作流的可靠性,而非数学推理能力;自主性水平显著影响表现;且不当评估曾严重阻碍基准测试进展。
原文摘要 · Abstract (English)
Modern work relies on an assortment of digital collaboration tools, yet routine processes continue to suffer from human error and delay. To address this gap, this dissertation extends TheAgentCompany with a finance-focused environment and investigates whether a general purpose LLM agent can complete representative wealth-management tasks both accurately and economically. This study introduces synthetic domain data, enriches colleague simulations, and prototypes an automatic task-generation pipeline. The study aims to create and assess an evaluation set that can meaningfully measure an agent's fitness for assistant-level wealth management work. We construct a benchmark of 12 task-pairs for wealth management assistants spanning retrieval, analysis, and synthesis/communication, with explicit acceptance criteria and deterministic graders. We seeded a set of new finance-specific data and introduced a high vs. low-autonomy variant of every task. The paper concluded that agents are limited less by mathematical reasoning and more so by end-to-end workflow reliability, and meaningfully affected by autonomy level, and that incorrect evaluation of models have hindered benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。