揭露大模型金融代理的盈利幻觉,提出新框架提升真实泛化能力
Profit Mirage: Revisiting Information Leakage in LLM-based Financial Agents
- 通过反事实扰动让模型学习因果关系而非记忆结果
- 在真实场景中表现优于所有基线,风险调整后收益显著提升
- 发布抗信息泄露基准,适合研究金融AI安全与可信性的人参考
基于大模型的金融代理虽能模拟人类交易员表现,但多数系统存在“盈利幻觉”:回测收益耀眼,一旦知识窗口到期即迅速消失,根源在于大模型固有的信息泄露。本文从四个维度系统量化该问题,并发布FinLake-Bench——一个抗信息泄露的评估基准。为缓解此问题,提出FactFin框架,通过反事实扰动迫使模型学习因果驱动因素而非记忆结果。该框架集成策略代码生成、检索增强生成、蒙特卡洛树搜索和反事实模拟四部分。大量实验表明,所提方法在样本外泛化能力上超越所有基线,实现更优的风险调整后绩效。
原文摘要 · Abstract (English)
LLM-based financial agents have attracted widespread excitement for their ability to trade like human experts. However, most systems exhibit a "profit mirage": dazzling back-tested returns evaporate once the model's knowledge window ends, because of the inherent information leakage in LLMs. In this paper, we systematically quantify this leakage issue across four dimensions and release FinLake-Bench, a leakage-robust evaluation benchmark. Furthermore, to mitigate this issue, we introduce FactFin, a framework that applies counterfactual perturbations to compel LLM-based agents to learn causal drivers instead of memorized outcomes. FactFin integrates four core components: Strategy Code Generator, Retrieval-Augmented Generation, Monte Carlo Tree Search, and Counterfactual Simulator. Extensive experiments show that our method surpasses all baselines in out-of-sample generalization, delivering superior risk-adjusted performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。