arXiv:2606.31522cs.CLcs.AI2026-06被引 1

提出金融代理心理稳定性评估基准,发现长期运行中行为指令会逐渐失效

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

论文配图:FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents
图 1 · 摘自论文原文
  • 构建仿真环境分离价格与价值,测试模型在不同市场下的行为稳定性
  • 18个模型均出现指令衰减,危机中重置指令使保守型代理表现提升4.4倍
  • 重置策略需根据代理类型和市场状态选择,盲目重置可能适得其反

大型语言模型被用作具有明确行为指令(如‘保本’或‘规避投机’)的自主金融代理,但随着市场环境长期积累,这些指令的行为影响力逐渐减弱,我们将其称为指令显著性衰减(MSD)。为此,我们提出FinPersona-Bench,一个模拟基准,通过合成市场使可观测价格与隐藏基本面价值分离,可验证三种失效模式:平静市场无信号交易、暴跌时恐慌抛售、泡沫期忽视基本面。评估18个前沿及开源大模型,每个分配三种行为类型(从严格保本到激进增长),结果显示MSD随时间累积且依赖模型。在危机场景中,静态代理与定期重置指令的代理间行为差距从首季度到末季度扩大4.4倍。重置效果并非普适:在低信号市场中持续帮助保守型代理,却显著恶化激进型代理行为。结果表明,长期可靠部署需基于代理类型与市场状态的精准、有选择的指令重置。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment. In practice, however, as market context accumulates over long horizons, these mandates gradually lose their behavioral influence, a phenomenon we formalize as Mandate Salience Decay (MSD). To measure MSD objectively, we introduce FinPersona-Bench, a simulation benchmark in which a synthetic market decouples observable price from hidden fundamental value, enabling falsifiable evaluation across three failure modes: trading without signal in calm markets, panic-selling during crashes, and ignoring fundamental value during speculative bubbles. Evaluating 18 leading frontier and open-source LLMs, each assigned one of three behavioral profiles ranging from strict capital preservation to aggressive growth, shows that MSD compounds over time and is model-dependent. In crash scenarios, the behavioral gap between static agents and those receiving periodic mandate re-grounding grows 4.4x from the first to the final quarter of the simulation. The effects of mandate re-grounding are not uniformly positive: it consistently helps conservative agents in low-signal markets but actively worsens behavior for aggressive agents in the same setting. These findings suggest that reliable long-horizon deployment requires selective, mandate-aware re-grounding based on agent profile and market regime.

金融AI行为稳定性大模型评估指令衰减

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。