arXiv:2605.28850cs.LGq-fin.CP2026-05

通过可审计的交易环境,发现大模型在市场压力下的行为异常前兆。

Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents

  • 构建可回放的交易测试平台,追踪模型决策与风险信号的动态演化。
  • 识别出预失败阶段的嵌入漂移、风险融合表征分离和局部流形秩收缩等信号。
  • 无需微调即可用结构化风险反馈作为对齐信号,适合研究模型可信性而非盈利。

我们研究大语言模型(LLM)代理在金融决策环境中的行为对齐与表征动态。TradeArena是一个可审计的交易代理测试平台,包含风险报告、执行模拟、记忆机制和可重播轨迹,使我们能够分析在市场压力下,模型的推理逻辑、持仓状态与干预行为如何演变。代码与数据可通过TradeArena仓库获取。我们发现预失败特征:计划嵌入偏离正常中心,融合的计划-风险表征能区分正常与预下跌状态,局部流形呈现有效秩收缩。在80个滚动失败锚点和8条LLM轨迹中,该模式在哈希、LSA、Transformer及白盒隐藏状态探测器中均持续存在。无思维链目标权重、词汇控制、OHLCV噪声和虚假审计的压力测试显示,若缺少推理逻辑,推理层面的收缩会消失,但意图空间与融合表征仍具信息量。结构化风险反馈可作为外部对齐信号,无需微调,但非万能增益:真实审计反馈提升部分模型的校准性,改善另一些模型的收益,同时暴露安慰剂或隐藏反馈虽短期回报高,但对齐诊断能力弱的情况。51只股票日内实验揭示一个相关性盲区:模型推理支持持有耦合资产,而风险层会剪裁此类暴露。最后,金融审计任务套件将比较焦点从“哪个模型交易最好”转向模型能否审计轨迹、遵守执行边界、复现成果并避免过度主张。这些结果支持一项研究主张,而非盈利主张:可审计的风险反馈与表征轨迹揭示了大模型金融推理是否对齐、漂移或失效。

原文摘要 · Abstract (English)

We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. TradeArena, an auditable trading-agent testbed with risk reports, execution simulation, memory, and replayable trajectories, lets us analyze how rationales, positions, and interventions evolve under market stress. Code and data artifacts are available through the \href{https://github.com/weich97/TreLLM-public}{TradeArena repository}. We find pre-failure signatures: planning embeddings drift from normal centroids, fused plan-risk representations separate normal from pre-drawdown states, and local manifolds exhibit effective-rank contraction. Across 80 rolling failure anchors and eight LLM trajectories, this pattern persists across hash, LSA, Transformer, and white-box hidden-state probes. Stress tests with CoT-free target weights, lexical controls, OHLCV noise, and false audits show that rationale-level contraction can vanish without rationales, while intent-space and fused signatures remain informative. Structured risk feedback can act as an external alignment signal without fine-tuning, but not as a universal performance enhancer: true audit feedback improves calibration for some models, returns for others, and exposes cases where placebo or hidden feedback has higher short-horizon return but weaker alignment diagnostics. A 51-stock intraday experiment reveals a correlation blind spot: LLM rationales justify exposure to coupled assets that the risk layer clips. Finally, a financial-audit task suite shifts comparison from ``which model trades best'' to whether models can audit trajectories, respect execution boundaries, reproduce artifacts, and avoid claim overreach. These results support a research claim, not a profitability claim: auditable risk feedback and representation trajectories reveal when LLM financial reasoning is aligning, drifting, or failing.

大模型金融决策可解释性风险对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。