arXiv:2605.06822cs.LG2026-05被引 2

让AI交易员像人类一样可审计地自我进化。

SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents

论文配图:SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents
图 1 · 摘自论文原文
  • 用可读的规则清单代替自由文本优化,约束AI推理过程。
  • 在三个股票板块测试中,小模型性能平均提升10至20个百分点。
  • 适合需要透明、可验证策略的金融机构使用。

大型语言模型(LLMs)正被越来越多地用于自主金融交易,该领域需持续适应噪声大、非平稳的市场。现有自进化代理通常通过无约束的自由文本提示优化来应对,但在信号弱、奖励延迟(如盈亏)的环境下,这种无结构方法加剧了信用分配难题:优化器无法可靠区分系统性逻辑错误与随机市场波动,最终导致策略漂移。为此,我们提出自进化可审计评分规则策略(SHARP),一种神经符号框架,将无约束的文本变异替换为结构化的符号策略优化。SHARP将代理推理限制在有限、可读的显式条件-动作规则清单中。当出现次优交易时,归因代理通过跨样本推理识别具体规则失效。这使得可靶向、原子级的策略修正,并经严格的向前滚动验证进行正则化。在三个不同股票板块和四种LLM骨干模型上评估,SHARP能持续将通用初始启发式转化为高度稳健的策略,使紧凑模型的实证性能平均提升10至20个百分点(例如,GPT-4o-mini)。最终证明,LLMs可在保持动态高效适应的同时,显著提升机构金融所需的结构性透明度与可审计性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed for autonomous financial trading, a domain requiring continuous adaptation to noisy, non-stationary markets. Existing self-improving agents typically address this through unbounded free-form prompt optimization. However, in low signal-to-noise environments with delayed scalar rewards (P\&L), this unstructured approach exacerbates the fundamental credit assignment problem: optimizers cannot reliably distinguish systematic logic flaws from stochastic market variance, inevitably leading to policy drift. To overcome this bottleneck, we introduce the Self-Evolving Human-Auditable Rubric Policy (SHARP), a neuro-symbolic framework that replaces unconstrained text mutation with structured, symbolic policy optimization. SHARP confines the agent's reasoning to a bounded, human-readable rubric of explicit condition-action rules. When sub-optimal trades occur, an attribution agent employs cross-sample reasoning across multiple samples to isolate specific rule failures. This enables targeted, atomic policy edits that are subsequently regularized through strict walk-forward validation. Evaluated across three diverse equity sectors and four LLM backbones, SHARP consistently transforms generic initial heuristics into highly robust strategies, lifting the empirical performance of compact models by 10 to 20 percentage points on average (e.g., GPT-4o-mini). Ultimately, SHARP demonstrates that LLMs can achieve dynamic and efficient adaptation while significantly enhancing the structural transparency and auditability demanded by institutional finance.

AI交易可解释性策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。