用大模型评估智能交易系统的决策过程,提升预测可靠性。
Multi-Dimensional Behavioral Evaluation of Agentic Stock Prediction Systems Using Large Language Model Judges with Closed-Loop Reinforcement Learning Feedback
- 通过大模型评委打分中间决策流,覆盖六维行为维度。
- 行为评分与20天夏普比率相关性达0.72,显著提升预测性能。
- 适用于任何可记录决策的智能系统,尤其适合高波动场景。
智能代理通过一系列相互依赖的自主决策生成输出,但传统评估仅关注结果,无法诊断决策过程。本文提出一种行为评估方法,通过评分每个决策点的中间行为轨迹来补充输出评估。将五日为单位的行为轨迹划分为片段,由三名大语言模型评委在六个领域特定维度(市场状态识别、路由、适应、风险校准、策略连贯性、错误恢复)进行评分。扰动实验验证各维度独立性,跨模型一致性Krippendorff's alpha达0.85。综合行为评分与20天实际夏普比率的斯皮尔曼相关系数为0.72。闭环反馈中,将各维度低分转化为信用分配惩罚并加入软演员-评论家(Soft Actor-Critic)奖励函数。三次微调循环(仅限验证集)使测试期(2017–2025)的一天平均绝对百分比误差从0.61%降至0.54%(相对下降11.5%,p<0.001,d=0.31),该改善在高波动率时段显著且可定位。该方法具有应用无关性,适用于任何能记录中间决策的智能系统。
原文摘要 · Abstract (English)
Agentic artificial intelligence systems produce outputs through sequences of interdependent autonomous decisions, yet standard evaluation assesses outputs alone and cannot diagnose the underlying process. We develop a behavioral evaluation methodology that complements output-level testing by scoring the intermediate decision process itself. Behavioral traces logged at each autonomous decision point are grouped into five-day episodes and scored along six domain-specific dimensions (regime detection, routing, adaptation, risk calibration, strategy coherence, error recovery) by an ensemble of three large language model (LLM) judges. A perturbation procedure that corrupts one dimension while leaving the other five intact confirms dimension specificity; cross-model agreement reaches Krippendorff's alpha = 0.85. The composite behavioral score correlates at Spearman rho = 0.72 with realized 20-day Sharpe ratio. Closing the loop, the framework converts deficient per-dimension scores into a credit-assigned penalty added to the Soft Actor-Critic reward. Three fine-tuning cycles, confined to validation data, reduce one-day MAPE from 0.61% to 0.54% (11.5% relative; p<0.001, d=0.31) on the held-out 2017 to 2025 test period, significant under Diebold-Mariano and localized by Giacomini-White to the high-volatility regime. The methodology is application-agnostic and applies to any agentic system whose intermediate decisions can be logged.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。