用一致性门控测试LLM投资理由可信度,避免过早发布虚假结论。
ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable
- 设计共识门控协议,检验LLM生成的投资理由是否稳定可靠。
- 实测显示100个对抗样本中多个维度不达标,最低一致性仅0.2022。
- 适合评估AI金融系统前的可靠性,防止盲目依赖模型判断。
基于大语言模型的金融代理在实际回报出现前就生成投资理由,导致评价滞后且噪声大。尽管使用LLM judges可提前评估,但未经验证的评判可能奖励冗长、自信或套用模板,而非真实金融判断。本文提出ValueBlindBench,一种预注册的一致性门控压力测试协议,用于判断投资理由是否可发表、合格或无效。在包含1,000次诚实决策和100个预注册对抗控制的原型中(共1,100条轨迹,5,500次评判),整体一致性达\(\barκ_w = 0.7168\),但多个过拟合主张被拦截。低秩系统陷入平局,一个维度(\texttt{constraint\_awareness})不通过,一致性仅\(\barκ_w = 0.2022\);单评委排名受家族依赖影响;简洁正确理由比诚实理由低\(Δ= -2.81\)分。锚定特异性探测进一步表明,如约束意识等金融概念具有操作关键性。研究目标并非构建排行榜或测量真实投资能力,而是作为AI金融评估前的校准层:决定一项基于LLM评判的投资理由是否足够稳定、一致且未被污染,方可报告。
原文摘要 · Abstract (English)
LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a delayed-ground-truth evaluation problem: realized returns remain the eventual arbiter of investment quality, but they arrive too late and are too noisy to guide many model-development and governance decisions. LLM judges offer a tempting shortcut for pre-deployment evaluation of AI-finance systems, but unvalidated judges may reward verbosity, confidence, or rubric mimicry rather than financial judgment. This paper introduces ValueBlindBench, a preregistered agreement-gated stress-test protocol for deciding when LLM-judged investment-rationale claims are publishable, qualified, or invalid. In a controlled market-state capital-allocation prototype with 1,000 honest decision cycles and 100 preregistered adversarial controls (1,100 trajectories, 5,500 judge calls), ValueBlindBench clears the aggregate agreement gate at \(\barκ_w = 0.7168\) but prevents several overclaims. Lower-rank systems collapse into a tie-class, one rubric dimension fails the per-dimension gate (\texttt{constraint\_awareness}, \(\barκ_w = 0.2022\)), single-judge rankings are family-dependent, and terse-correct rationales receive a \(Δ= -2.81\) rubric-point penalty relative to honest rationales. A targeted anchor-specificity probe further shows that financial constructs such as constraint awareness are operationally load-bearing. The scientific object is therefore not a leaderboard and not a claim to measure true investment skill. ValueBlindBench is a pre-calibration metrology layer for AI-finance evaluation: it governs whether a proposed LLM-judge-based investment-rationale claim is stable enough, agreed enough, and uncontaminated enough to be reported at all.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。