arXiv:2604.25224cs.AIq-fin.CP2026-04

用一致性门控测试LLM投资理由可信度,避免过早发布虚假结论。

ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable

  • 设计共识门控协议,检验LLM生成的投资理由是否稳定可靠。
  • 实测显示100个对抗样本中多个维度不达标,最低一致性仅0.2022。
  • 适合评估AI金融系统前的可靠性,防止盲目依赖模型判断。

基于大语言模型的金融代理在实际回报出现前就生成投资理由,导致评价滞后且噪声大。尽管使用LLM judges可提前评估,但未经验证的评判可能奖励冗长、自信或套用模板,而非真实金融判断。本文提出ValueBlindBench,一种预注册的一致性门控压力测试协议,用于判断投资理由是否可发表、合格或无效。在包含1,000次诚实决策和100个预注册对抗控制的原型中(共1,100条轨迹,5,500次评判),整体一致性达\(\barκ_w = 0.7168\),但多个过拟合主张被拦截。低秩系统陷入平局,一个维度(\texttt{constraint\_awareness})不通过,一致性仅\(\barκ_w = 0.2022\);单评委排名受家族依赖影响;简洁正确理由比诚实理由低\(Δ= -2.81\)分。锚定特异性探测进一步表明,如约束意识等金融概念具有操作关键性。研究目标并非构建排行榜或测量真实投资能力,而是作为AI金融评估前的校准层:决定一项基于LLM评判的投资理由是否足够稳定、一致且未被污染,方可报告。

原文摘要 · Abstract (English)

LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a delayed-ground-truth evaluation problem: realized returns remain the eventual arbiter of investment quality, but they arrive too late and are too noisy to guide many model-development and governance decisions. LLM judges offer a tempting shortcut for pre-deployment evaluation of AI-finance systems, but unvalidated judges may reward verbosity, confidence, or rubric mimicry rather than financial judgment. This paper introduces ValueBlindBench, a preregistered agreement-gated stress-test protocol for deciding when LLM-judged investment-rationale claims are publishable, qualified, or invalid. In a controlled market-state capital-allocation prototype with 1,000 honest decision cycles and 100 preregistered adversarial controls (1,100 trajectories, 5,500 judge calls), ValueBlindBench clears the aggregate agreement gate at \(\barκ_w = 0.7168\) but prevents several overclaims. Lower-rank systems collapse into a tie-class, one rubric dimension fails the per-dimension gate (\texttt{constraint\_awareness}, \(\barκ_w = 0.2022\)), single-judge rankings are family-dependent, and terse-correct rationales receive a \(Δ= -2.81\) rubric-point penalty relative to honest rationales. A targeted anchor-specificity probe further shows that financial constructs such as constraint awareness are operationally load-bearing. The scientific object is therefore not a leaderboard and not a claim to measure true investment skill. ValueBlindBench is a pre-calibration metrology layer for AI-finance evaluation: it governs whether a proposed LLM-judge-based investment-rationale claim is stable enough, agreed enough, and uncontaminated enough to be reported at all.

AI金融评测协议一致性测试大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。