AI金融分析中,能检索不等于能用,信息整合才是关键。
Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

- 设计结构化工作流,让信息从检索到判断无缝传递。
- 即使模型能准确检索,128,000词上下文仍使风险披露影响降为噪声。
- 适合关注AI投资决策可信度的研究者与从业者。
大型语言模型(LLMs)正被用作AI分析师处理财务披露并支持投资决策。然而,现有评估通常只关注其能否检索信息,而非这些信息是否真正影响判断。我们发现长上下文金融分析中存在检索-整合断层:在固定目标公司信息、仅变化无关上下文(2,000至128,000词)的实验中,风险披露对投资判断的影响降至实验噪声水平,而直接检索仍保持准确。该现象在多个模型家族和判断任务中复现,且在去除真实10-K文件中的披露内容后依然存在。更强大的模型仅延迟但未消除此断层。因果记忆干预表明,压缩摘要与原文查找共同将披露信息传入判断。工作流架构决定传输成败:分块-摘要流程会剔除相关信息,而紧邻决策的定向结构重述则可恢复其影响。因此,AI分析师表现由模型能力与工作流架构共同决定。基于检索的评估可能认证那些实际忽略所检索信息的系统。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。