LLM金融分析中,角色设定比检索内容更影响判断偏差。
The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

- 用角色提示控制模型解释同一财务文件的方式。
- 90%以上的判断差异源于解释偏差而非证据选择不同。
- 调整用户身份描述或分离输出可减轻但无法根除偏差。
大型语言模型(LLMs)越来越多地利用用户上下文(如记忆、档案和角色提示)来个性化回应。这种个性化可能影响基于证据的判断:相同的证据在不同用户上下文中可能导致不同结论。金融领域因决策依赖于解读长而复杂的文档,是研究此问题的理想场景。我们使用3,575份美国证券交易委员会(SEC)文件,在十二个LLM上进行了测试。通过对比角色条件检索、中性检索和记忆框架上下文,区分了证据选择与解释的影响。结果发现,多数用户上下文溢出效应来自模型在不同角色下对相同证据的解释差异,而非检索到不同证据。随后测试两种简单缓解策略:将相同的投资者心态以用户档案形式表达而非助手角色,以及分离基于证据与个性化的输出。两者均能减少溢出效应,但未完全消除,且效果在不同模型间差异显著。
原文摘要 · Abstract (English)
Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。