测试大模型在金融偏见下的决策能力,发现其易受人为偏见影响。
Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain
- 构建包含8868份长篇财报的基准数据集,模拟真实金融环境中的偏见干扰。
- 模型在有显性偏见时倾向盲目跟风,但通过去偏方法可超越人类预测表现。
- 适合关注AI金融决策可靠性的研究人员和投资者使用。
大型语言模型(LLMs)在金融领域的应用日益广泛,但其可靠性、对齐性及对抗性操纵的脆弱性引发关注。现有金融类基准多局限于小样本,未能充分展示模型在不确定性与潜在人为偏见下的表现。本文提出Fin-Bias,一个评估大模型在长期不确定金融情境下投资决策能力的基准,特别关注金融群体行为(herding)。该基准包含8868份企业特定分析师报告,涵盖不同行业,每份报告均附有由专业分析师给出的买入/中性/卖出评级。我们向LLMs提供带或不带评级的报告,甚至引入伪造评级,以生成其自身投资判断。结果表明,模型倾向于追随上下文中的显性偏见。我们还开发了一种检测人为观点的方法,促使模型独立思考,在部分任务中其对未来股价回报的预测甚至优于人类。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in financial contexts, raising critical concerns about reliability, alignment, and susceptibility to adversarial manipulation. While prior finance-related benchmarks assess LLMs' capabilities in stock trading, they are often restricted to small sample and fail to demonstrate LLM susceptibility to context with potential human bias. We introduce Fin-Bias (financial herding under long and uncertain financial context), a benchmark for evaluating LLM investment decision-making when faced with uncertainty and possible human-biased opinions. Fin-Bias includes 8868 long firm-specific analyst reports, including firm aspects summarized and analyzed by sophisticated analysts with investment ratings (Bullish/Neutral/Bearish) spanning from various industries. We present large language models with firm analyst reports with/without analyst investment ratings and even with 'fake' rating, to get investment ratings generated by LLMs. Our results reveal that LLMs tend to herd the explicit bias in context. We also develop a method to detect potential human opinions, which can encourage LLMs to think independently, some models even exceed human performance in predicting future stock return.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。