arXiv:2510.05702q-fin.CPcs.AI2025-10被引 2

发现开源大模型在投资决策中存在公司规模和估值偏见

Uncovering Representation Bias for Investment Decisions in Open-Source Large Language Models

  • 用轮转提示+约束解码,量化模型对150家美股的置信度
  • 模型对大公司和高估值企业更自信,风险越高越不自信
  • 技术行业置信度波动最大,适合做金融场景校准

大型语言模型正被广泛用于金融投资流程,但现有研究很少关注其在公司规模、行业或财务特征上的表征偏差,这些偏差可能显著影响决策。本文聚焦开源Qwen模型的表征偏差,针对约150家美国上市公司,采用平衡轮转提示法,结合约束解码与词元-对数几率聚合,生成不同金融语境下的企业级置信度评分。通过统计检验与方差分析发现,企业规模与估值持续提升模型置信度,而风险因素则降低置信度;行业间置信度差异显著,其中科技行业波动最大。当模型被提示特定金融类别时,其置信度排名与基本面数据最一致,其次为技术信号,与增长指标一致性最低。结果揭示了Qwen模型中的表征偏差,提示需采用行业敏感的校准与类别条件评估机制,以实现安全、公平的金融大模型部署。

原文摘要 · Abstract (English)

Large Language Models are increasingly adopted in financial applications to support investment workflows. However, prior studies have seldom examined how these models reflect biases related to firm size, sector, or financial characteristics, which can significantly impact decision-making. This paper addresses this gap by focusing on representation bias in open-source Qwen models. We propose a balanced round-robin prompting method over approximately 150 U.S. equities, applying constrained decoding and token-logit aggregation to derive firm-level confidence scores across financial contexts. Using statistical tests and variance analysis, we find that firm size and valuation consistently increase model confidence, while risk factors tend to decrease it. Confidence varies significantly across sectors, with the Technology sector showing the greatest variability. When models are prompted for specific financial categories, their confidence rankings best align with fundamental data, moderately with technical signals, and least with growth indicators. These results highlight representation bias in Qwen models and motivate sector-aware calibration and category-conditioned evaluation protocols for safe and fair financial LLM deployment.

大模型偏见金融AIQwen投资决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。