从行为金融学视角揭示视觉语言模型的决策偏见。
Behavioral Bias of Vision-Language Models: A Behavioral Finance View
- 构建端到端框架,评估模型在两类金融偏见下的表现
- 开源模型普遍存在近期性与权威性偏见,闭源模型影响小
- 为开源模型改进提供方向,适合关注AI可信性的研究者
大型视觉语言模型(LVLMs)随着大型语言模型(LLMs)接入视觉模块而迅速发展,形成更类人的能力。然而,其在不同领域的应用需谨慎评估,因其可能携带不可预期的偏见。本文从行为金融学视角出发,结合金融与心理学,研究LVLMs潜在的行为偏见。提出一个从数据收集到新评估指标的端到端框架,用于评估模型在两种经典人类金融行为偏见——近期性偏见和权威性偏见——中的动态表现。评估结果显示,诸如LLaVA-NeXT、MobileVLM-V2、Mini-Gemini、MiniCPM-Llama3-V 2.5和Phi-3-vision-128k等近期开源模型显著受这两种偏见影响,而专有模型GPT-4o几乎不受影响。研究结果指出了开源模型可改进的方向。代码已公开于https://github.com/mydcxiao/vlm_behavioral_fin。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) evolve rapidly as Large Language Models (LLMs) was equipped with vision modules to create more human-like models. However, we should carefully evaluate their applications in different domains, as they may possess undesired biases. Our work studies the potential behavioral biases of LVLMs from a behavioral finance perspective, an interdisciplinary subject that jointly considers finance and psychology. We propose an end-to-end framework, from data collection to new evaluation metrics, to assess LVLMs' reasoning capabilities and the dynamic behaviors manifested in two established human financial behavioral biases: recency bias and authority bias. Our evaluations find that recent open-source LVLMs such as LLaVA-NeXT, MobileVLM-V2, Mini-Gemini, MiniCPM-Llama3-V 2.5 and Phi-3-vision-128k suffer significantly from these two biases, while the proprietary model GPT-4o is negligibly impacted. Our observations highlight directions in which open-source models can improve. The code is available at https://github.com/mydcxiao/vlm_behavioral_fin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。