测试大模型在多语言金融假信息中的情境偏见,发现商业与开源模型均存明显偏差。
Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection
- 构建三类复杂金融情境,融合角色、地域、种族与信仰因素
- 覆盖英、中、希、孟加拉语的多语言假信息数据集,评估22个主流大模型
- 揭示模型在真实金融场景下普遍存在行为偏见,适合安全与合规研究者参考
大语言模型(LLMs)在金融领域广泛应用,但其训练数据源自人类文本,可能继承人类行为偏见,导致决策不稳定,尤其在处理金融信息时。现有研究多关注简单问答场景,缺乏对复杂现实金融环境和高风险、上下文敏感的多语言金融虚假信息检测(MFMD)任务的关注。本文提出MFMDScen,一个全面的基准测试框架,用于评估不同经济情境下LLMs在MFMD中的行为偏见。与金融专家合作,构建三类复杂情境:(i)基于角色与个性,(ii)基于角色与地区,(iii)融合种族与宗教信仰的角色情境。同时开发涵盖英语、中文、希腊语和孟加拉语的多语言金融虚假信息数据集。通过将这些情境与虚假声明结合,实现对22个主流大模型的系统性评估。结果表明,商业与开源模型均存在显著行为偏见。项目代码已开源:https://github.com/lzw108/FMD。
原文摘要 · Abstract (English)
Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-authored corpora, LLMs may inherit a range of human biases. Behavioral biases can lead to instability and uncertainty in decision-making, particularly when processing financial information. However, existing research on LLM bias has mainly focused on direct questioning or simplified, general-purpose settings, with limited consideration of the complex real-world financial environments and high-risk, context-sensitive, multilingual financial misinformation detection tasks MFMD. In this work, we propose MFMDScen, a comprehensive benchmark for evaluating behavioral biases of LLMs in MFMD across diverse economic scenarios. In collaboration with financial experts, we construct three types of complex financial scenarios: (i) role- and personality-based, (ii) role- and region-based, and (iii) role-based scenarios incorporating ethnicity and religious beliefs. We further develop a multilingual financial misinformation dataset covering English, Chinese, Greek, and Bengali. By integrating these scenarios with misinformation claims, MFMDScen enables a systematic evaluation of 22 mainstream LLMs. Our findings reveal that pronounced behavioral biases persist across both commercial and open-source models. This project is available at https://github.com/lzw108/FMD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。