测试主流AI评估投资风险能力,发现其结果受用户国籍性别影响,不符合金融监管要求。
Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?
- 使用真实用户画像测试GPT、Claude等主流AI和开源模型的风险评分机制
- 模型在尼日利亚、印度尼西亚等国用户上给出更高风险评分,违背公平性原则
- 无一模型在跨地区跨性别场景下保持评分一致性,不适合用于金融风险自动化
我们评估了AI系统在投资风险偏好评估中的可信度——这一任务在自动化前必须经过严格验证。研究基于专有模型(GPT、Claude、Gemini)和开源权重模型(LLaMA、DeepSeek、Mistral),采用精心设计的用户画像,涵盖不同国家与性别等属性。结果显示,当不应影响风险计算的用户属性(如国籍或性别)改变时,模型得分分布出现显著差异。例如,GPT-4o对尼日利亚和印度尼西亚用户的评分更高。尽管部分模型在低风险和中风险区间与预期得分接近,但均无法在区域和人口统计维度保持评分一致性,违反了人工智能与金融监管要求。
原文摘要 · Abstract (English)
We assess whether AI systems can credibly evaluate investment risk appetite-a task that must be thoroughly validated before automation. Our analysis was conducted on proprietary systems (GPT, Claude, Gemini) and open-weight models (LLaMA, DeepSeek, Mistral), using carefully curated user profiles that reflect real users with varying attributes such as country and gender. As a result, the models exhibit significant variance in score distributions when user attributes-such as country or gender-that should not influence risk computation are changed. For example, GPT-4o assigns higher risk scores to Nigerian and Indonesian profiles. While some models align closely with expected scores in the Low- and Mid-risk ranges, none maintain consistent scores across regions and demographics, thereby violating AI and finance regulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。