构建金融领域可信度评估基准,测试大模型在安全与合规上的表现
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
- 设计细粒度任务覆盖金融场景下的可信度维度
- 开源模型在行业公平性上优于闭源模型,但整体在法律意识上不足
- 适用于评估金融AI系统风险,适合研究者与从业者参考
近年来,大语言模型在解决金融问题方面展现出巨大潜力。然而,由于金融应用具有高风险、高重要性的特点,实际落地仍面临挑战。本文提出FinTrust,一个专为金融领域可信度评估设计的综合性基准。该基准基于真实业务场景,涵盖多维度对齐问题,并设置精细化任务以评估模型在安全性、公平性、合规性等方面的表现。我们在FinTrust上评估了十一款大模型,发现闭源模型如o4-mini在多数任务(如安全)中表现更优,而开源模型DeepSeek-V3在行业级公平性方面具有一定优势。但在涉及受托责任对齐和信息披露等复杂任务上,所有模型均表现不佳,暴露出明显的法律认知差距。我们认为FinTrust可成为评估金融领域大模型可信度的重要工具。
原文摘要 · Abstract (English)
Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property. This paper introduces FinTrust, a comprehensive benchmark specifically designed for evaluating the trustworthiness of LLMs in finance applications. Our benchmark focuses on a wide range of alignment issues based on practical context and features fine-grained tasks for each dimension of trustworthiness evaluation. We assess eleven LLMs on FinTrust and find that proprietary models like o4-mini outperforms in most tasks such as safety while open-source models like DeepSeek-V3 have advantage in specific areas like industry-level fairness. For challenging task like fiduciary alignment and disclosure, all LLMs fall short, showing a significant gap in legal awareness. We believe that FinTrust can be a valuable benchmark for LLMs' trustworthiness evaluation in finance domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。