arXiv:2505.24650cs.CEcs.LG2025-05被引 18

用可解释性技术破解金融大模型黑箱,提升透明度与合规性。

Beyond the Black Box: Interpretability of LLMs in Finance

  • 通过逆向分析模型内部激活机制,揭示金融任务中的决策逻辑。
  • 在交易策略、情绪分析等场景中验证了对幻觉和偏见的有效检测能力。
  • 为监管合规提供可追溯的解释路径,适合金融AI开发者与监管机构。

大型语言模型在金融领域展现出强大能力,涵盖报告生成、聊天机器人、情感分析、合规审查、投资建议、金融知识检索与摘要等任务。然而其内在复杂性与缺乏透明度,在高度监管的金融环境中带来显著挑战,尤其在可解释性、公平性与问责性方面。据我们所知,本文首次将机制可解释性应用于金融领域,通过逆向解析模型内部激活与电路结构,揭示特定特征如何影响预测结果,实现对模型行为的观察与调控。文中探讨了机制可解释性的理论基础,并通过一系列金融应用场景实验展示其实际价值,包括交易策略、情感分析、偏见与幻觉检测。尽管尚未广泛应用,但随着大模型在金融领域的普及,该技术预计将愈发关键。先进可解释工具有助于确保AI系统符合伦理规范、保持透明,并顺应不断演进的金融监管要求。本文特别强调这些技术如何满足监管与合规的可解释性需求,既应对当前挑战,也预判全球监管机构未来期望。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit remarkable capabilities across a spectrum of tasks in financial services, including report generation, chatbots, sentiment analysis, regulatory compliance, investment advisory, financial knowledge retrieval, and summarization. However, their intrinsic complexity and lack of transparency pose significant challenges, especially in the highly regulated financial sector, where interpretability, fairness, and accountability are critical. As far as we are aware, this paper presents the first application in the finance domain of understanding and utilizing the inner workings of LLMs through mechanistic interpretability, addressing the pressing need for transparency and control in AI systems. Mechanistic interpretability is the most intuitive and transparent way to understand LLM behavior by reverse-engineering their internal workings. By dissecting the activations and circuits within these models, it provides insights into how specific features or components influence predictions - making it possible not only to observe but also to modify model behavior. In this paper, we explore the theoretical aspects of mechanistic interpretability and demonstrate its practical relevance through a range of financial use cases and experiments, including applications in trading strategies, sentiment analysis, bias, and hallucination detection. While not yet widely adopted, mechanistic interpretability is expected to become increasingly vital as adoption of LLMs increases. Advanced interpretability tools can ensure AI systems remain ethical, transparent, and aligned with evolving financial regulations. In this paper, we have put special emphasis on how these techniques can help unlock interpretability requirements for regulatory and compliance purposes - addressing both current needs and anticipating future expectations from financial regulators globally.

可解释性金融AI大模型机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。