arXiv:2511.12154cs.LGcs.AI2025-11被引 3

用自监督学习融合金融交易的结构与文本信息,提升小数据场景下的金融分析能力

Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions

  • 将交易的结构化属性与非结构化描述统一建模为多模态表示
  • 在北美数千家机构的数据上验证了跨地域、跨机构的泛化能力
  • 适合金融风控、信用评估等需少样本学习的场景

我们提出了一种面向金融交易的多模态基础模型,将结构化属性与非结构化文本描述整合为统一表征。通过将掩码语言建模适配到交易序列,证明该方法不仅优于传统特征工程和离散事件序列方法,尤其在数据稀缺的开放银行场景中表现优异。据我们所知,这是首次在北美数千家金融机构上的大规模研究,表明多模态表征可跨地理区域和机构实现泛化。结果凸显自监督模型在欺诈检测、信用风险评估及客户洞察等金融应用中的潜力。

原文摘要 · Abstract (English)

We introduced a multimodal foundational model for financial transactions that integrates both structured attributes and unstructured textual descriptions into a unified representation. By adapting masked language modeling to transaction sequences, we demonstrated that our approach not only outperforms classical feature engineering and discrete event sequence methods but is also particularly effective in data-scarce Open Banking scenarios. To our knowledge, this is the first large-scale study across thousands of financial institutions in North America, providing evidence that multimodal representations can generalize across geographies and institutions. These results highlight the potential of self-supervised models to advance financial applications ranging from fraud prevention and credit risk to customer insights

金融AI多模态自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。