arXiv:2509.18405cs.CVcs.AI2025-09被引 1

用多模态大模型实现零样本支票字段检测,无需训练即可用。

Check Field Detection Agent (CFD-Agent) using Multimodal Large Language and Vision Language Models

  • 利用多模态大模型与视觉语言模型结合,无需训练直接识别支票关键字段。
  • 在110张不同格式支票上测试,零样本检测表现良好,泛化能力强。
  • 可生成高质量标注数据,适合金融机构快速部署和定制化检测系统。

支票仍是金融体系中的基础工具,支撑着大量跨机构交易,但其使用也使其长期成为欺诈目标,因此建立强大的支票欺诈检测机制至关重要。此类系统的核心在于准确识别和定位关键字段,如签名、磁墨字符识别(MICR)行、金额(草写/法定)、收款人和付款人等,这些字段需与同一客户的历史参考支票进行比对验证。传统方法依赖于在大规模、多样化且精标注数据集上训练的对象检测模型,但这类数据因隐私和专有性问题难以获取。本文提出一种全新的无训练框架,通过结合视觉语言模型(VLM)与多模态大语言模型(MLLM),实现支票组件的零样本检测,显著降低实际金融场景中的部署门槛。在人工精心整理的110张涵盖多种格式和布局的支票数据集上进行量化评估,结果表明该模型具备优异性能与泛化能力。此外,该框架还可作为生成高质量标注数据的启动机制,支持为特定机构需求定制实时对象检测模型。

原文摘要 · Abstract (English)

Checks remain a foundational instrument in the financial ecosystem, facilitating substantial transaction volumes across institutions. However, their continued use also renders them a persistent target for fraud, underscoring the importance of robust check fraud detection mechanisms. At the core of such systems lies the accurate identification and localization of critical fields, such as the signature, magnetic ink character recognition (MICR) line, courtesy amount, legal amount, payee, and payer, which are essential for subsequent verification against reference checks belonging to the same customer. This field-level detection is traditionally dependent on object detection models trained on large, diverse, and meticulously labeled datasets, a resource that is scarce due to proprietary and privacy concerns. In this paper, we introduce a novel, training-free framework for automated check field detection, leveraging the power of a vision language model (VLM) in conjunction with a multimodal large language model (MLLM). Our approach enables zero-shot detection of check components, significantly lowering the barrier to deployment in real-world financial settings. Quantitative evaluation of our model on a hand-curated dataset of 110 checks spanning multiple formats and layouts demonstrates strong performance and generalization capability. Furthermore, this framework can serve as a bootstrap mechanism for generating high-quality labeled datasets, enabling the development of specialized real-time object detection models tailored to institutional needs.

多模态支票检测零样本金融安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。