多模态大模型可有效识别文档篡改,但性能受微调影响大。
Can Multi-modal (reasoning) LLMs detect document manipulation?
- 用提示优化和推理分析评估多个多模态LLM的欺诈检测能力。
- 顶尖模型在零样本泛化上优于传统方法,但在分布外数据表现不一。
- 模型规模与推理能力对检测准确率影响有限,微调更关键。
文档欺诈对依赖安全可验证文档的行业构成重大威胁,亟需可靠的检测机制。本研究评估了包括OpenAI O1、OpenAI 4o、Gemini Flash(思考模式)、Deepseek Janus、Grok、Llama 3.2 和 4、Qwen 2 和 2.5 VL、Mistral Pixtral、Claude 3.5 和 3.7 Sonnet在内的多模态大语言模型在检测文档欺诈方面的有效性。在真实交易文档的标准数据集上,通过提示优化和对模型推理过程的详细分析,评估其识别文本篡改、格式错位及交易金额不一致等细微欺诈信号的能力。结果表明,表现最佳的多模态LLM展现出优越的零样本泛化能力,在分布外数据集上超越传统方法;而部分视觉LLM则表现不稳定或不佳。值得注意的是,模型规模与先进推理能力与检测准确率相关性较弱,表明任务特定微调至关重要。本研究揭示了多模态LLM在提升文档欺诈检测系统中的潜力,并为未来可解释、可扩展的欺诈缓解策略研究奠定基础。
原文摘要 · Abstract (English)
Document fraud poses a significant threat to industries reliant on secure and verifiable documentation, necessitating robust detection mechanisms. This study investigates the efficacy of state-of-the-art multi-modal large language models (LLMs)-including OpenAI O1, OpenAI 4o, Gemini Flash (thinking), Deepseek Janus, Grok, Llama 3.2 and 4, Qwen 2 and 2.5 VL, Mistral Pixtral, and Claude 3.5 and 3.7 Sonnet-in detecting fraudulent documents. We benchmark these models against each other and prior work on document fraud detection techniques using a standard dataset with real transactional documents. Through prompt optimization and detailed analysis of the models' reasoning processes, we evaluate their ability to identify subtle indicators of fraud, such as tampered text, misaligned formatting, and inconsistent transactional sums. Our results reveal that top-performing multi-modal LLMs demonstrate superior zero-shot generalization, outperforming conventional methods on out-of-distribution datasets, while several vision LLMs exhibit inconsistent or subpar performance. Notably, model size and advanced reasoning capabilities show limited correlation with detection accuracy, suggesting task-specific fine-tuning is critical. This study underscores the potential of multi-modal LLMs in enhancing document fraud detection systems and provides a foundation for future research into interpretable and scalable fraud mitigation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。