arXiv:2604.25186cs.CVcs.CE2026-04

构建首个文档视频智能评测基准,助力金融反欺诈与证据追溯。

FCMBench-Video: Benchmarking Document Video Intelligence

论文配图:FCMBench-Video: Benchmarking Document Video Intelligence
图 1 · 摘自论文原文
  • 设计原子化采集流程,生成可复用的多文档视频数据
  • 覆盖28类文档、1200段长视频,含1.1万条专家标注问答
  • 适配金融信贷场景,评估跨帧证据整合与真实性判断能力

文档理解在金融信贷审核、开户和远程验证中至关重要,既需决策准确,也需证据可追溯。相较于静态图像,文档视频具有时间冗余、顺序展开的证据流特征,需跨帧融合证据,并保留与真实性敏感和反欺诈相关的采集过程线索。我们提出FCMBench-Video,一个在真实拍摄条件下评估文档感知、时序定位与证据驱动推理的基准。为实现大规模隐私合规且真实的数据构建,采用原子级采集与组合流程:录制可复用的单文档片段,施加可控退化,按指定时间跨度组装成长视频。该基准由495个原子视频构成1,200段长视频,配有11,322条专家标注的问答对,涵盖28种文档类型,持续时间20–60秒,含5,960条中文和5,362条英文实例。对九个近期Video-MLLMs的评估显示,该基准能有效区分系统性能:计数任务最敏感于时长;跨文档验证与证据选择考验高层证据整合;视觉提示注入揭示互补鲁棒性。整体得分分布广且近似正态,表明基准既未饱和也无简单样本主导。综合结果表明,FCMBench-Video是追踪Video-MLLM在文档视频理解上进展、探索金融领域真实性敏感应用能力边界的可复现基准。

原文摘要 · Abstract (English)

Document understanding is a critical capability in financial credit review, onboarding, and remote verification, where both decision accuracy and evidence traceability matter. Compared with static document images, document videos present a temporally redundant and sequentially unfolding evidence stream, require evidence integration across frames, and preserve acquisition-process cues relevant to authenticity-sensitive and anti-fraud review. We introduce FCMBench-Video, a benchmark for document-video intelligence that evaluates document perception, temporal grounding, and evidence-grounded reasoning under realistic capture conditions. For privacy-compliant yet realistic data at scale, we organize construction as an atomic-acquisition and composition workflow that records reusable single-document clips, applies controlled degradations, and assembles long-form multi-document videos with prescribed temporal spans. FCMBench-Video is built from 495 atomic videos composed into 1,200 long-form videos paired with 11,322 expert-annotated question--answer instances, covering 28 document types over 20s--60s duration tiers and 5,960 Chinese / 5,362 English instances. Evaluations on nine recent Video-MLLMs show that FCMBench-Video provides meaningful separation across systems and capabilities: counting is the most duration-sensitive task, Cross-Document Validation and Evidence-Grounded Selection probe higher-level evidence integration, and Visual Prompt Injection provides a complementary robustness dimension. The overall score distribution is broad and approximately bell-shaped, indicating a benchmark that is neither saturated nor dominated by trivial cases. Together, these results position FCMBench-Video as a reproducible benchmark for tracking Video-MLLM progress on document-video understanding and probing capability boundaries in authenticity-sensitive credit-domain applications.

文档理解视频评测金融AI多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。