构建真实金融文档解析系统,突破传统评测局限
FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

- 采用视觉语言模型与强化学习结合的端到端解析框架
- 在真实场景下综合表现达81.43分,内部流程解析提升5.35点
- 适合需要高精度、强鲁棒性的金融自动化系统开发者
金融文档解析需兼顾准确性、结构一致性和可验证性,但现有评测基准常无法反映实际性能。本文提出FinixDoc,一个面向真实金融文档的端到端智能解析系统,核心为基于Qwen3-VL-4B的40亿参数视觉语言模型FinixDoc-VL。为揭示评测与部署间的差距,提出沿视觉质量与文档规模双轴划分的文档解析能力矩阵。据此,通过融合同形字符感知对比学习与多阶段强化学习(含领域特定奖励)进行域适应训练。为高效利用低质数据并支撑高质量数据生成,构建了带置信度感知专家评审的人工介入数据工厂流水线。评估方面,构建覆盖数字原生、摄像头拍摄、超大页面及内部工作流场景的FinixDocBench评测集,其中合规审查子集随报告发布。在主要子集上,FinixDoc-VL以81.43分位列所有基线最高,优于次优开源模型5.13分,尤其在内部财务流程任务(FinixInner: 84.08 vs. 78.73)中优势显著。
原文摘要 · Abstract (English)
Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We present FinixDoc, an end-to-end agentic parsing system for real-world financial documents, with FinixDoc-VL, a 4B-scale vision-language model built on Qwen3-VL-4B, as its core parser. To characterize the gap between benchmark and deployment performance, we introduce a Document Parsing Capability Matrix organized along two practical axes: visual quality and document scale. Guided by this matrix, FinixDoc-VL is trained with a domain-adapted recipe combining homoglyph-aware contrastive learning and multi-stage reinforcement learning with composite domain-specific rewards. To better leverage our accumulated advantage in low-quality financial-document data and support large-scale, high-quality data production, we further build a human-in-the-loop Data Factory pipeline with confidence-aware expert review. For evaluation, we construct FinixDocBench, a financial-domain evaluation suite covering digital-native, camera-captured, ultra-large-page, and internal-workflow scenarios, with a compliance-reviewed subset released alongside this technical report. On its main subsets, FinixDoc-VL achieves the highest overall score (81.43) among evaluated baselines, outperforming the next-best open-source model by 5.13 points, with the largest gains on internal financial workflows (FinixInner: 84.08 vs. 78.73).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。