评测大模型在法语金融文档上的表现,发现图表理解差且对话中错误会累积。
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
- 构建首个法语金融文档多模态基准Scribe Finance,含1204个专家验证问题。
- 模型对文本和表格准确率达85-90%,但图表理解仅34-62%。
- 多轮对话中早期错误导致整体准确率跌至约50%,适合金融领域研究者参考。
视觉语言模型(VLMs)在许多文档理解任务中表现良好,但在专业非英语领域的可靠性仍缺乏深入探索。这一差距在金融领域尤为关键,因为金融文档包含密集的监管文本、数值表格和视觉图表,提取错误可能带来实际后果。我们提出Scribe Finance,首个用于评估法语金融文档理解能力的多模态基准。该数据集包含1,204个由专家验证的问题,涵盖文本抽取、表格理解、图表解读及多轮对话推理,源自真实的投资说明书、KIDs和PRIIPs。我们采用LLM-as-judge协议,评估了六种开源权重的VLM(参数量8B-124B)。尽管模型在文本和表格任务上表现良好(准确率85-90%),但在图表理解方面表现较差(34-62%)。最显著的是,在多轮对话中,早期错误会持续传播,导致整体准确率降至约50%,与模型规模无关。结果表明,当前VLM在明确抽取任务中有效,但在交互式、多步金融分析中仍显脆弱。Scribe Finance为高风险场景下的性能评估与进步提供了一个挑战性基准。
原文摘要 · Abstract (English)
Vision-language models (VLMs) perform well on many document understanding tasks, yet their reliability in specialized, non-English domains remains underexplored. This gap is especially critical in finance, where documents mix dense regulatory text, numerical tables, and visual charts, and where extraction errors can have real-world consequences. We introduce Scribe Finance, the first multimodal benchmark for evaluating French financial document understanding. The dataset contains 1,204 expert-validated questions spanning text extraction, table comprehension, chart interpretation, and multi-turn conversational reasoning, drawn from real investment prospectuses, KIDs, and PRIIPs. We evaluate six open-weight VLMs (8B-124B parameters) using an LLM-as-judge protocol. While models achieve strong performance on text and table tasks (85-90% accuracy), they struggle with chart interpretation (34-62%). Most notably, multi-turn dialogue reveals a sharp failure mode: early mistakes propagate across turns, driving accuracy down to roughly 50% regardless of model size. These results show that current VLMs are effective for well-defined extraction tasks but remain brittle in interactive, multi-step financial analysis. Scribe Finance offers a challenging benchmark to measure and drive progress in this high-stakes setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。