提出可信赖的财务图文模型置信度评估方法,解决模型自信但错误的问题。
Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding

- 用训练过的内部探针估计置信度,比仅靠推理的方法更可靠
- 只有探针能生成可设阈值的分数,实现安全自动化决策
- 识别模型未读图表却给出流畅回答的情况,避免虚假自信
视觉语言模型(LVLMs)被越来越多地用于阅读财务图表、表格和文档,其中单个误读数据可能导致重大决策偏差,而模型最权威的回应有时恰恰是未查看图表时生成的。因此核心问题是信任而非准确:哪些答案可直接执行,哪些需提交审核。我们评估了七种置信度估计器——三种仅依赖推理,四种基于训练的内部探针——在五个开源LVLM和四个条件下的表现,涵盖三个财务视觉问答基准(一个双语)。所有探针均仅在自然图像上训练,未针对金融领域微调,以检验跨分布迁移能力。发现三方面:第一,校准性远比排序能力关键;推理基线虽能区分对错,但严重高估置信度,校准误差过大无法设置有效阈值,唯有训练探针能生成可用阈值的分数。第二,可靠性具有结构性,受模型与任务双重影响:最佳探针在20个(模型, 条件)组合中无一领先超过8次;双语对比揭示的‘语言鲁棒性’实为组合效应,在逐模型测试下消失。第三,若将自动处理视为在误差预算下的拒答(deferral),则可自动化的程度首先由模型能力决定,其次才受置信度限制:在简单条件下可安全自动化较高比例,但在困难条件下几乎为零(5%严格预算下接近零)。两种训练探针具备所需校准性,其中仅‘基于根基感知’的探针能在模型未使用图表时降低置信度,有效区分非基于证据的推测与合理回答。
原文摘要 · Abstract (English)
LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometimes one the model produced without reading the exhibit. The operational question is therefore trust, not accuracy: which answers can be acted on, and which escalated to a reviewer. We evaluate seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three financial visual question-answering benchmarks, one bilingual; every probe is trained only on natural images and applied to finance without adaptation, so the results measure out-of-distribution transfer. Three findings hold. First, the scarce property is calibration, not ranking: the inference baselines rank correct above incorrect answers competitively but are badly overconfident, calibration error far above what a threshold can tolerate, and only the trained probes produce a thresholdable score. Second, reliability is structured rather than global, along two axes a practitioner can read directly: the best estimator shifts with both model and task, none leading more than eight of twenty (model, condition) cells, and a controlled bilingual contrast exposes an apparent language robustness as a composition artifact that dissolves once models are read one at a time. Third, cast as deferral under an error budget, how much can be safely automated is set first by the model's competence and only narrowed by its confidence, so deferral clears a real share of the easiest condition and almost none of the hardest, near zero at a strict 5% budget. Two trained probes carry the calibration a deferral policy needs, and among them only the grounding-aware one lowers its confidence on answers a model gives without using the figure, separating detected non-grounding from a fluent guess.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。