用可视化矩阵帮金融分析师看清模型回答哪些有证据、哪些是瞎编。
EvidenceLens: A Claim-Evidence Matrix for Auditing Financial Question Answering

- 把回答拆成小主张,匹配原文、表格和图表证据
- 一眼看出证据覆盖不全或自相矛盾的地方
- 适合审计人员验证财报问答结果,提升可信度
大语言模型越来越多用于回答年报、业绩简报和分析师报告中的问题,但其输出在高风险金融场景中难以验证。流畅的回答可能混杂有依据的陈述、弱关联的合成内容以及无支持的断言,涵盖文本、表格和图表。我们提出 EvidenceLens,一个视觉分析原型,将金融问答视为主张-证据对齐问题。系统将答案分解为原子主张,汇总支持情况与置信度,识别支持缺口,并将主张级检查与原始段落、表格单元格和图表区域联动。其核心是多模态主张-证据矩阵,可即时呈现证据覆盖、矛盾关系及模态失衡。为保证可复现性,我们还定义了基于 JSON 的数据结构规范、轻量级多模态对齐流程,以及确定性的优先级排序机制,将后端信号映射为可审计的可视化结构。通过典型报告审计场景,我们展示 EvidenceLens 如何帮助分析师区分有依据的主张与过度自信的合成内容,而传统聊天界面会将其混淆。
原文摘要 · Abstract (English)
Large language models are increasingly used to answer questions over annual reports, earnings decks, and analyst notes, yet their outputs remain difficult to verify in high-stakes financial workflows. A fluent answer can blend directly grounded statements, weak synthesis, and unsupported claims across narrative text, tables, and charts. We present EvidenceLens, a visual analytics prototype that treats financial question answering as a claim-evidence alignment problem. The system decomposes an answer into atomic claims, summarizes support composition and confidence, support gaps, and coordinates claim-level inspection with source passages, table cells, and chart regions. Its core visual representation is a multimodal claim-evidence matrix that makes coverage, contradiction, and modality imbalance immediately visible. To support reproducibility, we also specify a JSON-based artifact schema, a lightweight multimodal alignment pipeline, and a deterministic review-priority ranking that maps backend signals into an auditable visual structure. Through representative report-auditing scenarios, we show how EvidenceLens helps analysts distinguish grounded claims from overconfident synthesis that conventional chat interfaces flatten.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。