九智能体框架提升金融问答可信度,逐条验证每项声明。
CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

- 将问题拆解为原子化声明,按类型分配证据可信度。
- 证据信任度提升至0.889,5.4%问题因证据不足主动放弃回答。
- 支持动态辩论与风险监控,适合高要求金融场景使用。
现有检索增强与多智能体系统的防幻觉机制存在缺陷:不同模态间矛盾时仍盲目信任证据,辩论仅针对整体报告而非具体声明,且验证滞后至生成后,导致智能体间错误无法及时发现。为此,我们提出 CLAIR-Fin,一个九智能体框架,将每个问题分解为在金融声明账本中维护的原子声明。通过不对称证据权威机制,根据声明类型决定证据可信度,而非等同对待所有模态;链式交接验证在撰写与对抗审查交接时检查事实依据,而非仅在流程末端;自适应反驳循环根据争议程度动态调节辩论深度;终端蕴含审计与持续幻觉风险指数区分通过审查与未被挑战的声明。我们在包含500个问题的跨模态金融评估集 BB-FinQA-X 上进行评测,该数据集基于孟加拉国银行年报构建,按查询类型、格式和难度分层。相较单次检索增强生成基线,其忠实度从0.780提升至0.889,并在证据不足时主动拒绝回答5.4%的问题,避免强行生成无依据回答;同时超越更强的 HyDE 和 Graph-RAG 等检索策略基线(≤0.874)。
原文摘要 · Abstract (English)
Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occurs only after drafting, leaving inter-agent errors undetected until the final text. To close this gap, we present CLAIR-Fin, a nine-agent framework that decomposes each question into atomic claims maintained in a typed Financial Claim Ledger. Each claim is resolved through Asymmetric Evidence Authority, which conditions evidence trust on claim type rather than treating all modalities as equally reliable; Chain-of-Custody Verification, which checks grounding at the hand-off between drafting and adversarial review rather than only at the pipeline's exit; an Adaptive Rebuttal Cycle, which routes contested claims through adversarial debate whose depth scales with what that debate finds; and a terminal entailment audit paired with a continuous Hallucination Risk Index that distinguishes claims that passed scrutiny from claims never contested. We evaluate CLAIR-Fin on BB-FinQA-X, a 500-question cross-modal financial evaluation set built from Bangladesh Bank Annual Report material, stratified by query type, format, and difficulty. Relative to a single-pass retrieval-augmented generation baseline, it raises faithfulness ($0.780 \rightarrow 0.889$) while abstaining on 5.4% of questions when evidence is insufficient rather than forcing an unsupported response, and it exceeds stronger retrieval-strategy baselines such as HyDE and Graph-RAG on faithfulness ($\leq 0.874$).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。