arXiv:2502.00989cs.CLcs.AI2025-02被引 2

让AI回答图表问题时能精准定位证据来源,提升可信度。

ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution

  • 用多智能体协作提取图表数据并定位支持答案的图像区域
  • 在多种图表类型上显著优于现有方法,实现细粒度证据标注
  • 适合需要可解释AI决策的专业人士和研究者

大型语言模型(LLMs)虽可回答图表问题,但常生成未经验证的幻觉内容。现有答案归因方法因视觉-语义上下文有限、视觉与文本对齐复杂及复杂布局中边界框预测困难,难以将回答准确锚定到原始图表。我们提出ChartCitor,一种多智能体框架,通过识别图表图像中的支持证据,实现细粒度边界框引用。系统协调多个LLM智能体完成图表到表格的提取、答案重述、表格增强、基于预筛选与重排序的证据检索,以及表格到图表的映射。ChartCitor在不同图表类型上均优于现有基线。定性用户研究表明,该系统通过提供更清晰的可解释性,增强了用户对生成式AI的信任,并帮助专业人士提升工作效率。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can perform chart question-answering tasks but often generate unverified hallucinated responses. Existing answer attribution methods struggle to ground responses in source charts due to limited visual-semantic context, complex visual-text alignment requirements, and difficulties in bounding box prediction across complex layouts. We present ChartCitor, a multi-agent framework that provides fine-grained bounding box citations by identifying supporting evidence within chart images. The system orchestrates LLM agents to perform chart-to-table extraction, answer reformulation, table augmentation, evidence retrieval through pre-filtering and re-ranking, and table-to-chart mapping. ChartCitor outperforms existing baselines across different chart types. Qualitative user studies show that ChartCitor helps increase user trust in Generative AI by providing enhanced explainability for LLM-assisted chart QA and enables professionals to be more productive.

图表理解可解释AI多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。