MuDoC让AI对话时能联动文档图文,实时定位来源增强可信度。
MuDoC: An Interactive Multimodal Document-grounded Conversational AI System
- 基于GPT-4o构建,支持文本与图表混合响应生成
- 可即时跳转原文本和图示,实现回答可验证
- 适合需要高可信度文档交互的科研、教育场景
多模态AI是提升人机交互效率的关键方向。现有系统在长文档的多模态交互上仍存在挑战。本文提出基于GPT-4o的交互式多模态文档对话系统MuDoC,能够结合文档中的文字内容与可视化元素(如图表)生成答案。系统通过智能教材界面支持用户实时回溯至原始文本与图表,提升回应可信度与可验证性。我们对MuDoC生成结果进行定性分析,揭示其在理解复杂图文关联方面的优势与局限。
原文摘要 · Abstract (English)
Multimodal AI is an important step towards building effective tools to leverage multiple modalities in human-AI communication. Building a multimodal document-grounded AI system to interact with long documents remains a challenge. Our work aims to fill the research gap of directly leveraging grounded visuals from documents alongside textual content in documents for response generation. We present an interactive conversational AI agent 'MuDoC' based on GPT-4o to generate document-grounded responses with interleaved text and figures. MuDoC's intelligent textbook interface promotes trustworthiness and enables verification of system responses by allowing instant navigation to source text and figures in the documents. We also discuss qualitative observations based on MuDoC responses highlighting its strengths and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。