用多智能体自动生成长文档问答数据,提升视觉语言模型理解能力
Multi-Agent Interactive Question Generation Framework for Long Document Understanding
- 设计多智能体交互框架,自动构建跨页长文档问答对
- 在英阿双语数百页文档上生成高质量问题,覆盖多个领域
- 适合需要长文本理解的视觉语言模型训练与评估
长文本场景下的文档理解在视觉-语言研究中仍具挑战性。尽管大型视觉-语言模型(LVLM)在短文本任务中表现优异,但在长文本环境下性能下降。主要瓶颈在于细粒度训练数据稀缺,尤其对阿拉伯语等低资源语言。现有技术严重依赖人工标注,成本高且效率低。本文提出一种完全自动化的多智能体交互框架,可高效生成英文和阿拉伯文长文档的单页与跨页问题。该方法覆盖多个领域的数百页文档,显著促进具备更强长文本理解能力的LVLM发展。实验表明,生成的英阿双语问题集(AraEngLongBench)对主流开源与闭源LVLM均构成显著挑战。代码与数据已公开于https://github.com/wangk0b/Multi_Agentic_QA_Long_Doc.git,附录提供样本问答对与结构化系统提示。
原文摘要 · Abstract (English)
Document Understanding (DU) in long-contextual scenarios with complex layouts remains a significant challenge in vision-language research. Although Large Vision-Language Models (LVLMs) excel at short-context DU tasks, their performance declines in long-context settings. A key limitation is the scarcity of fine-grained training data, particularly for low-resource languages such as Arabic. Existing state-of-the-art techniques rely heavily on human annotation, which is costly and inefficient. We propose a fully automated, multi-agent interactive framework to generate long-context questions efficiently. Our approach efficiently generates high-quality single- and multi-page questions for extensive English and Arabic documents, covering hundreds of pages across diverse domains. This facilitates the development of LVLMs with enhanced long-context understanding ability. Experimental results in this work have shown that our generated English and Arabic questions (\textbf{AraEngLongBench}) are quite challenging to major open- and close-source LVLMs. The code and data proposed in this work can be found in https://github.com/wangk0b/Multi_Agentic_QA_Long_Doc.git. Sample Question and Answer (QA) pairs and structured system prompts can be found in the Appendix.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。