自进化对抗流程自动生成高难度阿拉伯语长文档问答
A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation
- 多模型协同闭环迭代,无需人工干预
- 在阿语长文档上显著提升问答质量与难度
- 支持可调难度,适合构建高质量评测数据
我们提出一个端到端、自进化的对抗式工作流,用于阿拉伯语长上下文问答生成。通过协调多个专用大视觉语言模型(LVLM):问题生成器、评估器和答案生成集群,系统在无任何人工干预的情况下持续优化自身性能。从跨领域的多页阿拉伯语文档出发,问题生成器生成细粒度、上下文相关的查询,由答案生成集群处理,评估器则反馈质量指标。这一闭环循环实现持续学习:低置信度输出会触发自动重生成与模型更新,逐步提升问题的难度与相关性。此外,我们将质量指标设为可调超参数,实现可控且可定制的难度级别。我们发布了AraLongBench,一个涵盖数百页的大型阿拉伯语基准数据集,包含单页与多页挑战。实验表明,该自进化工作流显著优于静态流水线,大幅增强主流阿拉伯语大视觉语言模型(LVLM)的长文本理解能力。最后,我们还设计了全自动的智能体式长文档收集流程。
原文摘要 · Abstract (English)
We present an end-to-end, self-evolving adversarial workflow for long-context Question-Answer (QA) Generation in Arabic. By orchestrating multiple specialized LVLMs: a question generator, an evaluator, and a swarm of answer generators, our system iteratively refines its own performance without any human intervention. Starting from raw, multi-page Arabic documents across diverse domains, the question generator produces fine-grained, context-aware queries to be tackled by the answer generator swarm, and the evaluator assesses and feeds back quality metrics. This closed-loop cycle enables continuous learning: low-confidence outputs trigger automated re-generation and model updates, progressively enhancing question difficulty and relevance. Moreover, we set the quality metrics as a tunable hyperparameter, enabling question generation at controllable and customizable difficulty levels. We release AraLongBench, a large-scale Arabic benchmark of single- and multi-page challenges spanning hundreds of pages, and demonstrate that our self-evolving workflow substantially outperform static pipelines, markedly boosting the long-context comprehension capabilities of leading Arabic Large Vision Language Models (LVLMs). Lastly, we also meticulously architect a fully automated agentic workflow for long-context Arabic document collection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。