arXiv:2511.03656cs.CLcs.AI2025-11中稿 · ICANN 2025

构建首个面向多领域中文长文档问答的细粒度评测数据集。

ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation

  • 针对教育、金融、医疗等六大领域构建长文档QA数据集。
  • 包含6068个高质量问答对,覆盖10类细粒度问题类型。
  • 适合研究中文文档理解与智能问答系统的技术团队使用。

随着自然语言处理技术的快速发展,高质量中文文档问答数据集的需求持续增长。为此,我们提出中文多文档问答数据集(ChiMDQA),专为学术、教育、金融、法律、医疗和新闻等主流领域设计。该数据集涵盖六个不同领域的长文档,包含6068个经过严格筛选的高质量问答对,并细分为十类细粒度类别。通过严谨的文档筛选与系统化的问题设计方法,确保数据多样性和高质量,适用于文档理解、知识抽取及智能问答系统等多种NLP任务。本文还全面介绍了数据集的设计目标、构建方法与细粒度评估体系,为中文问答领域的未来研究与实际应用提供坚实基础。代码与数据已公开:https://anonymous.4open.science/r/Foxit-CHiMDQA/

原文摘要 · Abstract (English)

With the rapid advancement of natural language processing (NLP) technologies, the demand for high-quality Chinese document question-answering datasets is steadily growing. To address this issue, we present the Chinese Multi-Document Question Answering Dataset(ChiMDQA), specifically designed for downstream business scenarios across prevalent domains including academic, education, finance, law, medical treatment, and news. ChiMDQA encompasses long-form documents from six distinct fields, consisting of 6,068 rigorously curated, high-quality question-answer (QA) pairs further classified into ten fine-grained categories. Through meticulous document screening and a systematic question-design methodology, the dataset guarantees both diversity and high quality, rendering it applicable to various NLP tasks such as document comprehension, knowledge extraction, and intelligent QA systems. Additionally, this paper offers a comprehensive overview of the dataset's design objectives, construction methodologies, and fine-grained evaluation system, supplying a substantial foundation for future research and practical applications in Chinese QA. The code and data are available at: https://anonymous.4open.science/r/Foxit-CHiMDQA/.

中文问答多文档细粒度评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。