arXiv:2503.05049cs.CLcs.IR2025-03被引 8

动态生成问答数据集,防止模型记忆,提升评估可靠性

Dynamic-KGQA: A Scalable Framework for Generating Adaptive Question Answering Datasets

  • 基于知识图谱动态生成新问答数据,每轮迭代保持分布一致
  • 支持领域定制,生成紧凑语义连贯子图,利于模型训练与评测
  • 提供静态数据集划分,兼容现有评估方法,适合模型性能对比

随着基础模型快速发展,高质量、可适应且大规模的问答评估基准日益重要。传统问答基准多为静态公开数据,易被大语言模型记忆和污染,导致模型泛化能力被高估,难以真实反映实际表现。本文提出 Dynamic-KGQA,一个从知识图谱生成自适应问答数据的可扩展框架,可在每次运行时生成新数据集变体,同时保持底层分布一致,实现公平可复现的评估。该框架支持细粒度控制数据特征,可生成领域特定、主题聚焦的问答数据。此外,动态生成的紧凑语义连贯子图有助于提升模型对结构化知识的利用效率。为适配现有评估流程,我们还提供静态的大规模训练/测试/验证划分,确保与以往方法的可比性。Dynamic-KGQA 引入了动态可定制的评测范式,使问答系统评估更严谨、更具适应性。

原文摘要 · Abstract (English)

As question answering (QA) systems advance alongside the rapid evolution of foundation models, the need for robust, adaptable, and large-scale evaluation benchmarks becomes increasingly critical. Traditional QA benchmarks are often static and publicly available, making them susceptible to data contamination and memorization by large language models (LLMs). Consequently, static benchmarks may overestimate model generalization and hinder a reliable assessment of real-world performance. In this work, we introduce Dynamic-KGQA, a scalable framework for generating adaptive QA datasets from knowledge graphs (KGs), designed to mitigate memorization risks while maintaining statistical consistency across iterations. Unlike fixed benchmarks, Dynamic-KGQA generates a new dataset variant on every run while preserving the underlying distribution, enabling fair and reproducible evaluations. Furthermore, our framework provides fine-grained control over dataset characteristics, supporting domain-specific and topic-focused QA dataset generation. Additionally, Dynamic-KGQA produces compact, semantically coherent subgraphs that facilitate both training and evaluation of KGQA models, enhancing their ability to leverage structured knowledge effectively. To align with existing evaluation protocols, we also provide static large-scale train/test/validation splits, ensuring comparability with prior methods. By introducing a dynamic, customizable benchmarking paradigm, Dynamic-KGQA enables a more rigorous and adaptable evaluation of QA systems.

问答系统知识图谱动态评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。