构建超700万对中英知识图谱-文本数据,专用于评估中文大模型的可靠推理能力。
A Large-Scale Chinese Knowledge Graph-Text Alignment Dataset for Benchmarking Knowledge-Grounded LLMs
- 通过多阶段流程构建高质量中英文对齐数据,确保事实可靠性。
- 包含1500万条知识图谱三元组,覆盖四大领域,支持三大任务评测。
- 针对中文歧义、分词模糊等特性设计,适合中文知识推理研究者使用。
可靠评估中文知识增强型大语言模型需要明确对齐中文文本与可验证知识图谱事实的资源。现有中文基准主要评估通用语言理解,对中文特有语言现象下的结构化推理支持有限。本文提出中文数据-文本对(CDTP),一个涵盖四个宽泛领域的超大规模中文知识图谱-文本对齐数据集,包含超过700万对实例,共1500万条知识图谱三元组。通过多阶段构建流程——包括对齐过滤、人工验证和外部证据验证——提升语义一致性与事实可靠性。该数据集支持知识图谱补全(KGC)、问答(QA)和三元组到文本生成(T2T)三种任务。所有任务均考虑中文特有现象,如多义性、分词歧义和上下文依赖实体理解,可评估结构化推理、歧义感知的事实理解与知识引导生成能力。在多种开源与私有大模型上的实验表明,模型规模本身无法保证在这些中文知识密集型任务上的可靠表现;而基于CDTP的监督微调则持续提升域内性能与跨分布鲁棒性。该数据集、代码及评估协议已公开,为中文知识增强大模型的研发与评估提供可复用资源。
原文摘要 · Abstract (English)
Reliable evaluation of knowledge-grounded Large Language Models (LLMs) in Chinese requires resources that explicitly align Chinese-language text with verifiable Knowledge Graph (KG) facts. Yet existing Chinese benchmarks primarily assess general language understanding and offer limited support for structured reasoning under Chinese-specific linguistic phenomena. We introduce the Chinese Data-Text Pair (CDTP), a large-scale Chinese KG-text alignment dataset comprising more than 7 million aligned instances across four broad domains. Each instance pairs a Chinese-language text with one or more textually supported KG triples, totaling 15 million triples. A multi-stage construction pipeline combining alignment filtering, manual verification, and external evidence validation improves semantic consistency and factual reliability. CDTP supports Knowledge Graph Completion (KGC), Question Answering (QA), and Triple-to-Text Generation (T2T). Across all three tasks, the benchmark design accounts for Chinese-specific phenomena, including polysemy, word-segmentation ambiguity, and context-dependent entity interpretation, enabling the evaluation of structured reasoning, ambiguity-aware factual understanding, and knowledge-grounded generation. Experiments with diverse open-source and proprietary LLMs show that model scale alone does not guarantee reliable performance on these Chinese knowledge-intensive tasks, whereas supervised fine-tuning on CDTP consistently improves in-domain performance and out-of-distribution robustness. The publicly accessible dataset, code, and evaluation protocols provide a reusable resource for developing and evaluating knowledge-grounded LLMs in Chinese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。