构建中文多跳推理基准,评估大模型的常识推理能力
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
- 基于事实链生成多跳问题,覆盖中文特定知识
- 大模型在长尾知识推理上仍存明显短板
- 检索增强生成可显著提升性能,适合研究者参考
尽管大语言模型展现出强大的推理能力,但其在通用中文语境下的全面评估仍不足。为此,我们提出中文常识多跳推理(CCMOR)基准,用于评估大模型整合中文特有事实知识与多步逻辑推理的能力。首先从现有问答数据集构建领域均衡的种子集,再通过大模型驱动的流水线生成基于事实单元链的多跳问题。为确保数据质量,采用人机协同验证系统,由领域专家系统性地审核和优化生成的问题。利用CCMOR评估前沿大模型,结果表明其在处理长尾知识和执行知识密集型推理方面仍存在明显局限。值得注意的是,检索增强生成显著缓解了这些知识缺口,带来性能大幅提升。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have demonstrated advanced reasoning capabilities, their comprehensive evaluation in general Chinese-language contexts remains understudied. To bridge this gap, we propose Chinese Commonsense Multi-hop Reasoning (CCMOR), a novel benchmark designed to evaluate LLMs' ability to integrate Chinese-specific factual knowledge with multi-step logical reasoning. Specifically, we first construct a domain-balanced seed set from existing QA datasets, then develop an LLM-powered pipeline to generate multi-hop questions anchored on factual unit chains. To ensure the quality of resulting dataset, we implement a human-in-the-loop verification system, where domain experts systematically validate and refine the generated questions. Using CCMOR, we evaluate state-of-the-art LLMs, demonstrating persistent limitations in LLMs' ability to process long-tail knowledge and execute knowledge-intensive reasoning. Notably, retrieval-augmented generation substantially mitigates these knowledge gaps, yielding significant performance gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。