系统评估结构化知识提示的泛化能力,发现现有方法存在局限性。
Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking
- 从粒度、迁移性、可扩展性、通用性四方面评估结构化知识提示
- 构建包含9个任务的SUBARU多粒度多层级基准测试集
- 揭示当前SKP方法在泛化能力上的边界与不足,适合研究者参考
大语言模型在当前自然语言处理研究中展现出卓越的文本生成能力,但事实准确性缺失仍是其发展的主要瓶颈。结构化知识提示(SKP)通过引入结构化外部知识表示,显著提升了多个知识密集型任务的表现,成为主流范式。然而,现有方法多聚焦于特定问题,缺乏对SKP泛化能力与能力边界的全面探索。本文从粒度、迁移性、可扩展性、通用性四个维度系统评估SKP范式的泛化能力,并构建了一个新的多粒度、多层次基准SUBARU,涵盖9种不同粒度与难度的任务,以实现全面评估。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated exceptional performance in text generation within current NLP research. However, the lack of factual accuracy is still a dark cloud hanging over the LLM skyscraper. Structural knowledge prompting (SKP) is a prominent paradigm to integrate external knowledge into LLMs by incorporating structural representations, achieving state-of-the-art results in many knowledge-intensive tasks. However, existing methods often focus on specific problems, lacking a comprehensive exploration of the generalization and capability boundaries of SKP. This paper aims to evaluate and rethink the generalization capability of the SKP paradigm from four perspectives including Granularity, Transferability, Scalability, and Universality. To provide a thorough evaluation, we introduce a novel multi-granular, multi-level benchmark called SUBARU, consisting of 9 different tasks with varying levels of granularity and difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。