用认知复杂度框架重新评估大模型知识图谱任务
Characterizing Knowledge Graph Tasks in LLM Benchmarks Using Cognitive Complexity Frameworks
- 引入认知心理学复杂度框架分析任务难度
- 发现现有评测中价值分布不均且需求覆盖不足
- 适合研究评测基准设计与模型认知能力的学者
大型语言模型(LLMs)越来越多地用于涉及知识图谱(KGs)的任务,其评估通常聚焦于准确性和输出正确性。本文提出一种互补的任务表征方法,基于认知心理学中的三个复杂度框架。将其应用于LLM-KG-Bench框架后,揭示了价值分布特征,识别出被低估的任务需求,并推动评测任务在解释性与多样性方面的提升。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used for tasks involving Knowledge Graphs (KGs), whose evaluation typically focuses on accuracy and output correctness. We propose a complementary task characterization approach using three complexity frameworks from cognitive psychology. Applying this to the LLM-KG-Bench framework, we highlight value distributions, identify underrepresented demands and motivate richer interpretation and diversity for benchmark evaluation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。