用认知框架测评大模型对客家文化的理解能力
Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation
- 结合布鲁姆认知分类与检索增强生成,构建六层评估体系
- 在台湾客家数字档案上测试,准确率和文化相关性为关键指标
- 适合研究多文化AI、数字人文与认知评测的学者参考
本研究提出一种认知基准评估框架,用于检验大语言模型(LLMs)对特定文化知识的处理与应用能力。该框架融合布鲁姆认知分类法与检索增强生成(RAG),从六个层级的认知域——记忆、理解、应用、分析、评价和创造——评估模型表现。以精心整理的台湾客家数字文化档案为主要测试平台,评估模型生成内容在语义准确性和文化相关性方面的表现。
原文摘要 · Abstract (English)
This study proposes a cognitive benchmarking framework to evaluate how large language models (LLMs) process and apply culturally specific knowledge. The framework integrates Bloom's Taxonomy with Retrieval-Augmented Generation (RAG) to assess model performance across six hierarchical cognitive domains: Remembering, Understanding, Applying, Analyzing, Evaluating, and Creating. Using a curated Taiwanese Hakka digital cultural archive as the primary testbed, the evaluation measures LLM-generated responses' semantic accuracy and cultural relevance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。