arXiv:2409.01556cs.CLcs.AI2024-09中稿 · O-COCOSDA 2024被引 8

用客家文化评测大模型认知能力,发现检索增强能提效但难助创意。

Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture

  • 按布鲁姆认知六维度设计多层评测框架,覆盖记忆到创造。
  • RAG显著提升各认知层级准确率,尤其在事实检索与应用上。
  • 适合关注文化多样性、AI伦理与知识系统优化的研究者。

本研究提出一个综合性基准,用于评估大语言模型(LLMs)对文化知识的理解与处理能力,以客家文化为案例。基于布鲁姆分类法,构建涵盖记忆、理解、应用、分析、评价和创造六个认知维度的多维评估框架。该基准突破传统单一维度评价,深入分析模型在处理特定文化内容时的能力,从基础事实回忆到高阶认知任务如创造性综合。研究引入检索增强生成(RAG)技术,应对少数族裔文化知识在模型中表示不足的问题,实验证明RAG通过动态整合外部信息显著提升模型性能,尤其在需精确检索与应用的文化任务中表现突出。然而,结果也显示RAG在创造性任务中仍存局限,凸显进一步优化需求。该基准为评估与比较跨文化语境下的大模型提供了可靠工具,为人工智能驱动的文化知识保存与传播研究提供重要启示。

原文摘要 · Abstract (English)

This study introduces a comprehensive benchmark designed to evaluate the performance of large language models (LLMs) in understanding and processing cultural knowledge, with a specific focus on Hakka culture as a case study. Leveraging Bloom's Taxonomy, the study develops a multi-dimensional framework that systematically assesses LLMs across six cognitive domains: Remembering, Understanding, Applying, Analyzing, Evaluating, and Creating. This benchmark extends beyond traditional single-dimensional evaluations by providing a deeper analysis of LLMs' abilities to handle culturally specific content, ranging from basic recall of facts to higher-order cognitive tasks such as creative synthesis. Additionally, the study integrates Retrieval-Augmented Generation (RAG) technology to address the challenges of minority cultural knowledge representation in LLMs, demonstrating how RAG enhances the models' performance by dynamically incorporating relevant external information. The results highlight the effectiveness of RAG in improving accuracy across all cognitive domains, particularly in tasks requiring precise retrieval and application of cultural knowledge. However, the findings also reveal the limitations of RAG in creative tasks, underscoring the need for further optimization. This benchmark provides a robust tool for evaluating and comparing LLMs in culturally diverse contexts, offering valuable insights for future research and development in AI-driven cultural knowledge preservation and dissemination.

文化认知RAG评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。