用大模型构建跨文化常识知识图谱,让隐性文化知识可被结构化使用。
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
- 将大模型视为文化档案,通过迭代提示提取文化特有实体与关系
- 在5个国家验证,英文链表现最优,体现当前模型文化编码不均衡
- 结构化文化知识链能提升文化推理与故事生成,尤其适合英文场景
大语言模型(LLMs)从海量网络数据中学习了丰富的文化常识,为大规模建模文化常识提供了前所未有的机遇。然而这些知识大多隐含且无结构,限制了其可解释性与应用。本文提出一种基于提示的迭代框架,将LLMs视为文化档案,系统提取特定文化中的实体、关系与实践,并将其组织为跨语言的多步推理链条,构建文化常识知识图谱(CCKG)。我们在五个国家进行了人工评估,涵盖文化相关性、正确性及路径连贯性。结果显示,即使目标文化非英语(如中文、印尼语、阿拉伯语),英文链条的表现仍更优,表明当前模型的文化编码存在不平衡。将小型模型与CCKG结合后,在文化推理与故事生成任务中性能提升,其中英文链条带来的增益最大。结果表明,尽管大模型作为文化技术潜力巨大,但其局限性明显;而链式结构的文化知识是实现文化嵌入自然语言处理的重要基础。
原文摘要 · Abstract (English)
Large language models (LLMs) encode rich cultural knowledge learned from diverse web-scale data, offering an unprecedented opportunity to model cultural commonsense at scale. Yet this knowledge remains mostly implicit and unstructured, limiting its interpretability and use. We present an iterative, prompt-based framework for constructing a Cultural Commonsense Knowledge Graph (CCKG) that treats LLMs as cultural archives, systematically eliciting culture-specific entities, relations, and practices and composing them into multi-step inferential chains across languages. We evaluate CCKG on five countries with human judgments of cultural relevance, correctness, and path coherence. We find that the cultural knowledge graphs are better realized in English, even when the target culture is non-English (e.g., Chinese, Indonesian, Arabic), indicating uneven cultural encoding in current LLMs. Augmenting smaller LLMs with CCKG improves performance on cultural reasoning and story generation, with the largest gains from English chains. Our results show both the promise and limits of LLMs as cultural technologies and that chain-structured cultural knowledge is a practical substrate for culturally grounded NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。