用符号化图增强语言模型,让AI理解真实世界的深层含义
Neurosymbolic Graph Enrichment for Grounded World Models
- 将图像转为自然语言,再构建带逻辑模式的语义图
- 通过回流语言模型激活隐含知识,提升推理能力
- 适合需要深度上下文理解的AI系统开发
构建能够理解并推理复杂现实场景的人工智能系统仍是重大挑战。本文提出一种新方法,通过增强大型语言模型(LLM)的反应能力来应对复杂问题并解析深层次的现实语境意义。该方法结合最先进的大型语言模型与结构化语义表示,实现多模态、知识增强的形式化意义表征。流程从图像输入开始,利用先进大模型生成自然语言描述,再将其转化为抽象语义表示(AMR)图,并融合逻辑设计模式及来自语言学和事实知识库的分层语义。最终生成的图被送回语言模型,以激发由复杂启发式学习激活的隐含知识,包括语义蕴含、道德价值、具身认知和隐喻表达。该方法弥合了非结构化语言模型与形式化语义结构之间的鸿沟,为自然语言理解与推理中的复杂问题开辟新路径。
原文摘要 · Abstract (English)
The development of artificial intelligence systems capable of understanding and reasoning about complex real-world scenarios is a significant challenge. In this work we present a novel approach to enhance and exploit LLM reactive capability to address complex problems and interpret deeply contextual real-world meaning. We introduce a method and a tool for creating a multimodal, knowledge-augmented formal representation of meaning that combines the strengths of large language models with structured semantic representations. Our method begins with an image input, utilizing state-of-the-art large language models to generate a natural language description. This description is then transformed into an Abstract Meaning Representation (AMR) graph, which is formalized and enriched with logical design patterns, and layered semantics derived from linguistic and factual knowledge bases. The resulting graph is then fed back into the LLM to be extended with implicit knowledge activated by complex heuristic learning, including semantic implicatures, moral values, embodied cognition, and metaphorical representations. By bridging the gap between unstructured language models and formal semantic structures, our method opens new avenues for tackling intricate problems in natural language understanding and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。