用符号规则动态管理记忆,让智能体在信息不全时更聪明地记、查、删。
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
- 结合符号规则与神经网络,动态选择记忆策略
- 在512条记忆容量下表现最优,且决策过程可追踪
- 适合需要可解释记忆管理的复杂推理任务
部分可观测强化学习需决定何时保留、检索和遗忘信息。本文提出一种神经符号元策略,在保持执行过程符号化的同时,学习在每个决策点应用何种符号记忆启发式方法。实验基于RoomKG中的时序知识图谱记忆,其中隐藏状态和观测以资源描述框架(RDF)图表示,记忆通过时间戳标注的RDF三元组进行增强。模型融合记忆内容的知识图谱编码与问答、探索、遗忘三个价值头,生成兼具自适应性与可解释性的控制器。该方法通过基于RDF的表示、兼容注解的图语义及显式记忆状态上的图操作,实现直接的语义网基础支撑。在长期记忆容量为512的训练/测试房间分割上,具备限定条件感知的StarE-GNN配置在对比的符号、神经与神经符号系统中取得最佳保留性能,同时保持步级粒度的记忆管理决策可追溯性。
原文摘要 · Abstract (English)
Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execution symbolic. Our setting uses temporal knowledge-graph memory in RoomKG, where hidden state and observations are represented as Resource Description Framework (RDF) graphs and memory is augmented with temporal RDF triple annotations. The model combines knowledge-graph encoding of memory contents with value heads for question answering, exploration, and forgetting, yielding a controller that is both adaptive and inspectable. This gives the work a direct Semantic Web grounding through RDF-based representation, annotation-compatible graph semantics, and graph-based symbolic operations over explicit memory state. On train/test room splits at long-term memory capacity of 512, the qualifier-aware StarE-GNN configuration achieves the best held-out performance among the compared symbolic, neural, and neuro-symbolic systems while preserving step-level traceability of memory-management decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。