arXiv:2607.26072cs.IRcs.AI2026-07

测试大模型在BIM领域跨会话记忆能力,发现现有系统表现不佳。

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

  • 构建多会话BIM任务基准,模拟真实工程场景中的记忆需求
  • 最强系统仅32.4%答对率,暴露通用记忆系统的局限性
  • 适合研究专业领域智能体与结构化知识记忆的学者

长时记忆正成为基于大模型代理的核心能力,但现有评估主要聚焦开放域或人格驱动场景下的对话回忆。我们提出更强的测试标准:代理能否在实时、结构化、领域特定环境中复用先前会话的信息。研究聚焦建筑信息建模(BIM)这一专业工程流程,要求代理查询大型IFC模型的同时依赖项目规范、客户决策和工程惯例——这些信息常通过对话讨论,却未体现在模型中。为此,我们引入IFCMemoryBench,一个用于评估大模型在BIM信息检索中长时记忆能力的基准。该基准包含19个项目的143个跨会话任务和4,016次历史会话,源自IFC-Bench v2中的不完整信息问题。每个任务在早期对话中植入缺失项目上下文,并在后续以探测问题形式考察是否能结合记忆与实时IFC查询作答。评估框架将记忆性能分解为摄入、检索与利用三个阶段,采用专家验证的LLM评判答案质量与记忆质量。我们评测了代表性向量、图谱与文件型记忆系统。最强系统在部署现实的摄入范围内仅达32.4%的答案准确率,即使在理想过滤摄入或更强探测代理条件下也低于60%。分析显示,当前通用记忆系统常检索主题相关上下文,但存储项目知识为不完整或碎片化的事实。结果揭示了代理记忆在领域迁移中的差距,表明可靠的专业代理需具备关联对话、项目知识与结构化模型实体的领域感知记忆表征。

原文摘要 · Abstract (English)

Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or persona-grounded settings. We argue that a stronger test is whether an agent can reuse information from prior sessions while acting over a live, structured, domain-specific environment. We study this problem in Building Information Modelling (BIM), a professional engineering workflow where agents must query large IFC models while also relying on project specifications, client decisions, and engineering conventions often discussed in conversation but absent from the model. We introduce IFCMemoryBench, a benchmark for evaluating long-term memory in LLM-based BIM information retrieval. IFCMemoryBench contains 143 multi-session tasks across 19 projects and 4,016 prior sessions, derived from incomplete-information questions in IFC-Bench v2. Each task seeds missing project context across earlier conversations and later asks a probe question that can be answered only by combining remembered context with live IFC queries. Our evaluation framework decomposes memory performance into ingestion, retrieval, and utilization, and measures both answer quality and memory quality with expert-validated LLM judges. We evaluate representative vector-, graph-, and file-based memory systems. The strongest system achieves only 32.4% answer accuracy under a deployment-realistic ingestion scope, and remains below 60% under oracle-filtered ingestion or a stronger probe agent. Analysis shows that current general-purpose memory systems often retrieve topically relevant context but store project knowledge as incomplete or fragmented facts. These results reveal a domain-transfer gap in agent memory and suggest that reliable professional agents require domain-aware memory representations linking conversations, project knowledge, and structured model entities.

大模型记忆BIM信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。