让视频智能体自动修正记忆错误,避免混淆实体。
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

- 用跨模态绑定锁定实体身份,防止压缩丢失关键信息。
- 双向记忆修正提升历史记录一致性,使推理更准确。
- 多智能体交叉验证缺失证据时主动放弃回答,减少幻觉。
在长时序环境中运行的多模态智能体需构建并持续更新多媒体记忆,以支持一致的实体识别和时间定位推理。然而,现有方法在高压压缩与分段处理下常丢失细粒度实体线索,且依赖向量相似性检索,易召回语义相关但身份不符的证据,导致实体混淆、错误传播和幻觉回答。我们提出 ViSAGE,一种构建自修正、以实体为中心的记忆框架。具体而言,ViSAGE 通过跨模态绑定在长时程中锚定实体身份,并采用双向记忆精炼机制,回溯性地传播延迟的身份证据,统一历史记录并提升未来推理能力。同时引入多智能体交叉验证,在身份-证据对齐约束下评估检索结果,当证据缺失时选择拒绝回答而非给出无依据结论。大量实验表明,ViSAGE 持续优于最强基线,准确率提升 5.9%。
原文摘要 · Abstract (English)
Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive compression and segment-wise processing. They also rely heavily on vector similarity retrieval, which can surface semantically related yet identity-mismatched evidence, leading to entity confusion, error propagation, and hallucinated answers. We propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories. Specifically, ViSAGE anchors entity identity via cross-modal binding over long temporal ranges. It then applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning. We also introduce multi-agent cross-verification to assess retrieved evidence under an identity-evidence alignment onstraint, enabling abstention instead of unsupported answers when evidence is missing. Extensive results demonstrate that ViSAGE consistently outperforms the strongest baseline, achieving 5.9% higher accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。