让视觉语言模型学会自我总结物理知识,提升动态场景推理能力
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model

- 用自生成的‘知识笔记’外部化物理认知,构建可迭代的知识库
- 在PhysBench上达到56.68%准确率,比最优基线高4.96个百分点
- 适合需要长期物理推理与因果理解的智能系统研究者
视觉语言模型(VLM)在课本式物理题上表现优异,但在需要时序一致性和跨帧因果推理的动态真实场景中常失败。我们识别出两大核心问题:(1) 空间-时间身份漂移,即物体在连续帧中失去物理身份,破坏因果链;(2) 推理时洞察的不稳定性,模型偶有正确推理却无法固化复用。为此,我们提出PhysNote,一种代理框架,使VLM通过自生成的‘知识笔记’外化并精炼物理知识。PhysNote通过时空归一化稳定动态感知,将自生成洞察组织为分层知识库,并驱动一个迭代推理循环,在验证前基于视觉证据锚定假设,最终巩固可靠知识。在PhysBench上的实验表明,PhysNote总体准确率达56.68%,优于最佳多代理基线4.96%,并在所有四个物理推理领域均取得一致提升。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have demonstrated strong performance on textbook-style physics problems, yet they frequently fail when confronted with dynamic real-world scenarios that require temporal consistency and causal reasoning across frames. We identify two fundamental challenges underlying these failures: (1) spatio-temporal identity drift, where objects lose their physical identity across successive frames and break causal chains, and (2) volatility of inference-time insights, where a model may occasionally produce correct physical reasoning but never consolidates it for future reuse. To address these challenges, we propose PhysNote, an agentic framework that enables VLMs to externalize and refine physical knowledge through self-generated "Knowledge Notes." PhysNote stabilizes dynamic perception through spatio-temporal canonicalization, organizes self-generated insights into a hierarchical knowledge repository, and drives an iterative reasoning loop that grounds hypotheses in visual evidence before consolidating verified knowledge. Experiments on PhysBench demonstrate that PhysNote achieves 56.68% overall accuracy, a 4.96% improvement over the best multi-agent baseline, with consistent gains across all four physical reasoning domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。