用小模型生成物理场景描述,提升大模型推理能力
Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language Models
- 小模型专精物理场景描述,作为大模型的上下文输入
- 在复杂任务上提升平均准确率13.8%,实测跨模拟与真实场景迁移有效
- 适合需要物理推理的视觉语言模型研究者和开发者
视觉语言模型在物理推理方面仍面临挑战,主要源于难以将学习到的知识转化为对物理行为的预测。尽管持续微调可缓解此问题,但对大模型成本过高,且难以针对每个任务重复执行。为此,我们提出物理上下文构建器(Physics Context Builders, PCBs),一种模块化框架:通过微调小型专用视觉语言模型,生成详细的物理场景描述,用于增强大型模型的推理能力。PCBs实现了感知与推理的分离,便于分析二者对物理理解的贡献。我们在CLEVRER和Falling Tower(一个包含模拟与真实场景的稳定性检测数据集)上进行实验,结果表明,PCBs在复杂物理推理任务中显著提升性能,平均准确率最高提升13.8%。值得注意的是,PCBs在模拟到真实场景的迁移(Sim2Real)中表现良好,成功从仿真数据推广至真实世界。
原文摘要 · Abstract (English)
Physical reasoning remains a significant challenge for Vision-Language Models (VLMs). This limitation arises from an inability to translate learned knowledge into predictions about physical behavior. Although continual fine-tuning can mitigate this issue, it is expensive for large models and impractical to perform repeatedly for every task. This necessitates the creation of modular and scalable ways to teach VLMs about physical reasoning. To that end, we introduce Physics Context Builders (PCBs), a modular framework where specialized smaller VLMs are fine-tuned to generate detailed physical scene descriptions. These can be used as physical contexts to enhance the reasoning capabilities of larger VLMs. PCBs enable the separation of visual perception from reasoning, allowing us to analyze their relative contributions to physical understanding. We perform experiments on CLEVRER and on Falling Tower, a stability detection dataset with both simulated and real-world scenes, to demonstrate that PCBs provide substantial performance improvements, increasing average accuracy by up to 13.8% on complex physical reasoning tasks. Notably, PCBs also show strong Sim2Real transfer, successfully generalizing from simulated training data to real-world scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。