arXiv:2511.20085cs.AIcs.MA2025-11被引 4

让大模型像人一样一步步看图推理,还更省资源。

VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis

  • 用视觉工具穿插在思考链中,实现多轮图文交互推理。
  • 在多个遥感数据集上,推理透明度和生成质量均超当前最优方法。
  • 可将复杂推理能力迁移到小模型,适合部署在资源受限场景。

当前遥感图像分析正从传统目标识别转向复杂智能推理,对模型的推理能力与工具调用灵活性提出更高要求。为此,我们提出一种新型多模态智能体框架——视觉交错思维链(VICoT),通过动态将视觉工具融入思维链,实现显式的多轮推理。基于栈式推理结构与模块化MCP兼容工具套件,VICoT使大语言模型能高效完成多轮、交错的视觉-语言推理任务,具备强泛化性与灵活性。我们还提出推理栈蒸馏方法,将复杂智能体行为迁移至小型轻量模型,在显著降低复杂度的同时保持推理能力。在多个遥感基准测试中,VICoT在推理透明度、执行效率与生成质量方面均显著优于现有最先进框架。

原文摘要 · Abstract (English)

The current remote sensing image analysis task is increasingly evolving from traditional object recognition to complex intelligence reasoning, which places higher requirements on the model's reasoning ability and the flexibility of tool invocation. To this end, we propose a new multimodal agent framework, Vision-Interleaved Chain-of-Thought Framework (VICoT), which implements explicit multi-round reasoning by dynamically incorporating visual tools into the chain of thought. Through a stack-based reasoning structure and a modular MCP-compatible tool suite, VICoT enables LLMs to efficiently perform multi-round, interleaved vision-language reasoning tasks with strong generalization and flexibility.We also propose the Reasoning Stack distillation method to migrate complex Agent behaviors to small, lightweight models, which ensures the reasoning capability while significantly reducing complexity. Experiments on multiple remote sensing benchmarks demonstrate that VICoT significantly outperforms existing SOTA frameworks in reasoning transparency, execution efficiency, and generation quality.

多模态推理遥感分析智能体框架思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。