arXiv:2509.16343cs.CVcs.AI2025-09中稿 · MORS 2026 Artifici…

用推理时扩展提升遥感视觉模型,不重训也能更强。

Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling

  • 用迭代思考-批评-行动循环让多个模型互验纠错
  • 在复杂问题上准确率提升40.67%,整体准确率达78.8%
  • 无需训练即可增强现有大模型,适合遥感领域应用

构建高风险领域如遥感的鲁棒视觉系统需要超越单次推理的更强视觉推理能力;然而,重新训练大型模型通常计算成本高且依赖大量数据。我们提出视觉推理代理(VRA),一种无需训练的代理式视觉推理框架,通过迭代的思考-批评-行动循环,协调现成的大规模视觉语言模型(LVLMs)与大规模推理模型(LRM),实现跨模型验证、自我批评和递归优化。在遥感基准数据集VRSBench VQA上,VRA持续优于多个独立的LVLM基线,在涵盖感知与推理任务的挑战性问题类型上最高提升40.67%。此外,集成三个LVLM与VRA后,整体准确率从52.8%提升至78.8%,证明了增加推理时算力的代理式推理的有效性。

原文摘要 · Abstract (English)

Building robust vision systems for high-stakes domains such as remote sensing requires stronger visual reasoning than what single-pass inference typically provides; yet, retraining large models is often computationally expensive and data intensive. We present Visual Reasoning Agent (VRA), a training-free agentic visual reasoning framework that orchestrates off-the-shelf large vision-language models (LVLMs) with a large reasoning model (LRM) through an iterative Think-Critique-Act loop for cross-model verification, self-critique, and recursive refinement. On the remote sensing benchmark VRSBench VQA dataset, VRA consistently outperforms multiple standalone LVLM baselines and achieves up to 40.67\% improvement on challenging question types spanning both perception and reasoning tasks. In addition, integrating three LVLMs with VRA improves the overall accuracy of the standalone LVLMs from 52.8% to 78.8%, demonstrating the effectiveness of agentic reasoning with increased inference-time compute.

视觉推理遥感代理框架推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。