首个面向无人机影像的多模态数学推理基准,检验视觉语言模型在真实场景下的算术与逻辑能力。
Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration
- 构建包含3773题的无人机影像数学推理数据集,覆盖6类学科20个主题。
- 14个主流视觉语言模型在该基准上表现不佳,平均准确率不足40%。
- 链式思考提示和微调可提升性能,为未来研究提供方向。
数学推理对无人机遥感中的精确距离与面积计算、轨迹估计和空间分析至关重要,但现有视觉语言模型(VLMs)在此领域尚未得到充分评估。为此,我们提出AVI-Math,首个系统评估无人机影像中多模态数学推理的基准,超越简单计数任务,涵盖几何、逻辑、代数等领域的专业知识。数据集包含3,773道高质量车辆相关问题,源于不同高度与视角的无人机视图,真实反映实际应用场景,确保问题的多样性与复杂性。本文对14个主流VLM进行了全面评测,结果表明:尽管这些模型在以往多模态基准上表现良好,但在AVI-Math任务中仍面临显著挑战。详细分析揭示当前VLM在数学推理方面存在明显局限,并指明未来研究方向。此外,我们探索了链式思考提示(Chain-of-Thought prompting)与微调技术,显示出改善推理能力的潜力。研究不仅暴露了VLM在数学推理上的不足,也为实现可信的无人机应用视觉语言模型提供了关键洞见。代码与数据集将公开于https://github.com/VisionXLab/avi-math。
原文摘要 · Abstract (English)
Mathematical reasoning is critical for tasks such as precise distance and area computations, trajectory estimations, and spatial analysis in unmanned aerial vehicle (UAV) based remote sensing, yet current vision-language models (VLMs) have not been adequately tested in this domain. To address this gap, we introduce AVI-Math, the first benchmark to rigorously evaluate multimodal mathematical reasoning in aerial vehicle imagery, moving beyond simple counting tasks to include domain-specific knowledge in areas such as geometry, logic, and algebra. The dataset comprises 3,773 high-quality vehicle-related questions captured from UAV views, covering 6 mathematical subjects and 20 topics. The data, collected at varying altitudes and from multiple UAV angles, reflects real-world UAV scenarios, ensuring the diversity and complexity of the constructed mathematical problems. In this paper, we benchmark 14 prominent VLMs through a comprehensive evaluation and demonstrate that, despite their success on previous multimodal benchmarks, these models struggle with the reasoning tasks in AVI-Math. Our detailed analysis highlights significant limitations in the mathematical reasoning capabilities of current VLMs and suggests avenues for future research. Furthermore, we explore the use of Chain-of-Thought prompting and fine-tuning techniques, which show promise in addressing the reasoning challenges in AVI-Math. Our findings not only expose the limitations of VLMs in mathematical reasoning but also offer valuable insights for advancing UAV-based trustworthy VLMs in real-world applications. The code, and datasets will be released at https://github.com/VisionXLab/avi-math
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。