arXiv:2604.12772cs.CVcs.MA2026-04

用多智能体系统自动发现卫星图像中的新闻事件并生成描述。

A Multi-Agent Feedback System for Detecting and Describing News Events in Satellite Imagery

论文配图:A Multi-Agent Feedback System for Detecting and Describing News Events in Satellite Imagery
图 1 · 摘自论文原文
  • 设计多智能体迭代流程,结合新闻文章地理编码与图像序列合成。
  • 相比传统方法发现5倍更多事件,构建了包含5000个序列的新数据集。
  • 适合遥感监测、新闻自动化和灾害响应等场景使用。

卫星影像的变化通常跨越多个时间步。尽管已有双时相变化描述数据集,但遥感领域仍缺乏至少包含两幅图像的多时相事件描述数据集,原因在于(1)在卫星影像中搜索可见事件,以及(2)标注多时相序列均需大量时间和人力。为解决这一问题,我们提出SkyScraper——一种迭代式多智能体工作流,通过地理编码新闻文章并为对应的卫星图像序列生成描述。实验表明,SkyScraper比传统地理编码方法成功发现5倍更多的事件,证明智能体反馈是有效挖掘多时相事件的策略。我们将该框架应用于全球新闻文章数据库,构建了一个包含5000个序列的新型多时相描述数据集。通过自动识别与新闻事件相关的影像,本研究亦可支持新闻报道与公共信息传播。

原文摘要 · Abstract (English)

Changes in satellite imagery often occur over multiple time steps. Despite the emergence of bi-temporal change captioning datasets, there is a lack of multi-temporal event captioning datasets (at least two images per sequence) in remote sensing. This gap exists because (1) searching for visible events in satellite imagery and (2) labeling multi-temporal sequences require significant time and labor. To address these challenges, we present SkyScraper, an iterative multi-agent workflow that geocodes news articles and synthesizes captions for corresponding satellite image sequences. Our experiments show that SkyScraper successfully finds 5x more events than traditional geocoding methods, demonstrating that agentic feedback is an effective strategy for surfacing new multi-temporal events in satellite imagery. We apply our framework to a large database of global news articles, curating a new multi-temporal captioning dataset with 5,000 sequences. By automatically identifying imagery related to news events, our work also supports journalism and reporting efforts.

遥感监测多智能体新闻生成图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。