arXiv:2608.08874cs.CV2026-08

首个面向航拍图像的语音查询指代分割研究,实现自然语言精准定位目标。

AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images

论文配图:AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images
图 1 · 摘自论文原文
  • 构建双路径网络,融合语音与视觉特征,保留边界与语义细节。
  • 在纯净测试集上达到62.09% mIoU,优于最强基线5.38个百分点。
  • 支持带噪声干扰的严苛场景,适合无人机导航等真实应用。

语音为密集遥感图像中任意目标的指定提供了自然、免手持的交互方式,但现有遥感指代分割基准仅支持文本表达。为此,我们引入 exttt{dataset},一个基于RISBench构建的语音查询基准,增加了带有口音和不同语音特征的语音数据,同时保留原始图像、掩码及数据划分。其严苛评估集结合了旋翼、风噪和混合干扰,以及三种信噪比水平。我们还提出 exttt{model},一种高效的双路网络,包含保持边界的视觉路径、保持词元的语音编码、核线性跨模态注意力机制和分辨率精修头。该设计在不生成稠密语音-视觉关联矩阵的前提下,在两个尺度上对视觉特征进行条件化,再利用高分辨率视觉特征恢复精细边界。在干净测试集上, exttt{model} 使用Swin-Base模型获得62.09%的平均交并比(mIoU)和68.22%的整体交并比(oIoU),分别优于最强的音频适配遥感基线5.38和2.08个百分点;在最严苛设置下仍保持54.09%的最优mIoU。据我们所知,这是首个针对遥感图像全句语音查询指代分割的基准与模型研究。代码将公开发布。

原文摘要 · Abstract (English)

Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote-sensing image segmentation benchmarks accept only written expressions. To bridge this gap, we introduce \dataset, a spoken-query benchmark derived from RISBench that adds accent- and voice-diverse speech while preserving the original image, mask, and data splits. Its hard evaluation sets combine rotor, wind, and mixed interference with three signal-to-noise levels. We also propose \model, an efficient bilateral network that combines a boundary-preserving visual path with token-preserving speech encoding, kernel linear cross-modal attention, and a resolution refinement head. The design conditions visual features at two scales without materializing a dense speech--visual affinity matrix, then restores fine boundaries using high-resolution visual features. On the clean test split, \model with Swin-Base achieves 62.09\% mean intersection over union (mIoU) and 68.22\% overall intersection over union (oIoU), outperforming the strongest audio-adapted remote-sensing baseline by 5.38 and 2.08 percentage points, respectively. It retains the best hard-set mIoU at 54.09\%. To the best of our knowledge, this is the first benchmark and model study of full-sentence spoken-query referring segmentation for remote-sensing imagery. The code will be made publicly available.

语音理解遥感分割跨模态无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。