针对高分辨率遥感图像,实现智能去冗余,兼顾语义与几何信息。
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
- 按任务动态分路压缩:语义流保小目标,几何流补空间结构。
- 在XLRS-Bench上计算量降低80%以上,精度不降反升。
- 无需训练,适配遥感语义与场景理解两类任务。
多模态大语言模型在地球观测中潜力巨大,但超高清(UHR)影像生成的海量视觉标记带来巨大计算开销,严重制约推理效率。现有压缩方法多采用静态统一策略,忽视遥感任务中“语义-几何双重性”:语义任务关注对象抽象含义,适合大幅剔除背景;而场景几何任务依赖空间拓扑完整性。为此,我们提出DualComp——一种任务自适应的双流压缩框架。由轻量级预训练路由器动态引导,将特征处理拆分为两条专用路径。在对象语义流中,空间连续语义聚合器(SCSA)采用大小自适应聚类,保留小目标的同时聚合冗余背景;在场景几何流中,指令引导结构恢复器(IGSR)引入贪心路径追踪拓扑重建机制,恢复空间骨架。在超高清遥感基准测试XLRS-Bench上的实验表明,DualComp以极低计算成本实现高保真遥感理解,在效率与精度上均取得显著提升。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated immense potential in Earth observation. However, the massive visual tokens generated when processing Ultra-High-Resolution (UHR) imagery introduce prohibitive computational overhead, severely bottlenecking their inference efficiency. Existing visual token compression methods predominantly adopt static and uniform compression strategies, neglecting the inherent "Semantic-Geometric Duality" in remote sensing interpretation tasks. Specifically, object semantic tasks focus on the abstract semantics of objects and benefit from aggressive background pruning, whereas scene geometric tasks critically rely on the integrity of spatial topology. To address this challenge, we propose DualComp, a task-adaptive dual-stream token compression framework. Dynamically guided by a lightweight pre-trained router, DualComp decouples feature processing into two dedicated pathways. In the object semantic stream, the Spatially-Contiguous Semantic Aggregator (SCSA) utilizes size-adaptive clustering to aggregates redundant background while protecting small object. In the scene geometric stream, the Instruction-Guided Structure Recoverer (IGSR) introduces a greedy path-tracing topology completion mechanism to reconstruct spatial skeletons. Experiments on the UHR remote sensing benchmark XLRS-Bench demonstrate that DualComp accomplishes high-fidelity remote sensing interpretation at an exceptionally low computational cost, achieving simultaneous improvements in both efficiency and accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。