arXiv:2411.19913cs.CVcs.AI2024-11中稿 · publication in the…被引 3

量化航拍图像中真实与合成数据的差异,提升模型泛化能力。

Quantifying the synthetic and real domain gap in aerial scene understanding

  • 用多模型共识和深度结构指标评估场景复杂度
  • 真实数据一致性更高,合成数据更难适应模型
  • 适合做航拍场景仿真优化与域自适应的研究者

量化真实与合成图像之间的差距对提升基于Transformer的模型及数据集至关重要,尤其在航拍场景理解这一潜力巨大的未充分探索领域。本文提出一种新方法,利用多模型共识度量(MMCM)和基于深度的结构指标,实现对感知与结构差异的稳健评估。实验基于真实数据集Dronescapes和合成数据集Skyscenes,结果表明:真实场景下先进视觉变压器的一致性更高,而合成场景则表现出更大变异性,对模型适应性构成更大挑战。研究揭示了内在复杂性与域间差距,强调需提高仿真保真度与模型泛化能力。该工作为航拍场景理解中的域特性与模型表现关系提供了关键洞见,指明了改进域自适应策略的路径。

原文摘要 · Abstract (English)

Quantifying the gap between synthetic and real-world imagery is essential for improving both transformer-based models - that rely on large volumes of data - and datasets, especially in underexplored domains like aerial scene understanding where the potential impact is significant. This paper introduces a novel methodology for scene complexity assessment using Multi-Model Consensus Metric (MMCM) and depth-based structural metrics, enabling a robust evaluation of perceptual and structural disparities between domains. Our experimental analysis, utilizing real-world (Dronescapes) and synthetic (Skyscenes) datasets, demonstrates that real-world scenes generally exhibit higher consensus among state-of-the-art vision transformers, while synthetic scenes show greater variability and challenge model adaptability. The results underline the inherent complexities and domain gaps, emphasizing the need for enhanced simulation fidelity and model generalization. This work provides critical insights into the interplay between domain characteristics and model performance, offering a pathway for improved domain adaptation strategies in aerial scene understanding.

航拍理解域差距合成数据视觉变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。