arXiv:2605.19624cs.CVcs.AI2026-05被引 1

让卫星仿真图更像真实影像,同时保留精确标注。

Component-Aware Structure-Preserving Style Transfer for Satellite Visual Sim2Real Data Construction

论文配图:Component-Aware Structure-Preserving Style Transfer for Satellite Visual Sim2Real Data Construction
图 1 · 摘自论文原文
  • 按部件分离风格,用真实图像提取样式并注入合成图对应区域。
  • 生成图像与真实域差距小,FID 54.32,KID 0.048,优于现有方法。
  • 适合需要高精度姿态估计的卫星视觉仿真转实测数据场景。

针对基于相机的卫星视觉感知,模拟到现实的数据构建需在保留仿真几何标注的同时逼近真实传感器外观。真实卫星目标的可靠姿态标签与部件级掩码难以大规模获取,而合成渲染虽提供精确几何信息,但存在显著外观差距。本文提出一种面向部件的结构保持式风格迁移框架,用于卫星视觉的仿真到现实数据构建。通过校准的真实采集、ArUco标定的相机位姿测量、CAD渲染及部件掩码,构建弱配对的实-仿样本。从无标注真实图像中提取部件级真实域风格码,并通过掩码对齐调制注入合成卫星区域。为确保生成图像适用于下游传感器监督任务,结合对抗训练、局部对比一致性、自正则化与边缘保持约束。实验在5,000张渲染卫星图像和100张校准环境下拍摄的真实图像上进行。真实图像作为目标域外观参考与最终评估依据,下游GDRNet姿态估计算法仅在合成或转换后的合成图像上训练。相比代表性图像翻译基线,所提方法实现最低图像分布差异(FID 54.32,KID 0.048)。使用转换数据训练时,ADD通过率提升至0.260,AUC达0.611。结果表明,在该校准设置下,部件级外观迁移可有效提升标注保留的卫星视觉模拟转现实数据生成质量。

原文摘要 · Abstract (English)

For camera-based satellite visual sensing, Sim2Real data construction requires images that approach real-domain sensor appearance while retaining the annotations inherited from simulation. Real sensor images of satellite targets with reliable pose labels and component-level masks are difficult to acquire at scale, whereas synthetic rendering provides exact geometric annotations but suffers from a visible appearance gap. This paper presents a component-aware structure-preserving style transfer framework for satellite visual synthetic-to-real data construction. The method builds weakly paired real--synthetic samples from calibrated real acquisition, ArUco-based camera-pose measurement, CAD rendering, and component masks. It then extracts part-wise real-domain style codes from unlabeled real images and injects them into corresponding synthetic satellite regions through mask-aligned modulation. To keep the generated images usable for downstream sensor-data supervision, adversarial training is combined with local contrastive consistency, self-regularization, and edge-preserving constraints. Experiments are conducted on 5,000 rendered satellite images and 100 real images captured in a calibrated setup. The real images provide target-domain appearance references and final evaluation images, while the downstream GDRNet pose estimator is trained only on synthetic or translated synthetic images. Compared with representative image-translation baselines, the proposed method achieves the lowest image distribution discrepancy, with an FID of 54.32 and a KID of 0.048. When the translated data are used to train GDRNet in this target-domain adaptation setting, the ADD pass rate improves to 0.260 and the AUC improves to 0.611. These results indicate that component-level appearance transfer can improve annotation-preserving satellite visual Sim2Real data generation in the considered calibrated setup.

卫星视觉风格迁移仿真转现实姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。