用虚拟机场生成假数据,让行李车检测少花35%标注成本
Evaluating Synthetic Data for Baggage Trolley Detection in Airport Logistics
- 基于阿尔及尔机场数字孪生构建高保真合成数据
- 混合训练仅用40%真实数据就达94% mAP@50
- 适合需要降本提效的机场智能监控系统
高效行李车管理对减少机场拥堵、保障资产可用性至关重要。自动检测系统面临两大挑战:一是安全与隐私法规限制大规模数据采集;二是现有公开数据集缺乏多样性、规模和标注质量,难以应对实际运营中密集重叠的车组形态。为此,我们基于高保真数字孪生(NVIDIA Omniverse)构建合成数据生成流程,产出带方向边界框的丰富标注数据,可捕捉紧密嵌套等复杂车组结构。评估YOLO-OBB五种训练策略:纯真实、纯合成、线性探测、全微调、混合训练。结果表明,混合训练仅使用40%真实数据即可达到或超过全真实数据基线,mAP@50达0.94,mAP@50-95达0.77,同时降低25%至35%标注工作量。多种子实验显示mAP@50标准差低于0.01,验证了合成数据在自动化行李车检测中的实用有效性。
原文摘要 · Abstract (English)
Efficient luggage trolley management is critical for reducing congestion and ensuring asset availability in modern airports. Automated detection systems face two main challenges. First, strict security and privacy regulations limit large-scale data collection. Second, existing public datasets lack the diversity, scale, and annotation quality needed to handle dense, overlapping trolley arrangements typical of real-world operations. To address these limitations, we introduce a synthetic data generation pipeline based on a high-fidelity Digital Twin of Algiers International Airport using NVIDIA Omniverse. The pipeline produces richly annotated data with oriented bounding boxes, capturing complex trolley formations, including tightly nested chains. We evaluate YOLO-OBB using five training strategies: real-only, synthetic-only, linear probing, full fine-tuning, and mixed training. This allows us to assess how synthetic data can complement limited real-world annotations. Our results show that mixed training with synthetic data and only 40 percent of real annotations matches or exceeds the full real-data baseline, achieving 0.94 mAP@50 and 0.77 mAP@50-95, while reducing annotation effort by 25 to 35 percent. Multi-seed experiments confirm strong reproducibility with a standard deviation below 0.01 on mAP@50, demonstrating the practical effectiveness of synthetic data for automated trolley detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。