arXiv:2509.17707cs.CV2025-09综述被引 1

用计算机视觉自动识别集装箱等运输单元,提升港口效率

Automatic Intermodal Loading Unit Identification using Computer Vision: A Scoping Review

  • 梳理63篇论文,总结视觉识别方法演进路径
  • 端到端识别准确率从5%到96%,动态摄像头成新趋势
  • 适合物流、港口智能化研究者参考,推动技术落地

背景:集装箱、半挂车和换装体等标准运输单元(ILUs)的标准化推动了全球贸易发展,但其高效可靠的识别仍是港口与码头运营中的瓶颈。目标:系统梳理基于计算机视觉(CV)的ILU识别方法,厘清术语定义,总结技术演进脉络,揭示研究空白与未来方向对场站作业的影响。方法:遵循PRISMA-ScR流程,在Google Scholar与dblp中检索英文文献,经双人筛选后,对研究方法、数据集与评估指标进行归纳分析。结果:共纳入63项实证研究,时间跨度为1990至2025年。发现识别技术由静态(如OCR门禁)向车载摄像头系统转变,实现更精准监控;报告的端到端识别准确率范围为5%至96%。结论:建议统一术语,倡导开放数据集、代码库与模型权重以促进公平评估,并提出未来应聚焦车载摄像头带来的新挑战、探索合成数据生成、将多阶段方法整合为统一端到端模型,以及攻克无上下文文本识别难题。

原文摘要 · Abstract (English)

Background: The standardisation of Intermodal Loading Units (ILUs), including containers, semi-trailers, and swap bodies, has transformed global trade, yet efficient and robust identification remains an operational bottleneck in ports and terminals. Objective: To map Computer Vision (CV) methods for ILU identification, clarify terminology, summarise the evolution of proposed approaches, and highlight research gaps, future directions and their potential effects on terminal operations. Methods: Following PRISMA-ScR, we searched Google Scholar and dblp for English-language studies with quantitative results. After dual reviewer screening, the studies were charted across methods, datasets, and evaluation metrics. Results: 63 empirical studies on CV-based solutions for the ILU identification task, published between 1990 and 2025 were reviewed. Methodological evolution of ILU identification solutions, datasets, evaluation of the proposed methods and future research directions are summarised. A shift from static (e.g. OCR-gates) to vehicle mounted camera setups, which enables precise monitoring is observed. The reported results for end-to-end accuracy range from 5% to 96%. Conclusions: We propose standardised terminology, advocate for open-access datasets, codebases and model weights to enable fair evaluation and define future work directions. The shift from static to dynamic camera settings introduces new challenges that have transformative potential for transportation and logistics. However, the lack of public benchmark datasets, open-access code, and standardised terminology hinders the advancements in this field. As for the future work, we suggest addressing the new challenges emerged from vehicle mounted cameras, exploring synthetic data generation, refining the multi-stage methods into unified end-to-end models to reduce complexity, and focusing on contextless text recognition.

计算机视觉智能港口运输单元识别动态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。