arXiv:2501.14510cs.CVcs.GR2025-01被引 9

用合成数据训练的深度模型,单图就能预测相机参数。

Deep-BrownConrady: Prediction of Camera Calibration and Distortion Parameters Using Deep Learning and Synthetic Data

  • 用AI模拟生成带参数变化的图像数据,训练网络预测镜头参数。
  • 在真实图像上测试,预测误差小于1.5像素,接近传统方法精度。
  • 适合自动驾驶、机器人等需快速标定的场景使用。

本研究针对仅凭单张图像进行相机标定与畸变参数预测的难题,提出一种基于深度学习的方法。核心贡献包括:(1) 验证了在真实与合成图像混合数据上训练的深度学习模型,可准确从单图中预测相机与镜头参数;(2) 基于AILiveSim仿真平台构建了全面的合成数据集,涵盖焦距与镜头畸变参数的变化,为模型训练与测试提供坚实基础。训练主要依赖合成数据,辅以少量真实图像,评估合成数据训练模型在真实图像上的泛化能力。传统标定方法需多角度校准图,但公开数据集常缺乏此类图像。采用基于ResNet架构的回归网络,依据Brown-Conrady模型预测相机参数,适用于自动驾驶、机器人与增强现实等需高精度标定的应用场景。

原文摘要 · Abstract (English)

This research addresses the challenge of camera calibration and distortion parameter prediction from a single image using deep learning models. The main contributions of this work are: (1) demonstrating that a deep learning model, trained on a mix of real and synthetic images, can accurately predict camera and lens parameters from a single image, and (2) developing a comprehensive synthetic dataset using the AILiveSim simulation platform. This dataset includes variations in focal length and lens distortion parameters, providing a robust foundation for model training and testing. The training process predominantly relied on these synthetic images, complemented by a small subset of real images, to explore how well models trained on synthetic data can perform calibration tasks on real-world images. Traditional calibration methods require multiple images of a calibration object from various orientations, which is often not feasible due to the lack of such images in publicly available datasets. A deep learning network based on the ResNet architecture was trained on this synthetic dataset to predict camera calibration parameters following the Brown-Conrady lens model. The ResNet architecture, adapted for regression tasks, is capable of predicting continuous values essential for accurate camera calibration in applications such as autonomous driving, robotics, and augmented reality. Keywords: Camera calibration, distortion, synthetic data, deep learning, residual networks (ResNet), AILiveSim, horizontal field-of-view, principal point, Brown-Conrady Model.

相机标定合成数据深度学习ResNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。