用语音指令精准采集高质量校准图像,提升机器人视觉系统标定效率。
Acquisition of high-quality images for camera calibration in robotics applications via speech prompts
- 通过语音指令触发拍摄,实现高精度时间对齐的图像采集。
- 实验表明该方法可成功校准复杂多相机系统,避免模糊与运动伪影。
- 适合需要快速、无接触标定的机器人应用场景,提升操作便捷性。
精确的相机内参与外参标定是依赖视觉输入的机器人应用的重要前提。尽管已有研究探索使用自然图像进行标定,但实际系统仍多依赖特定标定靶,如棋盘格或AprilTag网格。在获取不同视角的标定图像并检测特征描述符后,通常通过优化过程最小化几何重投影误差。为使优化收敛,输入图像需具备足够质量,尤其要求清晰度高;不应包含运动模糊或滚动快门伪影,这些可能源于标定板在拍摄过程中未保持静止。本文提出一种通过夹式麦克风录制语音指令控制的新型标定图像采集技术,相比遥控触发或事后从视频序列中过滤模糊帧更具鲁棒性和用户友好性。为此,我们采用先进的语音转文字模型,并利用其精确的逐词时间戳来实现触发词的精准时序对齐。实验结果表明,该方法显著提升了用户体验,实现快速高效采集,成功校准复杂多相机系统。
原文摘要 · Abstract (English)
Accurate intrinsic and extrinsic camera calibration can be an important prerequisite for robotic applications that rely on vision as input. While there is ongoing research on enabling camera calibration using natural images, many systems in practice still rely on using designated calibration targets with e.g. checkerboard patterns or April tag grids. Once calibration images from different perspectives have been acquired and feature descriptors detected, those are typically used in an optimization process to minimize the geometric reprojection error. For this optimization to converge, input images need to be of sufficient quality and particularly sharpness; they should neither contain motion blur nor rolling-shutter artifacts that can arise when the calibration board was not static during image capture. In this work, we present a novel calibration image acquisition technique controlled via voice commands recorded with a clip-on microphone, that can be more robust and user-friendly than e.g. triggering capture with a remote control, or filtering out blurry frames from a video sequence in postprocessing. To achieve this, we use a state-of-the-art speech-to-text transcription model with accurate per-word timestamping to capture trigger words with precise temporal alignment. Our experiments show that the proposed method improves user experience by being fast and efficient, allowing us to successfully calibrate complex multi-camera setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。