用语音实时播报激光雷达检测结果,提升自动驾驶可读性。
LiDAR-based Object Detection with Real-time Voice Specifications
- 融合点云与图像的多模态网络实现三维目标检测
- 在3000样本上达87.0%准确率,较基线提升19.5个百分点
- 支持印度男声语音反馈,适合无障碍导航与人机交互
本文提出一种基于激光雷达的实时语音提示目标检测系统,通过多模态PointNet框架融合KITTI数据集中的3D点云与RGB图像。在3000样本子集上达到87.0%的验证准确率,相较200样本基线的67.5%显著提升,得益于空间与视觉信息融合、加权损失缓解类别不平衡,以及自适应训练优化。系统集成Tkinter原型,使用Edge TTS(en-IN-PrabhatNeural)生成自然印度男性语音输出,结合3D可视化与实时反馈,增强自主导航与辅助技术中的可访问性与安全性。研究提供完整方法、详尽实验分析及应用挑战综述,推动人机交互与环境感知的可扩展发展,契合当前研究趋势。
原文摘要 · Abstract (English)
This paper presents a LiDAR-based object detection system with real-time voice specifications, integrating KITTI's 3D point clouds and RGB images through a multi-modal PointNet framework. It achieves 87.0% validation accuracy on a 3000-sample subset, surpassing a 200-sample baseline of 67.5% by combining spatial and visual data, addressing class imbalance with weighted loss, and refining training via adaptive techniques. A Tkinter prototype provides natural Indian male voice output using Edge TTS (en-IN-PrabhatNeural), alongside 3D visualizations and real-time feedback, enhancing accessibility and safety in autonomous navigation, assistive technology, and beyond. The study offers a detailed methodology, comprehensive experimental analysis, and a broad review of applications and challenges, establishing this work as a scalable advancement in human-computer interaction and environmental perception, aligned with current research trends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。