用非凸方法优化机器人视觉导航的特征选择,兼顾精度与实时性。
Non-submodular Visual Attention for Robot Navigation
- 设计基于MSE的非子模目标函数,动态选择关键视觉特征
- 提出四种近似算法,计算复杂度从多项式到近常数时间
- 在真实平台验证,适合资源受限的机器人实时导航
本文提出一种面向任务的计算框架,用于提升机器人在时间与能源受限条件下的视觉惯性导航(VIN)性能。通过基于均方误差(MSE)的非子模目标函数和简化动态预测模型,智能选择视觉特征。针对该问题的NP难性质,设计四种多项式时间近似算法:具有常数保证的经典贪心法;显著降低复杂度的低秩贪心变体;平衡效率与解质量的随机贪心采样器;以及基于一阶泰勒展开的线性化选择器,可实现近常数时间运行。通过子模率、曲率及逐元素曲率分析,建立严格的性能边界。在标准基准与自研控制感知平台上的大量实验验证了理论结果,表明所提方法在保证强近似性能的同时支持实时部署。
原文摘要 · Abstract (English)
This paper presents a task-oriented computational framework to enhance Visual-Inertial Navigation (VIN) in robots, addressing challenges such as limited time and energy resources. The framework strategically selects visual features using a Mean Squared Error (MSE)-based, non-submodular objective function and a simplified dynamic anticipation model. To address the NP-hardness of this problem, we introduce four polynomial-time approximation algorithms: a classic greedy method with constant-factor guarantees; a low-rank greedy variant that significantly reduces computational complexity; a randomized greedy sampler that balances efficiency and solution quality; and a linearization-based selector based on a first-order Taylor expansion for near-constant-time execution. We establish rigorous performance bounds by leveraging submodularity ratios, curvature, and element-wise curvature analyses. Extensive experiments on both standardized benchmarks and a custom control-aware platform validate our theoretical results, demonstrating that these methods achieve strong approximation guarantees while enabling real-time deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。