让机器人根据任务和视觉观察自动配置扫描参数,提升检测精度。
Task-Aware Scanning Parameter Configuration for Robotic Inspection Using Vision Language Embeddings and Hyperdimensional Computing

- 用语言指令和视觉图像生成任务感知的扫描参数组合
- 在5个参数上准确率超92%,比传统方法更稳定快速
- 适合工业质检场景,无需人工调参,可实时部署
机器人激光剖面扫描广泛用于尺寸验证与表面检测,但测量精度常受传感器配置影响远大于机械运动。工业级剖面仪包含采样频率、测量范围、曝光时间、接收动态范围和照明等多个耦合参数,仍依赖试错法调节;配置不当会导致饱和、截断或漏检,无法后期修复。本文提出指令条件化的传感参数推荐:给定预扫描的RGB图像和自然语言检测指令,推断机器人安装剖面仪的关键参数离散配置。为评估该问题,构建了真实世界多模态数据集Instruct-Obs2Param,涵盖16个物体在多视角姿态与光照变化下的检测意图与标准参数配置。随后提出ScanHD框架,基于高维计算将指令与观测编码为任务感知代码,通过紧凑记忆实现参数级关联推理,在匹配离散扫描模式的同时提供稳定、可解释、低延迟决策。在Instruct-Obs2Param上,ScanHD在5个参数上平均精确率达92.7%,平均Win@1准确率达98.1%,具备强跨划分泛化能力,推理延迟低,优于规则启发式方法、传统多模态模型及多模态大语言模型。本工作实现了从任务意图与场景上下文出发的自主指令感知传感配置,消除了人工调参需求,使传感器配置从静态设置变为可适应的决策变量。
原文摘要 · Abstract (English)
Robotic laser profiling is widely used for dimensional verification and surface inspection, yet measurement fidelity is often dominated by sensor configuration rather than robot motion. Industrial profilers expose multiple coupled parameters, including sampling frequency, measurement range, exposure time, receiver dynamic range, and illumination, that are still tuned by trial-and-error; mismatches can cause saturation, clipping, or missing returns that cannot be recovered downstream. We formulate instruction-conditioned sensing parameter recommendation; given a pre-scan RGB observation and a natural-language inspection instruction, infer a discrete configuration over key parameters of a robot-mounted profiler. To benchmark this problem, we develop Instruct-Obs2Param, a real-world multimodal dataset linking inspection intents and multi-view pose and illumination variation across 16 objects to canonical parameter regimes. We then propose ScanHD, a hyperdimensional computing framework that binds instruction and observation into a task-aware code and performs parameter-wise associative reasoning with compact memories, matching discrete scanner regimes while yielding stable, interpretable, low-latency decisions. On Instruct-Obs2Param, ScanHD achieves 92.7% average exact accuracy and 98.1% average Win@1 accuracy across the five parameters, with strong cross-split generalization and low-latency inference suitable for deployment, outperforming rule-based heuristics, conventional multimodal models, and multimodal large language models. This work enables autonomous, instruction-conditioned sensing configuration from task intent and scene context, eliminating manual tuning and elevating sensor configuration from a static setting to an adaptive decision variable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。