轻量级多模态推理框架提升机器人临床场景理解能力
Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
- 结合小模型与代理架构,实现思维链推理与视听融合
- 在视频理解任务中达到领先准确率,且推理更鲁棒
- 适合手术辅助、患者监护等需要结构化决策的医疗场景
医疗机器人需在动态临床环境中具备稳健的多模态感知与推理能力以保障安全。现有视觉语言模型(VLMs)虽具通用性强的优势,但在时间推理、不确定性估计及结构化输出方面仍不足,难以满足机器人规划需求。本文提出一种轻量级智能体式多模态框架,基于Qwen2.5-VL-3B-Instruct模型与SmolAgent协同层,支持思维链推理、语音-视觉融合及动态工具调用。框架生成结构化场景图,并采用混合检索模块实现可解释、自适应推理。在Video-MME基准和自建临床数据集上的评估显示,其准确率具有竞争力,且鲁棒性优于当前最优VLMs,展现出在机器人辅助手术、患者监测与决策支持中的应用潜力。
原文摘要 · Abstract (English)
Healthcare robotics requires robust multimodal perception and reasoning to ensure safety in dynamic clinical environments. Current Vision-Language Models (VLMs) demonstrate strong general-purpose capabilities but remain limited in temporal reasoning, uncertainty estimation, and structured outputs needed for robotic planning. We present a lightweight agentic multimodal framework for video-based scene understanding. Combining the Qwen2.5-VL-3B-Instruct model with a SmolAgent-based orchestration layer, it supports chain-of-thought reasoning, speech-vision fusion, and dynamic tool invocation. The framework generates structured scene graphs and leverages a hybrid retrieval module for interpretable and adaptive reasoning. Evaluations on the Video-MME benchmark and a custom clinical dataset show competitive accuracy and improved robustness compared to state-of-the-art VLMs, demonstrating its potential for applications in robot-assisted surgery, patient monitoring, and decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。