用带记忆的AI助手,让外骨骼更懂工人动作意图。
Interpretable Locomotion Prediction in Construction Using a Memory-Driven LLM Agent With Chain-of-Thought Reasoning
- 用大模型+记忆系统分析语音和视觉数据预测动作
- 加入长期记忆后准确率提升至0.90,误判减少
- 适合动态施工场景中需高安全性的智能辅助
建筑任务具有高度不确定性,环境动态且对安全要求严苛,给工人带来显著风险。外骨骼可提供助力,但若无法准确识别多样化的动作意图则难以发挥作用。本文提出一种基于大语言模型(LLM)与记忆系统的运动模式预测代理,旨在提升外骨骼在该类场景中的辅助能力。该代理融合多模态输入——来自智能眼镜的视觉数据与语音指令,包含感知模块、短期记忆(STM)、长期记忆(LTM)及优化模块,以有效预测运动模式。评估显示,无记忆时基准加权F1得分为0.73,加入STM后升至0.81,同时使用STM与LTM时达到0.90,尤其在模糊或高安全要求指令下表现优异。校准指标方面,布里尔得分从0.244降至0.090,期望校准误差(ECE)从0.222降至0.044,表明预测可靠性显著提高。该框架支持更安全、更高层次的人机协同,为动态产业中的自适应辅助系统提供了前景。
原文摘要 · Abstract (English)
Construction tasks are inherently unpredictable, with dynamic environments and safety-critical demands posing significant risks to workers. Exoskeletons offer potential assistance but falter without accurate intent recognition across diverse locomotion modes. This paper presents a locomotion prediction agent leveraging Large Language Models (LLMs) augmented with memory systems, aimed at improving exoskeleton assistance in such settings. Using multimodal inputs - spoken commands and visual data from smart glasses - the agent integrates a Perception Module, Short-Term Memory (STM), Long-Term Memory (LTM), and Refinement Module to predict locomotion modes effectively. Evaluation reveals a baseline weighted F1-score of 0.73 without memory, rising to 0.81 with STM, and reaching 0.90 with both STM and LTM, excelling with vague and safety-critical commands. Calibration metrics, including a Brier Score drop from 0.244 to 0.090 and ECE from 0.222 to 0.044, affirm improved reliability. This framework supports safer, high-level human-exoskeleton collaboration, with promise for adaptive assistive systems in dynamic industries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。