用结构化文本增强视觉语言模型,实现高精度室内动作识别
KRAST: Knowledge-Augmented Robotic Action Recognition with Structured Text for Vision-Language Models
- 将动作类别文本转为可学习提示注入冻结的视觉语言模型
- 在ETRI-Activity3D数据集上达95%以上准确率,优于现有方法
- 适合需要少标注、强泛化的机器人感知场景
准确的基于视觉的动作识别对开发能在复杂现实环境中安全可靠运行的自主机器人至关重要。本文通过引入领域特定知识,提升视觉语言模型(VLMs)在视频中识别室内日常动作的能力。我们采用提示学习框架,将每个动作类别的文本描述作为可学习提示嵌入到冻结的预训练VLM主干中,并设计和评估了多种文本结构化与编码策略。在ETRI-Activity3D数据集上的实验表明,该方法仅使用测试时的RGB视频输入,即可达到95%以上的准确率,显著优于当前最优方法。结果证明,知识增强提示能以最小监督实现鲁棒的动作识别。
原文摘要 · Abstract (English)
Accurate vision-based action recognition is crucial for developing autonomous robots that can operate safely and reliably in complex, real-world environments. In this work, we advance video-based recognition of indoor daily actions for robotic perception by leveraging vision-language models (VLMs) enriched with domain-specific knowledge. We adapt a prompt-learning framework in which class-level textual descriptions of each action are embedded as learnable prompts into a frozen pre-trained VLM backbone. Several strategies for structuring and encoding these textual descriptions are designed and evaluated. Experiments on the ETRI-Activity3D dataset demonstrate that our method, using only RGB video inputs at test time, achieves over 95\% accuracy and outperforms state-of-the-art approaches. These results highlight the effectiveness of knowledge-augmented prompts in enabling robust action recognition with minimal supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。