arXiv:2512.19742cs.LG2025-12

用大模型让手机识别动作更准还懂解释。

On-device Large Multi-modal Agent for Human Activity Recognition

  • 融合大模型与多模态数据,实现动作识别与推理
  • 在多个数据集上达到顶尖准确率,且可解释性强
  • 适合需要交互式健康监测的场景

人类活动识别(HAR)在医疗与智能环境中有广泛应用。近期大型语言模型(LLMs)的发展为HAR带来了新可能,不仅支持活动分类,还可提供可解释性与类人交互。本文提出一种面向HAR的大规模多模态智能体框架,整合LLM能力以提升性能与用户体验。该框架不仅能完成活动分类,还通过推理与问答功能,弥合技术输出与用户友好洞察之间的差距。我们在多个常用数据集(包括HHAR、Shoaib、Motionsense)上进行了广泛评估,结果表明,该模型在分类准确率上接近当前最优水平,同时显著提升了可解释性与交互能力。

原文摘要 · Abstract (English)

Human Activity Recognition (HAR) has been an active area of research, with applications ranging from healthcare to smart environments. The recent advancements in Large Language Models (LLMs) have opened new possibilities to leverage their capabilities in HAR, enabling not just activity classification but also interpretability and human-like interaction. In this paper, we present a Large Multi-Modal Agent designed for HAR, which integrates the power of LLMs to enhance both performance and user engagement. The proposed framework not only delivers activity classification but also bridges the gap between technical outputs and user-friendly insights through its reasoning and question-answering capabilities. We conduct extensive evaluations using widely adopted HAR datasets, including HHAR, Shoaib, Motionsense to assess the performance of our framework. The results demonstrate that our model achieves high classification accuracy comparable to state-of-the-art methods while significantly improving interpretability through its reasoning and Q&A capabilities.

多模态动作识别大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。