arXiv:2510.21424cs.CLcs.AI2025-10被引 4

用视觉语言模型提升医疗场景下动态人体行为识别能力

Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings

  • 构建描述性字幕数据集,支持动态行为的文本化表达
  • 在医疗行为识别任务中,视觉语言模型性能媲美甚至超过传统深度学习模型
  • 为远程健康监测中的智能系统提供新范式,适合关注AI医疗落地的研究者

随着生成式AI的发展,视觉语言模型(VLMs)在医疗应用中展现出巨大潜力。然而,在远程健康监测中的人体活动识别(HAR)领域,其应用仍相对有限。VLMs具备更强的灵活性,可克服传统深度学习模型的部分局限性。但关键挑战在于如何评估其动态、非确定性的输出结果。为此,本文构建了一个描述性字幕数据集,并提出全面的评估方法来衡量VLM在HAR中的表现。与先进深度学习模型的对比实验表明,VLM在多数情况下达到相当甚至更高的准确率。该研究建立了有力的基准,为将VLM融入智能医疗系统开辟了新路径。

原文摘要 · Abstract (English)

As generative AI continues to evolve, Vision Language Models (VLMs) have emerged as promising tools in various healthcare applications. One area that remains relatively underexplored is their use in human activity recognition (HAR) for remote health monitoring. VLMs offer notable strengths, including greater flexibility and the ability to overcome some of the constraints of traditional deep learning models. However, a key challenge in applying VLMs to HAR lies in the difficulty of evaluating their dynamic and often non-deterministic outputs. To address this gap, we introduce a descriptive caption data set and propose comprehensive evaluation methods to evaluate VLMs in HAR. Through comparative experiments with state-of-the-art deep learning models, our findings demonstrate that VLMs achieve comparable performance and, in some cases, even surpass conventional approaches in terms of accuracy. This work contributes a strong benchmark and opens new possibilities for the integration of VLMs into intelligent healthcare systems.

视觉语言模型行为识别医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。