arXiv:2412.12990cs.CVcs.ET2024-12

探索新型传感器与生成技术,提升动作识别的效率与伦理合规性。

Future Aspects in Human Action Recognition: Exploring Emerging Techniques and Ethical Influences

  • 利用下一代传感器捕捉图像间动态信息,优化时序分析。
  • 通过强化学习生成带标签的合成视频,缓解数据不足问题。
  • 关注技术伦理,推动可信赖的人机交互研究。

基于视觉的人类动作识别广泛应用于监控系统、体育分析、医疗辅助技术及人机交互等领域,旨在识别和分类视频中个体的行为。由于动作发生在连续帧序列中,时序分析带来额外复杂性,现有方法在计算成本和适应性方面仍存挑战。下一代硬件传感器提供的含过渡信息的视觉数据,有望助力机器人领域解决该问题。然而,尽管存在大量静态图像数据集,用于训练人工智能模型,真实人类活动视频数据却受限于规模小、分布不均或来源控制不足。为此,可通过生成带标注的逼真合成视频,结合强化学习技术减少对大规模数据集的依赖。同时,人类因素的介入引发伦理关切,现有技术争议需引起研究者重视。

原文摘要 · Abstract (English)

Visual-based human action recognition can be found in various application fields, e.g., surveillance systems, sports analytics, medical assistive technologies, or human-robot interaction frameworks, and it concerns the identification and classification of individuals' activities within a video. Since actions typically occur over a sequence of consecutive images, it is particularly challenging due to the inclusion of temporal analysis, which introduces an extra layer of complexity. However, although multiple approaches try to handle temporal analysis, there are still difficulties because of their computational cost and lack of adaptability. Therefore, different types of vision data, containing transition information between consecutive images, provided by next-generation hardware sensors will guide the robotics community in tackling the problem of human action recognition. On the other hand, while there is a plethora of still-image datasets, that researchers can adopt to train new artificial intelligence models, videos representing human activities are of limited capabilities, e.g., small and unbalanced datasets or selected without control from multiple sources. To this end, generating new and realistic synthetic videos is possible since labeling is performed throughout the data creation process, while reinforcement learning techniques can permit the avoidance of considerable dataset dependence. At the same time, human factors' involvement raises ethical issues for the research community, as doubts and concerns about new technologies already exist.

动作识别生成模型伦理考量传感器融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。