arXiv:2605.24488cs.CVcs.GR2026-05SIGGRAPH

用舞蹈分析法识别虚拟世界中的暗示性动作,不依赖外观信息。

Appearance-Invariant Detection of Suggestive Motion via Laban Movement Descriptors

  • 基于拉班动作分析提取61维运动特征,仅用骨骼轨迹分类
  • 在17小时数据上达到68%准确率,媲美复杂视频模型
  • 特征可解释,适合实时虚拟环境内容审核

在线多人3D虚拟环境中内容审核日益自动化,但现有方法集中于图像、视频和音频,忽视了暗示性动作。本文提出仅基于运动的分类流程,利用拉班动作分析(LMA)描述符从SMPL骨骼轨迹中检测暗示性与明确性动作。在涵盖日常、艺术、暗示及明确动作的17+小时数据集上,使用61维LMA特征的逻辑回归模型在无泄露评估下达到68%的二元SFW/NSFW准确率(随机森林为70%)。该性能与一个在同一运动重渲染为无外观视频(无衣物、皮肤、场景)的深度学习模型相当。各关节轨迹的间接性(迂回度,即路径长度与净位移之比)在暗示性层级达到峰值,表明拉班空间因素中的直接-间接极性可作为功能动作用向暗示性动作转变的可解释标记。最终,基于拉班的运动描述符提供了一种轻量、可解释的暗示性动作检测方法:每个决策均可分解为命名且理论支持的特征。由于分类器仅处理姿态轨迹,审核可直接在虚拟环境的化身姿态上运行,无需任何外观数据。

原文摘要 · Abstract (English)

Content moderation in online multiplayer 3D virtual environments is increasingly automated, yet detection has focused on images, video, and audio, leaving suggestive motion a blind spot. We present a motion-only classification pipeline that detects suggestive and explicit movement from SMPL skeleton trajectories using Laban Movement Analysis (LMA) descriptors. On a dataset spanning everyday, artistic, suggestive, and explicit movement (17+ hours of video), a logistic regression trained on 61-feature LMA descriptors reaches 68% binary SFW/NSFW accuracy (70% random forest) under a leak-free evaluation protocol. At this level, our descriptor performs comparably to a learned video model trained on the same motion re-rendered as appearance-free video, a gray figure with no clothing, skin, or scene. The indirectness (tortuosity) of each joint's trajectory, measured as the ratio of the joint's path length to its net displacement, peaks at the suggestive tier, showing that the Direct-to-Indirect polarity of Laban's Space factor provides an interpretable marker of the shift from functional to suggestive motion. Ultimately, Laban-based kinematic descriptors offer a lightweight, interpretable approach to suggestive-motion detection: every decision decomposes into named, theory-grounded features. Because the classifier operates on pose trajectories alone, moderation can run directly on avatar poses in virtual environments, with no appearance data.

动作识别内容审核可解释性虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。