让AI实时模仿人类肢体语言,发现时序连贯性比画面清晰度更重要
Non-verbal Real-time Human-AI Interaction in Constrained Robotic Environments
- 基于2D人体关键点构建实时人机非语言交互框架
- 在NVIDIA Orin Nano上达100FPS,预训练可显著降低运动误差
- 时序连贯性影响真实表现,适用于机器人环境中的自然交互
我们研究当代生成模型在全身动作背景下,生成数据与人类数据的统计保真度问题。核心问题是:生成模型是否超越表面模仿,真正参与无声但富有表现力的肢体语言对话?为此,我们提出首个从2D人体关键点实时生成自然人机非语言交互的框架。实验采用四种轻量级架构,在NVIDIA Orin Nano上最高运行于100 FPS,有效闭环感知-动作循环。基于437段人类视频训练,结果显示预训练于合成序列可显著减少运动误差,且不牺牲速度。然而,仍存在明显真实差距:当最佳模型在Sora和VEO生成的关键点上评估时,对Sora数据性能下降明显,而对VEO下降较轻,表明时序连贯性而非图像保真度主导真实表现。结果表明,人类与AI动作间仍存在可统计区分的差异。
原文摘要 · Abstract (English)
We study the ongoing debate regarding the statistical fidelity of AI-generated data compared to human-generated data in the context of non-verbal communication using full body motion. Concretely, we ask if contemporary generative models move beyond surface mimicry to participate in the silent, but expressive dialogue of body language. We tackle this question by introducing the first framework that generates a natural non-verbal interaction between Human and AI in real-time from 2D body keypoints. Our experiments utilize four lightweight architectures which run at up to 100 FPS on an NVIDIA Orin Nano, effectively closing the perception-action loop needed for natural Human-AI interaction. We trained on 437 human video clips and demonstrated that pretraining on synthetically-generated sequences reduces motion errors significantly, without sacrificing speed. Yet, a measurable reality gap persists. When the best model is evaluated on keypoints extracted from cutting-edge text-to-video systems, such as SORA and VEO, we observe that performance drops on SORA-generated clips. However, it degrades far less on VEO, suggesting that temporal coherence, not image fidelity, drives real-world performance. Our results demonstrate that statistically distinguishable differences persist between Human and AI motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。