arXiv:2504.19460cs.HCcs.AI2025-04被引 1

舞者动起来,音乐随之改变,实时互动更沉浸。

A Real-Time Gesture-Based Control Framework

  • 通过视频分析人体动作,实时转换为音乐控制信号。
  • 仅需50至80个样本即可实现用户无关的姿势识别。
  • 适合现场表演、互动装置与个人创作,体验自然流畅。

我们提出一种实时、人机协同的手势控制框架,通过分析实时视频输入,动态调整音频与音乐以响应人体运动。该系统建立视觉与听觉刺激之间的响应式连接,使舞者与表演者不仅能跟随音乐,还能通过动作影响音乐。适用于现场演出、互动装置和个人使用,带来沉浸式体验。框架融合计算机视觉与机器学习技术,追踪并解析运动行为,支持对节拍、音高、效果和播放顺序等音频元素的操控。经持续训练后可实现用户无关功能,仅需50至80个样本即可标注简单手势。系统结合手势训练、触发映射与音频操控,构建出动态交互体验。手势作为输入信号,被映射为声音控制指令,自然调节音乐要素,展现人机协作的无缝融合。

原文摘要 · Abstract (English)

We introduce a real-time, human-in-the-loop gesture control framework that can dynamically adapt audio and music based on human movement by analyzing live video input. By creating a responsive connection between visual and auditory stimuli, this system enables dancers and performers to not only respond to music but also influence it through their movements. Designed for live performances, interactive installations, and personal use, it offers an immersive experience where users can shape the music in real time. The framework integrates computer vision and machine learning techniques to track and interpret motion, allowing users to manipulate audio elements such as tempo, pitch, effects, and playback sequence. With ongoing training, it achieves user-independent functionality, requiring as few as 50 to 80 samples to label simple gestures. This framework combines gesture training, cue mapping, and audio manipulation to create a dynamic, interactive experience. Gestures are interpreted as input signals, mapped to sound control commands, and used to naturally adjust music elements, showcasing the seamless interplay between human interaction and machine response.

手势控制实时交互音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。