用事件序列建模音乐游戏关卡,精准捕捉节奏与动作时序关系。
Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling

- 将关卡生成转为音符与节拍偏移交替的符号序列建模
- 在事件级评估中优于传统帧式基线模型
- 适合研究音乐如何引导节奏对齐的动作设计
程序化生成音乐游戏关卡是一个充满挑战的任务,需将音乐结构转化为时间对齐的游戏交互事件序列。现有方法多采用帧级表示,将音频划分为均匀时间网格并在每帧预测事件,导致游戏事件在多个帧中隐式分布,难以刻画事件间的时间关系及人类创作关卡中的长程结构。本文以程序化生成为场景,研究音乐线索如何映射到交互事件序列。受事件驱动符号音乐建模启发,提出一种基于标记的序列建模方法,将关卡生成视为多模态序列到序列问题:给定音频片段与关卡元数据,模型生成交替的玩法事件与节拍偏移标记序列,显式表达动作及其在节拍空间中的相对时间。基于此框架构建Transformer模型,在事件级评估中优于代表性帧级基线模型,并支持系统分析音频如何在元数据之外增强节奏对齐事件预测能力。
原文摘要 · Abstract (English)
Procedural generation of music game levels is an exciting yet challenging problem, as levels must translate musical structure into interactive sequences of timed gameplay events. Most existing approaches formulate this task by frame-based representations, dividing audio into uniform time grids and predicting events at each frame. This makes gameplay events implicit across many frames. As a result, it is hard to describe event-level timing relations and longer-range structure found in human-authored levels. We use procedural generation as a practical setting to study how musical cues map to interactive event sequences. Inspired by event-based symbolic music modeling, we propose a token-level sequence formulation that casts level generation as a multimodal sequence-to-sequence problem. Conditioned on an audio excerpt and level metadata, the model generates a token sequence alternating gameplay-event and beat-shift tokens. This explicitly represents actions and their relative timing in beat space. Based on this formulation, we build a Transformer model. It outperforms representative frame-level baselines under event-level evaluation. It also enables systematic analysis of how audio supports rhythm-aligned event prediction beyond metadata conditioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。