arXiv:2608.10360cs.HCcs.AI2026-08

用自然语言实时控制阿拉伯马卡姆音乐生成,精准保留微音程特色。

MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

论文配图:MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model
图 1 · 摘自论文原文
  • 将实时演奏转为文本提示,动态驱动生成模型
  • 子秒级延迟下提升微音程使用率,显著优于基线
  • 适合跨文化音乐共创与自适应音乐教育场景

阿拉伯马卡姆音乐以微音程、调式和装饰性对答为特征,是当前生成音乐模型最缺乏支持的传统之一,因训练体系仍以西方平均律为主。实时伴奏更凸显此差距:AI需实时聆听、动态调整并尊重微音程结构。流式文本到音乐模型虽具强大生成能力,却缺乏精确控制接口。本文提出 MazzikaAI,一个基于知识的系统,利用自然语言作为实时控制环的执行器。通过将实时 MIDI、手势和推断和声编译为持续更新的文本提示,该系统在不微调模型的前提下,操控未修改的流式生成器 Google Lyria RealTime。系统嵌入六种核心马卡姆的知识、典型装饰音及合奏动态,实现实时响应,关键到可听延迟低于1秒。实证评估显示,动态提示编译能可靠地将生成内容锚定于微音程体系,显著提高离格四分音含量。除核心实现外,MazzikaAI 展示了确定性知识规则如何有效连接专家性的非西方音乐传统与未微调的基础模型。该架构为实时人机共创建立可扩展范式,为交互式伴奏、自适应音乐教育及跨文化生成音频提供通用蓝图。

原文摘要 · Abstract (English)

Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative music models, whose training frameworks remain predominantly Western and equaltempered. Real time accompaniment sharpens this gap: an AI partner must listen, adapt dynamically, and respect idiomatic microtonal structures. Streaming text to music models provide strong generative capabilities but lack precise control interfaces. We present MazzikaAI, a knowledge based system that uses natural language as the actuator of a realtime control loop. By compiling live MIDI, gesture, and inferred harmony into continuously updated text prompts, MazzikaAI steers an unmodified streaming generator, Google Lyria RealTime, without requiring model finetuning. The system embeds expert knowledge of six core maqamat, characteristic ornaments, and ensemble dynamics, maintaining realtime responsiveness with subsecond keytoaudibleupdate latency. Empirical evaluations demonstrate that dynamic prompt compilation reliably grounds generation in microtonal scales, significantly increasing offgrid quartertone content over baseline generation. Beyond its core implementation, MazzikaAI illustrates how deterministic knowledgebased rules can effectively bridge expert, nonWestern musical traditions and unfinetuned foundation models. This architecture establishes a scalable paradigm for realtime humanAI cocreation, offering a generalizable blueprint for interactive accompaniment, adaptive music education, and culturally inclusive generative audio across diverse global idioms.

音乐生成实时控制微音程跨文化AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。