探索AI生成印度古典唱腔时音乐家的互动体验与挑战
Exploratory Study Of Human-AI Interaction For Hindustani Music
- 用三层生成模型让音乐家通过三种方式与AI交互
- 发现输出无约束且旋律不连贯是主要问题
- 适合研究音乐生成与人机协作的学者和开发者
本文研究了三位参与者在真实场景中使用GaMaDHaNi——一种用于印度古典声乐旋律的新型分层生成模型的交互过程。通过三种预设交互模式开展用户研究,尽管模型未针对真实世界使用进行适配(训练数据与实际应用存在差异),仍作为试点帮助理解实践音乐家对这类模型的期望、反应与偏好。研究发现两大核心挑战:模型输出缺乏约束,以及生成旋律存在不连贯现象。结合印度古典音乐的语境,论文旨在为未来模型设计提供改进方向。
原文摘要 · Abstract (English)
This paper presents a study of participants interacting with and using GaMaDHaNi, a novel hierarchical generative model for Hindustani vocal contours. To explore possible use cases in human-AI interaction, we conducted a user study with three participants, each engaging with the model through three predefined interaction modes. Although this study was conducted "in the wild"- with the model unadapted for the shift from the training data to real-world interaction - we use it as a pilot to better understand the expectations, reactions, and preferences of practicing musicians when engaging with such a model. We note their challenges as (1) the lack of restrictions in model output, and (2) the incoherence of model output. We situate these challenges in the context of Hindustani music and aim to suggest future directions for the model design to address these gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。