用音乐混音器的操控方式控制大模型,让抽象交互更直观可感。
Mixer Metaphors: audio interfaces for non-musical applications
- 借鉴模拟合成器和混音台的物理控制逻辑,设计语言模型操控界面。
- 艺术家实验表明,音频化控制使大模型操作更直接、更具身体参与感。
- 适合对创意交互、跨感官设计感兴趣的开发者与艺术家。
传统NIME会议聚焦于音乐表达的交互设计。本文反其道而行之,探讨音乐类接口能否成功应用于非音乐场景。为此,我们设计并实现了一款新设备,采用源自模拟合成器和音频混音的界面隐喻,用于物理操控大型语言模型(LLM)的抽象参数。通过对比带与不带音频增强功能的两种版本,邀请一组艺术家在为期一周的使用中分别体验。结果表明,具有音频启发式设计的版本提供了更即时、直接且具身体感的控制体验,显著提升了用户在探索与创作中的实验意愿。研究证明,跨感官隐喻有助于激发创造性思维与具身实践,为新型人机交互设计提供新范式。
原文摘要 · Abstract (English)
The NIME conference traditionally focuses on interfaces for music and musical expression. In this paper we reverse this tradition to ask, can interfaces developed for music be successfully appropriated to non-musical applications? To help answer this question we designed and developed a new device, which uses interface metaphors borrowed from analogue synthesisers and audio mixing to physically control the intangible aspects of a Large Language Model. We compared two versions of the device, with and without the audio-inspired augmentations, with a group of artists who used each version over a one week period. Our results show that the use of audio-like controls afforded more immediate, direct and embodied control over the LLM, allowing users to creatively experiment and play with the device over its non-mixer counterpart. Our project demonstrates how cross-sensory metaphors can support creative thinking and embodied practice when designing new technological interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。