用AI把味觉描述转成音乐,让味觉与声音产生跨模态共鸣。
A Multimodal Symphony: Integrating Taste and Sound through Generative AI
- 用微调的MusicGEN模型,将详细味觉描述生成对应音乐。
- 111名参与者评价显示,微调模型生成的音乐更贴合味觉描述。
- 适合对跨模态生成、感官交互感兴趣的开发者与研究者。
近几十年来,神经科学和心理学研究揭示了味觉与听觉感知之间的直接关联。本文探索了能够将味觉信息转化为音乐的多模态生成模型,基于这一基础研究展开。文章简要回顾了该领域的最新进展,强调关键发现与方法。实验中,我们使用微调后的生成音乐模型(MusicGEN),根据每首曲目对应的详细味觉描述生成音乐。结果显示:根据111名参与者的评估,微调模型生成的音乐在整体上更一致地反映了输入的味觉描述,优于未微调模型。本研究标志着在理解与开发人工智能、声音与味觉之间具身交互方面的重要进展,为生成式AI领域开辟了新可能。数据集、代码及预训练模型已公开:https://osf.io/xs5jy/。
原文摘要 · Abstract (English)
In recent decades, neuroscientific and psychological research has traced direct relationships between taste and auditory perceptions. This article explores multimodal generative models capable of converting taste information into music, building on this foundational research. We provide a brief review of the state of the art in this field, highlighting key findings and methodologies. We present an experiment in which a fine-tuned version of a generative music model (MusicGEN) is used to generate music based on detailed taste descriptions provided for each musical piece. The results are promising: according the participants' ($n=111$) evaluation, the fine-tuned model produces music that more coherently reflects the input taste descriptions compared to the non-fine-tuned model. This study represents a significant step towards understanding and developing embodied interactions between AI, sound, and taste, opening new possibilities in the field of generative AI. We release our dataset, code and pre-trained model at: https://osf.io/xs5jy/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。