arXiv:2608.12615cs.SDcs.LG2026-08

根据驾驶场景实时生成匹配音乐,提升行车体验

Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences

论文配图:Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences
图 1 · 摘自论文原文
  • 融合行车画面与车辆数据,提取驾驶情境并映射为音乐特征
  • 实现低延迟音乐生成,支持场景变化时的自然过渡
  • 内置安全约束,适合车载系统实际部署

车内音乐可作为自适应界面,提升驾驶体验、注意力与身心健康。我们提出Drive-to-Music,一个基于多模态驾驶信号实时生成音乐的上下文感知系统。该系统利用行车记录仪图像和车辆遥测数据,提取场景语义与驾驶情境,将其映射为高层音乐描述符,并以此调控生成音频模型,合成符合情境的音轨。架构结合感知与生成模块,将视觉与运动输入转化为结构化音乐属性,并实现低延迟音频合成。系统支持驾驶条件演变时的平滑过渡,通过在生成全流程中引入基于约束的控制与安全检查,确保鲁棒性与部署可行性。结果表明,在汽车环境中实现实时、上下文感知的音乐生成是可行的,为个性化、自适应的车载音频体验奠定基础。

原文摘要 · Abstract (English)

In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being. We present Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals. Using dashcam imagery and vehicle telemetry, the system extracts scene semantics and driving context, maps them to high-level musical descriptors, and conditions generative audio models to produce contextually aligned soundtracks. The architecture combines perception and generative components to translate visual and kinematic inputs into structured musical attributes and synthesize audio with low latency. It supports smooth transitions as driving conditions evolve, and to ensure robustness and deployment readiness, we incorporate constraint-based controls and safety checks across the generation pipeline. Our results demonstrate the feasibility of real-time, context-aware music generation in automotive settings, providing a foundation for personalized and adaptive in-vehicle audio experiences.

车载音乐生成音频上下文感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。