arXiv:2511.10218cs.AI2025-11被引 3

通过多模态增强与频谱融合,提升城市交通状态建模的准确性。

MTP: Exploring Multimodal Urban Traffic Profiling with Modality Augmentation and Spectrum Fusion

  • 融合数值、视觉、文本三模态数据,从频域学习交通信号特征。
  • 在六个真实数据集上优于现有方法,显著提升预测性能。
  • 适合关注城市交通感知与智能决策的研究者和工程师。

随着现代城市化进程加速,来自各类传感器的交通信号在监测城市状态方面发挥着重要作用,为保障出行安全、缓解交通拥堵和优化城市出行提供了坚实基础。现有交通信号建模方法通常依赖原始数据模态,即城市传感器的直接数值读数,但这种单模态方法忽略了多源异构城市数据中蕴含的语义信息,限制了对交通信号的全面理解及复杂交通动态的精准预测。为此,我们提出一种新型多模态框架MTP,用于城市交通状态建模,通过数值、视觉和文本三个视角学习多模态特征。三个分支在频域内驱动城市交通信号的多视角学习,频率学习策略精细地提炼信息。具体而言,我们对交通信号进行视觉增强,将其转化为频域图像与周期性图像以支持视觉学习;基于特定主题、背景信息和项目描述,对交通信号生成描述性文本以支持文本学习;同时,采用频域多层感知机对原始数值模态进行学习。设计分层对比学习机制,在三个分支间融合多模态频谱信息。在六个真实世界数据集上的大量实验表明,该方法显著优于当前最优模型。

原文摘要 · Abstract (English)

With rapid urbanization in the modern era, traffic signals from various sensors have been playing a significant role in monitoring the states of cities, which provides a strong foundation in ensuring safe travel, reducing traffic congestion and optimizing urban mobility. Most existing methods for traffic signal modeling often rely on the original data modality, i.e., numerical direct readings from the sensors in cities. However, this unimodal approach overlooks the semantic information existing in multimodal heterogeneous urban data in different perspectives, which hinders a comprehensive understanding of traffic signals and limits the accurate prediction of complex traffic dynamics. To address this problem, we propose a novel Multimodal framework, MTP, for urban Traffic Profiling, which learns multimodal features through numeric, visual, and textual perspectives. The three branches drive for a multimodal perspective of urban traffic signal learning in the frequency domain, while the frequency learning strategies delicately refine the information for extraction. Specifically, we first conduct the visual augmentation for the traffic signals, which transforms the original modality into frequency images and periodicity images for visual learning. Also, we augment descriptive texts for the traffic signals based on the specific topic, background information and item description for textual learning. To complement the numeric information, we utilize frequency multilayer perceptrons for learning on the original modality. We design a hierarchical contrastive learning on the three branches to fuse the spectrum of three modalities. Finally, extensive experiments on six real-world datasets demonstrate superior performance compared with the state-of-the-art approaches.

多模态交通建模频谱融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。