arXiv:2409.08543cs.IRcs.AI2024-09被引 4

用音频文本融合与低秩适配提升推荐系统性能

ATFLRec: A Multimodal Recommender System with Audio-Text Fusion and Low-Rank Adaptation via Instruction-Tuned Large Language Model

  • 将音频与文本融合输入指令微调大模型,通过低秩适配高效微调
  • 在多个数据集上达到更高AUC,优于传统及图神经网络基线
  • 独立适配音文模态效果最佳,适合多模态推荐研究者

推荐系统在电商和娱乐等领域通过个性化推荐提升用户满意度。本研究探索将文本与音频多模态数据融入大语言模型以增强推荐性能。传统文本与音频推荐系统存在冷启动问题,而现有大模型方法计算开销大。为此引入低秩适配(LoRA),在不损失性能的前提下提升效率。提出ATFLRec框架,整合音频与文本模态,采用多种LoRA配置与模态融合策略。实验表明,ATFLRec优于传统及图神经网络基线模型,取得更高AUC分数;分别对音频与文本数据使用独立的LoRA模块可实现最优性能,不同池化方法与梅尔滤波器组数量显著影响结果。该研究为优化多模态推荐系统提供了重要参考。

原文摘要 · Abstract (English)

Recommender Systems (RS) play a pivotal role in boosting user satisfaction by providing personalized product suggestions in domains such as e-commerce and entertainment. This study examines the integration of multimodal data text and audio into large language models (LLMs) with the aim of enhancing recommendation performance. Traditional text and audio recommenders encounter limitations such as the cold-start problem, and recent advancements in LLMs, while promising, are computationally expensive. To address these issues, Low-Rank Adaptation (LoRA) is introduced, which enhances efficiency without compromising performance. The ATFLRec framework is proposed to integrate audio and text modalities into a multimodal recommendation system, utilizing various LoRA configurations and modality fusion techniques. Results indicate that ATFLRec outperforms baseline models, including traditional and graph neural network-based approaches, achieving higher AUC scores. Furthermore, separate fine-tuning of audio and text data with distinct LoRA modules yields optimal performance, with different pooling methods and Mel filter bank numbers significantly impacting performance. This research offers valuable insights into optimizing multimodal recommender systems and advancing the integration of diverse data modalities in LLMs.

多模态推荐低秩适配大模型应用音频融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。