图像化用户行为数据能让大模型预测更准,且零成本。
To See or To Read: User Behavior Reasoning in Multimodal LLMs
- 用文本、散点图、流程图三种方式表示用户购买序列
- 图像表示使预测准确率提升87.5%,无额外计算开销
- 适合研究多模态模型推理与用户行为分析的学者
多模态大语言模型(MLLMs)正在重塑现代智能系统对序列用户行为数据的推理方式。然而,文本还是图像形式的用户行为数据更能提升MLLM性能,这一问题仍待深入探索。我们提出 exttt{BehaviorLens},一个系统性的基准框架,通过将交易数据以(1)文字段落、(2)散点图、(3)流程图三种形式呈现,评估六种MLLM在用户行为推理中的模态权衡。基于真实世界购买序列数据集,发现当数据以图像形式表示时,MLLM的下一笔购买预测准确率相比等效文本表示提升了87.5%,且无需额外计算成本。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are reshaping how modern agentic systems reason over sequential user-behavior data. However, whether textual or image representations of user behavior data are more effective for maximizing MLLM performance remains underexplored. We present \texttt{BehaviorLens}, a systematic benchmarking framework for assessing modality trade-offs in user-behavior reasoning across six MLLMs by representing transaction data as (1) a text paragraph, (2) a scatter plot, and (3) a flowchart. Using a real-world purchase-sequence dataset, we find that when data is represented as images, MLLMs next-purchase prediction accuracy is improved by 87.5% compared with an equivalent textual representation without any additional computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。