用社交媒体数据评测大模型推断用户偏好的能力。
SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context

- 基于171人长期社交媒体数据构建多模态偏好基准
- 模型能识别宽泛兴趣,但对细粒度和近期兴趣效果差
- 适合研究长时序用户建模与个性化对话的学者
个性化语言模型常通过记忆能力评估:能否回忆用户明确陈述的偏好?更全面的个性化需要更强的能力——从用户自然留下的多模态痕迹中推断其关心的事。我们提出SocialPersona,一个用于评估多模态大语言模型(MLLMs)能否从长期社交媒体时间线中恢复显性偏好并用于对话的基准。该基准基于171位普通用户的非推广类社交媒体长期记录,包含文本、图像、时间戳及2,597个经人工验证的偏好标签,覆盖七个兴趣领域,区分稳定兴趣与近期兴趣。支持两项任务:从多模态上下文构建结构化用户画像,以及生成符合推断画像的回应。在专有和开源的MLLM上实验显示,模型可识别宽泛兴趣领域,但在细粒度和近期兴趣上表现下降,且当需使用推断画像进行对话个性化时性能进一步恶化。结合文本与图像提供互补偏好信号的证据,结果表明跨模态、长时程用户建模仍是关键挑战,而SocialPersona有助于衡量并推动向能推断并响应揭示偏好助手的进展。
原文摘要 · Abstract (English)
Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogue? More comprehensive personalization demands a harder capability -- inferring what users care about from the multimodal traces they naturally leave behind. We introduce SocialPersona, a benchmark for evaluating whether multimodal large language models (MLLMs) can recover revealed preferences from longitudinal social-media timelines and use them in dialogue. Built from longitudinal timelines of 171 everyday, non-promotional social-media users, SocialPersona contains text, images, timestamps, and 2,597 human-verified preference tags across seven interest domains, separating stable interests from recent interests. It supports two tasks: constructing structured user profiles from multimodal context and generating responses aligned with inferred profiles. Experiments with proprietary and open-weight MLLMs show that models can identify broad interest domains, yet their performance drops on fine-grained and recent interests and degrades further when inferred profiles must be used to personalize dialogue. Together with evidence that text and images provide complementary preference signals, these results indicate that robust cross-modal, long-horizon user modeling remains a key challenge, and that SocialPersona can help measure and advance progress toward assistants that infer and act on revealed preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。