arXiv:2501.06189cs.AIcs.CL2025-01

用多模态大模型打造能理解社交内容的智能代理

A Multimodal Social Agent

  • 基于多模态大模型构建社会内容分析代理,支持规划与反思
  • 在问答、标题生成等任务上显著优于基线模型
  • 适合需要自动化社交内容理解的应用场景

近年来,大语言模型(LLMs)在常识推理任务中展现出显著进展,而这一能力对理解社会动态、互动与沟通至关重要。然而,将计算机与这些社会认知能力相结合的潜力仍待探索。本文提出MuSA,一种基于多模态大模型的社会内容分析代理,可处理文本丰富的社交内容,完成问答、视觉问答、标题生成和分类等以人类为中心的任务。该代理采用规划、推理、行动、优化、批评与精炼策略完成任务。实验表明,MuSA在问答、标题生成和内容分类任务中表现显著优于基线模型,具备自动化并提升社会内容分析的能力,有助于各类应用场景中的决策支持。

原文摘要 · Abstract (English)

In recent years, large language models (LLMs) have demonstrated remarkable progress in common-sense reasoning tasks. This ability is fundamental to understanding social dynamics, interactions, and communication. However, the potential of integrating computers with these social capabilities is still relatively unexplored. However, the potential of integrating computers with these social capabilities is still relatively unexplored. This paper introduces MuSA, a multimodal LLM-based agent that analyzes text-rich social content tailored to address selected human-centric content analysis tasks, such as question answering, visual question answering, title generation, and categorization. It uses planning, reasoning, acting, optimizing, criticizing, and refining strategies to complete a task. Our approach demonstrates that MuSA can automate and improve social content analysis, helping decision-making processes across various applications. We have evaluated our agent's capabilities in question answering, title generation, and content categorization tasks. MuSA performs substantially better than our baselines.

多模态社会智能大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。