用大模型从推文生成可解释的用户画像,跨领域适应性强。
From Millions of Tweets to Actionable Insights: Leveraging LLMs for User Profiling
- 基于领域关键陈述,分两阶段生成抽象与提取式用户画像
- 在波斯语政治推特数据上比顶尖方法提升9.8%准确率
- 无需大量标注数据,适合需可解释性的人机协同场景
通过内容分析进行社交媒体用户画像对虚假信息检测、互动预测、仇恨言论监控和用户行为建模至关重要。现有方法如推文摘要、属性画像和潜在表征学习存在可迁移性差、特征不可解释、依赖大规模标注数据或预设类别僵化等问题。本文提出一种基于大语言模型(LLM)的新方法,利用领域关键陈述作为核心特征,构建画像基础。该两阶段方法先通过特定领域知识库进行半监督过滤,再生成抽象(合成描述)和提取(代表性推文选取)两类用户画像。借助LLM的内生知识,仅需少量人工验证,即可实现跨领域适应,减少对大规模标注数据的依赖。所生成的画像为自然语言形式,能有效压缩用户数据规模,释放LLM的推理与知识能力,适用于下游社交网络任务。我们贡献了一个波斯语政治推特(X)数据集及基于LLM的评估框架,并经人工验证。实验表明,本方法相比最先进的LLM和传统方法显著提升9.8%,证明其在灵活性、适应性和可解释性上的优势。
原文摘要 · Abstract (English)
Social media user profiling through content analysis is crucial for tasks like misinformation detection, engagement prediction, hate speech monitoring, and user behavior modeling. However, existing profiling techniques, including tweet summarization, attribute-based profiling, and latent representation learning, face significant limitations: they often lack transferability, produce non-interpretable features, require large labeled datasets, or rely on rigid predefined categories that limit adaptability. We introduce a novel large language model (LLM)-based approach that leverages domain-defining statements, which serve as key characteristics outlining the important pillars of a domain as foundations for profiling. Our two-stage method first employs semi-supervised filtering with a domain-specific knowledge base, then generates both abstractive (synthesized descriptions) and extractive (representative tweet selections) user profiles. By harnessing LLMs' inherent knowledge with minimal human validation, our approach is adaptable across domains while reducing the need for large labeled datasets. Our method generates interpretable natural language user profiles, condensing extensive user data into a scale that unlocks LLMs' reasoning and knowledge capabilities for downstream social network tasks. We contribute a Persian political Twitter (X) dataset and an LLM-based evaluation framework with human validation. Experimental results show our method significantly outperforms state-of-the-art LLM-based and traditional methods by 9.8%, demonstrating its effectiveness in creating flexible, adaptable, and interpretable user profiles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。