arXiv:2605.21338cs.CL2026-05

测试大模型处理海量社交媒体文本的能力,发现其分析性能随数据量和任务复杂度显著下降。

Text Analytics Evaluation Framework: A Case Study on LLMs and Social Media

论文配图:Text Analytics Evaluation Framework: A Case Study on LLMs and Social Media
图 1 · 摘自论文原文
  • 设计470个人工标注问题,评估大模型对多篇社交媒体文本的语义理解与推理能力。
  • 输入超500条文本时,模型在数值计算等任务上性能大幅下滑,尤其开放权重模型更明显。
  • 任务越复杂(如比较、计数),模型表现越差,凸显当前大模型在量化分析上的局限。

大模型在多种自然语言任务中表现出色,但在实际数据分析场景中仍存在显著差距,尤其是在处理长序列非结构化文档(如新闻流或社交媒体帖子)时。为实证评估大模型在此类场景下的有效性,本文提出一个基于问题的评估框架,包含470个手动构建的问题,用于检验大模型对聚合文本数据的语义理解和推理能力。我们在涵盖多种NLP任务(包括情感分析、仇恨言论检测和情绪识别)的多样化Twitter数据集上应用该基准。结果表明,模型表现高度依赖输入规模与数据源复杂性,在多标签或目标依赖场景下明显下降。随着任务复杂度提升,性能从基础语义存在识别逐步退化至对比、计数、计算等高阶操作。此外,当输入规模超过500条时,各类大模型(特别是开放权重模型)普遍出现性能严重下降,尤其在数值任务上。这些发现揭示了当前大模型在大规模文本集合上进行严谨定量分析时的关键架构瓶颈。

原文摘要 · Abstract (English)

LLMs have demonstrated exceptional proficiency in a wide range of NLP tasks. However, a notable gap remains in practical data analysis scenarios, particularly when LLMs are required to process long sequences of unstructured documents, such as news feeds or, as specifically addressed in this paper, social media posts. To empirically assess the effectiveness of LLMs in this setting, we introduce a question-based evaluation framework comprising 470 manually curated questions designed to evaluate LLMs' semantic understanding and reasoning abilities over aggregated text data. We apply our benchmark on diverse Twitter datasets covering various NLP tasks, including sentiment analysis, hate speech detection, and emotion recognition. Our results reveal that the performance depends heavily on input scale and the complexity of the data sources, declining noticeably in multi-label or target-dependent scenarios. In addition, as task complexity increases, performance drops progressively from basic semantic existence identification to more demanding operations such as comparison, counting, and calculation. Furthermore, as the input size grows beyond 500 instances, we identify a common limitation across LLMs, particularly Open-weights models: performance degrades substantially, especially on numerical tasks. These findings highlight critical architectural bottlenecks in current LLMs for performing rigorous quantitative analysis over large text collections.

大模型评估社交媒体分析量化推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。