arXiv:2505.15695cs.CL2025-05EMNLP被引 6

测试大模型在真实网络评论中挖掘观点的有效性。

Can Large Language Models be Effective Online Opinion Miners?

  • 构建新基准OOMB,评估大模型从复杂内容中提取观点的能力
  • 验证大模型在实体-特征-观点三元组抽取上表现良好
  • 适合对舆情分析、大模型应用感兴趣的读者

用户生成的在线内容蕴含丰富的客户偏好与市场趋势信息,但其高度多样化、复杂且上下文丰富,给传统观点挖掘方法带来挑战。为此,我们提出在线观点挖掘基准(OOMB),一个全新的数据集与评估协议,用于衡量大语言模型(LLMs)在复杂多样的在线环境中有效挖掘观点的能力。OOMB 提供了全面的(实体、特征、观点)三元组标注以及以观点为中心的摘要,突出每段内容中的关键观点主题,从而支持对模型抽取与摘要能力的双重评估。通过该基准,我们系统分析了当前仍具挑战性的方面及大模型的适应性,探讨其在真实在线场景中作为观点挖掘工具的可行性。本研究为基于大模型的观点挖掘奠定基础,并指明未来研究方向。

原文摘要 · Abstract (English)

The surge of user-generated online content presents a wealth of insights into customer preferences and market trends. However, the highly diverse, complex, and context-rich nature of such contents poses significant challenges to traditional opinion mining approaches. To address this, we introduce Online Opinion Mining Benchmark (OOMB), a novel dataset and evaluation protocol designed to assess the ability of large language models (LLMs) to mine opinions effectively from diverse and intricate online environments. OOMB provides extensive (entity, feature, opinion) tuple annotations and a comprehensive opinion-centric summary that highlights key opinion topics within each content, thereby enabling the evaluation of both the extractive and abstractive capabilities of models. Through our proposed benchmark, we conduct a comprehensive analysis of which aspects remain challenging and where LLMs exhibit adaptability, to explore whether they can effectively serve as opinion miners in realistic online scenarios. This study lays the foundation for LLM-based opinion mining and discusses directions for future research in this field.

观点挖掘大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。