构建首个基于同用户评论的隐式比较挖掘数据集,解决真实评论中缺乏显式对比的问题。
Comparing Without Saying: A Dataset and Benchmark for Implicit Comparative Opinion Mining from Same-User Reviews
- 从同一用户的多条独立评论中挖掘隐式偏好关系,不依赖显式比较词汇
- 数据集含4150对评论(15191句),标注了细粒度属性提及与整体偏好
- 首次建立该任务基准,为未来研究提供挑战性评估标准
现有比较意见挖掘研究主要关注显式比较表达,但这类表达在真实评论中较少见。本文提出SUDO,一个针对同用户评论的隐式比较意见挖掘新数据集,可无需显式比较线索即可靠推断用户偏好。SUDO包含4,150个标注的评论对(共15,191句),具有双层结构,涵盖属性级提及与评论级偏好。我们采用传统机器学习和语言模型两类基线方法进行评测。实验表明,尽管语言模型表现优于传统方法,但整体性能仍处于中等水平,揭示了该任务的固有难度,并确立SUDO作为未来研究的挑战性基准。
原文摘要 · Abstract (English)
Existing studies on comparative opinion mining have mainly focused on explicit comparative expressions, which are uncommon in real-world reviews. This leaves implicit comparisons - here users express preferences across separate reviews - largely underexplored. We introduce SUDO, a novel dataset for implicit comparative opinion mining from same-user reviews, allowing reliable inference of user preferences even without explicit comparative cues. SUDO comprises 4,150 annotated review pairs (15,191 sentences) with a bi-level structure capturing aspect-level mentions and review-level preferences. We benchmark this task using two baseline architectures: traditional machine learning- and language model-based baselines. Experimental results show that while the latter outperforms the former, overall performance remains moderate, revealing the inherent difficulty of the task and establishing SUDO as a challenging and valuable benchmark for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。