arXiv:2508.17258cs.CLcs.IR2025-08被引 1

用不确定性量化整合多个思维链代理,提升少标注场景下的情感分析效果

Are You Sure You're Positive? Consolidating Chain-of-Thought Agents with Uncertainty Quantification for Aspect-Category Sentiment Analysis

  • 基于大模型的分词级不确定度评分,融合多个思维链代理
  • 在3B和70B+参数模型上验证,少数据场景表现稳定
  • 适合标注稀缺、需可复现结果的新领域应用

方面类别情感分析通过识别产品评论中特定主题及其关联观点,提供细粒度洞察。当前主流方法依赖监督学习,但新领域数据稀缺且标注成本高。我们主张在零样本设置下利用大语言模型,以节省时间和资源。此外,标注偏差可能导致监督方法在特定场景下表现良好,但在缺乏标注的新领域中泛化能力差,影响可复现性。本文提出新方法,通过大模型的分词级不确定度得分,整合多个思维链代理。我们在Llama和Qwen的3B与70B+参数版本上进行实验,验证了该方法在实际需求中的有效性,并探讨了在标签稀缺条件下如何评估准确性。

原文摘要 · Abstract (English)

Aspect-category sentiment analysis provides granular insights by identifying specific themes within product reviews that are associated with particular opinions. Supervised learning approaches dominate the field. However, data is scarce and expensive to annotate for new domains. We argue that leveraging large language models in a zero-shot setting is beneficial where the time and resources required for dataset annotation are limited. Furthermore, annotation bias may lead to strong results using supervised methods but transfer poorly to new domains in contexts that lack annotations and demand reproducibility. In our work, we propose novel techniques that combine multiple chain-of-thought agents by leveraging large language models' token-level uncertainty scores. We experiment with the 3B and 70B+ parameter size variants of Llama and Qwen models, demonstrating how these approaches can fulfil practical needs and opening a discussion on how to gauge accuracy in label-scarce conditions.

情感分析大模型不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。