用NLP检测企业环保谎言,揭示现有方法的局限与改进方向
Detecting Greenwashing: A Natural Language Processing Literature Survey
- 将绿色洗白拆解为多个可操作的NLP任务,如气候话题识别与欺骗性语言检测
- 当前部分任务在控制条件下表现接近完美,但涉及模糊性与主观判断的任务仍难突破
- 缺乏经验证的绿色洗白数据集,建议结合媒体报告和监管记录提升标注可靠性
绿色洗白指企业或政府故意误导公众对其环境影响的认知。本文系统综述了自然语言处理(NLP)在文本数据中检测绿色洗白的方法,重点关注企业气候沟通。不将绿色洗白视为单一任务,而是分析研究者用于近似该问题的一系列气候NLP任务,涵盖气候主题识别到欺骗性沟通模式识别。重点考察这些方法的理论基础:任务定义、数据集构建及模型评估如何影响可靠性。研究发现,当前领域碎片化严重:若干子任务在受控条件下已接近完美性能,但涉及歧义、主观性或推理的任务仍具挑战。关键问题是尚无经验证的绿色洗白案例数据集。我们主张推进自动化检测需采用严谨的NLP方法,结合可靠标注与可解释模型设计。未来工作应利用第三方权威判断(如媒体报道、监管记录)以降低标注主观性和法律风险,并采用分步式流程支持人工监督、可追溯推理与高效模型设计。
原文摘要 · Abstract (English)
Greenwashing refers to practices by corporations or governments that intentionally mislead the public about their environmental impact. This paper provides a comprehensive and methodologically grounded survey of natural language processing (NLP) approaches for detecting greenwashing in textual data, with a focus on corporate climate communication. Rather than treating greenwashing as a single, monolithic task, we examine the set of NLP problems, also known as climate NLP tasks, that researchers have used to approximate it, ranging from climate topic detection to the identification of deceptive communication patterns. Our focus is on the methodological foundations of these approaches: how tasks are formulated, how datasets are constructed, and how model evaluation influences reliability. Our review reveals a fragmented landscape: several subtasks now exhibit near-perfect performance under controlled settings, yet tasks involving ambiguity, subjectivity, or reasoning remain challenging. Crucially, no dataset of verified greenwashing cases currently exists. We argue that advancing automated greenwashing detection requires principled NLP methodologies that combine reliable data annotations with interpretable model design. Future work should leverage third-party judgments, such as verified media reports or regulatory records, to mitigate annotation subjectivity and legal risk, and adopt decomposed pipelines that support human oversight, traceable reasoning, and efficient model design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。