arXiv:2608.09280cs.CL2026-08

分析EMNLP 2025清单响应,发现45%理由敷衍,半数作者忽视潜在社会风险。

Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025

论文配图:Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025
图 1 · 摘自论文原文
  • 统计7.4万条清单回复,发现伦理问题常被孤立处理
  • 44.9%的否定回答理由简短空洞,6%存在逻辑矛盾
  • 建议设最低字数、加强风险审查,避免形式主义

负责任的自然语言处理包含透明性、伦理与社会影响。为推动这一目标,ACL发布了EMNLP 2025检查清单。本文首次构建并公开两个数据集:一是主会议与发现赛道全部清单响应及理由;二是清单项与论文段落的映射关系。我们分析了73,922条响应与理由。在主赛道中,作者将伦理问题与正文分离,呈现“事后补救”趋势。针对否定回答,44.9%的理由属于质量差或恶意敷衍,内容简短或空白。还发现6%的清单存在父级与子级回答间的逻辑矛盾。此外,53%的作者声称其工作无潜在风险或社会影响,缺乏实质评估。该现象在发现赛道中同样显著。研究提出改进建议:强制最低字数、强化对应用风险的审查,以提升清单实效性。

原文摘要 · Abstract (English)

Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promote responsible practice. Recently, ACL released the EMNLP 2025 Checklists to aid transparency on the current research practice, which we focus on. We curate and release the first two datasets of: a) all the checklist responses and justifications from the EMNLP 2025 Main and Finding tracks; b) checklist reference linking to paper sections. We also provide the first analysis of recent EMNLP Checklists, by examining $73,922$ responses and justifications to them. For the Main track, we find that authors isolate ethics questions of the Checklist from the paper's bulk, mimicking the trend of ethics being an afterthought. We then examine \texttt{NO} responses. We find $44.9\%$ of justifications are poor or bad-faith, being brief or empty. Then, we find significant issues with the checklist design and effort of authors, namely that $6\%$ of all checklists contained logical contradictions between parent and child responses. We also find evidence of surface compliance for responsible ethics, with $53\%$ authors dismissing potential risks or social impacts of their work, for which there should be none. We compare this to the Findings track, noticing a similar trend in both tracks. Lastly, we discuss the implications of the checklist design and provide recommendations for future checklist iterations. Including: a) enforcing a minimum word count, b) enforcing more scrutiny on the risks of appliances.

负责任AI伦理检查形式主义清单分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。