arXiv:2608.11171cs.CLcs.AI2026-08

六年追踪可信NLP研讨会,揭示模型从解释到可控的演进路径。

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

  • 基于144篇论文分类分析,梳理可信NLP六大维度演变。
  • 2025-2026年真相性研究占比达37%,成增长最快方向。
  • 机械可解释性在2026年回升,适合关注AI安全与可控性的研究者。

自2021年起,与主要ACL会议同期举办的可信自然语言处理研讨会(TrustNLP)已从8篇论文发展至41篇,六届共收录144篇论文,记录了领域从静态模型后验解释向生成系统机制理解与主动控制的转变。我们基于可信框架(TrustLLM、DecodingTrust)对论文进行六维分类,发现能力涌现与信任维度存在共现。首个高影响力聊天模型发布时,所有信任维度同时激活;后续模型迭代则聚焦于真实性与安全对齐。分析显示:真实性在2025-2026年占论文37%,为增速最快维度;公平性始终是核心主题;可解释性呈倒U型轨迹——后验方法衰退后,2026年因机械可解释性复苏。跨会议对比(ACL、NAACL、EACL、EMNLP,约2000篇)表明,TrustNLP议题分布与领域平均高度一致。研究提炼四项结构洞见,并提出可操作的研究方向。

原文摘要 · Abstract (English)

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (~2K papers) in the same period shows that TrustNLP's topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.

可信AI可解释性安全对齐趋势分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。