arXiv:2511.18499cs.CL2025-11

呼吁文学学者参与大模型可解释性研究,以突破单一工具标准的局限。

For Those Who May Find Themselves on the Red Team

  • 主张文学学者介入大模型可解释性研究
  • 强调当前解释方法过于依赖工具性标准
  • 建议通过红队测试等场景开展跨学科对话

本文主张,文学学者必须参与大型语言模型(LLM)可解释性研究。尽管这一参与可能涉及意识形态冲突,甚至隐含共谋风险,但其必要性显而易见:当前解释方法的工具性本质不应成为衡量模型解释的唯一标准。我提出,红队(red team)是开展这一思想交锋的一个可行场所。

原文摘要 · Abstract (English)

This position paper argues that literary scholars must engage with large language model (LLM) interpretability research. While doing so will involve ideological struggle, if not out-right complicity, the necessity of this engagement is clear: the abiding instrumentality of current approaches to interpretability cannot be the only standard by which we measure interpretation with LLMs. One site at which this struggle could take place, I suggest, is the red team.

可解释性文学研究大模型红队

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。