arXiv:2607.11977cs.AIcs.CL2026-07

机器生成文本的权威正被优化指标取代,但它们无法分辨错误与创新。

Optimization Is Not All You Need

论文配图:Optimization Is Not All You Need
图 1 · 摘自论文原文
  • 将语言生成的权威从人类评判转向损失函数与奖励模型。
  • 优化过程能测文本罕见度,却无法判断是错误还是创造。
  • 适合关注AI伦理、语言生成机制的读者。

2019年,OpenAI发布两百万条未经修正的GPT-2生成文本,用于辅助检测机器生成内容。其后续更流畅的输出通常被视为工程进步;我们则将其视为优化文化的新表现:一种比技术更早存在的信念,即在预设维度上可测量的进步便足以定义价值。通过追溯这一信念在预训练、解码、偏好调优、基准测试与界面设计中的体现,并回溯其在审计社会中的渊源,我们发现其极限所在:优化程序可衡量生成文本的稀有性,却无法判断这种罕见是错误还是创造。尽管缺乏判断能力,此类系统却在五年内取得了合法语言表达的制定权。过去由学院、课堂、语法书和考官掌握的语言权威,如今交由损失函数、奖励模型、基准测试与系统提示所构成的工具体系执行,而这一体系虽能行使裁决之职,却无真正评判之能。

原文摘要 · Abstract (English)

In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it instead as the newest expression of optimization culture: the conviction, older than the technology, that measurable improvement along predefined axes exhausts the question of value. Tracing that conviction through the stack-pretraining, decoding, preference tuning, benchmarking, interface-and back through its genealogy in the audit society, we arrive at the limit: an optimization procedure can measure how improbable a piece of generated text is; it cannot tell whether that unlikelihood is error or invention. A procedure that cannot make that distinction has nonetheless, within half a decade, assumed the authority to set the protocols of legitimate language. Held for centuries by academies and schoolrooms, grammars and examiners, this authority has been given over to loss functions, reward models, benchmarks, and system prompts: an apparatus that executes the office of judgment with no capacity for judging.

AI伦理语言生成优化文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。