让大模型玩概念定义与反例博弈,看它能否像哲学家一样越辩越明。
The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models

- 用大模型生成反例,再由另一模型修复定义,循环迭代。
- 模型判别反例时接受率是人类的两倍,但整体准确性未提升。
- 反复迭代只会让定义变啰嗦,对某些概念根本修不出稳定定义。
概念分析——通过提出定义并以反例不断修正——是哲学方法的核心。我们研究大语言模型能否通过迭代的分析与修复链条完成这一任务:一个模型实例生成对已有定义的反例,另一个模型负责修复定义,过程重复进行。在20个概念和数千次反例-修复循环中,尽管多数模型生成的反例被专家人类及模型裁判判定为无效,但模型裁判接受的比例约为人类的两倍。然而,人类与模型在单个反例有效性判断上具中等一致性。进一步发现,延长迭代次数仅导致定义愈发冗长,准确率未提高;部分概念本身难以形成稳定定义。这些结果表明,尽管大模型可参与哲学推理,但反例-修复循环很快进入边际效益递减阶段,可作为评估大模型是否具备持续高水平哲学思辨能力的重要测试案例。
原文摘要 · Abstract (English)
Conceptual analysis -- proposing definitions and refining them through counterexamples -- is central to philosophical methodology. We study whether language models can perform this task through iterated analysis and repair chains: one model instance generates counterexamples to a proposed definition, another repairs the definition, and the process repeats. Across 20 concepts and thousands of counterexample-repair cycles, we find that, although many LM-generated counterexamples are judged invalid by both expert humans and an LM judge, the LM judge accepts roughly twice as many as humans do. Nonetheless, per-item validity judgments are moderately consistent across humans and between humans and the LM. We further find that extended iteration produces increasingly verbose definitions without improving accuracy. We also see that some concepts resist stable definitions in general. These findings suggest that while LMs can engage in philosophical reasoning, the counterexample-repair loop hits diminishing returns quickly and could be a fruitful test case for evaluating whether LMs can sustain high-level iterated philosophical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。