让AI通过对话自我进化,在无标准答案的复杂问题中提升判断力。
Self-evolving expertise in complex non-verifiable subject domains: dialogue as implicit meta-RL
- 用对话+反思+记忆机制模拟元强化学习,让AI从经验中积累专业能力
- 在两种模型上测试,对话训练后评分显著优于基线(Elo/AlphaRank等指标)
- 适合研究开放性难题的AI系统设计者,或需要持续进化的决策支持工具
所谓“棘手问题”涉及多维复杂场景、结果不可验证、影响异质且无唯一正确答案,人类长期面临此类挑战。现代例子包括司法框架选择、环境污染治理、疫情韧性规划与粮食安全。当前正探索使用先进人工智能系统(特别是基于大语言模型的智能体)与人类协作解决这些问题。尽管可通过微调、人工提示工程或外部工具增强大模型能力,但其缺乏在复杂领域中通过经验自主发展专业能力的内在机制。本文提出Dialectica框架,让智能体围绕特定主题进行结构化对话,结合记忆、自我反思和策略约束的上下文编辑。形式上,讨论被视为一种隐式元强化学习过程。通过评判生成回应的成对比较进行事后评估。在两种模型架构(本地运行的Qwen3:30b和OpenAI的o4-mini)上,结果显示:在对话中启用基于反思的上下文编辑,能使智能体在Elo评分、归一化Bradley-Terry-Davidson能力与AlphaRank质量分布上全面超越基线。定性分析显示,反思日志能识别自身弱点并有效引导后续陈述,定量与定性证据高度一致,支持对话驱动的上下文演化是开放非验证领域中目标专业化提升的可行路径。
原文摘要 · Abstract (English)
So-called `wicked problems', those involving complex multi-dimensional settings, non-verifiable outcomes, heterogeneous impacts and a lack of single objectively correct answers, have plagued humans throughout history. Modern examples include decisions over justice frameworks, solving environmental pollution, planning for pandemic resilience and food security. The use of state-of-the-art artificial intelligence systems (notably Large Language Model-based agents) collaborating with humans on solving such problems is being actively explored. While the abilities of LLMs can be improved by, for example, fine-tuning, hand-crafted system prompts and scaffolding with external tools, LLMs lack endogenous mechanisms to develop expertise through experience in such settings. This work address this gap with Dialectica, a framework where agents engage in structured dialogue on defined topics, augmented by memory, self-reflection, and policy-constrained context editing. Formally, discussion is viewed as an implicit meta-reinforcement learning process. The `dialogue-trained' agents are evaluated post-hoc using judged pairwise comparisons of elicited responses. Across two model architectures (locally run Qwen3:30b and OpenAI's o4-mini) results show that enabling reflection-based context editing during discussion produces agents which dominate their baseline counterparts on Elo scores, normalized Bradley-Terry-Davidson ability, and AlphaRank mass. The predicted signatures of learning are observed qualitatively in statement and reflection logs, where reflections identify weaknesses and reliably shape subsequent statements. Agreement between quantitative and qualitative evidence supports dialogue-driven context evolution as a practical path to targeted expertise amplification in open non-verifiable domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。