arXiv:2602.11516cs.AI2026-02被引 1

让AI像人一样持续优化思考方式,提升自适应能力。

Human-Inspired Continuous Learning of Internal Reasoning Processes: Learning How to Think for Adaptive AI Systems

  • 将思考过程本身作为学习对象,动态优化推理结构
  • 实测运行时间减少23.9%,系统边执行边进化认知架构
  • 适合需要长期自适应的智能系统开发人员

学习内部推理过程对于构建能够在动态现实环境中持续适应的AI系统至关重要。然而,现有方法大多侧重于任务特定输出或静态知识表示,忽视了对内部推理结构、动作调度策略和学习机制本身的持续优化。本文提出一种受人类启发的持续学习框架,将推理、行动、反思与验证统一于一个序列化推理模型中,并通过并行学习增强。该框架将内部思维过程明确作为主要学习目标,系统性记录内部推理轨迹与环境交互,使系统不仅能优化任务内容,还能改进推理组织、调度与演化机制。这一设计实现边处理边学习,使认知结构在执行过程中持续优化。此外,框架支持用学习到的程序替代预设逻辑,并引入分层的学习-学习机制,联合调整任务参数与学习策略。实验结果表明,在温度传感器异常检测任务中,引入内部过程学习后平均运行时间降低23.9%。

原文摘要 · Abstract (English)

Learning internal reasoning processes is crucial for developing AI systems capable of sustained adaptation in dynamic real-world environments. However, most existing approaches primarily emphasize learning task-specific outputs or static knowledge representations, while overlooking the continuous refinement of internal reasoning structures, action scheduling policies, and learning mechanisms themselves. In this paper, we propose a human-inspired continuous learning framework that unifies reasoning, action, reflection, and verification within a sequential reasoning model enhanced by parallel learning. The framework explicitly treats internal thinking processes as primary learning objects. It systematically records internal reasoning trajectories and environmental interactions as structured learning material, enabling the system to optimize not only task-level content but also the organization, scheduling, and evolution of reasoning activities. This design realizes learning alongside processing, allowing cognitive structures to improve during execution. Furthermore, the framework supports controlled replacement of predefined logic with learned procedures and introduces a hierarchical learning-to-learn mechanism that jointly adapts task-level parameters and learning strategies. As a result, the system progressively evolves its internal cognitive architecture while preserving operational stability. Experimental results on a temperature sensor abnormality detection task show that incorporating internal-process learning reduces average runtime by 23.9%.

持续学习推理优化自适应系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。