arXiv:2605.23987cs.AIcs.RO2026-05

让机器人自主发现新特征、新类别并优化动作,实现持续进化。

Beyond Predefined Learning Objects: A Thinking-Learning Interaction Model for Up-to-Date Autonomous Robot Learning

论文配图:Beyond Predefined Learning Objects: A Thinking-Learning Interaction Model for Up-to-Date Autonomous Robot Learning
图 1 · 摘自论文原文
  • 思考与学习双向互动,动态调整输入、输出和动作策略。
  • 识别准确率从41.9%提升至84.5%,动作长度由13.0降至4.0。
  • 适合长期运行的开放环境机器人,突破预设任务限制。

在开放变化环境中,自主机器人难以依赖预定义的输入、输出和动作流程。现有学习方法虽能通过环境交互提升性能,但学习对象常固定不变,如特征、识别结果、网络结构或任务目标,限制了对新特征、新类别或更优任务流程的适应能力。为此,本文提出一种思维-学习交互模型:思考通过识别潜在变化、筛选有用证据、组织训练材料和规划验证动作来引导学习;学习则通过更新任务知识、特征选择经验、动作策略和未来推理过程反向促进思考。基于该双向机制,机器人可逐步突破预设学习框架,在持续环境交互中自适应调整识别关系与动作关系。具体支持自适应输入特征发现、输出类别扩展、学习模型更新与动作流程重构。实验表明,该模型使特征适应下的识别准确率从0.419提升至0.845,新类别生成准确率更高,模型更新成功率提升,动作长度平均从13.0降至4.0;在学习增强思维方面,有效证据选择率从0.272升至0.965,表明学习成果显著改善了未来的证据选择与推理能力。

原文摘要 · Abstract (English)

Autonomous robots operating in open and changing environments cannot always rely on predefined inputs, outputs, and action routines. Although existing learning methods enable robots to improve their performance through environmental interaction, the objects of learning are often fixed in advance, such as input features, recognition outputs, network structures, task goals, or action sequences. This limits their ability to adapt when new features, new categories, or more efficient task routines appear during long-term operation. To address this problem, this paper proposes a thinking-learning interaction model for autonomous robots. The core idea is that thinking guides learning by identifying potential changes, selecting useful evidence, organizing training materials, and planning verification actions, while learning promotes thinking by updating task knowledge, feature-selection experience, action strategies, and future reasoning processes. Based on this bidirectional mechanism, the robot can gradually move beyond predefined learning settings and adapt its recognition relations and action relations through continuous interaction with the environment. Specifically, the proposed model supports adaptive input feature discovery, output category expansion, learning model update, and action routine reconstruction. Experimental results show that the proposed model improves the final recognition accuracy from 0.419 to 0.845 in feature adaptation, achieves higher new-category formation accuracy and model-update success rate, and reduces the average action length from 13.0 to 4.0 in action routine reconstruction. In learning-enhanced thinking, the useful evidence selection rate increases from 0.272 to 0.965, indicating that learning results can effectively improve future evidence selection and reasoning.

自主机器人持续学习思维-学习自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。