让大模型扮演机器学习任务角色并自我纠错,测试其通用能力。
Mockingbird: How does LLM perform in general machine learning tasks?
- 设计角色扮演+自我反思机制,让LLM适应通用机器学习任务。
- 在常见任务上表现可接受,但无法超越专业文档和人工反馈。
- 适合对LLM泛化能力感兴趣的开发者与研究者参考。
大语言模型(LLMs)正越来越多地用于聊天机器人、信息摘要和代码生成等任务。随着推理能力与速度的快速提升,它们在聊天机器人之外的通用机器学习任务中展现出巨大潜力。本文基于对此潜力的好奇,提出框架Mockingbird,将LLMs适配至通用机器学习任务,并在多个任务上评估其性能与可扩展性。该框架核心思想是让LLM扮演特定功能角色,并通过反思自身错误来改进。评估与分析表明,由LLM驱动的方法(如Mockingbird)在常见机器学习任务上可达到可接受效果;然而,仅靠自我反思尚无法超越领域专用文档或人类专家反馈的效果。
原文摘要 · Abstract (English)
Large language models (LLMs) are now being used with increasing frequency as chat bots, tasked with the summarizing information or generating text and code in accordance with user instructions. The rapid increase in reasoning capabilities and inference speed of LLMs has revealed their remarkable potential for applications extending beyond the domain of chat bots to general machine learning tasks. This work is conducted out of the curiosity about such potential. In this work, we propose a framework Mockingbird to adapt LLMs to general machine learning tasks and evaluate its performance and scalability on several general machine learning tasks. The core concept of this framework is instructing LLMs to role-play functions and reflect on its mistakes to improve itself. Our evaluation and analysis result shows that LLM-driven machine learning methods, such as Mockingbird, can achieve acceptable results on common machine learning tasks; however, solely reflecting on its own currently cannot outperform the effect of domain-specific documents and feedback from human experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。