LLM生成的解释虽质量高,却无法提升人类学习,难达超强机器学习目标
LLM-Generated Explanations Do Not Suffice for Ultra-Strong Machine Learning
- 用神经符号框架自动生成逻辑程序的自然语言解释
- 人工评估显示其解释质量优于直接提示和手工模板
- 但对人类学习无帮助,说明需符合认知规律的教法
超强机器学习(USML)指符号学习系统不仅能自我优化,还能通过可量化方式提升人类表现。本文提出LENS(基于神经摘要的逻辑编程解释),将符号程序合成与大语言模型结合,自动生成逻辑程序的自然语言解释,取代以往依赖人工设计的模板。通过LLM作为评判者和专家验证,LENS生成的解释质量高于直接提示和手工模板。随后在三个相关领域的人类实验中,测试这些解释是否能有效传授主动学习策略。探索性分析表明,简洁的专家写作解释对初始水平较高的学习者有益,而LLM生成解释并未带来优于人类自主学习的效果,尽管其评分更高。该案例揭示:实现USML需基于人类学习机制,当前的LLM生成解释未反映人类认知约束,且‘以LLM为评判’的评估体系无法真实反映支持人类学习的有效性。
原文摘要 · Abstract (English)
Ultra Strong Machine Learning (USML) refers to symbolic learning systems that not only improve their own performance but can also teach their acquired knowledge to quantifiably improve human performance. We introduce LENS (Logic Programming Explanation via Neural Summarisation), a neuro-symbolic framework that combines symbolic program synthesis with large language models (LLMs). This framework automatically generates natural language explanations of learned logic programs, replacing hand-crafted templates used in prior USML work. Using LLMs-as-judges evaluation and expert validation, we show that LENS produces higher-quality explanations than both direct LLM prompting and hand-crafted templates. We then examine whether LENS explanations suffice for achieving USML in a human trial teaching active learning strategies across three related domains. Our exploratory analysis suggests that concise, expert-written explanations may benefit learners with higher initial performance, while LLM-generated explanations provide no advantage over human self learning despite being rated as higher quality. This case study reveals that achieving USML requires methods grounded in human learning, where current LLM-generated explanations do not capture human cognitive constraints and LLMs-as-judges evaluations do not reflect what effectively supports human learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。