用低成本自动生成高质量英语辅导反馈数据,提升AI导师效果。
FEAT: A Preference Feedback Dataset through a Cost-Effective Auto-Generation and Labeling Framework for English AI Tutoring
- 结合人工与大模型协作生成反馈,降低标注成本
- 仅5%-10%人工数据可显著优于全人工数据
- 适合教育AI、自动反馈系统研究者使用
在英语教学辅导中,教师反馈对学生成长至关重要。近年来,基于AI的辅导系统兴起,但其训练需要大量高质量的教师反馈数据,而人工生成耗时且昂贵。本研究提出FEAT框架,构建三个互补数据集:(1) DIRECT-Manual (DM),由人类与大语言模型协作生成高质量反馈,成本较高;(2) DIRECT-Generated (DG),纯大模型生成,成本低但质量较低;(3) DIRECT-Augmented (DA),以DG为主,加入少量DM数据以提升质量并保持成本优势。实验表明,将5%-10%的DM数据融入DG,性能优于使用100% DM的情况。
原文摘要 · Abstract (English)
In English education tutoring, teacher feedback is essential for guiding students. Recently, AI-based tutoring systems have emerged to assist teachers; however, these systems require high-quality and large-scale teacher feedback data, which is both time-consuming and costly to generate manually. In this study, we propose FEAT, a cost-effective framework for generating teacher feedback, and have constructed three complementary datasets: (1) DIRECT-Manual (DM), where both humans and large language models (LLMs) collaboratively generate high-quality teacher feedback, albeit at a higher cost; (2) DIRECT-Generated (DG), an LLM-only generated, cost-effective dataset with lower quality;, and (3) DIRECT-Augmented (DA), primarily based on DG with a small portion of DM added to enhance quality while maintaining cost-efficiency. Experimental results showed that incorporating a small portion of DM (5-10%) into DG leads to superior performance compared to using 100% DM alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。