arXiv:2601.13075cs.LGcs.AI2026-01

AI导师助本科生从想法到成稿,效果优于顶级大模型。

METIS: Mentoring Engine for Thoughtful Inquiry & Solutions

  • 分阶段引导+工具增强,动态匹配研究流程
  • 在90个任务中71%胜过Claude Sonnet 4.5,54%胜过GPT-5
  • 适合需要系统指导的本科生科研入门

许多学生缺乏专家研究指导。本文探讨能否用AI导师帮助本科生将一个想法发展为一篇论文。为此构建了METIS——一种工具增强、阶段感知的助手,具备文献检索、指南推荐、方法审查和记忆功能。通过LLM作为裁判的成对偏好、学生角色评分表、短回合辅导及证据合规性检查,在六个写作阶段评估METIS与GPT-5、Claude Sonnet 4.5的表现。在90个单回合提示中,LLM裁判更青睐METIS(胜过Claude Sonnet 4.5达71%,胜过GPT-5达54%)。学生评分(清晰度/可操作性/约束契合度;90提示×3裁判)在各阶段均更高。在五种多轮情景下,METIS最终产出质量略高于GPT-5。提升集中在文档依赖阶段(D-F),与阶段感知路由设计一致;失败模式包括过早调用工具、浅层知识嵌入及偶尔阶段误判。

原文摘要 · Abstract (English)

Many students lack access to expert research mentorship. We ask whether an AI mentor can move undergraduates from an idea to a paper. We build METIS, a tool-augmented, stage-aware assistant with literature search, curated guidelines, methodology checks, and memory. We evaluate METIS against GPT-5 and Claude Sonnet 4.5 across six writing stages using LLM-as-a-judge pairwise preferences, student-persona rubrics, short multi-turn tutoring, and evidence/compliance checks. On 90 single-turn prompts, LLM judges preferred METIS to Claude Sonnet 4.5 in 71% and to GPT-5 in 54%. Student scores (clarity/actionability/constraint-fit; 90 prompts x 3 judges) are higher across stages. In multi-turn sessions (five scenarios/agent), METIS yields slightly higher final quality than GPT-5. Gains concentrate in document-grounded stages (D-F), consistent with stage-aware routing and groundings failure modes include premature tool routing, shallow grounding, and occasional stage misclassification.

AI导师科研辅助阶段感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。