构建首个多语言长视频导师式问答数据集,推动AI从查事实转向给指导。
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
- 基于180小时多语言长视频构建9000组问答对,定义清晰度、契合度、学习价值等新评估维度
- 多智能体架构在复杂话题和低资源语言上表现最优,显著优于单/双智能体与RAG
- 揭示大模型自动评价与人工判断存在较大差异,强调需谨慎设计评估方法
问答系统通常以事实正确性为评估标准,但教育与职业指导等真实场景需要的是能提供反思与引导的导师式回答。现有评测基准极少捕捉这一差异,尤其在多语言与长篇内容场景中。本文提出MentorQA,首个面向长视频的多语言导师式问答数据集与评估框架,包含近9000个问答对,覆盖180小时内容及四种语言。我们定义了超越事实准确性的评估维度,涵盖清晰度、契合度与学习价值。在控制条件下对比单智能体、双智能体、RAG与多智能体架构,发现多智能体管道在复杂主题与低资源语言中表现最佳。进一步分析表明,基于大模型的自动评估存在显著偏差,与人工判断不一致。本工作确立导师式问答为独立研究方向,并提供多语言基准,用于研究智能体架构与评估设计在教育AI中的应用。数据集与评估框架已开源:https://github.com/AIM-SCU/MentorQA。
原文摘要 · Abstract (English)
Question answering systems are typically evaluated on factual correctness, yet many real-world applications-such as education and career guidance-require mentorship: responses that provide reflection and guidance. Existing QA benchmarks rarely capture this distinction, particularly in multilingual and long-form settings. We introduce MentorQA, the first multilingual dataset and evaluation framework for mentorship-focused question answering from long-form videos, comprising nearly 9,000 QA pairs from 180 hours of content across four languages. We define mentorship-focused evaluation dimensions that go beyond factual accuracy, capturing clarity, alignment, and learning value. Using MentorQA, we compare Single-Agent, Dual-Agent, RAG, and Multi-Agent QA architectures under controlled conditions. Multi-Agent pipelines consistently produce higher-quality mentorship responses, with especially strong gains for complex topics and lower-resource languages. We further analyze the reliability of automated LLM-based evaluation, observing substantial variation in alignment with human judgments. Overall, this work establishes mentorship-focused QA as a distinct research problem and provides a multilingual benchmark for studying agentic architectures and evaluation design in educational AI. The dataset and evaluation framework are released at https://github.com/AIM-SCU/MentorQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。