评测大模型生成课件、播客等非传统学术输出的能力与效果
Advancing Academic Chatbots: Evaluation of Non Traditional Outputs
- 对比图谱RAG与混合关键词语义搜索,发现后者更准且少幻觉
- GPT 4o mini在问答和内容生成上表现最优,LLaMA 3在叙事连贯性上具潜力
- 强调人工评审对布局与风格缺陷的识别至关重要,适合教育与科研场景
现有大模型评估多集中于事实问答或短摘要等常规任务。本研究从两方面拓展评估维度:一是比较基于知识图谱的Graph RAG与混合关键词-语义搜索的Advanced RAG在问答中的表现;二是评估大模型生成高质量非传统学术输出(如课件、播客脚本)的能力。实验采用Meta LLaMA 3 70B与OpenAI GPT 4o mini API构建原型系统。问答性能通过人类评分(11个质量维度)与大模型裁判进行双验证。结果显示,GPT 4o mini搭配Advanced RAG生成答案最准确,而Graph RAG虽有提升但易引入幻觉,可能因结构复杂与人工配置所致。课件与播客生成基于文档增强检索,结果仍以GPT 4o mini最优,但LLaMA 3在叙事连贯性方面表现良好。人工评审在发现排版与风格缺陷中起关键作用,凸显结合人机评估对新兴学术输出评估的必要性。
原文摘要 · Abstract (English)
Most evaluations of large language models focus on standard tasks such as factual question answering or short summarization. This research expands that scope in two directions: first, by comparing two retrieval strategies, Graph RAG, structured knowledge-graph based, and Advanced RAG, hybrid keyword-semantic search, for QA; and second, by evaluating whether LLMs can generate high quality non-traditional academic outputs, specifically slide decks and podcast scripts. We implemented a prototype combining Meta's LLaMA 3 70B open weight and OpenAI's GPT 4o mini API based. QA performance was evaluated using both human ratings across eleven quality dimensions and large language model judges for scalable cross validation. GPT 4o mini with Advanced RAG produced the most accurate responses. Graph RAG offered limited improvements and led to more hallucinations, partly due to its structural complexity and manual setup. Slide and podcast generation was tested with document grounded retrieval. GPT 4o mini again performed best, though LLaMA 3 showed promise in narrative coherence. Human reviewers were crucial for detecting layout and stylistic flaws, highlighting the need for combined human LLM evaluation in assessing emerging academic outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。