构建土耳其语教育视频摘要数据集与自动共识框架
TR-EduVSum: A Turkish-Focused Dataset and Consensus Framework for Educational Video Summarization
- 基于多人摘要聚类生成共识内容,用嵌入模型捕捉语义单元
- 自动生成的黄金标准摘要与大模型摘要重合度高,最高达0.81
- 方法可低成本推广至其他突厥语系语言,适合多语言摘要研究者
本研究提出一种全自动、可复现的框架,基于土耳其语教育视频的多份人工摘要生成黄金标准摘要。研究构建了名为TR-EduVSum的新数据集,包含82个“数据结构与算法”领域课程视频,共3281条独立人工摘要。受金字塔评估方法启发,提出AutoMUP(自动意义单元金字塔)方法:通过嵌入对人工摘要中的意义单元进行聚类,统计参与者间一致性,根据共识权重生成分级摘要。黄金标准摘要对应最高共识配置,由最频繁被支持的意义单元构成。实验表明,AutoMUP摘要与Flash 2.5和GPT-5.1等大模型摘要具有高度语义重叠(最高0.81)。消融实验证明共识权重与聚类对摘要质量起决定性作用。该方法可低代价推广至其他突厥语系语言。
原文摘要 · Abstract (English)
This study presents a framework for generating the gold-standard summary fully automatically and reproducibly based on multiple human summaries of Turkish educational videos. Within the scope of the study, a new dataset called TR-EduVSum was created, encompassing 82 Turkish course videos in the field of "Data Structures and Algorithms" and containing a total of 3281 independent human summaries. Inspired by existing pyramid-based evaluation approaches, the AutoMUP (Automatic Meaning Unit Pyramid) method is proposed, which extracts consensus-based content from multiple human summaries. AutoMUP clusters the meaning units extracted from human summaries using embedding, statistically models inter-participant agreement, and generates graded summaries based on consensus weight. In this framework, the gold summary corresponds to the highest-consensus AutoMUP configuration, constructed from the most frequently supported meaning units across human summaries. Experimental results show that AutoMUP summaries exhibit high semantic overlap with robust LLM (Large Language Model) summaries such as Flash 2.5 and GPT-5.1. Furthermore, ablation studies clearly demonstrate the decisive role of consensus weight and clustering in determining summary quality. The proposed approach can be generalized to other Turkic languages at low cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。