构建首个可控教育领域情感分析合成数据集,助力课程改进研究。
A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis
- 用三轮人工校验生成1万条真实感强的课程评论,含20个教学维度标签
- 最强模型微调后仅达0.2930微F1,表明任务具有挑战性
- 适合做教育数据分析、自监督学习或合成数据评估的研究者使用
教育领域的情感分析可辅助课程优化,但公开标注的学生反馈稀缺,因教育评论具有隐私性、机构特异性且标注成本高。本研究构建了一个受控的合成教育情感分析基准,包含10,000条合成课程评论,具备明确的训练-验证-测试划分,采用覆盖教学品质、评估与管理、学习需求、学习环境及参与度的20维教学框架。语料通过采样目标标签与语义细微差异属性生成,并经三轮评委-编辑流程优化提示词以提升真实性。在该基准上,基于TF-IDF、两阶段Transformer和联合编码器的基线模型表明任务非平凡:未调参的BERT达到0.2760的保留检测微F1,而小幅降低学习率的版本提升至0.2930。全批次GPT-based推理在零样本下得0.2519微F1,检索增强少样本提示下为0.2501,优于传统基线并接近紧凑型联合编码器。对2,829条来自Herath等人的映射真实反馈进行保守外部评估,BERT在9维重叠上获得0.4593微F1,显示部分合成到真实迁移能力。通过真实感与忠实度分析作为生成器诊断,揭示了基准稳定性及仍存在的标签噪声。本研究贡献了一个合成教育情感分析语料库、可复现的生成流程与可重复的基准设置,填补了公共标注数据匮乏领域的空白。
原文摘要 · Abstract (English)
Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because educational reviews are private, institution-specific, and expensive to annotate. This study introduces a controlled synthetic benchmark for educational ABSA built from 10,000 synthetic course reviews with explicit train-validation-test splits and a 20-aspect pedagogical schema spanning instructional quality, assessment and course management, learning demand, learning environment, and engagement. The corpus is generated with sampled target labels, sampled nuance attributes, and a realism-tuned prompt refined through a three-cycle judge-editor procedure. On the resulting benchmark, local baselines with TF-IDF, two-step transformers, and joint encoders show that the task is nontrivial; the strongest untuned model, BERT, reaches a held-out detection micro-F1 of 0.2760, while a modest lower-rate BERT schedule improves this to 0.2930. Full-test GPT-based inference with gpt-5.2 reaches 0.2519 micro-F1 in zero-shot mode and 0.2501 with retrieval-based few-shot prompting, placing batch inference above the classical baseline and close to the compact joint encoders. A conservative external evaluation on 2,829 mapped student-feedback reviews from Herath et al. yields a micro-F1 of 0.4593 for BERT on a 9-aspect overlap, indicating partial synthetic-to-real transfer. Realism and faithfulness analyses are reported as generator diagnostics that clarify how the benchmark was stabilized and where label noise remains. The study therefore contributes a synthetic educational ABSA corpus, a documented generation procedure, and a reproducible benchmark setting for a domain in which public labeled data remain difficult to obtain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。