用智能代理架构检测教材历史偏见,避免误判且成本可控。
An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks
- 构建多代理评估系统,区分教材叙述与引用史料,减少误判。
- 罗马尼亚中学教材测试中,83.3%内容被判定为教学可接受,严重度均值2.9/7。
- 相比零样本基线,人类盲评更偏好该系统,适合教育监管与审核场景。
历史教材常隐含偏见、民族主义叙事和选择性省略,难以规模化审计。本文提出一种智能体评估架构,包含多模态筛查代理、五种异质评估代理组成的评审团,以及负责结论合成与人工升级的元代理。核心贡献是“来源溯源协议”,能区分教材叙述与引用的历史资料,防止单模型评估中的系统性误报。在对罗马尼亚高中历史教材的实证研究中,270个筛选片段中有83.3%被判定为教学可接受(平均严重度2.9/7),远优于零样本基线的5.4/7,表明智能体协商可缓解过度惩罚。在18名评估者参与的盲测中(54次对比),独立协商配置在64.8%情况下胜过启发式变体和零样本基线。每本教材约2美元成本,显示该架构在教育治理中具备经济可行性。
原文摘要 · Abstract (English)
History textbooks often contain implicit biases, nationalist framing, and selective omissions that are difficult to audit at scale. We propose an agentic evaluation architecture comprising a multimodal screening agent, a heterogeneous jury of five evaluative agents, and a meta-agent for verdict synthesis and human escalation. A central contribution is a Source Attribution Protocol that distinguishes textbook narrative from quoted historical sources, preventing the misattribution that causes systematic false positives in single-model evaluators. In an empirical study on Romanian upper-secondary history textbooks, 83.3\% of 270 screened excerpts were classified as pedagogically acceptable (mean severity 2.9/7), versus 5.4/7 under a zero-shot baseline, demonstrating that agentic deliberation mitigates over-penalization. In a blind human evaluation (18 evaluators, 54 comparisons), the Independent Deliberation configuration was preferred in 64.8\% of cases over both a heuristic variant and the zero-shot baseline. At approximately \$2 per textbook, these results position agentic evaluation architectures as economically viable decision-support tools for educational governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。