让AI从多重解读视角预测故事相似性,提升判断准确性。
Multiperspectivity as a Resource for Narrative Similarity Prediction
- 构建31种不同风格的LLM角色,模拟多元解读视角。
- 在SemEval-2026数据集上达0.705准确率,规模越大越准。
- 关注性别解读的模型反而更不准确,提醒评估需包容多元理解。
预测叙事相似性本质上是解释性任务:同一文本可有多种合理解读,导致相似性判断差异,给依赖单一标准的语义评估带来根本挑战。我们不试图消除这种多视角现象,而是将其纳入预测系统的决策过程。为此,我们构建了31个LLM人格化角色,涵盖遵循诠释框架的从业者到更直觉化的普通用户风格。实验基于SemEval-2026 Task 4数据集,系统取得0.705的准确率。准确率随集成规模提升,符合弱化独立性条件下的康多塞陪审团定理动态。从业者角色个体表现较差,但错误相关性低,使多数投票时集成增益更大。误差分析显示,所有角色类别中,聚焦性别解读的词汇与准确率呈稳定负相关,表明可能过度关注基准未涵盖的维度,或存在未被标注的有效解读。这一发现强调评估框架需容纳诠释多样性。
原文摘要 · Abstract (English)
Predicting narrative similarity can be understood as an inherently interpretive task: different, equally valid readings of the same text can produce divergent interpretations and thus different similarity judgments, posing a fundamental challenge for semantic evaluation benchmarks that encode a single ground truth. Rather than treating this multiperspectivity as a challenge to overcome, we propose to incorporate it in the decision making process of predictive systems. To explore this strategy, we created an ensemble of 31 LLM personas. These range from practitioners following interpretive frameworks to more intuitive, lay-style characters. Our experiments were conducted on the SemEval-2026 Task 4 dataset, where the system achieved an accuracy score of 0.705. Accuracy improves with ensemble size, consistent with Condorcet Jury Theorem-like dynamics under weakened independence. Practitioner personas perform worse individually but produce less correlated errors, yielding larger ensemble gains under majority voting. Our error analysis reveals a consistent negative association between gender-focused interpretive vocabulary and accuracy across all persona categories, suggesting either attention to dimensions not relevant for the benchmark or valid interpretations absent from the ground truth. This finding underscores the need for evaluation frameworks that account for interpretive plurality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。