构建长篇小说多维度语义相似度数据集,助力文学类语言模型评估
FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction
- 基于作者元数据与学者验证,构建12维语义相似度标注的长篇小说数据集
- 测试多种嵌入模型发现其普遍关注表层特征而非深层文学语义
- 强调作者知情同意,保障创作主体权益,适合文学计算研究者使用
随着语言模型处理长篇复杂文本能力提升,其在计算文学研究中的应用日益受到关注。然而,由于长篇文本细粒度标注成本高,且公开文献存在数据污染问题,评估模型在该领域的有效性仍具挑战。现有嵌入相似度数据集因侧重粗粒度、短文本,难以适用于文学任务。本文构建并发布FICSIM,一个包含近期创作长篇小说的数据集,涵盖由作者提供元数据支持的12个维度相似度评分,并经数字人文学者验证。我们评估了多类嵌入模型,发现其普遍倾向关注表面特征,而忽略对文学研究有用的语义类别。数据收集全程注重作者自主权,依赖持续且知情的作者同意。
原文摘要 · Abstract (English)
As language models become capable of processing increasingly long and complex texts, there has been growing interest in their application within computational literary studies. However, evaluating the usefulness of these models for such tasks remains challenging due to the cost of fine-grained annotation for long-form texts and the data contamination concerns inherent in using public-domain literature. Current embedding similarity datasets are not suitable for evaluating literary-domain tasks because of a focus on coarse-grained similarity and primarily on very short text. We assemble and release FICSIM, a dataset of long-form, recently written fiction, including scores along 12 axes of similarity informed by author-produced metadata and validated by digital humanities scholars. We evaluate a suite of embedding models on this task, demonstrating a tendency across models to focus on surface-level features over semantic categories that would be useful for computational literary studies tasks. Throughout our data-collection process, we prioritize author agency and rely on continual, informed author consent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。