无需标注数据,自动识别法语移民叙事中的共通经历
Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives
- 用零样本大模型和结构特征检测跨路线移民故事的相似经历
- 最佳方法相关性达0.38,融合31个特征后提升至0.45
- 出发阶段的故事最易出现经历共鸣,适合研究社会创伤与叙事模式
穿越撒哈拉和巴尔干等地理迥异路线的移民常讲述惊人相似的经历:警察暴力、走私贩剥削、危险渡海及家庭分离。本文提出体验性互文性检测任务:在无标注数据情况下,自动识别迁移叙事间的共通经历线索。基于108篇涵盖两条路线的法语移民叙述,我们自动生成句子对,并使用多种无注释方法进行评分:词汇基线、句向量、词性结构特征、移民主题词典、上下文感知叙事特征,以及采用Qwen2.5-7B和Mistral-7B的零样本大模型评分(三种提示策略)。所有方法均在816条专家标注的互文性判断上验证(评分者间一致性Krippendorff's α=0.27)。结果表明,所有表面、结构和嵌入方法与专家判断的相关性均较弱(r ≤ 0.30);其中Qwen2.5-7B零样本表现最佳(r = 0.38);少样本示例虽降低Qwen性能,但显著提升Mistral表现;叙事位置显著预测互文性,出发阶段的句子对具有最高经历共鸣;通过融合全部31个特征的监督混合模型,相关性达到r = 0.45,较最优单方法提升21%。
原文摘要 · Abstract (English)
Migrants traversing geographically distinct routes such as the Trans-Saharan and Balkan corridors often recount strikingly parallel lived experiences: police violence, smuggler exploitation, dangerous crossings, and family separation. We introduce the task of experiential intertextuality detection: automatically identifying shared experiential echoes across migration narratives without requiring annotated training data. From 108 French migration narratives spanning both corridors, we automatically generate sentence pairs and score them using annotation-free methods: lexical baselines, sentence embeddings, POS-based structural features, a migration-specific theme lexicon, context-aware narrative features, and zero-shot LLM scoring with Qwen2.5-7B and Mistral-7B under three prompting strategies. We validate all methods against 816 expert-annotated intertextuality judgments (inter-annotator Krippendorff's $α= 0.27$). Our results reveal that all surface, structural, and embedding methods correlate only weakly with expert judgments ($r \leq 0.30$); Qwen2.5-7B zero-shot achieves the best single-method correlation ($r = 0.38$); few-shot examples degrade Qwen but dramatically improve Mistral; narrative position significantly predicts intertextuality, with departure-phase pairs showing the highest experiential echoes; and a supervised hybrid combining all 31 features achieves $r = 0.45$, a 21% improvement over the best individual method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。