arXiv:2504.20679cs.CLcs.IR2025-04中稿 · SIGIR 2025综述被引 1

用信息检索方法自动匹配跨时间调查问题,提升社会科学研究的可比性。

Are Information Retrieval Approaches Good at Harmonising Longitudinal Survey Questions in Social Science?

  • 提出新任务:识别调查中概念层面等价的问题与选项
  • 专用神经模型在1946-2020数据上表现最优,F1最高
  • 适合做长期社会科学研究的数据整合人员参考

自动化检测纵向社会科学研究中语义等价的问题对长期研究至关重要。由于理论构念(如住房、工作)在不同研究间表达不一致,且随时间演变,这一任务面临双重挑战。本文由计算机科学家与调查专家合作,提出一种新的信息检索任务:在问题与回答选项间识别概念等价性,以统一纵向人口研究。我们在覆盖1946-2020年的调查数据集上测试多种无监督方法,包括概率模型、语言模型线性探测和专用于信息检索的预训练神经网络。结果显示,专用于IR的神经模型整体性能最佳,其他方法表现相当。此外,用神经模型重排序概率模型结果,仅带来最多0.07的F1分数提升。调查专家的定性评估表明,模型对高词汇重叠问题敏感度低,尤其在子概念错配时表现不佳。该分析为社会科学研究中的长期数据调和提供了重要参考。

原文摘要 · Abstract (English)

Automated detection of semantically equivalent questions in longitudinal social science surveys is crucial for long-term studies informing empirical research in the social, economic, and health sciences. Retrieving equivalent questions faces dual challenges: inconsistent representation of theoretical constructs (i.e. concept/sub-concept) across studies as well as between question and response options, and the evolution of vocabulary and structure in longitudinal text. To address these challenges, our multi-disciplinary collaboration of computer scientists and survey specialists presents a new information retrieval (IR) task of identifying concept (e.g. Housing, Job, etc.) equivalence across question and response options to harmonise longitudinal population studies. This paper investigates multiple unsupervised approaches on a survey dataset spanning 1946-2020, including probabilistic models, linear probing of language models, and pre-trained neural networks specialised for IR. We show that IR-specialised neural models achieve the highest overall performance with other approaches performing comparably. Additionally, the re-ranking of the probabilistic model's results with neural models only introduces modest improvements of 0.07 at most in F1-score. Qualitative post-hoc evaluation by survey specialists shows that models generally have a low sensitivity to questions with high lexical overlap, particularly in cases where sub-concepts are mismatched. Altogether, our analysis serves to further research on harmonising longitudinal studies in social science.

信息检索调查分析数据融合自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。