arXiv:2604.12099cs.IRcs.CL2026-04

对比七种选文策略对文本分析结果的影响,发现语义检索最实用。

The Effect of Document Selection on Query-focused Text Analysis

  • 系统比较随机、语义与混合检索等七种选文方法
  • 语义或混合检索在多数场景下效果最佳,计算开销适中
  • 为文本分析提供可复用的数据筛选框架,适合研究者参考

文本分析常需从文档集合中选择数据,因并非所有文档都与研究问题相关,且计算资源有限。然而,现有研究极少探讨不同选择策略的影响。本文系统评估了七种选择方法(从随机选择到混合检索)在两个数据集上对四种文本分析方法(LDA、BERTopic、TopicGPT、HiCode)的影响,覆盖26个开放性问题。结果表明,语义或混合检索是强推荐的通用策略,既能避免弱策略的缺陷,又无需复杂方法带来的额外计算开销。本研究将数据选择提升为方法论决策,而非单纯的技术限制,推动新策略的发展。

原文摘要 · Abstract (English)

Analyses of document collections often require selecting what data to analyze, as not all documents are relevant to a particular research question and computational constraints preclude analyzing all documents, yet little work has examined effects of selection strategy choices. We systematically evaluate seven selection methods (from random selection to hybrid retrieval) on outputs from four text analyses methods (LDA, BERTopic, TopicGPT, HiCode) over two datasets with 26 open-ended queries. Our evaluation reveals practice guidance: semantic or hybrid retrieval offer strong go-to approaches that avoid the pitfalls of weaker selection strategies and the unnecessary compute overhead of more complicated ones. Overall, our evaluation framework establishes data selection as a methodological decision, rather than a practical necessity, inviting the development of new strategies.

文本分析数据选择检索策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。