arXiv:2608.28965cs.AIcs.LG2026-08

让搜索更准:把口语化地点转成精准地理实体

From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People Search

  • 用提示不对称双编码器,动态处理地名变体和同名歧义
  • 在非标准地名上召回率提升18个百分点,达46%
  • 适合需要精准地理检索的搜索系统,尤其处理复杂表达时

人员搜索需将自由表述的地名映射为结构化检索所用的地理实体。词法标准化器虽能处理标准名称,但在别名、拼写错误、都市表达及同名歧义上表现脆弱。本文将该任务定义为基于固定本体的分级、集合式实体检索。提出三项耦合设计需求:区分保持身份的变体与依赖知识的别名,控制有效同名实体中的漏检,分离稳定转换与可变实体知识。实现方案包括带校准别名支持的提示不对称双编码器、受控歧义负样本、可局部更新的实体文档,无需重训练。在生产环境开发基准与公开GeoNames迁移任务上,任务适配模型显著优于冻结编码器与标准分词基线。控制性消融实验显示,专用监督作用超越标准微调与编码器扩展。在GeoNames上,适配模型在零到中等字符重叠下均提升已知目标的Recall@1,而字符n-gram在总召回@5上仍略胜一筹。盲测人类评估在分层生产挑战集上,相关性P@1从28.0%提升至46.0%(p=0.012)。固定查询端点估计显示,对非标准查询表现更好,且在高频查询上接近对照组;随机上线实验未发现用户参与度下降。结果表明,任务适配的地理实体检索是现有基于分类树标准器的实际替代方案,尤其在非标准查询中收益最大。

原文摘要 · Abstract (English)

People search must map free-form location phrases to geographic entities used as structured retrieval filters. Lexical standardizers handle canonical names well but are brittle to aliases, misspellings, metropolitan expressions, and same-name ambiguity. We formulate this task as graded, set-valued entity retrieval over a fixed ontology. We identify three coupled design requirements: distinguishing identity-preserving variation from knowledge-dependent aliases, controlling false negatives among valid same-name entities, and separating stable transformations from mutable entity knowledge. We realize them in a prompt-asymmetric bi-encoder with calibrated alias support, bounded ambiguity-aware negatives, and editable entity documents that support localized updates without retraining. Across a fixed production-derived development benchmark and a public GeoNames transfer task, task adaptation improves substantially over frozen encoders and standard token baselines. Controlled development ablations show that specialized supervision contributes beyond standard task fine-tuning and encoder scaling. On GeoNames, the adapted model improves known-target Recall@1 throughout zero-to-moderate character overlap, while character n-grams retain a small aggregate Target Recall@5 advantage. In a blinded human comparison on a stratified production challenge set, our model raises relevant P@1 from 28.0% to 46.0% (p=0.012). Fixed-query endpoint estimates improve on non-canonical queries and remain close to control on frequent queries; a randomized live experiment detects no engagement regression. These results support task-adapted geographic entity retrieval as a practical replacement for the incumbent taxonomy-based standardizer, with the largest relevance gains on non-canonical queries.

地理检索实体匹配搜索优化自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。