针对乌尔都语低资源特性,提出自适应语义推荐架构ULTRA。
ULTRA:Urdu Language Transformer-based Recommendation Architecture
- 根据查询长度动态分配至标题级或全文级语义管道。
- 在乌尔都新闻数据集上精度超90%,优于单一管道基线。
- 适合低资源语言的个性化内容推荐场景。
乌尔都语作为低资源语言,缺乏有效的语义内容推荐系统,尤其在个性化新闻检索领域。现有方法多依赖词法匹配或语言无关技术,难以捕捉语义意图,在不同查询长度和信息需求下表现不佳,导致推荐相关性与适应性下降。为此,本文提出ULTRA(Urdu Language Transformer-based Recommendation Architecture),一种自适应语义推荐框架。ULTRA采用双嵌入架构,结合查询长度感知路由机制,动态区分短查询(意图聚焦)与长查询(上下文丰富)。基于阈值决策,用户查询被引导至优化于标题级或全文级表示的专用语义管道,确保检索时的语义粒度恰当。系统利用Transformer嵌入与优化池化策略,突破表面关键词匹配,实现上下文感知的相似性搜索。在大规模乌尔都新闻语料库上的实验证明,该架构在多样化查询类型下持续提升推荐相关性,相比单管道基线精度提升超过90%,验证了查询自适应语义对齐的有效性。研究结果表明,ULTRA是一种稳健且可泛化的低资源语言内容推荐架构,为语义检索系统设计提供实用洞见。
原文摘要 · Abstract (English)
Urdu, as a low-resource language, lacks effective semantic content recommendation systems, particularly in the domain of personalized news retrieval. Existing approaches largely rely on lexical matching or language-agnostic techniques, which struggle to capture semantic intent and perform poorly under varying query lengths and information needs. This limitation results in reduced relevance and adaptability in Urdu content recommendation. We propose ULTRA (Urdu Language Transformer-based Recommendation Architecture),an adaptive semantic recommendation framework designed to address these challenges. ULTRA introduces a dual-embedding architecture with a query-length aware routing mechanism that dynamically distinguishes between short, intent-focused queries and longer, context-rich queries. Based on a threshold-driven decision process, user queries are routed to specialized semantic pipelines optimized for either title/headline-level or full-content/document level representations, ensuring appropriate semantic granularity during retrieval. The proposed system leverages transformer-based embeddings and optimized pooling strategies to move beyond surface-level keyword matching and enable context-aware similarity search. Extensive experiments conducted on a large-scale Urdu news corpus demonstrate that the proposed architecture consistently improves recommendation relevance across diverse query types. Results show gains in precision above 90% compared to single-pipeline baselines, highlighting the effectiveness of query-adaptive semantic alignment for low-resource languages. The findings establish ULTRA as a robust and generalizable content recommendation architecture, offering practical design insights for semantic retrieval systems in low-resource language settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。