CareerX: A Retrieval-Augmented Generation Framework for Personalized AI-Driven Career Guidance
Asai, Akari, Zeqiu Wu, Wang, Yizhong, Sil, Avirup, Hannaneh Hajishirzi
Why it has this license class
AChecked 30 Sept 2026. Open license (CC-BY, CC-BY-SA, CC0, public domain): full text indexed and used in synthesis.
| Source | License | Open-access status | Read as |
|---|---|---|---|
| openalex | cc-by | green | Green |
| arxiv | http://creativecommons.org/licenses/by/4.0/ | — | Green |
Abstract
BCareer guidance systems must adapt to rapidly shifting labor markets, yet traditional platforms rely on static databases while standalone large language models (LLMs) are limited by training data cutoffs and prone to hallucination. This paper presents CareerX, a web-based career guidance system that augments LLM generation with live web retrieval through a SearXNG-based meta-search pipeline. The system dynamically generates personalized intake questionnaires via LLM-driven schema generation, formulates targeted search queries, scrapes and cleans live web sources, and synthesizes structured career dashboards through Zod-enforced schema-constrained generation. Unlike canonical RAG architectures that retrieve from pre-built vector stores, CareerX performs direct web retrieval — live search results are scraped and injected into the LLM context window, trading embedding-based semantic precision for real-time data currency. The system produces location-specific salary data, source-attributed university recommendations, and traceable URL citations that standalone LLMs cannot provide without retrieval augmentation. Initial measurements across three successful test profiles indicate a mean end-to-end pipeline time of 66.7 seconds (SD = 6.2s), of which LLM synthesis accounts for approximately 15–20 seconds, with an average of 37.7 web sources retrieved per session and 79% schema completeness.
Claims built on this paper
D- Hallucination and factual inaccuracy are the shared motivation for grounding LLMs across this body of work: Akari et al. 2023 blame errors on sole reliance on parametric knowledge, Matsumoto et al. 2024 cite hallucinated or irrelevant content and noisy data, and Sušnjak et al. 2025 explicitly design hallucination-mitigation solutions. Trace →
- Both Akari et al. 2023 and Matsumoto et al. 2024 warn that naive retrieval creates its own reliability problems—unhelpful output from indiscriminately injected passages, and difficulty selecting appropriate knowledge from large, noisy sources—motivating their respective self-reflection and graph-of-thoughts mechanisms. Trace →