I'm building a resume screening/ranking system (matching resumes to job descriptions using pretrained sentence embeddings + cosine similarity, no fine-tuning at this stage) as a learning project aimed at becoming a market-ready NLP practitioner. Problem: I could not find a trustworthy, publicly available English dataset with genuine human-labeled resume-job match scores. I checked several Kaggle/HuggingFace options and found labels that were either AI-generated (e.g., GPT-4o) or fully synthetic with demographic columns (race/ethnicity/gender) tied to the match label, which raised bias concerns. Academic literature (ConFit, PJFNN papers) confirms that the only broadly public dataset for this exact task (Person-Job Fit) is the 2019 Alibaba matching competition dataset, which is Chinese-only; most published research instead uses private company-provided data. My current approach: Use two real (non-synthetic) datasets: a scraped resume corpus and a real LinkedIn job postings corpus, both verified for low duplication and cleaned of PII. Build a small manual evaluation set myself (~24 resume-job pairs, selected to cover clear matches, clear non-matches, and ambiguous cases), scoring them on a 0-3 relevance scale with a confidence flag. Use this manual set as ground truth to compute ranking metrics (Precision@K, MRR, NDCG) once the embedding-based matching pipeline is built. Question: Is this a sound methodology given the lack of reliable public ground truth, or is there a better-e…

Full article content could not be extracted automatically. Read the original below.