# Semantic Similarity & False Positives ## Problem When searching for **Python** and **Machine Learning**, a candidate with only **JavaScript, TypeScript, Node.js, Vue.js, PWA, CI/CD, Web Architectures** receives a score of **0.392** (39%). This seems too high for someone who has **none of the searched competences**. ## Root Cause: Semantic Embeddings The `per_skill` strategy works as follows: 1. For each **required** competence, find the **best match** from candidate competences 2. Average those best-match scores The Azure OpenAI embedding model (`text-embedding-3-large`) finds **semantic similarity** between concepts, not just exact matches: ### Example Similarities | Required | Best Candidate Match | Why Similar? | Approx. Score | |----------|---------------------|--------------|---------------| | Python | JavaScript / TypeScript / Node.js | All programming languages | ~0.50-0.60 | | Machine Learning | Progressive Web App / general software dev | Broader technical domain | ~0.20-0.30 | **Average:** `(0.55 + 0.25) / 2 = 0.40` → explains the 0.392 score! ## Why This Happens Embedding models are trained to understand: - **Python** and **JavaScript** are both programming languages - They share many concepts (variables, functions, loops, OOP) - Semantically, they're closer to each other than to "accounting" or "project management" This is **by design** - embeddings capture semantic relationships, not just exact string matches. ## Solution Options ### Option 1: Raise the `partial` Threshold ⭐ RECOMMENDED **Change in `config.toml`:** ```toml [matching.thresholds] top = 0.8 good = 0.65 partial = 0.5 # ← was 0.4, now 0.5 low = 0.3 # ← was 0.2, now 0.3 ``` **Effect:** - Score 0.392 → category "Low" (instead of "Partial") - Clearer distinction between "somewhat related" and "actually matching" - More balanced category distribution **Pros:** - Simple one-line config change - Makes categories more meaningful - Doesn't affect matching, only categorization **Cons:** - Might push some legitimate partial matches into "Low" --- ### Option 2: Stricter Competence Inference **Change in `config.toml`:** ```toml [matching.inference] max_competences = 16 min_similarity = 0.6 # ← was 0.4, raise to 0.6 or 0.7 ``` **Effect:** - When inferring competences from descriptions, only closer semantic matches count - Doesn't affect manual competence lists **Pros:** - Reduces noise in inferred competences - More precise matching **Cons:** - Only affects inference, not manual searches - Might miss some valid related competences --- ### Option 3: More Specific Competences **Instead of:** ``` Required: ["Python", "Machine Learning"] ``` **Use:** ``` Required: ["Python", "scikit-learn", "TensorFlow", "PyTorch", "pandas", "NumPy"] ``` **Effect:** - More specific competences are harder to match accidentally - "scikit-learn" is semantically very different from "Vue.js" **Pros:** - More precise matching - Better reflects actual skill requirements **Cons:** - Requires more detailed competence lists - Might miss candidates who have Python but use different ML libraries --- ### Option 4: Switch to `aggregate` Strategy **Change in `config.toml`:** ```toml [matching.similarity] strategy = "aggregate" # ← was "per_skill" ``` **How it works:** - Computes **mean embedding** of all required competences - Computes **mean embedding** of all candidate competences - Returns cosine similarity between the two means **Effect with your example:** - Required mean: average of [Python, Machine Learning] - Candidate mean: average of [JavaScript, TypeScript, Node.js, Vue.js, PWA, CI/CD, Web Arch] - The diluted candidate mean has **lower similarity** to Python/ML than per-skill **Pros:** - Better for "overall technical fit" vs "specific skill coverage" - Reduces false positives from cross-domain matches **Cons:** - Extra candidate competences dilute the score (see SCORE_BUG_FIX.md) - All required competences get the same score - Less granular than per-skill --- ### Option 5: Filter by Matched Competences **Use MCP tools:** ```python # After search, filter to only show results where specific competences matched filter_search_results( search_id="...", filter_id="python-ml-only", min_similarity=0.65, # High threshold for "matched" ) ``` **Effect:** - Only shows results where similarity >= 0.65 for Python or Machine Learning - Filters out the JavaScript/TypeScript candidates **Pros:** - Most precise control - Can experiment with different thresholds interactively **Cons:** - Requires post-search filtering - Adds an extra step to workflow --- ## Recommended Approach **Combination of Option 1 + Option 5:** 1. **Raise `partial` threshold to 0.5:** ```toml [matching.thresholds] partial = 0.5 ``` 2. **Use interactive filtering when needed:** - For broad exploration: accept lower scores - For precise matching: apply `filter_search_results()` with higher thresholds This gives you: - ✅ Clearer category boundaries (0.392 → Low, not Partial) - ✅ Flexibility to tighten results when needed - ✅ No loss of recall (candidates still appear, just in correct category) --- ### Option 6: Switch to BM25 + RRF (lexical matching) ⭐ BEST for exact skill names **Change in `config.toml`:** ```toml [matching.similarity] use_bm25_search = true ``` **How it works:** - For each required competence, ranks candidate competences using BM25 (term-frequency/ inverse-document-frequency lexical scoring) and normalizes via Reciprocal Rank Fusion. - Candidates with **no shared token** with the required competence receive score **0.0** — the false-positive problem is completely eliminated. **Effect with your example:** - Required: `["Python", "Machine Learning"]` - Candidate: `["JavaScript", "TypeScript", "Node.js", "Vue.js"]` - Score: **0.0** (no token overlap at all) **Pros:** - ✅ Eliminates false positives for lexically disjoint skills - ✅ No Azure embedding calls for competence matching (faster, lower cost) - ✅ Deterministic — same query always returns same ranked order **Cons:** - ❌ Misses synonyms and cross-language pairs: - "ML" ≠ "Machine Learning" (no token overlap) - "Softwarearchitektur" ≠ "Software Architecture" - "Data Science" ≠ "Machine Learning" - ❌ Role similarity still uses embeddings (unchanged) **Mitigation:** combine with `use_auto_tagging = true` (see Option 7 below). --- ### Option 7: BM25 + RRF + Auto-Tagging ⭐ BEST for mixed-language / abbreviation-heavy datasets **Change in `config.toml`:** ```toml [matching.similarity] use_bm25_search = true use_auto_tagging = true [azure_openai] chat_deployment = "gpt-4.1" # your chat deployment name ``` **And set the env variable:** ``` AZURE_OPENAI_LLM_API_KEY= ``` **How it works:** - Before BM25 scoring, the LLM is asked: "which of these *required* competences are already covered by the candidate's *existing* competences (via synonym, abbreviation, or cross-language equivalence)?" - The canonical required-competence names identified by the LLM are appended to the candidate's working list for this BM25 call only (result is **never persisted**). - BM25 then finds token matches including the LLM-expanded terms. **Effect:** - "ML" → LLM identifies "Machine Learning" as covered → added → BM25 matches ✅ - "CI" → LLM identifies "CI/CD" as covered → BM25 matches ✅ - "Softwarearchitektur" → LLM identifies "Software Architecture" → BM25 matches ✅ **Pros:** - ✅ Eliminates false positives (BM25 zero-out rule) - ✅ Bridges synonyms, abbreviations, and cross-language pairs (LLM) - ✅ Expansion is ephemeral — no database pollution **Cons:** - ❌ One additional LLM call per candidate per matching run (latency + cost) - ❌ Requires `AZURE_OPENAI_LLM_API_KEY` environment variable --- ## Strategy Comparison | Strategy | False positives | Synonyms/Cross-language | Cost | Speed | |----------|----------------|------------------------|------|-------| | `per_skill` (default) | ⚠️ Moderate | ✅ Good | Azure embeddings | Fast (cached) | | `aggregate` | ⚠️ Moderate | ✅ Good | Azure embeddings | Fast (cached) | | `use_bm25_search` | ✅ Eliminated | ❌ Misses | No embedding calls | Very fast | | `use_bm25_search` + `use_auto_tagging` | ✅ Eliminated | ✅ Good | + 1 LLM call/candidate | Moderate | --- ## Technical Background ### Why Embeddings Find Semantic Similarity The `text-embedding-3-large` model is trained on massive amounts of text including: - Programming language documentation - Technical tutorials and courses - Stack Overflow discussions - GitHub repositories It learns that: - **Python** and **JavaScript** both appear in programming contexts - They share similar syntax concepts - Developers often know both - They're used for similar tasks (web development, data processing) This makes the embeddings **semantically aware**, which is powerful for fuzzy matching but can cause false positives when you need exact skill matches. ### Per-Skill vs Aggregate **Per-Skill:** - Finds best match for each required competence - Good for: "Must have Python AND Machine Learning" - Risk: Cross-domain false positives (Python ↔ JavaScript) **Aggregate:** - Computes overall semantic fit - Good for: "General technical background in this domain" - Risk: Diluted by extra competences Choose based on your use case! --- ## See Also - `SCORE_BUG_FIX.md` - Detailed test coverage for scoring behavior - `docs/architecture.md` - Similarity strategy documentation - `config.toml.example` - Configuration options