Files
Orchestrator/bahn/teamlandkarte-mcp/docs/semantic_similarity_tradeoffs.md
T
ankn a5f8fb49ab Migrate all repos into monorepo context folders
Bahn: aisupport, Analyse-O2C-C2S, awesome-bahn-mcp-servers, beam-mcp,
      Confluence_Bot, db-planet-mcp-server, O2C-Harness, project-audit,
      Projekt-KIQ-HP, teamlandkarte-mcp
Dhive: Jury-Voting
Privat: CV, NoteGraph (NOTE: NoteGraph needs complete redo after consolidation)
Shared: AI-Orchestrator, OrgMyLife, power_skills_and_more
Shared/references: symphony (read-only)

Bahn repos remain available as independent remotes - this monorepo
pulls them in via subtree, the originals are untouched.
2026-06-30 20:39:52 +02:00

315 lines
9.4 KiB
Markdown

# Semantic Similarity & False Positives
## Problem
When searching for **Python** and **Machine Learning**, a candidate with only **JavaScript, TypeScript, Node.js, Vue.js, PWA, CI/CD, Web Architectures** receives a score of **0.392** (39%).
This seems too high for someone who has **none of the searched competences**.
## Root Cause: Semantic Embeddings
The `per_skill` strategy works as follows:
1. For each **required** competence, find the **best match** from candidate competences
2. Average those best-match scores
The Azure OpenAI embedding model (`text-embedding-3-large`) finds **semantic similarity** between concepts, not just exact matches:
### Example Similarities
| Required | Best Candidate Match | Why Similar? | Approx. Score |
|----------|---------------------|--------------|---------------|
| Python | JavaScript / TypeScript / Node.js | All programming languages | ~0.50-0.60 |
| Machine Learning | Progressive Web App / general software dev | Broader technical domain | ~0.20-0.30 |
**Average:** `(0.55 + 0.25) / 2 = 0.40` → explains the 0.392 score!
## Why This Happens
Embedding models are trained to understand:
- **Python** and **JavaScript** are both programming languages
- They share many concepts (variables, functions, loops, OOP)
- Semantically, they're closer to each other than to "accounting" or "project management"
This is **by design** - embeddings capture semantic relationships, not just exact string matches.
## Solution Options
### Option 1: Raise the `partial` Threshold ⭐ RECOMMENDED
**Change in `config.toml`:**
```toml
[matching.thresholds]
top = 0.8
good = 0.65
partial = 0.5 # ← was 0.4, now 0.5
low = 0.3 # ← was 0.2, now 0.3
```
**Effect:**
- Score 0.392 → category "Low" (instead of "Partial")
- Clearer distinction between "somewhat related" and "actually matching"
- More balanced category distribution
**Pros:**
- Simple one-line config change
- Makes categories more meaningful
- Doesn't affect matching, only categorization
**Cons:**
- Might push some legitimate partial matches into "Low"
---
### Option 2: Stricter Competence Inference
**Change in `config.toml`:**
```toml
[matching.inference]
max_competences = 16
min_similarity = 0.6 # ← was 0.4, raise to 0.6 or 0.7
```
**Effect:**
- When inferring competences from descriptions, only closer semantic matches count
- Doesn't affect manual competence lists
**Pros:**
- Reduces noise in inferred competences
- More precise matching
**Cons:**
- Only affects inference, not manual searches
- Might miss some valid related competences
---
### Option 3: More Specific Competences
**Instead of:**
```
Required: ["Python", "Machine Learning"]
```
**Use:**
```
Required: ["Python", "scikit-learn", "TensorFlow", "PyTorch", "pandas", "NumPy"]
```
**Effect:**
- More specific competences are harder to match accidentally
- "scikit-learn" is semantically very different from "Vue.js"
**Pros:**
- More precise matching
- Better reflects actual skill requirements
**Cons:**
- Requires more detailed competence lists
- Might miss candidates who have Python but use different ML libraries
---
### Option 4: Switch to `aggregate` Strategy
**Change in `config.toml`:**
```toml
[matching.similarity]
strategy = "aggregate" # ← was "per_skill"
```
**How it works:**
- Computes **mean embedding** of all required competences
- Computes **mean embedding** of all candidate competences
- Returns cosine similarity between the two means
**Effect with your example:**
- Required mean: average of [Python, Machine Learning]
- Candidate mean: average of [JavaScript, TypeScript, Node.js, Vue.js, PWA, CI/CD, Web Arch]
- The diluted candidate mean has **lower similarity** to Python/ML than per-skill
**Pros:**
- Better for "overall technical fit" vs "specific skill coverage"
- Reduces false positives from cross-domain matches
**Cons:**
- Extra candidate competences dilute the score (see SCORE_BUG_FIX.md)
- All required competences get the same score
- Less granular than per-skill
---
### Option 5: Filter by Matched Competences
**Use MCP tools:**
```python
# After search, filter to only show results where specific competences matched
filter_search_results(
search_id="...",
filter_id="python-ml-only",
min_similarity=0.65, # High threshold for "matched"
)
```
**Effect:**
- Only shows results where similarity >= 0.65 for Python or Machine Learning
- Filters out the JavaScript/TypeScript candidates
**Pros:**
- Most precise control
- Can experiment with different thresholds interactively
**Cons:**
- Requires post-search filtering
- Adds an extra step to workflow
---
## Recommended Approach
**Combination of Option 1 + Option 5:**
1. **Raise `partial` threshold to 0.5:**
```toml
[matching.thresholds]
partial = 0.5
```
2. **Use interactive filtering when needed:**
- For broad exploration: accept lower scores
- For precise matching: apply `filter_search_results()` with higher thresholds
This gives you:
- ✅ Clearer category boundaries (0.392 → Low, not Partial)
- ✅ Flexibility to tighten results when needed
- ✅ No loss of recall (candidates still appear, just in correct category)
---
### Option 6: Switch to BM25 + RRF (lexical matching) ⭐ BEST for exact skill names
**Change in `config.toml`:**
```toml
[matching.similarity]
use_bm25_search = true
```
**How it works:**
- For each required competence, ranks candidate competences using BM25 (term-frequency/
inverse-document-frequency lexical scoring) and normalizes via Reciprocal Rank Fusion.
- Candidates with **no shared token** with the required competence receive score **0.0**
— the false-positive problem is completely eliminated.
**Effect with your example:**
- Required: `["Python", "Machine Learning"]`
- Candidate: `["JavaScript", "TypeScript", "Node.js", "Vue.js"]`
- Score: **0.0** (no token overlap at all)
**Pros:**
- ✅ Eliminates false positives for lexically disjoint skills
- ✅ No Azure embedding calls for competence matching (faster, lower cost)
- ✅ Deterministic — same query always returns same ranked order
**Cons:**
- ❌ Misses synonyms and cross-language pairs:
- "ML" ≠ "Machine Learning" (no token overlap)
- "Softwarearchitektur" ≠ "Software Architecture"
- "Data Science" ≠ "Machine Learning"
- ❌ Role similarity still uses embeddings (unchanged)
**Mitigation:** combine with `use_auto_tagging = true` (see Option 7 below).
---
### Option 7: BM25 + RRF + Auto-Tagging ⭐ BEST for mixed-language / abbreviation-heavy datasets
**Change in `config.toml`:**
```toml
[matching.similarity]
use_bm25_search = true
use_auto_tagging = true
[azure_openai]
chat_deployment = "gpt-4.1" # your chat deployment name
```
**And set the env variable:**
```
AZURE_OPENAI_LLM_API_KEY=<your-chat-api-key>
```
**How it works:**
- Before BM25 scoring, the LLM is asked: "which of these *required* competences are
already covered by the candidate's *existing* competences (via synonym, abbreviation,
or cross-language equivalence)?"
- The canonical required-competence names identified by the LLM are appended to the
candidate's working list for this BM25 call only (result is **never persisted**).
- BM25 then finds token matches including the LLM-expanded terms.
**Effect:**
- "ML" → LLM identifies "Machine Learning" as covered → added → BM25 matches ✅
- "CI" → LLM identifies "CI/CD" as covered → BM25 matches ✅
- "Softwarearchitektur" → LLM identifies "Software Architecture" → BM25 matches ✅
**Pros:**
- ✅ Eliminates false positives (BM25 zero-out rule)
- ✅ Bridges synonyms, abbreviations, and cross-language pairs (LLM)
- ✅ Expansion is ephemeral — no database pollution
**Cons:**
- ❌ One additional LLM call per candidate per matching run (latency + cost)
- ❌ Requires `AZURE_OPENAI_LLM_API_KEY` environment variable
---
## Strategy Comparison
| Strategy | False positives | Synonyms/Cross-language | Cost | Speed |
|----------|----------------|------------------------|------|-------|
| `per_skill` (default) | ⚠️ Moderate | ✅ Good | Azure embeddings | Fast (cached) |
| `aggregate` | ⚠️ Moderate | ✅ Good | Azure embeddings | Fast (cached) |
| `use_bm25_search` | ✅ Eliminated | ❌ Misses | No embedding calls | Very fast |
| `use_bm25_search` + `use_auto_tagging` | ✅ Eliminated | ✅ Good | + 1 LLM call/candidate | Moderate |
---
## Technical Background
### Why Embeddings Find Semantic Similarity
The `text-embedding-3-large` model is trained on massive amounts of text including:
- Programming language documentation
- Technical tutorials and courses
- Stack Overflow discussions
- GitHub repositories
It learns that:
- **Python** and **JavaScript** both appear in programming contexts
- They share similar syntax concepts
- Developers often know both
- They're used for similar tasks (web development, data processing)
This makes the embeddings **semantically aware**, which is powerful for fuzzy matching but can cause false positives when you need exact skill matches.
### Per-Skill vs Aggregate
**Per-Skill:**
- Finds best match for each required competence
- Good for: "Must have Python AND Machine Learning"
- Risk: Cross-domain false positives (Python ↔ JavaScript)
**Aggregate:**
- Computes overall semantic fit
- Good for: "General technical background in this domain"
- Risk: Diluted by extra competences
Choose based on your use case!
---
## See Also
- `SCORE_BUG_FIX.md` - Detailed test coverage for scoring behavior
- `docs/architecture.md` - Similarity strategy documentation
- `config.toml.example` - Configuration options