Files
Orchestrator/bahn/teamlandkarte-mcp/openspec/changes/add-bm25-rrf-competence-matching/design.md
T
ankn a5f8fb49ab Migrate all repos into monorepo context folders
Bahn: aisupport, Analyse-O2C-C2S, awesome-bahn-mcp-servers, beam-mcp,
      Confluence_Bot, db-planet-mcp-server, O2C-Harness, project-audit,
      Projekt-KIQ-HP, teamlandkarte-mcp
Dhive: Jury-Voting
Privat: CV, NoteGraph (NOTE: NoteGraph needs complete redo after consolidation)
Shared: AI-Orchestrator, OrgMyLife, power_skills_and_more
Shared/references: symphony (read-only)

Bahn repos remain available as independent remotes - this monorepo
pulls them in via subtree, the originals are untouched.
2026-06-30 20:39:52 +02:00

667 lines
25 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Design: BM25 + RRF Competence Matching
## Context
This design specifies the internal architecture for the BM25 + Reciprocal Rank Fusion
(RRF) competence matching strategy with optional LLM-based Auto-Tagging, controlled by
`matching.similarity.use_bm25_search` and `matching.similarity.use_auto_tagging`.
The goal is to provide a lexical alternative to embedding-based similarity that assigns
near-zero scores when candidate competences do not literally overlap with the required
competences, while still producing a normalized score compatible with the existing
`Matcher` and `scorer` contracts.
Auto-Tagging optionally extends BM25 by having an LLM pre-expand each candidate's
competence list with canonical forms of any required competences it already covers
(synonyms, abbreviations, cross-language equivalents). This bridges the lexical gap for
known skill aliases without persisting any data.
---
## Key Decisions
### 1. BM25 is per-skill only, not aggregate
**Decision:** `_bm25_rrf_similarity()` operates per required competence (like the existing
`per_skill` strategy), not as an aggregate over all required competences.
**Rationale:**
- Aggregate-style BM25 would require concatenating all required competences into a single
query document, which loses the per-skill granularity needed for `matched_competences`
and `missing_competences` diagnostic output.
- Per-skill BM25 + RRF preserves the same output shape as `_per_skill_similarity()`,
making it a drop-in replacement.
---
### 2. BM25 index is built once globally over all filtered candidates
**Decision:** One `Bm25Index` is built over the **union of all filtered candidates'
competences** before the per-candidate scoring loop in `Matcher.match()`, not fresh per
`compute_competence_similarity()` call. The index is passed in as `global_index` and each
per-candidate call filters BM25 results down to that candidate's own competence set.
**Rationale:**
- **IDF pathology fix:** With a per-person corpus (N = that person's competence count),
BM25 IDF is `log((N df + 0.5) / (df + 0.5))`. For N = 530 and any term that appears
in most or all documents, this is 0 or negative. `rank_bm25` therefore returns zero or
negative raw scores for typical real-world skill lists — making every BM25 match appear
as zero before `max(0.0, raw)` clamping. A global corpus of all candidates' competences
ensures N is large enough for IDF to carry meaningful signal.
- **Still cheap:** Building one index over the deduplicated union of all candidates'
competences is O(M) where M is the total unique competence count — still milliseconds.
- **Correctness over caching:** The global index is not cached across MCP tool calls; it is
rebuilt fresh for each `match_candidates` invocation, so stale state is impossible.
---
### 3. RRF with single ranked list (BM25 only)
**Decision:** For the initial implementation, `reciprocal_rank_fusion()` fuses a single
ranked list (BM25 ranks only), effectively normalizing BM25 scores via the RRF formula.
**Formula (single list):**
```
rrf_score(candidate) = 1 / (k + rank(candidate))
```
where `k = 60` (standard constant, Cormack et al. 2009), and `rank` is 1-based.
**Output normalization:**
The raw RRF scores sum to a value > 1 across all candidates. To map to [0.0, 1.0]:
```
normalized = rrf_score(best_candidate) / rrf_score_at_rank_1
= 1 / (k + rank) * (k + 1)
= (k + 1) / (k + rank)
```
For `k=60`, rank 1 → 1.0, rank 2 → 61/62 ≈ 0.984, rank 10 → 61/70 ≈ 0.871.
This means BM25 scores do NOT inherently drop to near-zero for adjacent candidates.
To address this, RRF scores are **zeroed** if the underlying BM25 score is 0.0, because
a BM25 score of 0 means **no token overlap at all**. This is the critical property that
fixes the false-positive problem.
**Implementation:**
```python
def reciprocal_rank_fusion(
ranked_lists: list[list[tuple[str, float]]],
*,
k: int = 60,
) -> dict[str, float]:
...
# Skip lists with no BM25 signal at all
if all(score == 0.0 for _, score in ranked_list):
continue
# IMPORTANT: Keep (candidate, 0.0) entries in the ranked list for transparency,
# but treat them as "no signal" for fusion. I.e., candidates with BM25==0.0 MUST
# NOT receive a positive fused score just because they have a (low) rank at the
# end of the list.
for rank, (candidate, score) in enumerate(ranked_list, start=1):
if score == 0.0:
continue
fused[candidate] += 1 / (k + rank)
...
```
---
### 4. `best_match` reported per required competence
**Decision:** `best_match` in the output dict is the candidate with the highest BM25 score
(rank 1 in the BM25 list). If no candidate has any token overlap (BM25 = 0 for all),
`best_match = None` and `score = 0.0`.
---
### 5. Role similarity unchanged
**Decision:** `compute_role_similarity()` is always embedding-based, regardless of
`use_bm25_search`.
**Rationale:**
- Role names are short, single-phrase labels ("Backend Developer", "Data Engineer").
- Embeddings handle multilingual and synonym variants better for roles.
- BM25's exact-match strength is less valuable for role inference.
---
### 6. `rank_bm25` as the BM25 library
**Decision:** Use the `rank-bm25` PyPI package (`BM25Okapi` class).
**Rationale:**
- Pure Python, no C extensions or native code.
- Actively maintained, MIT licensed.
- Implements BM25 Okapi (the standard variant) with configurable `k1` and `b` parameters.
- Minimal footprint, no additional build steps.
**Tokenization:**
Candidate and query texts are tokenized by lowercasing and splitting on non-word
characters (whitespace, hyphens, slashes, parentheses, dots):
```python
import re
def _tokenize(text: str) -> list[str]:
return [t for t in re.split(r"[\W_]+", text.lower()) if t]
```
This ensures "Progressive Web App (PWA)" tokenizes to
`["progressive", "web", "app", "pwa"]`, and "CI/CD Pipeline" becomes
`["ci", "cd", "pipeline"]`.
---
### 7. No changes to `Matcher`, `scorer`, or output format
**Decision:** `_bm25_rrf_similarity()` returns `dict[str, dict[str, object]]` in the
identical shape to `_per_skill_similarity()`:
```python
{
"Python": {
"score": 0.0,
"best_match": None,
"rationale": "BM25: no token overlap with any candidate competence.",
},
"Machine Learning": {
"score": 0.0,
"best_match": None,
"rationale": "BM25: no token overlap with any candidate competence.",
},
}
```
The `Matcher.match()` loop, the `ScoredCapacity` model, the `compute_overall()` scorer,
and all MCP tool output formatting remain unchanged.
---
### 8. LLM Auto-Tagging: ephemeral competence expansion
**Decision:** When `use_auto_tagging = true`, the `AutoTagger` issues a single LLM
chat-completion call per candidate per scoring invocation to identify which required
competences are already covered by the candidate's existing competences (via synonym,
abbreviation, or cross-language equivalence). The response is a JSON list of canonical
required-competence names to add to the candidate's working list. The expanded list is
used only for that BM25 scoring call and is **never persisted**.
**Rationale:**
- Auto-tagging directly addresses BM25's known limitation with lexical disjoint synonyms
(e.g., "ML" for "Machine Learning", "Softwarearchitektur" for "Software Architecture").
- Making the expansion ephemeral keeps the data model clean and audit-friendly.
- Restricting additions to "names already covered by the candidate's existing entries"
means the LLM cannot invent skills — it can only surface canonical equivalents.
- LLM call cost is bounded: one call per candidate per `compute_competence_similarity()`
invocation, with a compact prompt (required list + candidate list).
**LLM prompt schema:**
```
System: You are a skill-taxonomy assistant. Given a list of REQUIRED competences and a
candidate's EXISTING competences, output ONLY the canonical names from REQUIRED
that are already covered by one or more entries in EXISTING (via synonym,
abbreviation, or cross-language equivalence). Do not invent new skills.
User: REQUIRED: ["Machine Learning", "Python", "CI/CD"]
EXISTING: ["ML", "Python", "Jenkins", "Continuous Integration"]
Response (JSON): {"additions": ["Machine Learning", "CI/CD"]}
```
**Azure OpenAI `response_format`:** `{"type": "json_object"}` is used to enforce
structured JSON output, eliminating parsing failures.
**Error handling:** If the LLM call fails (timeout, API error, malformed JSON), the
`AutoTagger` returns the original unmodified candidate list. BM25 scoring continues with
no expansion — graceful degradation.
---
## Data Flow
```
find_matching_capacities(...) / find_matching_tasks(...)
└─► Matcher.match(capacities, requirements)
├─ [use_bm25_search=True]
│ Build ONE Bm25Index over all filtered candidates' competences (global corpus)
│ → global_bm25_index (passed into every compute_competence_similarity call)
└─► for each candidate cap in filtered:
SimilarityEngine.compute_competence_similarity(
required=requirements.competences,
candidate=cap.competences,
global_index=global_bm25_index,
)
├─ [use_bm25_search=False]
│ _per_skill_similarity() / _aggregate_similarity()
│ (embedding-based, unchanged)
└─ [use_bm25_search=True]
├─ [use_auto_tagging=True]
│ AutoTagger.expand_competences(
│ required=required,
│ existing=candidate,
│ )
│ → expanded_candidate (ephemeral, not persisted)
└─► _bm25_rrf_similarity(required, expanded_candidate, global_index)
├─► global_index.rank(query=required_comp)
│ → filter to candidate_set
│ → list[(candidate_text, bm25_score)]
│ (falls back to bm25_rank_competences() if no global_index)
└─► reciprocal_rank_fusion(ranked_lists)
→ dict[required_comp → rrf_score]
```
---
## Module Structure
```
src/teamlandkarte_mcp/matching/
├── auto_tagger.py # AutoTagger class + expand_competences()
├── bm25.py # Bm25Index dataclass + bm25_rank_competences()
├── rrf.py # reciprocal_rank_fusion()
├── matcher.py # unchanged
├── scorer.py # unchanged
└── similarity.py # extended: _bm25_rrf_similarity(), auto-tag branch, flag routing
src/teamlandkarte_mcp/azure/
└── openai_client.py # extended: chat_completion(), chat_deployment, llm_api_key
```
---
## `bm25.py` — Public Interface
```python
@dataclass
class Bm25Index:
"""In-memory BM25 index over a corpus of competence strings.
Intended to be built once over the union of all candidates' competences
(the global corpus) so that IDF values are meaningful. Each per-candidate
scoring call queries the global index and filters results to that candidate's
own competence set.
Args:
corpus: List of competence strings to index.
Notes:
Tokenization lowercases and splits on non-word characters.
An empty corpus produces an index that scores all queries as 0.0.
Raw BM25 scores are clamped to ``max(0.0, raw)`` — ubiquitous terms
whose IDF is negative simply produce 0.0.
"""
corpus: list[str]
def rank(self, query: str) -> list[tuple[str, float]]:
"""Rank corpus documents for a single query string.
Args:
query: Required competence text used as the BM25 query.
Returns:
List of (candidate_text, bm25_score) sorted by score descending.
Entries with score 0.0 are included (caller decides cutoff).
Scores are clamped to max(0.0, raw); negative IDF yields 0.0.
"""
...
def bm25_rank_competences(
required: str,
candidates: list[str],
) -> list[tuple[str, float]]:
"""Rank candidate competences against a single required competence.
Convenience wrapper: builds a temporary ``Bm25Index`` and ranks the query.
Suitable for standalone use and tests; in production the global index built
by ``Matcher`` is preferred (see ``Bm25Index`` notes on IDF correctness).
Args:
required: The required competence text (BM25 query).
candidates: Candidate competence strings to rank.
Returns:
Ranked list of (candidate_text, bm25_score), descending by score.
"""
...
```
---
## `rrf.py` — Public Interface
```python
def reciprocal_rank_fusion(
ranked_lists: list[list[tuple[str, float]]],
*,
k: int = 60,
) -> dict[str, float]:
"""Fuse multiple ranked lists using Reciprocal Rank Fusion.
For each candidate present in any list, the fused score is:
score(c) = Σ_i 1 / (k + rank_i(c))
where rank_i(c) is the 1-based rank of candidate c in list i.
Candidates with zero BM25 score in ALL contributing lists receive
a fused score of 0.0 (no token overlap signal).
Scores are normalized so the top-ranked candidate receives 1.0.
Args:
ranked_lists: One or more ranked lists of (text, score) pairs.
Each list MUST be sorted by score descending.
Lists with all-zero scores are skipped (no signal).
k: RRF smoothing constant (default 60, Cormack et al. 2009).
Higher k flattens rank differences; lower k amplifies them.
Returns:
Mapping of candidate text → normalized fused score in [0.0, 1.0].
Returns empty dict if all input lists are empty or all-zero.
"""
...
```
---
## `auto_tagger.py` — Public Interface
```python
class AutoTagger:
"""LLM-based competence expander for BM25 pre-processing.
Calls an Azure OpenAI chat model to identify which required competences
are already covered by the candidate's existing competences (via synonym,
abbreviation, or cross-language equivalence) and returns the canonical
required-competence names as additions.
All LLM calls use ``response_format={"type": "json_object"}`` to ensure
parseable output. On any LLM or parsing error, the original candidate list
is returned unchanged (graceful degradation).
Args:
client: An ``AzureOpenAIClient`` instance configured with a chat
deployment (``chat_deployment`` and ``llm_api_key``).
"""
def __init__(self, client: AzureOpenAIClient) -> None: ...
def expand_competences(
self,
required: list[str],
existing: list[str],
) -> list[str]:
"""Expand a candidate's competence list with covered required terms.
Sends a prompt to the LLM asking: given ``required`` competences and
the candidate's ``existing`` competences, which required competences
are already implicitly covered by the existing entries?
The returned list is ``existing + additions`` where ``additions``
are the canonical required-competence names identified by the LLM.
The result is **ephemeral**: it is used only for the current BM25
scoring call and is never written back to the database.
Args:
required: List of required competence names (BM25 query terms).
existing: Candidate's current competence list.
Returns:
Combined list ``existing + additions``. If the LLM call fails or
returns no additions, returns ``existing`` unchanged.
"""
...
```
---
## Configuration
### `SimilarityConfig` (updated)
```python
@dataclass(frozen=True)
class SimilarityConfig:
"""Similarity engine configuration."""
embedding_model: str = "text-embedding-3-large"
embedding_dimensions: int = 3072
strategy: str = "per_skill"
use_bm25_search: bool = False
"""When True, BM25 + RRF replaces embedding-based competence similarity.
Role similarity is always embedding-based.
Default: False (preserves existing behavior).
"""
use_auto_tagging: bool = False
"""When True (and use_bm25_search is also True), calls the LLM AutoTagger
before each BM25 scoring call to expand the candidate's competence list
with canonical equivalents of required competences it already covers.
The expansion is ephemeral and never persisted.
Default: False.
"""
```
### `AzureOpenAIConfig` (updated)
```python
@dataclass(frozen=True)
class AzureOpenAIConfig:
"""Azure OpenAI client configuration."""
# ... existing fields ...
chat_deployment: str = ""
"""Azure deployment name of the chat model used for auto-tagging
(e.g. ``"gpt-4.1"``). Required only when ``use_auto_tagging = true``.
"""
llm_api_key: str = ""
"""Azure OpenAI API key for the chat deployment. Read from the
``AZURE_OPENAI_LLM_API_KEY`` environment variable.
Required when ``use_auto_tagging = true``; raises ``ConfigurationError``
at startup if absent.
"""
```
### Config parser (`config.toml` keys)
```toml
[matching.similarity]
strategy = "per_skill"
use_bm25_search = false # new — enables BM25+RRF competence matching
use_auto_tagging = false # new — enables LLM pre-expansion before BM25
[azure_openai]
# ... existing keys ...
chat_deployment = "gpt-4.1" # new — required when use_auto_tagging = true
```
### Environment variables
| Variable | Required when | Description |
|----------|--------------|-------------|
| `AZURE_OPENAI_LLM_API_KEY` | `use_auto_tagging = true` | API key for the Azure OpenAI chat deployment used in auto-tagging. Server raises `ConfigurationError` at startup if absent and `use_auto_tagging = true`. |
---
## BM25 Score Behavior
### Example: Query "Python" against candidate competences
| Candidate | Tokens | BM25 Score | RRF Score (k=60) |
|-----------|--------|------------|-----------------|
| Python | ["python"] | ~3.5 | 1.000 |
| Python (advanced) | ["python", "advanced"] | ~2.1 | 0.984 |
| JavaScript | ["javascript"] | 0.0 | **0.000** |
| TypeScript | ["typescript"] | 0.0 | **0.000** |
| Node.js | ["node", "js"] | 0.0 | **0.000** |
Candidates without any shared token receive **0.0** — the false-positive problem is
eliminated.
### Example: Query "Machine Learning" against candidate competences
| Candidate | Tokens | BM25 Score | RRF Score |
|-----------|--------|------------|-----------|
| Machine Learning | ["machine", "learning"] | ~4.0 | 1.000 |
| Deep Learning | ["deep", "learning"] | ~1.8 | 0.984 |
| Machine Learning (Python) | ["machine", "learning", "python"] | ~3.2 | 0.492 |
| JavaScript | ["javascript"] | 0.0 | **0.000** |
---
## Known Limitations
### 1. Conceptual synonyms with zero token overlap score 0.0
BM25 requires at least one shared token between query and document. Skill pairs that are
semantically equivalent but lexically disjoint will always score 0.0:
| Required competence | Candidate competence | Shared tokens | BM25 score |
|---------------------|----------------------|---------------|------------|
| Data Science | Machine Learning | none | **0.000** |
| Data Science | Artificial Intelligence | none | **0.000** |
| Frontend Development | UI Engineering | none | **0.000** |
| Agile | Scrum | none | **0.000** |
| Softwarearchitektur | Software Architecture | none | **0.000** |
These are **false negatives** — real skill matches that BM25 misses entirely. Embeddings
handle all of these correctly via semantic similarity.
**Consequence:** BM25 mode is best suited to datasets where competence names are
**canonical, consistent, and in a single language**. If profiles mix German and English
skill names, or use informal abbreviations ("ML" for "Machine Learning"), BM25 will
produce false negatives that embeddings would have caught.
**Mitigation:** Enable `use_auto_tagging = true`. The LLM will bridge synonym and
cross-language gaps by pre-expanding the candidate list before BM25 scoring. This
addresses limitations 1, 2, and 3 at the cost of one LLM call per candidate per scoring
invocation.
---
### 2. Abbreviations and acronyms
Short forms that do not share tokens with the expanded form score 0.0:
| Required | Candidate | BM25 score |
|----------|-----------|------------|
| Machine Learning | ML | **0.000** |
| Continuous Integration | CI | **0.000** |
| Natural Language Processing | NLP | **0.000** |
Partial exception: "CI/CD" tokenizes to `["ci", "cd"]`, so a query "CI/CD Pipeline"
would find partial overlap with a candidate "CI/CD", but not with "Continuous
Integration/Continuous Delivery".
**Mitigation:** `use_auto_tagging = true` (see limitation 1).
---
### 3. Cross-language skill names
A required competence in German and a candidate competence in English (or vice versa)
score 0.0 even when they describe the same skill:
| Required | Candidate | BM25 score |
|----------|-----------|------------|
| Softwarearchitektur | Software Architecture | **0.000** |
| Maschinelles Lernen | Machine Learning | **0.000** |
| Datenwissenschaft | Data Science | **0.000** |
This is a significant limitation for German-language skill databases. Embeddings handle
cross-language pairs well; BM25 does not.
**Mitigation:** `use_auto_tagging = true` (see limitation 1).
---
### 4. RRF score compression near 1.0
Because the normalization formula is `(k+1) / (k+rank)` with `k=60`, ranks 15 all
compress into the range `[0.871, 1.0]`. There is little score differentiation between
the top-ranked candidates. This is intentional (rank differences matter less than
presence/absence of token overlap) but means BM25 mode should not be used to produce
fine-grained competence similarity scores — only to separate matching (score > 0) from
non-matching (score = 0) candidates.
---
### 5. Small per-person corpus causes IDF pathology — mitigated by global index
BM25's IDF term rewards tokens that are rare across the corpus. When the BM25 index was
built per candidate (N = that person's competence count ≈ 530), any term appearing in
most documents had IDF ≤ 0, causing `rank_bm25` to return negative raw scores. After
`max(0.0, raw)` clamping, all candidates scored 0.0 — making the BM25 signal useless.
**Fix (implemented):** The `Bm25Index` is now built once over the **union of all filtered
candidates' competences** in `Matcher.match()` before the per-candidate loop (see Key
Decision #2). This makes N equal to the total unique competence count across all
candidates, giving IDF proper signal. Per-candidate scoring filters the global results to
only that candidate's competence set.
**Residual limitation:** If the filtered candidate pool itself is very small (e.g., 23
candidates with few unique competences), IDF may still be weak. In practice this is
negligible since matching runs are always invoked against a meaningful pool of candidates.
---
## Test Strategy
### Unit tests
1. **`tests/test_bm25.py`**
- `Bm25Index` with empty corpus → all scores 0.0
- Exact match → score > 0
- No token overlap → score == 0.0
- Partial token overlap (multi-token queries) → partial score
- Case-insensitive matching ("Python" == "python")
- Tokenization edge cases: hyphens, parentheses, slashes
2. **`tests/test_rrf.py`**
- Single list, single candidate → score 1.0
- Single list, multiple candidates → scores normalized to [0, 1]
- All-zero list → returns empty / all zeros
- `k` parameter effect on score distribution
3. **`tests/test_auto_tagger.py`** (new)
- LLM returns valid additions → `expand_competences()` appends them to existing list
- LLM returns empty additions → original list returned unchanged
- LLM call raises exception → original list returned unchanged (graceful degradation)
- LLM returns malformed JSON → original list returned unchanged
- Additions from LLM that are NOT in `required` are ignored
- `expand_competences()` never duplicates entries already in `existing`
- All tests use a mocked `AzureOpenAIClient` (no real LLM calls)
4. **`tests/test_similarity.py`** (extended)
- `compute_competence_similarity()` with `use_bm25_search=True`
- Required "Python", candidate ["JavaScript"] → score 0.0
- Required "Python", candidate ["Python"] → score 1.0
- Required "Python" + "Machine Learning", candidate ["JavaScript", "TypeScript"]
→ both scores 0.0, overall competence_score 0.0
- With `use_bm25_search=True, use_auto_tagging=True`: mocked `AutoTagger` injects
"Machine Learning" into candidate ["ML"] → score > 0
### Integration test
5. **`tests/test_matcher_integration.py`** (extended)
- With `use_bm25_search=True`: candidate with exact skills scores higher than
candidate with only semantically adjacent skills.
- The false-positive scenario (Python/ML vs JavaScript/TypeScript) is explicitly
tested.
- With `use_bm25_search=True, use_auto_tagging=True` (mocked LLM): candidate with
"ML" matches required "Machine Learning" after auto-tag expansion.