Bahn: aisupport, Analyse-O2C-C2S, awesome-bahn-mcp-servers, beam-mcp,
Confluence_Bot, db-planet-mcp-server, O2C-Harness, project-audit,
Projekt-KIQ-HP, teamlandkarte-mcp
Dhive: Jury-Voting
Privat: CV, NoteGraph (NOTE: NoteGraph needs complete redo after consolidation)
Shared: AI-Orchestrator, OrgMyLife, power_skills_and_more
Shared/references: symphony (read-only)
Bahn repos remain available as independent remotes - this monorepo
pulls them in via subtree, the originals are untouched.
13 KiB
Spec Delta: BM25 + RRF Competence Matching
ADDED Requirements
Requirement: BM25 + RRF optional competence matching strategy
The system SHALL support an optional BM25-based competence matching strategy controlled
by matching.similarity.use_bm25_search in config.toml. When enabled, BM25 + RRF
replaces embedding-based competence similarity for find_matching_capacities and
find_matching_tasks. Role similarity SHALL always remain embedding-based.
Scenario: Candidate without matching skills receives zero competence score
- GIVEN
use_bm25_search = trueinconfig.toml - AND a task requires competences
["Python", "Machine Learning"] - AND a candidate has competences
["JavaScript", "TypeScript", "Node.js", "Vue.js"] - WHEN the system computes
competence_scorefor that candidate - THEN
competence_score == 0.0 - AND the candidate is categorized as "Irrelevant" or "Low" based on
overall_score
Scenario: Candidate with exact matching skills receives high competence score
- GIVEN
use_bm25_search = trueinconfig.toml - AND a task requires competences
["Python", "Machine Learning"] - AND a candidate has competences
["Python", "Machine Learning", "Pandas"] - WHEN the system computes
competence_scorefor that candidate - THEN
competence_score > 0.8 - AND the candidate scores significantly higher than a candidate with no token overlap
Scenario: BM25 mode is disabled by default
- GIVEN no
use_bm25_searchkey inconfig.toml(oruse_bm25_search = false) - WHEN the system computes competence similarity
- THEN the existing embedding-based strategy is used
- AND behavior is bit-for-bit identical to the pre-change implementation
Requirement: BM25 index ranks candidates by token overlap per required competence
The Bm25Index SHALL build an in-memory BM25 index over a corpus of competence strings
and rank them for each required competence query. The index is built once over the global
candidate pool per matching run (see "Global BM25 index" requirement below) and is queried
per candidate with results filtered to that candidate's competence set.
Scenario: BM25 index ranks exact token match highest
- GIVEN a candidate corpus
["Python", "JavaScript", "TypeScript"] - AND query
"Python" - WHEN the
Bm25Index.rank()method is called - THEN
"Python"is ranked first with the highest BM25 score - AND
"JavaScript"and"TypeScript"have BM25 score 0.0 (no shared tokens)
Scenario: BM25 index handles empty corpus without error
- GIVEN a candidate with no competences (empty corpus)
- WHEN the system computes
competence_score - THEN
competence_score == 0.0for every required competence - AND no exception is raised
Scenario: Tokenization is case-insensitive and splits on non-word characters
- GIVEN a candidate corpus
["Progressive Web App (PWA)", "CI/CD Pipeline"] - AND query
"ci cd" - WHEN the
Bm25Index.rank()method is called - THEN
"CI/CD Pipeline"receives a score > 0 - AND
"Progressive Web App (PWA)"receives score 0.0
Requirement: Reciprocal Rank Fusion normalizes BM25 ranks into scores in [0, 1]
The reciprocal_rank_fusion() function SHALL fuse ranked lists into a normalized score
dict. Candidates with zero BM25 score in all contributing lists SHALL receive a fused
score of 0.0. The top-ranked candidate SHALL receive score 1.0.
Scenario: Single-list fusion normalizes top candidate to 1.0
- GIVEN a single ranked list
[("Python", 3.5), ("Python (advanced)", 2.1)] - WHEN
reciprocal_rank_fusion([ranked_list], k=60)is called - THEN
"Python"receives normalized score1.0 - AND
"Python (advanced)"receives a score in(0, 1)
Scenario: All-zero BM25 list produces empty result
- GIVEN a ranked list
[("JavaScript", 0.0), ("TypeScript", 0.0)] - WHEN
reciprocal_rank_fusion([ranked_list])is called - THEN the function returns an empty dict
{} - AND the caller maps this to
score = 0.0andbest_match = None
Requirement: Global BM25 index is built once over all filtered candidates
When use_bm25_search = true, the Matcher SHALL build one Bm25Index over the
deduplicated union of all filtered candidates' competences before the per-candidate
scoring loop. The same index SHALL be passed into every compute_competence_similarity()
call and results SHALL be filtered to the current candidate's competence set. This ensures
BM25 IDF values reflect the full candidate pool rather than a single person's small corpus.
Scenario: Candidate with unique skill scores above zero in global index
- GIVEN
use_bm25_search = trueinconfig.toml - AND two candidates: one with
["Python"]only, another with 6 unrelated skills - AND the task requires
["Python"] - WHEN the system computes
competence_scorefor each candidate - THEN the candidate with
["Python"]receivescompetence_score > 0.0 - AND the candidate with no Python token receives
competence_score == 0.0
Scenario: Ubiquitous term present in every candidate scores 0.0
- GIVEN a global corpus where a token appears in every candidate's competences
- WHEN
Bm25Index.rank()is called with that token as the query - THEN the raw BM25 score is clamped to
max(0.0, raw)(IDF is negative → score is 0.0) - AND no exception is raised
Scenario: Global index filters results to current candidate's competence set
- GIVEN a global index built over candidates A (
["Python", "Django"]) and B (["Java", "Spring"]) - AND the system scores candidate A for required
["Python"] - WHEN
_bm25_rrf_similarity()is called with candidate A's competence set - THEN only
"Python"and"Django"appear in the filtered results - AND
"Java"and"Spring"do NOT appear in candidate A's result
Requirement: Optional LLM-based Auto-Tagging expands candidate competences before BM25
When use_auto_tagging = true (and use_bm25_search = true), the system SHALL call the
AutoTagger before each BM25 scoring call to identify which required competences are
already covered by the candidate's existing entries via synonym, abbreviation, or
cross-language equivalence. The LLM returns canonical required-competence names as
additions. The expanded list is ephemeral — it SHALL be used only for the current BM25
scoring call and SHALL NOT be written back to any database or persisted between calls.
Scenario: Auto-tagging bridges synonym gap so candidate matches required competence
- GIVEN
use_bm25_search = trueanduse_auto_tagging = trueinconfig.toml - AND a task requires competences
["Machine Learning"] - AND a candidate has competences
["ML", "Python"] - AND the LLM identifies
"ML"as covering"Machine Learning" - WHEN the system computes
competence_score - THEN
competence_score > 0.0(auto-tagging added"Machine Learning"to working list) - AND the candidate's stored competences remain
["ML", "Python"](not mutated)
Scenario: Auto-tagging bridges cross-language gap
- GIVEN
use_bm25_search = trueanduse_auto_tagging = trueinconfig.toml - AND a task requires competences
["Software Architecture"] - AND a candidate has competences
["Softwarearchitektur"] - AND the LLM identifies
"Softwarearchitektur"as covering"Software Architecture" - WHEN the system computes
competence_score - THEN
competence_score > 0.0
Scenario: Auto-tagging graceful degradation on LLM failure
- GIVEN
use_bm25_search = trueanduse_auto_tagging = trueinconfig.toml - AND the LLM call raises an exception (timeout, API error, or malformed JSON)
- WHEN the system computes
competence_score - THEN no exception is propagated to the caller
- AND BM25 scoring continues on the original (unexpanded) candidate list
- AND a warning is logged
Scenario: Auto-tagging LLM cannot invent skills not in required list
- GIVEN
use_auto_tagging = true - AND the LLM response includes an addition that is NOT in the
requiredlist - WHEN
AutoTagger.expand_competences()processes the response - THEN the hallucinated addition is silently dropped
- AND only valid required-competence names are appended to the working list
Scenario: Auto-tagging is no-op when use_auto_tagging is false
- GIVEN
use_bm25_search = trueanduse_auto_tagging = false(default) inconfig.toml - WHEN the system computes
competence_score - THEN no LLM call is made
- AND BM25 runs on the raw candidate competence list without expansion
Requirement: use_auto_tagging configuration key with default false
config.toml SHALL support a matching.similarity.use_auto_tagging boolean key.
The default value SHALL be false. The key SHALL be documented in config.toml.example.
use_auto_tagging = true is only effective when use_bm25_search = true.
Scenario: Config key defaults to false when omitted
- GIVEN a
config.tomlthat does not containuse_auto_tagging - WHEN the server starts and loads configuration
- THEN
SimilarityConfig.use_auto_tagging == False - AND no
AutoTaggeris constructed - AND no LLM client is constructed for auto-tagging
Scenario: Config key enables auto-tagging when set to true
- GIVEN
config.tomlcontainsuse_auto_tagging = trueunder[matching.similarity] - AND
chat_deploymentis set andAZURE_OPENAI_LLM_API_KEYis present - WHEN the server starts and loads configuration
- THEN
SimilarityConfig.use_auto_tagging == True - AND an
AutoTaggeris constructed and passed toSimilarityEngine
Requirement: azure_openai.chat_deployment configuration key for auto-tagging
config.toml SHALL support an azure_openai.chat_deployment string key naming the Azure
OpenAI chat model deployment used by the AutoTagger. This key is required when
use_auto_tagging = true. The AZURE_OPENAI_LLM_API_KEY environment variable SHALL
supply the API key for that deployment. The server SHALL raise a ConfigurationError at
startup if use_auto_tagging = true and AZURE_OPENAI_LLM_API_KEY is absent.
Scenario: Server raises ConfigurationError when API key missing with auto-tagging enabled
- GIVEN
use_auto_tagging = trueinconfig.toml - AND the environment variable
AZURE_OPENAI_LLM_API_KEYis not set - WHEN the server starts and loads configuration
- THEN a
ConfigurationErroris raised before any MCP tool is registered - AND the error message indicates the missing environment variable
Scenario: chat_deployment key is read from config when auto-tagging enabled
- GIVEN
config.tomlcontainschat_deployment = "gpt-4.1"under[azure_openai] - AND
use_auto_tagging = trueandAZURE_OPENAI_LLM_API_KEYis set - WHEN the server starts
- THEN the
AzureOpenAIClientfor auto-tagging is initialized with deployment"gpt-4.1"
Requirement: use_bm25_search configuration key with default false
config.toml SHALL support a matching.similarity.use_bm25_search boolean key.
The default value SHALL be false. The key SHALL be documented in config.toml.example.
Scenario: Config key defaults to false when omitted
- GIVEN a
config.tomlthat does not containuse_bm25_search - WHEN the server starts and loads configuration
- THEN
SimilarityConfig.use_bm25_search == False - AND the embedding-based strategy is used unchanged
Scenario: Config key enables BM25 mode when set to true
- GIVEN
config.tomlcontainsuse_bm25_search = trueunder[matching.similarity] - WHEN the server starts and loads configuration
- THEN
SimilarityConfig.use_bm25_search == True - AND
compute_competence_similarity()delegates to_bm25_rrf_similarity()
MODIFIED Requirements
Requirement: Role similarity is always embedding-based regardless of use_bm25_search
Role similarity computation SHALL remain embedding-based even when use_bm25_search = true.
The BM25 strategy applies only to competence similarity.
Scenario: Role scoring unaffected by use_bm25_search flag
- GIVEN
use_bm25_search = trueinconfig.toml - AND a task requiring role "Backend Developer"
- AND a capacity with role "Backend Developer"
- WHEN the system computes
role_score - THEN
role_scoreis computed via embedding similarity (unchanged behavior) - AND embedding cache and Azure OpenAI calls for role embeddings proceed as before
REMOVED Requirements
Requirement: Legacy debug scripts are removed from the repository root
The files debug_similarity_analysis.py, debug_categorization.py,
SEMANTIC_SIMILARITY_ANALYSIS.md, and SCORE_BUG_FIX.md SHALL be removed from the
repository root. They are superseded by the test suite and docs/semantic_similarity_tradeoffs.md.
Scenario: Legacy files do not exist after the change is applied
- GIVEN the change
add-bm25-rrf-competence-matchingis fully applied - WHEN the repository root is listed
- THEN
debug_similarity_analysis.pydoes not exist - AND
debug_categorization.pydoes not exist - AND
SEMANTIC_SIMILARITY_ANALYSIS.mddoes not exist - AND
SCORE_BUG_FIX.mddoes not exist