Bahn: aisupport, Analyse-O2C-C2S, awesome-bahn-mcp-servers, beam-mcp,
Confluence_Bot, db-planet-mcp-server, O2C-Harness, project-audit,
Projekt-KIQ-HP, teamlandkarte-mcp
Dhive: Jury-Voting
Privat: CV, NoteGraph (NOTE: NoteGraph needs complete redo after consolidation)
Shared: AI-Orchestrator, OrgMyLife, power_skills_and_more
Shared/references: symphony (read-only)
Bahn repos remain available as independent remotes - this monorepo
pulls them in via subtree, the originals are untouched.
18 KiB
Tasks: Replace Heuristics with Azure OpenAI
Change:
replace-heuristics-with-azure-openaiStatus: Approved / Implemented
Overview
Replace legacy client-side sampling/heuristic logic with direct Azure OpenAI integration:
- Embeddings (
text-embedding-3-large) for competence/role similarity - LLM (
gpt-4.1) for requirements extraction and validation - Persistent SQLite embedding cache
- Two configurable similarity strategies
Removed: migration note about database.toml rename (now a completed internal detail).
Estimated Effort
Total: 28-35 hours (1 week full-time or 2 weeks part-time)
Task Breakdown
Phase 1: Configuration & Infrastructure (3-4 hours)
- 1.1 Rename
database.toml→config.toml - 1.2 Rename
database.toml.example→config.toml.example - 1.3 Update all file references in code (
README.md,docs/,src/config.py, etc.) - 1.4 Extend
src/config.py: AddAzureOpenAIConfigmodelendpoint,api_version_embeddings,api_version_llmembeddings.model,embeddings.deployment(deployment must equal model name)llm.model,llm.deployment(deployment must equal model name),llm.temperature,llm.max_tokens
- 1.5 Extend
src/config.py: AddEmbeddingCacheConfigmodelenabled,db_path,ttl_days
- 1.6 Extend
src/config.py: Addmatching.similarity.strategyfield ("per_skill"or"aggregate") - 1.7 Update
config.toml.examplewith new[azure_openai],[embedding_cache],[matching.similarity]sections - 1.8 Load
AZURE_OPENAI_EMBEDDING_API_KEYfrom environment inload_config() - 1.9 Load
AZURE_OPENAI_LLM_API_KEYfrom environment inload_config() - 1.10 Add
openaipackage topyproject.tomldependencies - 1.11 Run
uv syncto install dependencies - 1.12 Add
config.tomlto.gitignore(if not already there) - 1.13 Configure pytest integration marker in
pytest.ini(if not already present):integration: marks tests that call real Azure OpenAI API (deselect with '-m "not integration"')
Phase 2: Embedding Cache (2-3 hours)
- 2.1 Create
src/teamlandkarte_mcp/cache/embedding_cache.py - 2.2 Implement
EmbeddingCacheclass:__init__(db_path: str, ttl_days: int)- Create SQLite database and schema if not exists:
CREATE TABLE IF NOT EXISTS embeddings ( id INTEGER PRIMARY KEY AUTOINCREMENT, text TEXT UNIQUE NOT NULL, model TEXT NOT NULL, embedding BLOB NOT NULL, created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP ) CREATE INDEX IF NOT EXISTS idx_text_model ON embeddings(text, model)
- 2.3 Implement
EmbeddingCache.get(text: str, model: str) -> Optional[list[float]]- Query by text + model
- Check if
created_atis within TTL - Deserialize BLOB to
list[float] - Return
Noneif missing or expired
- 2.4 Implement
EmbeddingCache.put(text: str, model: str, embedding: list[float]) -> None- Serialize embedding as UTF-8 bytes of JSON:
json.dumps(embedding).encode("utf-8") - INSERT OR REPLACE into database
- Serialize embedding as UTF-8 bytes of JSON:
- 2.5 Implement
EmbeddingCache.cleanup_expired() -> int- DELETE embeddings older than TTL
- Return count of deleted rows
- 2.6 Add graceful handling for concurrent access (SQLite locks)
- 2.7 Create
tests/test_embedding_cache.py- Test cache hit/miss
- Test TTL expiration
- Test cleanup
- Test concurrent access
- 2.8 Add
embeddings_cache.dbto.gitignore
Phase 3: Azure OpenAI Client (3-4 hours)
- 3.1 Create
src/teamlandkarte_mcp/azure/__init__.py - 3.2 Create
src/teamlandkarte_mcp/azure/openai_client.py - 3.3 Define
AzureAPIError(RuntimeError)exception class - 3.4 Implement
AzureOpenAIClientclass:__init__(config: AzureOpenAIConfig, embedding_cache: Optional[EmbeddingCache])- Initialize two
AzureOpenAIclients (embeddings + LLM) usingAzureKeyCredential
- 3.5 Implement
AzureOpenAIClient.get_embedding(text: str) -> list[float]- Check cache first (if enabled)
- If cache miss: call Azure API
- Store in cache (if enabled)
- Raise
AzureAPIErroron failure (no fallback)
- 3.6 Implement
AzureOpenAIClient.get_embeddings_batch(texts: list[str]) -> list[list[float]]- Call
get_embedding()for each text individually (per user requirement #3) - Return list of embeddings in same order
- Call
- 3.7 Implement
AzureOpenAIClient.chat_completion(system: str, user: str, response_format: Optional[dict] = None) -> str- Call Azure OpenAI chat completions API
- Use
gpt-4.1model with configured temperature/max_tokens - Support
response_format={"type": "json_object"}for structured JSON - Implement retry logic (3 attempts, exponential backoff)
- Raise
AzureAPIErroron final failure (no fallback)
- 3.8 Add logging for API calls (cache hit/miss, request/response sizes)
- 3.9 Create
tests/test_azure_client.py- Mock Azure API with
unittest.mock - Test embedding retrieval with cache hit/miss
- Test LLM completion
- Test error handling and retries
- Mock Azure API with
- 3.10 Create
tests/test_azure_client_integration.py(marked@pytest.mark.integration)- Real API calls (requires valid credentials)
- Test embedding generation
- Test LLM structured JSON response
Phase 4: Similarity Engine (4-5 hours)
- 4.1 Create
src/teamlandkarte_mcp/matching/similarity.py - 4.2 Implement
cosine_similarity(vec_a: list[float], vec_b: list[float]) -> float- Compute dot product and magnitudes
- Return cosine similarity in [0.0, 1.0]
- Handle zero vectors gracefully
- 4.3 Implement
SimilarityEngineclass:__init__(client: AzureOpenAIClient, strategy: str = "per_skill")- Validate strategy is
"per_skill"or"aggregate"
- 4.4 Implement
SimilarityEngine._per_skill_similarity(required: list[str], candidate: list[str]) -> dict[str, dict[str, Any]]- For each required skill:
- Get embedding
- Compare with all candidate skill embeddings
- Find best match (highest cosine similarity)
- Return
{required: {best_match: str|None, score: float, rationale: str}}
- Rationale example:
"Cosine similarity: 0.92 (best match: 'React.js')"
- For each required skill:
- 4.5 Implement
SimilarityEngine._aggregate_similarity(required: list[str], candidate: list[str]) -> dict[str, dict[str, Any]]- Compute average embedding of all required skills
- Compute average embedding of all candidate skills
- Compute cosine similarity between averages
- Return same format as
_per_skill_similarity(each required skill gets same aggregate score) - Rationale example:
"Aggregate embedding similarity: 0.85"
- 4.6 Implement
SimilarityEngine.compute_competence_similarity(required: list[str], candidate: list[str]) -> dict[str, dict[str, Any]]- Route to
_per_skill_similarity()or_aggregate_similarity()based on configured strategy - Maintain same output format as current
TaskAnalyzer.semantic_competence_similarity() - Guarantee
rationaleis always present (even when best_match is None) - Edge cases:
required == []→ return{}candidate == []→ return per-required entry withbest_match=None,score=0.0, and rationale
- Route to
- 4.7 Implement
SimilarityEngine.compute_role_similarity(required_role: str, candidate_role: str) -> float- Edge cases:
- If either role is empty/None → return 0.0
- If either role is "(unknown)" (case-insensitive, trimmed) → return 0.0
- Otherwise compute cosine similarity of role embeddings
- Edge cases:
- 4.8 Create
tests/test_similarity.py- Test
cosine_similarity()with known vectors - Mock embeddings and test both strategies
- Test edge cases (empty lists, identical skills, no matches)
- Test
- 4.9 Create
tests/test_similarity_integration.py(marked@pytest.mark.integration)- Compare both strategies with real embeddings
- Validate scores are reasonable (e.g., "Python" vs "Python 3" → high similarity)
Phase 5: TaskAnalyzer Refactoring (3-4 hours)
- 5.1 Update
src/teamlandkarte_mcp/matching/task_analyzer.py - 5.2 Change
TaskAnalyzer.__init__signature:- Remove:
__init__(self, mcp: FastMCP) - Add:
__init__(self, azure_client: AzureOpenAIClient)
- Remove:
- 5.3 Delete
_sample_json()method completely - 5.4 Delete
_heuristic_competences_from_text()static method completely - 5.5 Refactor
extract_ranked_roles(description: str, limit: int = 5) -> list[RankedRole]:- Remove all heuristic fallback code
- Build prompt for structured JSON response:
system: "Extract a ranked list of possible roles from a task description. Return STRICT JSON: {roles:[{rank:int, role:str, rationale:str}]}" user: "Limit: {limit}\n\nTask description:\n{description}" - Call
self._client.chat_completion(system, user, response_format={"type": "json_object"}) - Parse JSON response and build
list[RankedRole] - Raise
AzureAPIErroron failure (no fallback)
- 5.6 Refactor
_extract_requirements_from_description(description: str) -> ExtractedRequirements:- Remove all heuristic fallback code
- Build prompt for structured JSON:
system: "Extract structured requirements from a task description. Return STRICT JSON with keys: competences (list[str]), date_start (YYYY-MM-DD|null), date_end (YYYY-MM-DD|null), roles ([{rank, role, rationale}])." user: "Task description:\n{description}" - Call
self._client.chat_completion(system, user, response_format={"type": "json_object"}) - Parse JSON and build
ExtractedRequirements - Raise
AzureAPIErroron failure (no fallback)
- 5.7 Refactor
apply_requirement_update(current: Requirements, change_description: str) -> Requirements:- Remove all heuristic fallback code
- Build prompt with current requirements + change request
- Call
self._client.chat_completion(system, user, response_format={"type": "json_object"}) - Parse JSON and build updated
Requirements - Raise
AzureAPIErroron failure (no fallback)
- 5.8 Delete
semantic_competence_similarity()method completely (replaced bySimilarityEngine) - 5.9 Update
ExtractedRequirementsdataclass docstring (no longer mentions heuristics) - 5.10 Update
TaskAnalyzerclass docstring (remove FastMCP sampling references, mention Azure OpenAI) - 5.11 Remove unused imports (
FastMCP,random,asyncioif no longer needed) - 5.12 Update all
tests/test_task_analyzer.py(or similar):- Mock
AzureOpenAIClientinstead ofFastMCP - Remove heuristic fallback tests
- Add tests for Azure API error propagation
- Mock
- 5.13 Update
AnalysisErrorsemantics (or rename/remove if no longer needed):- Previously: heuristic fallback masked missing sampling
- Now: Azure failures must surface as exceptions (no heuristics)
Phase 6: Matcher & Scorer Integration (2-3 hours)
- 6.1 Identify where
Matcheror scoring logic callsTaskAnalyzer.semantic_competence_similarity() - 6.2 Update
Matcher.__init__to acceptSimilarityEngineas parameter - 6.3 Replace
analyzer.semantic_competence_similarity(required, candidate)withsimilarity_engine.compute_competence_similarity(required, candidate) - 6.4 Update role matching to use
similarity_engine.compute_role_similarity(required_role, candidate_role) - 6.5 Ensure scorer aggregation logic works with new similarity output format (should be identical)
- 6.6 Update
tests/test_scorer.py(and/or addtests/test_matcher.pyif introduced):- Mock
SimilarityEngineinstead ofTaskAnalyzer.semantic_competence_similarity - Test that both similarity strategies produce valid scores
- Mock
- 6.7 Update matcher/scorer tests to reflect new edge-case contracts:
- role similarity returns 0.0 for empty/"(unknown)" roles
- competence similarity returns 0.0 when candidate list is empty
Phase 7: Server Initialization (2 hours)
- 7.1 Update
build_server()insrc/teamlandkarte_mcp/mcp_server.py - 7.2 Update server build to use top-level embedding cache config (not nested under azure):
cfg.embedding_cache.enabled/db_path/ttl_days
- 7.3 Initialize
EmbeddingCache(if enabled) usingcfg.embedding_cache - 7.4 Initialize
AzureOpenAIClientusingcfg.azure_openai - 7.5 Initialize
SimilarityEngineusingcfg.matching.similarity.strategy - 7.6 Update
TaskAnalyzerinstantiation - 7.7 Update
Matcherinstantiation - 7.8 Add startup logging (stderr):
- Log Azure OpenAI endpoint
- Log embedding cache status (enabled/disabled, path)
- Log similarity strategy
- 7.9 Add graceful error handling if Azure credentials missing:
- Log error to stderr
- Exit with clear message: "Azure OpenAI credentials not found in environment"
Phase 8: Cost Tracking & Monitoring (2 hours)
- 8.1 Create
src/teamlandkarte_mcp/azure/cost_tracker.py - 8.2 Implement
CostTrackerclass:- Class-level constants for pricing
log_embedding_request(text: str, cached: bool)log_llm_request(input_tokens: int, output_tokens: int)get_session_costs() -> dict(return totals)
- 8.3 Integrate
CostTrackerintoAzureOpenAIClient:- Track embedding requests (cache hit/miss)
- Track LLM requests (token counts from API response)
- 8.4 Add periodic cost logging to stderr:
- Every 10 API calls or every 5 minutes
- Log:
"Azure OpenAI costs this session: embeddings=$X.XX, llm=$Y.YY, total=$Z.ZZ"
Phase 9: Documentation (2-3 hours)
- 9.1 Update
README.md:- Add "Azure OpenAI Setup" section with prerequisites
- Document environment variables (
AZURE_OPENAI_EMBEDDING_API_KEY,AZURE_OPENAI_LLM_API_KEY) - Document
config.tomlAzure sections - Add cost estimation table
- Add embedding cache explanation
- Add similarity strategy comparison
- 9.2 Add migration guide to
README.md:- Step 1: Rename
database.toml→config.toml - Step 2: Add Azure sections to config
- Step 3: Add Azure API keys to
.env - Step 4: Run
uv sync - Step 5: Start server (embedding cache will be created automatically)
- Step 1: Rename
- 9.3 Update
docs/troubleshooting.md:- Add section: "Azure OpenAI API Errors"
- Document common errors (rate limits, invalid credentials, network issues)
- Document
AzureAPIErrorand how to debug
- 9.4 Update
config.toml.example:- Add comprehensive comments for all Azure sections
- Document both similarity strategies with examples
- Document embedding cache TTL trade-offs
- 9.5 Update OpenSpec architecture docs:
openspec/changes/add-capacity-matching-mcp-server/architecture.mdopenspec/changes/add-capacity-matching-mcp-server/design.md- Update diagrams to show Azure OpenAI dependency
- Update component descriptions
- 9.6 Create
docs/azure_openai_setup.md(detailed setup guide):- Prerequisites (Azure subscription, OpenAI resource)
- API key generation steps
- Cost monitoring in Azure portal
- Troubleshooting deployment/model availability
Phase 10: Testing & Validation (3-4 hours)
- 10.1 Run fast unit tests in CI mode:
pytest -m "not integration" -q - 10.2 Run integration tests with real Azure API:
pytest -m integration - 10.3 Test embedding cache behavior:
- Start server, run search → check cache DB created
- Restart server, run same search → verify cache hit
- Wait past TTL, run cleanup → verify expired entries removed
- 10.4 Test both similarity strategies:
- Set
matching.similarity.strategy = "per_skill"→ run search - Set
matching.similarity.strategy = "aggregate"→ run search - Compare results and validate scoring differences
- Set
- 10.5 Test cost tracking:
- Run several searches
- Check stderr for cost logs
- Validate cost estimates are reasonable
- 10.6 Test error handling:
- Comment out Azure API keys → verify graceful startup error
- Mock API failure → verify
AzureAPIErrorpropagates correctly - Mock rate limit error → verify retry logic works
- 10.7 Performance benchmark:
- Cold cache: measure time for first search (100 candidates)
- Warm cache: measure time for repeated search
- Document speedup ratio
- 10.8 Manual validation in Cherry Studio:
- Browse tasks
- Extract requirements from complex task description
- Run matching workflow
- Validate similarity scores look reasonable
- Test guided capture with Azure LLM
- Test update_requirements
Phase 11: Cleanup & Final Review (2-3 hours)
- 11.1 Final sweep: ensure there are no remaining FastMCP sampling references
- 11.2 Final sweep: ensure no heuristic fallback paths remain
- 11.3 Tighten type hints, docstrings, and error semantics
- 11.4 Run
ruff check(viauv) and fix all issues - 11.5 Run
mypy(viauv) and fix all issues - 11.6 Final code review: validate all modules, remove stray prints/debug logs
- 11.7 Git commit review: ensure no secrets (config/.env/cache DB) are tracked
- 11.7a Cleanup: remove obsolete
scripts/trino_smoke_check.py - 11.7b Dependency hygiene: keep
trinoas a required runtime dependency - 11.8 Merge to main and tag release
Risk Mitigation
Risk: Azure API downtime → System unusable
Mitigation: Document SLA expectations; consider adding a simple health check tool/endpoint and a clear user-facing error message (AzureAPIError).
Risk: Unexpected high costs
Mitigation: Enable embedding cache (default); monitor costs weekly; set Azure budget alerts.
Risk: Embedding cache grows too large
Mitigation: Run periodic cleanup; set reasonable TTL (default 30 days); ensure cache DB files stay ignored.
Risk: Secrets accidentally committed (config.toml, .env, caches)
Mitigation: Keep config.toml and .env ignored; review git status before commit; run a secret scan locally (e.g. git grep for key patterns) before merging.
Risk: Migration breaks existing deployments
Mitigation: Provide clear migration guide; test migration path manually; run integration tests against Azure in at least one real environment.
Success Criteria
- Unit tests pass (non-integration)
- Integration tests pass (
pytest -m integration) in an environment with valid Azure credentials - Server starts successfully with Azure credentials
- Matching results show improved semantic understanding vs. heuristics
- Embedding cache demonstrates significant performance improvement on repeated searches
- Cost tracking shows accurate estimates
- Documentation complete and clear
- Repo clean: no secrets or local config files tracked (verify before merge)
Total estimated effort: 28-35 hours
Priority: Medium-High (significant quality improvement, but breaking change)
Approval status: ✅ Approved