Files
Orchestrator/bahn/teamlandkarte-mcp/openspec/changes/replace-heuristics-with-azure-openai/proposal.md
T
ankn a5f8fb49ab Migrate all repos into monorepo context folders
Bahn: aisupport, Analyse-O2C-C2S, awesome-bahn-mcp-servers, beam-mcp,
      Confluence_Bot, db-planet-mcp-server, O2C-Harness, project-audit,
      Projekt-KIQ-HP, teamlandkarte-mcp
Dhive: Jury-Voting
Privat: CV, NoteGraph (NOTE: NoteGraph needs complete redo after consolidation)
Shared: AI-Orchestrator, OrgMyLife, power_skills_and_more
Shared/references: symphony (read-only)

Bahn repos remain available as independent remotes - this monorepo
pulls them in via subtree, the originals are untouched.
2026-06-30 20:39:52 +02:00

28 KiB
Raw Blame History

Change Proposal: Replace Heuristics with Azure OpenAI Embeddings and LLM

  • Change ID: replace-heuristics-with-azure-openai
  • Status: Implemented
  • Target: teamlandkarte-mcp
  • Author: Thomas Handke
  • Date: 2026-02-13

Summary

Replace the current FastMCP sampling (unavailable) and deterministic heuristic fallbacks with direct Azure OpenAI API integration for:

  1. Competence and role similarity matching via embeddings (text-embedding-3-large, 3072 dimensions)
  2. Requirements extraction and validation via LLM (gpt-4.1)

Additionally:

  • Rename database.tomlconfig.toml (and template accordingly)
  • Implement persistent embedding cache using a local SQLite database
  • Support two similarity scoring strategies (per-skill vs. aggregate) via configuration
  • Remove all FastMCP sampling code and heuristic fallbacks completely

Motivation

Current State Problems

  1. FastMCP sampling is unavailable: The mcp package does not expose a FastMCP.sampling API, forcing the system to rely entirely on weak deterministic heuristics.

  2. Heuristic limitations:

    • Role inference: simple keyword matching ("cloud" → "Cloud Engineer")
    • Competence extraction: small hardcoded keyword list (~15 terms)
    • Competence similarity: exact normalized match (1.0) or substring (0.7), otherwise 0.0
    • No true semantic understanding
  3. Poor matching quality:

    • Synonyms not recognized (e.g., "React.js" vs "ReactJS")
    • Related skills missed (e.g., "FastAPI" vs "REST API Development")
    • Role extraction frequently returns "(unknown)"
  4. Maintenance burden: Heuristic keyword lists require manual updates

Proposed Benefits

  1. Semantic matching: True similarity via embeddings (e.g., "Kubernetes" ↔ "K8s", "Python" ↔ "Python 3")
  2. Robust extraction: LLM-backed role and competence extraction from free text
  3. Better user experience: More accurate Top/Good/Partial/Low categorization
  4. Reduced maintenance: No manual keyword list updates
  5. Production-ready: Direct API integration, no dependency on MCP sampling feature

Goals

  1. Replace TaskAnalyzer.semantic_competence_similarity with embedding-based similarity
  2. Replace TaskAnalyzer.extract_ranked_roles with Azure OpenAI LLM
  3. Replace TaskAnalyzer._extract_requirements_from_description with Azure OpenAI LLM
  4. Replace TaskAnalyzer.apply_requirement_update with Azure OpenAI LLM
  5. Implement persistent embedding cache (SQLite)
  6. Support two similarity strategies via config
  7. Remove all FastMCP sampling and heuristic code
  8. Rename database.tomlconfig.toml

Non-Goals

  • Changing the overall matching workflow (confirmation gate, guided capture, etc.)
  • Changing the scoring weights (competence 0.8, role 0.2)
  • Changing the categorical thresholds (Top/Good/Partial/Low)
  • Supporting multiple embedding models
  • Supporting other LLM providers (only Azure OpenAI)

Proposed Changes

1. Configuration Changes

1.1 Rename configuration file

  • database.tomlconfig.toml
  • database.toml.exampleconfig.toml.example
  • Update all references in code, docs, README

1.2 New configuration sections in config.toml

[azure_openai]
endpoint = "https://aiservice-ca00361106.cognitiveservices.azure.com/"
api_version_embeddings = "2024-02-01"
api_version_llm = "2024-12-01-preview"

[azure_openai.embeddings]
model = "text-embedding-3-large"
# Azure deployment name is equal to model name
deployment = "text-embedding-3-large"

[azure_openai.llm]
model = "gpt-4.1"
# Azure deployment name is equal to model name
deployment = "gpt-4.1"
temperature = 0.2
max_tokens = 2000

[embedding_cache]
enabled = true
db_path = "embeddings_cache.db"
ttl_days = 30

[matching.similarity]
# Strategy: "per_skill" or "aggregate"
# - per_skill: Match each required skill to best candidate skill (current behavior)
# - aggregate: Compare average embedding of all required vs all candidate skills
strategy = "per_skill"
# Chat model/deployment name (Azure OpenAI)
chat_model = "gpt-4.1"

1.3 Environment variables (.env)

AZURE_OPENAI_EMBEDDING_API_KEY="..."
AZURE_OPENAI_LLM_API_KEY="..."

2. Embedding Cache Implementation

Create src/teamlandkarte_mcp/cache/embedding_cache.py:

class EmbeddingCache:
    """Persistent embedding cache using SQLite.
    
    Schema:
        embeddings(
            id INTEGER PRIMARY KEY,
            text TEXT UNIQUE NOT NULL,
            model TEXT NOT NULL,
            embedding BLOB NOT NULL,  -- UTF-8 bytes of JSON-serialized list[float]
            created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
        )
    
    Notes:
        - Store embeddings as UTF-8 encoded JSON in a BLOB column for portability.
        - Serialize with: json.dumps(embedding).encode("utf-8")
        - Deserialize with: json.loads(blob.decode("utf-8"))
    """
    
    def __init__(self, db_path: str, ttl_days: int = 30):
        """Initialize cache, create DB if not exists."""
        
    def get(self, text: str, model: str) -> Optional[list[float]]:
        """Retrieve cached embedding, return None if missing/expired."""
        
    def put(self, text: str, model: str, embedding: list[float]) -> None:
        """Store embedding in cache."""
        
    def cleanup_expired(self) -> int:
        """Remove embeddings older than TTL, return count deleted."""

3. Azure OpenAI Client Layer

Create src/teamlandkarte_mcp/azure/openai_client.py:

class AzureOpenAIClient:
    """Wrapper for Azure OpenAI API calls using AsyncAzureOpenAI client."""
    
    def __init__(self, config: AzureOpenAIConfig, embedding_cache: EmbeddingCache):
        """Initialize AsyncAzureOpenAI clients for embeddings and LLM.
        
        Uses openai.AsyncAzureOpenAI for all async API calls.
        """
        
    async def get_embedding(self, text: str) -> list[float]:
        """Get embedding for text, using cache if available.
        
        Raises:
            AzureAPIError: If API call fails (no fallback).
        """
        
    async def get_embeddings_batch(self, texts: list[str]) -> list[list[float]]:
        """Get embeddings for multiple texts (not batched per user requirement)."""
        
    async def chat_completion(
        self, 
        system: str, 
        user: str,
        response_format: dict = None
    ) -> str:
        """LLM completion with structured JSON response.
        
        Raises:
            AzureAPIError: If API call fails (no fallback).
        """

4. Similarity Computation

Create src/teamlandkarte_mcp/matching/similarity.py:

def cosine_similarity(vec_a: list[float], vec_b: list[float]) -> float:
    """Compute cosine similarity between two vectors."""

class SimilarityEngine:
    """Compute semantic similarity using embeddings."""
    
    def __init__(
        self, 
        client: AzureOpenAIClient, 
        strategy: str = "per_skill"
    ):
        """Initialize with Azure client and strategy."""
        
    async def compute_competence_similarity(
        self,
        required: list[str],
        candidate: list[str],
    ) -> dict[str, dict[str, Any]]:
        """Compute similarity using configured strategy.
        
        Returns same format as current semantic_competence_similarity:
            {required_comp: {best_match: str|null, score: float, rationale: str}}

        Behavioral guarantees:
        - `rationale` is always included (even when best_match is null).
        - If `required` is empty, return {}.
        - If `candidate` is empty, return an entry per required competence with:
            best_match=null, score=0.0, and a short rationale.
        """
        
    async def _per_skill_similarity(
        self, required: list[str], candidate: list[str]
    ) -> dict[str, dict[str, Any]]:
        """For each required skill, find best matching candidate skill."""
        
    async def _aggregate_similarity(
        self, required: list[str], candidate: list[str]
    ) -> dict[str, dict[str, Any]]:
        """Compare average embeddings of required vs candidate skills."""
        
    async def compute_role_similarity(
        self, required_role: str, candidate_role: str
    ) -> float:
        """Compute role similarity (0.0 to 1.0).

        Behavioral guarantees:
        - If either role is empty/None, return 0.0.
        - If either role is "(unknown)" (case-insensitive, trimmed), return 0.0.
        - Otherwise compute cosine similarity of the role embeddings.

        Note:
        - Role similarity returns only a float score. Rationales are provided by
          `compute_competence_similarity()` results.
        """

5. TaskAnalyzer Refactoring

Remove:

  • _sample_json() method
  • _heuristic_competences_from_text() static method
  • All heuristic fallback code in:
    • extract_ranked_roles()
    • _extract_requirements_from_description()
    • apply_requirement_update()
    • semantic_competence_similarity()

Replace with:

class TaskAnalyzer:
    """Task text analyzer using Azure OpenAI."""
    
    def __init__(self, azure_client: AzureOpenAIClient):
        self._client = azure_client
        
    async def extract_ranked_roles(
        self, description: str, limit: int = 5
    ) -> list[RankedRole]:
        """Extract ranked roles using Azure OpenAI LLM.
        
        Raises:
            AzureAPIError: If API call fails.
        """
        
    async def _extract_requirements_from_description(
        self, description: str
    ) -> ExtractedRequirements:
        """Extract requirements using Azure OpenAI LLM.
        
        Raises:
            AzureAPIError: If API call fails.
        """
        
    async def apply_requirement_update(
        self, current: Requirements, change_description: str
    ) -> Requirements:
        """Update requirements using Azure OpenAI LLM.
        
        Raises:
            AzureAPIError: If API call fails.
        """

Additional behavior change:

  • The current TaskAnalyzer raises AnalysisError only when MCP sampling is present but fails/returns invalid JSON, and otherwise silently falls back to heuristics.
  • After this change, heuristics are removed. Any Azure OpenAI failure (HTTP error, timeout after retries, invalid JSON/schema) will result in an analysis exception (e.g. AzureAPIError and/or a narrower AnalysisError) and should be surfaced to callers.

Testing impact:

  • Existing tests that currently assert heuristic fallback behavior in TaskAnalyzer must be updated.
  • Replace “heuristic fallback expected” assertions with either:
    • mocked Azure responses (unit tests), or
    • explicit error expectations when Azure is unavailable.

6. Matcher & Scorer Integration

Update src/teamlandkarte_mcp/matching/matcher.py:

Key Changes:

  1. Accept SimilarityEngine in constructor
  2. Replace TaskAnalyzer.semantic_competence_similarity() calls with SimilarityEngine.compute_competence_similarity()
  3. Replace the existing _role_similarity() function with SimilarityEngine.compute_role_similarity()
  4. Maintain current scoring weights (competence 0.8, role 0.2)

Details:

  • The Matcher currently uses TaskAnalyzer.semantic_competence_similarity() for competence matching (per-skill, best-match strategy)
  • It also has a standalone _role_similarity() function for role matching (currently uses simple string comparison)
  • Both will be replaced by SimilarityEngine methods, which provide semantic similarity via embeddings
  • The SimilarityEngine.compute_competence_similarity() returns the same dict structure as current semantic_competence_similarity(), ensuring drop-in compatibility
  • The SimilarityEngine.compute_role_similarity() returns a float (0.0 to 1.0), matching current _role_similarity() signature
class Matcher:
    """Capacity matcher with semantic similarity."""
    
    def __init__(
        self, 
        analyzer: TaskAnalyzer, 
        similarity_engine: SimilarityEngine,
        config: MatchingConfig
    ):
        """Initialize matcher with analyzer, similarity engine, and config."""
        self._analyzer = analyzer
        self._similarity = similarity_engine
        self._config = config
        
    async def _compute_match_score(
        self, 
        required: Requirements, 
        capacity: Capacity
    ) -> MatchScore:
        """Compute match score using SimilarityEngine.
        
        Competence matching:
            comp_sim = await self._similarity.compute_competence_similarity(
                required.competences, capacity.competences
            )
        
        Role matching:
            role_score = await self._similarity.compute_role_similarity(
                required.role, capacity.role
            )
        
        Final score: (0.8 * avg_comp_score) + (0.2 * role_score)
        """

Owning module note:

  • The end-to-end match scoring aggregation currently lives in src/teamlandkarte_mcp/matching/matcher.py (with supporting score shaping asserted by tests/test_scorer.py).

7. Server Initialization Changes

Update src/teamlandkarte_mcp/mcp_server.py:

# Initialize Azure OpenAI components
embedding_cache = EmbeddingCache(
    db_path=cfg.embedding_cache.db_path,
    ttl_days=cfg.embedding_cache.ttl_days,
) if cfg.embedding_cache.enabled else None

azure_client = AzureOpenAIClient(cfg.azure_openai, embedding_cache)
similarity_engine = SimilarityEngine(
    azure_client,
    strategy=cfg.matching.similarity.strategy,
)

# Initialize analyzer with Azure client
analyzer = TaskAnalyzer(azure_client)

# Initialize matcher with similarity engine
matcher = Matcher(analyzer, similarity_engine, cfg.matching)

8. Cost Estimation & Monitoring

Add cost tracking module src/teamlandkarte_mcp/azure/cost_tracker.py:

class CostTracker:
    """Track and estimate Azure OpenAI API costs."""
    
    # Pricing (as of 2026-02, may change)
    EMBEDDING_COST_PER_1K_TOKENS = 0.00013  # text-embedding-3-large
    LLM_INPUT_COST_PER_1K_TOKENS = 0.03     # gpt-4.1
    LLM_OUTPUT_COST_PER_1K_TOKENS = 0.06    # gpt-4.1
    
    def log_embedding_request(self, text: str, cached: bool):
        """Log embedding request for cost tracking."""
        
    def log_llm_request(self, input_tokens: int, output_tokens: int):
        """Log LLM request for cost tracking."""
        
    def get_session_costs(self) -> dict:
        """Return estimated costs for current session."""

Add to README.md:

## Azure OpenAI Cost Estimation

Typical matching workflow costs (estimated):

| Operation | API Calls | Estimated Cost |
|-----------|-----------|----------------|
| Single capacity match (100 candidates) | ~200 embeddings (mostly cached), 0 LLM | $0.001 - $0.01 |
| Task requirements extraction | 0 embeddings, 1 LLM call (~1000 tokens) | $0.03 - $0.06 |
| Role inference | 0 embeddings, 1 LLM call (~500 tokens) | $0.015 - $0.03 |

With embedding caching enabled (default), repeated searches are significantly cheaper.

**Monthly cost estimate** (100 searches/day, 20 days):
- Without cache: ~$60-120/month
- With cache (90% hit rate): ~$10-20/month

Testing Strategy

Unit Tests (Mocked Azure)

All unit tests should mock Azure OpenAI API responses to avoid API costs and ensure deterministic behavior:

Embedding Mocks:

  • Use fixed-dimension vectors matching text-embedding-3-large (3072 dimensions)
  • Create realistic patterns:
    • Similar terms: [0.9, 0.8, 0.1, ..., 0.0] and [0.85, 0.82, 0.15, ..., 0.0] → high cosine similarity (~0.95)
    • Different terms: [0.9, 0.1, 0.0, ..., 0.0] and [0.1, 0.9, 0.0, ..., 0.0] → low cosine similarity (~0.1)
  • Use unittest.mock.AsyncMock or pytest-mock for AsyncAzureOpenAI client

LLM Mocks:

  • Mock chat.completions.create() to return structured JSON responses
  • Example: {"roles": [{"role": "Backend Developer", "confidence": 0.9, "rationale": "..."}], ...}

Example Mock Pattern:

@pytest.mark.asyncio
async def test_compute_competence_similarity():
    mock_client = AsyncMock()
    mock_client.get_embedding.side_effect = [
        [0.9, 0.8, 0.1] + [0.0] * 3069,  # "Python"
        [0.85, 0.82, 0.15] + [0.0] * 3069,  # "Python 3"
    ]
    engine = SimilarityEngine(mock_client, strategy="per_skill")
    result = await engine.compute_competence_similarity(["Python"], ["Python 3"])
    assert result["Python"]["score"] > 0.9

Integration Tests (Real Azure)

  • Mark all integration tests with @pytest.mark.integration
  • Configure pytest.ini:
    [pytest]
    markers =
        integration: marks tests that call real Azure OpenAI API (deselect with '-m "not integration"')
    
  • Run integration tests manually before deployment: pytest -m integration
  • Run fast tests in CI: pytest -m "not integration"
  • Integration tests should use small, cheap prompts to minimize costs

Test Coverage Goals

  • Unit tests: ≥90% coverage for all new code (cache, client, similarity, analyzer)
  • Integration tests: Cover happy path + common error scenarios (API timeout, invalid credentials)
  • Manual tests: Full end-to-end workflow in Cherry Studio (search → filter → confirm)

Implementation Task List

Phase 1: Configuration & Infrastructure (3-4 hours)

  • 1.1 Rename database.tomlconfig.toml and template
  • 1.2 Update all file references in code
  • 1.3 Add [azure_openai] section to config model (src/config.py)
  • 1.4 Add [embedding_cache] section to config model
  • 1.5 Add [matching.similarity] section to config model
  • 1.6 Update config.toml.example with new sections
  • 1.7 Load AZURE_OPENAI_EMBEDDING_API_KEY from .env
  • 1.8 Load AZURE_OPENAI_LLM_API_KEY from .env
  • 1.9 Add openai package to pyproject.toml dependencies

Phase 2: Embedding Cache (2-3 hours)

  • 2.1 Create src/teamlandkarte_mcp/cache/embedding_cache.py
  • 2.2 Implement EmbeddingCache.__init__ (create DB/schema if not exists)
  • 2.3 Implement EmbeddingCache.get() (with TTL check)
  • 2.4 Implement EmbeddingCache.put()
  • 2.5 Implement EmbeddingCache.cleanup_expired()
  • 2.6 Add SQLite schema migration support (if DB exists but schema old)
  • 2.7 Add unit tests for EmbeddingCache
  • 2.8 Add .gitignore entry for embeddings_cache.db

Phase 3: Azure OpenAI Client (3-4 hours)

Testing Strategy:

  • Unit tests should mock Azure API responses with realistic embedding vectors (3072 dimensions for text-embedding-3-large)

  • Use fixed seed embeddings for deterministic test behavior (e.g., [0.1, 0.2, ..., 0.0] patterns)

  • Integration tests marked with @pytest.mark.integration use real API calls

  • 3.1 Create src/teamlandkarte_mcp/azure/__init__.py

  • 3.2 Create src/teamlandkarte_mcp/azure/openai_client.py

  • 3.3 Implement AzureOpenAIClient.__init__

  • 3.4 Implement AzureOpenAIClient.get_embedding() with cache integration

  • 3.5 Implement AzureOpenAIClient.get_embeddings_batch() (individual calls)

  • 3.6 Implement AzureOpenAIClient.chat_completion() with retry logic

  • 3.7 Add AzureAPIError exception class

  • 3.8 Add unit tests for client (mocked API)

  • 3.9 Add integration test with real API (marked as @pytest.mark.integration)

Phase 4: Similarity Engine (4-5 hours)

  • 4.1 Create src/teamlandkarte_mcp/matching/similarity.py
  • 4.2 Implement cosine_similarity() function
  • 4.3 Implement SimilarityEngine.__init__
  • 4.4 Implement SimilarityEngine._per_skill_similarity()
  • 4.5 Implement SimilarityEngine._aggregate_similarity()
  • 4.6 Implement SimilarityEngine.compute_competence_similarity() (router)
  • 4.7 Implement SimilarityEngine.compute_role_similarity()
  • 4.8 Add unit tests for similarity computation (mocked embeddings)
  • 4.9 Add integration test comparing both strategies

Phase 5: TaskAnalyzer Refactoring (3-4 hours)

  • 5.1 Remove TaskAnalyzer.__init__(self, mcp: FastMCP) signature
  • 5.2 Add new TaskAnalyzer.__init__(self, azure_client: AzureOpenAIClient)
  • 5.3 Remove _sample_json() method completely
  • 5.4 Remove _heuristic_competences_from_text() method completely
  • 5.5 Refactor extract_ranked_roles() to use azure_client.chat_completion()
  • 5.6 Remove heuristic fallback from extract_ranked_roles()
  • 5.7 Refactor _extract_requirements_from_description() to use LLM
  • 5.8 Remove heuristic fallback from _extract_requirements_from_description()
  • 5.9 Refactor apply_requirement_update() to use LLM
  • 5.10 Remove heuristic fallback from apply_requirement_update()
  • 5.11 Remove semantic_competence_similarity() method (replaced by SimilarityEngine)
  • 5.12 Update all TaskAnalyzer unit tests

Phase 6: Matcher & Scorer Integration (2-3 hours)

  • 6.1 Update Matcher.__init__ to accept SimilarityEngine
  • 6.2 Replace competence similarity calls with SimilarityEngine.compute_competence_similarity()
  • 6.3 Replace role similarity calls with SimilarityEngine.compute_role_similarity()
  • 6.4 Update scorer aggregation logic if needed
  • 6.5 Update integration tests for matching pipeline

Phase 7: Server Initialization (2 hours)

  • 7.1 Update build_server() in mcp_server.py to load Azure config
  • 7.2 Initialize EmbeddingCache instance
  • 7.3 Initialize AzureOpenAIClient instance
  • 7.4 Initialize SimilarityEngine instance
  • 7.5 Update TaskAnalyzer instantiation
  • 7.6 Update Matcher instantiation
  • 7.7 Add startup logging for Azure OpenAI connection
  • 7.8 Add graceful error handling if Azure credentials missing

Phase 8: Cost Tracking & Monitoring (2 hours)

  • 8.1 Create src/teamlandkarte_mcp/azure/cost_tracker.py
  • 8.2 Implement CostTracker class with logging methods
  • 8.3 Integrate cost tracking into AzureOpenAIClient
  • 8.4 Add periodic cost report logging (stderr)
  • 8.5 Add cost summary to tool outputs (optional, via config)

Phase 9: Documentation (2-3 hours)

  • 9.1 Update README.md with Azure OpenAI setup instructions
  • 9.2 Add Azure cost estimation section to README.md
  • 9.3 Update docs/troubleshooting.md with Azure API error guidance
  • 9.4 Document similarity strategies (per_skill vs aggregate)
  • 9.5 Update config.toml.example with comprehensive comments
  • 9.6 Update OpenSpec architecture docs
  • 9.7 Add migration guide from old database.toml to new config.toml

Phase 10: Testing & Validation (3-4 hours)

  • 10.1 Configure pytest.ini with integration marker: markers = integration: marks tests that call real Azure OpenAI API (deselect with '-m "not integration"')
  • 10.2 Run full test suite after refactoring (pytest -m "not integration" for fast CI)
  • 10.3 Add new integration tests for Azure API paths (marked with @pytest.mark.integration)
  • 10.4 Test embedding cache persistence across server restarts
  • 10.5 Test both similarity strategies with real data
  • 10.6 Validate cost tracking accuracy
  • 10.7 Test error handling when Azure API unavailable
  • 10.8 Performance benchmark: compare cache hit/miss scenarios
  • 10.9 Manual validation in Cherry Studio

Phase 11: Cleanup (1 hour)

  • 11.1 Remove all commented-out FastMCP sampling code
  • 11.2 Remove unused imports (from mcp.server.fastmcp import FastMCP from TaskAnalyzer)
  • 11.3 Update type hints and docstrings
  • 11.4 Run linter and fix style issues
  • 11.5 Final code review

Impact Analysis

Changed Files

New files:

  • config.toml (renamed from database.toml)
  • config.toml.example (renamed from database.toml.example)
  • src/teamlandkarte_mcp/cache/embedding_cache.py
  • src/teamlandkarte_mcp/azure/__init__.py
  • src/teamlandkarte_mcp/azure/openai_client.py
  • src/teamlandkarte_mcp/azure/cost_tracker.py
  • src/teamlandkarte_mcp/matching/similarity.py
  • embeddings_cache.db (generated at runtime, gitignored)

Modified files:

  • src/teamlandkarte_mcp/config.py (new config sections)
  • src/teamlandkarte_mcp/matching/task_analyzer.py (complete refactor)
  • src/teamlandkarte_mcp/matching/matcher.py (similarity engine integration)
  • src/teamlandkarte_mcp/mcp_server.py (initialization changes)
  • README.md (setup instructions, cost docs)
  • docs/troubleshooting.md (Azure error guidance)
  • pyproject.toml (add openai dependency)
  • .gitignore (add embeddings_cache.db, config.toml)
  • All OpenSpec architecture/design docs

Deleted code:

  • All FastMCP sampling code in TaskAnalyzer
  • All heuristic fallback code in TaskAnalyzer
  • _sample_json(), _heuristic_competences_from_text() methods

Breaking Changes

  1. Configuration file renamed: Users must rename database.tomlconfig.toml
  2. New required environment variables: AZURE_OPENAI_EMBEDDING_API_KEY, AZURE_OPENAI_LLM_API_KEY
  3. Hard dependency on Azure OpenAI: No offline/fallback mode
  4. New dependency: openai Python package
  5. TaskAnalyzer constructor signature changed: TaskAnalyzer.__init__(self, mcp: FastMCP)TaskAnalyzer.__init__(self, azure_client: AzureOpenAIClient)
    • This affects any code that instantiates TaskAnalyzer directly
    • Server initialization in mcp_server.py must be updated
  6. Matcher constructor signature changed: Matcher.__init__(self, analyzer: TaskAnalyzer, config: MatchingConfig)Matcher.__init__(self, analyzer: TaskAnalyzer, similarity_engine: SimilarityEngine, config: MatchingConfig)
    • Adds new required similarity_engine parameter
    • Server initialization must pass SimilarityEngine instance

Migration Path

  1. Rename database.tomlconfig.toml
  2. Add new Azure OpenAI sections to config.toml
  3. Add Azure API keys to .env
  4. Install updated dependencies: uv sync
  5. First run will create embeddings_cache.db automatically

Security & Constraints

Security

  • API keys in .env: Never commit .env or config.toml with credentials
  • Embedding cache: Contains only embeddings, not raw capacity data (safe to persist)
  • Network: All API calls over HTTPS to Azure OpenAI endpoint
  • Error messages: Do not log API keys in error messages

Operational Constraints

  • Azure OpenAI dependency: System will fail if Azure API unavailable (no fallback)
  • API rate limits: Azure OpenAI enforces rate limits (TPM/RPM); implement retry with backoff
  • Cost: Real monetary cost per API call (mitigated by caching)
  • Latency: First search for a capacity will be slower (embedding generation); subsequent searches fast (cached)

Configuration Defaults

[embedding_cache]
enabled = true
db_path = "embeddings_cache.db"
ttl_days = 30

[matching.similarity]
strategy = "per_skill"  # Maintains current behavior

Cost Estimation Details

API Pricing (as of 2026-02)

Service Model Cost
Embeddings text-embedding-3-large $0.00013 per 1K tokens
LLM (input) gpt-4.1 $0.03 per 1K tokens
LLM (output) gpt-4.1 $0.06 per 1K tokens

Typical Workflow Costs

Scenario 1: Task requirements extraction

  • Input: ~500 tokens (task description)
  • Output: ~300 tokens (JSON with roles/competences/dates)
  • Cost: (500 × $0.03 / 1000) + (300 × $0.06 / 1000) = $0.033

Scenario 2: Capacity matching (100 candidates, 5 required skills)

  • First run (cold cache):
    • Required skills: 5 × ~10 tokens = 50 tokens → ~$0.0000065
    • Candidate skills: 100 candidates × 3 skills avg × 10 tokens = 3000 tokens → ~$0.00039
    • Total: ~$0.0004 (negligible)
  • Subsequent runs (warm cache): $0 (all cached)

Scenario 3: Role inference

  • Input: ~400 tokens
  • Output: ~200 tokens
  • Cost: ~$0.024

Monthly Estimates

Light usage (20 searches/month, 10 extractions):

  • Embeddings: ~$0.05
  • LLM: ~$0.50
  • Total: ~$0.55/month

Moderate usage (100 searches/month, 50 extractions):

  • Embeddings: ~$0.20 (with 90% cache hit rate)
  • LLM: ~$2.50
  • Total: ~$2.70/month

Heavy usage (500 searches/month, 200 extractions):

  • Embeddings: ~$1.00 (with 90% cache hit rate)
  • LLM: ~$10.00
  • Total: ~$11/month

Open Questions

None (all clarifications provided by user).

Approval & Timeline

  • Estimated effort: 28-35 hours (1 week full-time or 2 weeks part-time)
  • Risk level: Medium (API dependency, cost implications, major refactor)
  • Approval required: Yes (architecture change, new external dependency)

Next Steps

  1. Review and approve this proposal
  2. Create detailed implementation branch
  3. Implement phases 1-11 sequentially
  4. Conduct thorough testing (unit + integration + manual)
  5. Document migration guide
  6. Deploy to staging environment
  7. Monitor costs and performance
  8. Deploy to production