# Design: Accelerate Embedding Similarity ## Context The server uses Azure OpenAI embeddings for competence and role similarity. The matching pipeline can be slow because embeddings are currently requested on-demand and repeatedly from inner loops. The system already has a persistent SQLite embedding cache (`EmbeddingCache`). However, without bulk prefetch and without a run-scoped memoization layer, the similarity engine still: - repeatedly normalizes and hashes text keys - performs many repeated SQLite reads - performs sequential Azure calls for cache misses ## Goals - Bulk prefetch embeddings for a matching run. - Maintain deterministic behavior and strict error semantics. - Preserve existing similarity contracts. ## Architecture Changes ### Components impacted - `SimilarityEngine` (`src/teamlandkarte_mcp/matching/similarity.py`) - Add instance variable for in-memory cache (lives with the engine; not created/cleared per invocation) - Add public method `prefetch_embeddings(required_texts, candidate_texts, required_roles, candidate_roles) -> dict[str, list[float]]` - `AzureOpenAIClient` (`src/teamlandkarte_mcp/azure/openai_client.py`) - Modify `get_embeddings_batch()` to use true batch API with `input=[...]` - `CostTracker` (`src/teamlandkarte_mcp/azure/cost_tracker.py`) - May need adjustment to log batch requests (log once per chunk with count) - `EmbeddingCache` (no schema changes expected) ### Data flow (proposed) 1. **Collect** all texts that will be embedded in the current run. - This is a **single global collection pass** over the inputs and **all free capacities** (roles + competences), not a per-person loop that embeds incrementally. 2. **Normalize + deduplicate** using existing `_normalize_text()` logic. - Skip empty/whitespace-only texts - Emit a **Python logger warning** if all candidate competences normalize to empty 3. **Resolve embeddings** via: - in-memory cache (instance variable in `SimilarityEngine`) - SQLite cache (`EmbeddingCache`) - Azure embeddings API (**batch only the cache-missing texts**, **chunked** by batch size) 4. **Compute** cosine similarities locally using the prefetched mapping. ### Normalization and keys Current cache key behavior: - normalization: trim + collapse whitespace - key: SHA256 of `model|dims|normalized(text).lower()` This proposal keeps the same normalization and key derivation to avoid invalidating the on-disk cache. ## Batch embeddings API The OpenAI/Azure embeddings API supports embedding multiple inputs per request (`input=[...]`). Implementation requirements: - preserve stable mapping from input texts → returned embedding vectors - rely on `index` field in API response if available - otherwise assume stable ordering - chunk large requests to avoid request size limits (configurable batch size, default 128) - strict failure semantics: - entire matching operation fails immediately on error - log the failing **chunk input list** via Python logging to aid diagnosis - cost tracking: log once per batch chunk with a count of embeddings requested ## Trade-offs - Batch calls reduce network requests dramatically but may increase blast radius: one failing request could affect multiple inputs. - Mitigation: keep retry logic at the chunk level. - Prefetch requires holding larger embedding dictionaries in memory. - Mitigation: scope to a single invocation, and only store vectors required for that run. ## Test Strategy - Unit test prefetch/dedup: ensure the Azure client is called once per unique normalized text (only for cache misses). - Unit test cache layering: ensure hits are served from in-memory cache (first), then SQLite, and do not call Azure. - Unit test chunking: ensure multiple Azure calls are made when `len(cache_missing_texts) > batch_size`. - Unit test empty input handling: verify warning when all candidate competences normalize to empty; verify prefetch skips empty texts. - Unit test error identification: verify that batch failures report which specific text(s) caused the error. - Integration test (manual/gated): validate real Azure batch API behavior with live credentials.