Bahn: aisupport, Analyse-O2C-C2S, awesome-bahn-mcp-servers, beam-mcp,
Confluence_Bot, db-planet-mcp-server, O2C-Harness, project-audit,
Projekt-KIQ-HP, teamlandkarte-mcp
Dhive: Jury-Voting
Privat: CV, NoteGraph (NOTE: NoteGraph needs complete redo after consolidation)
Shared: AI-Orchestrator, OrgMyLife, power_skills_and_more
Shared/references: symphony (read-only)
Bahn repos remain available as independent remotes - this monorepo
pulls them in via subtree, the originals are untouched.
4.4 KiB
4.4 KiB
Change: Add capacity→task matching tools and embedding-based inference
Why
The server currently supports task→capacity matching (list_open_tasks, get_task_details, find_matching_capacities, find_capacities_for_task). There is no symmetrical workflow to start from a capacity/person and discover relevant tasks.
Additionally, role/requirements inference and task validation still rely on the Azure chat API. This introduces latency/cost, complicates failure semantics, and is redundant given the system already uses embeddings + cosine similarity for matching.
What changes
New MCP tools
list_free_capacities(limit=20)- Show the X most recent active capacities in a table-first Markdown output.
get_capacity_details(capacity_id)- Show a single capacity/person entry (same fields as the list, table-first).
find_matching_tasks(capacity_id)- Match a capacity against all open tasks based on role + competences.
- Returns paginated Markdown table with task_id, title, required competences, availability, score, and category.
Rework role + competence inference (remove chat API)
- Replace chat-based role inference with embedding-based role inference:
- infer roles from a combined text: task title (if present) + task description
- embed this combined text once and match against a pre-embedded role vocabulary from the Data Lake.
- Replace chat-based competence extraction/validation with embedding-based competence matching:
- embed the task text and match against a pre-embedded competence vocabulary from the Data Lake.
- Store the embedding for
task_id(combined title+description text) in the local SQLite embedding cache to minimize repeated embedding costs.
Data Lake sourcing
- Competence vocabulary:
hive.tier1_open_lake.beschaffungstool_kmp_skill_latest(existing source) - Role vocabulary:
teamlandkarte_v_capacity_roles_latestvia columns:- role name column:
name - filters:
active=trueandstaffing_board_relevant=true
- role name column:
Azure OpenAI usage
- Azure embeddings remain required.
- Azure chat API usage is removed completely once all dependent features are migrated.
Impact
- MCP API surface: Adds 3 tools (capacity browsing + capacity→task matching).
- Business logic: TaskAnalyzer and validation logic migrate to embeddings-only.
- Data access: Requires new read queries to fetch role/competence vocabularies.
- Caching: Reuses existing SQLite embedding cache; adds stable keys for task_id text embeddings.
- Docs: README, assistant prompt, and architecture must be updated.
Non-goals
- Changing scoring weights or category thresholds.
- Building a full-text search for tasks/capacities.
Open questions (need confirmation)
- “Most recent capacities”: should ordering be by
begin_datedescending, byend_datedescending, or by a DB “created/updated” timestamp (if available)? list_free_capacities“owner/team”: do you want owner_name only, or bothowner_nameand an explicitteamfield (if available in the view)?- Do we keep the existing
infer_roles(...)tool name/shape, but change its implementation to embeddings-only (recommended), or introduce a new tool and deprecate the old one? - For embedding-based competence inference for validation: should the output be a ranked list with similarity scores, or a binary chosen set (top-N)? If top-N, what default N?
Confirmed decisions
list_free_capacities: order by most recentcreation_date.- Capacity “Owner/Team”: use the
owner_namecolumn (contains the team name). - Role inference tool:
- introduce a new tool and deprecate
infer_roles - the new tool returns only the single closest role (no rank/rationale)
- introduce a new tool and deprecate
- Competence inference:
- return a maximum of X competences (configurable, default: 8)
- apply a configurable similarity score threshold (default: 0.4)
- if
max_competencesis not specified: no limit on number of competences - if
min_similarityis not specified: no threshold filtering - at least one parameter must be specified to avoid unbounded results
- if both are set: select top X among those >= threshold
- Refresh strategy:
- on startup only
- only cache-missing roles/competences are embedded
- embed all open tasks (combined title+description) on startup
- if something is still missing at runtime, embed on-demand (strict semantics)