Files
ankn a5f8fb49ab Migrate all repos into monorepo context folders
Bahn: aisupport, Analyse-O2C-C2S, awesome-bahn-mcp-servers, beam-mcp,
      Confluence_Bot, db-planet-mcp-server, O2C-Harness, project-audit,
      Projekt-KIQ-HP, teamlandkarte-mcp
Dhive: Jury-Voting
Privat: CV, NoteGraph (NOTE: NoteGraph needs complete redo after consolidation)
Shared: AI-Orchestrator, OrgMyLife, power_skills_and_more
Shared/references: symphony (read-only)

Bahn repos remain available as independent remotes - this monorepo
pulls them in via subtree, the originals are untouched.
2026-06-30 20:39:52 +02:00

10 KiB
Raw Blame History

Change Proposal: Improve Task Details, Role Inference, Guided Requirements Capture, and Search Reliability

  • Change ID: improve-task-details-role-inference-and-guided-flow
  • Status: Implemented
  • Target: teamlandkarte-mcp
  • Author: (to fill)
  • Date: 2026-02-12

Summary

This proposal refines several MCP tools and interaction flows to improve usability in Cherry Studio and other MCP clients:

  1. get_task_details should return a concise markdown summary (table) of the task plus description below.
  2. Introduce a dedicated tool infer_roles for role inference from either a DB task ID or free-text description.
  3. validate_task_requirements should return a single comparison table plus a short written summary.
  4. When a user requests a person search without providing project/task details, switch to a step-by-step guided capture (no "all questions at once").
  5. Enforce a mandatory confirmation step before triggering matching, for both free-text and guided capture / structured entry.
  6. Fix a suspected competence matching regression (cloud expert + AWS scoring ends up Low), likely caused by extraction/matching issues.
  7. Improve search session robustness for pagination/filtering:
    • search tools emit a JSON block containing search_id/filter_id
    • non-UUID search_id inputs are rejected early with a clear message
    • internal tracking of the latest search id (session state) for diagnostics only; no "get last search id" tool is provided
  8. Update documentation, architecture, and design markdown accordingly.

Motivation

  • Current outputs are verbose and inconsistent across tools.
  • Role inference exists but is not clearly available as an explicit tool.
  • Guided capture and confirmation are required to prevent accidental matching runs.
  • Cache invalidation/expiry breaks multi-step flows (filter_search_results, pagination, etc.).

Goals

  • Produce stable, markdown-first outputs that are easy to consume in chat UIs.
  • Make role inference available as an explicit, reusable capability for both DB-backed tasks and free-text.
  • Improve correctness and robustness of role/requirements extraction and semantic matching via the servers Azure OpenAI integration.
  • Ensure multi-step search workflows remain functional for realistic time spans.
  • Remove artificial constraints on the number of exposed MCP tools.

Non-Goals

  • Re-introducing Hive/PyHive support.
  • Changing the underlying DB schema.
  • Replacing LLM-based role inference with a numeric-only scorer (categorical matching output should remain the main UX).

Proposed Changes

1) get_task_details output format

Change get_task_details(task_id) to return:

  1. A markdown table with columns:
    • Task ID
    • Title
    • Created (date only, no time)
    • Start
    • End
    • Competences (comma-separated, no numbering)
    • Inferred Role
  2. Below the table:
    • The task description (plain markdown, no table)

Notes:

  • "Inferred Role" is derived using infer_roles (see section 2). If no role is inferred, display (none).

2) Dedicated role inference tool: infer_roles

Add a tool to infer roles from either:

  • A DB task via task_id, or
  • A free-text task_description

Signature:

  • infer_roles(task_id: Optional[str] = None, task_description: Optional[str] = None, limit: int = 5) -> str

Output:

  • A markdown table containing all recognized roles, each with:
    • Rank
    • Role
    • Rationale

Behavior:

  • Exactly one of task_id or task_description must be provided.
  • Uses Azure OpenAI chat completion for role inference (no heuristic fallback).

3) validate_task_requirements output format

Change output to:

  1. A single markdown comparison table that contrasts DB vs extracted information.
  2. Table must include:
    • Task ID
    • Title
    • Time ranges (DB vs extracted)
    • Competences (DB vs extracted)
  3. Below the table, a short written summary, including:
    • Skill overlap and notable differences
    • Any missing DB fields (dates/skills)

4) Step-by-step guided capture mode (step-specific tools)

Introduce a guided capture workflow for the case:

User requests people/capacity search but provides insufficient task/project details.

The guided flow should collect input sequentially using step-specific tools (to avoid free-text ambiguity):

  1. Task description
  2. Role
  3. Time range
  4. Competences

Implementation idea:

  • Add a server-side state machine stored in SessionState.
  • Add tools such as:
    • start_guided_capture()
    • guided_set_description(description: str)
    • guided_set_role(role_name: str)
    • guided_set_time_range(date_start: Optional[str] = None, date_end: Optional[str] = None)
    • guided_set_competences(competences: list[str])

5) Enforce confirmation before matching (shared mechanism, hard-gated)

extract_requirements and structured/guided capture must not directly trigger matching.

Instead, both should:

  • produce a "pending" Requirements object, and
  • require a confirmation step before matching tools can execute.

Hard enforcement:

  • find_matching_capacities (and any DB-backed matching tool) must refuse to run unless requirements have been explicitly confirmed or confirmation is auto-skipped by configuration (see below).

Config key (name + location):

  • Add a matching section flag in database.toml:

    [matching]
    require_confirmation = true
    
  • Default: true (strict).

  • When false: pending requirements are treated as auto-confirmed (skip confirmation).

Implementation idea:

  • Keep confirm_requirements parameters for backwards compatibility, but route the behavior through a shared internal helper and a shared tool:
    • internal: _set_pending_requirements(req)
    • tool: confirm_requirements(confirm: bool = True)

6) Fix competence matching producing incorrect Low results

Investigate and fix why "Cloud expert" + "AWS" competences yields Low across all capacities.

Hypotheses:

  • Requirements extraction produces empty or wrong competences.
  • Competence normalization / matching thresholds could be misapplied.

Deliverables:

  • Add a regression test reproducing the AWS case.
  • Ensure competence scoring uses the extracted competences correctly.

Capacity scoring UX note:

  • Keep numeric scores (overall_score, competence_score, role_score) visible.
  • Emphasize categorical results (Top/Good/Partial/Low) as the primary interpretation in docs and tables.

7) Fix search cache expiry / invalid IDs in multi-step tools

Observed:

  • filter_search_results and get_results_by_category sometimes return "Unknown or expired search_id".

Additional observed behavior in chat clients:

  • Some clients shorten, reformat, or accidentally reuse a previous search_id when the user requests another category (e.g. "show Good results").

Assumption:

  • The server process is continuous (no restart between tool calls).

Likely causes to verify (not only TTL):

  • Eviction due to [cache].max_size (entry removed even within TTL).
  • Search IDs not stored on all paths (serialization error, early return, etc.).
  • ID formatting/copy issues (backticks/whitespace) in client UIs.
  • Hidden process restarts / multiple instances in the client environment.

Proposed fixes:

  • Verify SearchCache TTL configuration and access.
  • Verify and test behavior under max-size eviction.
  • Ensure IDs are stored on every store_search call.
  • Add troubleshooting guidance to copy IDs without surrounding formatting.
  • Ensure search tools return a machine-readable JSON block containing IDs:
    • find_matching_capacities returns {search_id, filter_id=null}
    • filter_search_results returns {search_id, filter_id}
    • get_results_by_category echoes {search_id, filter_id, category, page, ...}
  • Reject non-UUID search_id values early (clear message: expected UUID).
  • Track latest search_id server-side (session state) for diagnostics only; no "get last search id" tool is provided.

8) Documentation updates

Update:

  • README.md
  • docs/troubleshooting.md
  • OpenSpec design/architecture docs

…to reflect:

  • new/changed tool outputs
  • infer_roles
  • guided capture step tools
  • confirmation requirement
  • role inference and capacity scoring UX (categorical)
  • cache behavior expectations

Implementation Task List

A. Tool output refactors

  1. Update get_task_details to produce the new table-first format.
  2. Implement infer_roles and reuse it from get_task_details.
  3. Update validate_task_requirements output as specified.
  4. Add/adjust markdown table helpers if needed (column wrapping, newline handling).

B. Role inference and extraction robustness

  1. Review role/requirements extraction prompts and output validation.
  2. Add regression tests for:
    • role inference
    • AWS/cloud expert competence matching not collapsing to Low

C. Guided capture + confirmation flow

  1. Design a minimal state machine for guided capture in SessionState.
  2. Implement step-specific guided tools:
    • start_guided_capture
    • guided_set_description
    • guided_set_role
    • guided_set_time_range
    • guided_set_competences
  3. Add matching.require_confirmation to the configuration model + loader (default: true).
  4. Implement shared confirmation tool confirm_requirements(confirm: bool = True).
  5. Refactor extract_requirements and structured/guided collection tools to set pending requirements and instruct confirmation.
  6. Refactor matching tools to hard-enforce confirmation unless matching.require_confirmation = false.

D. Search cache reliability

  1. Diagnose and fix SearchCache invalidation/expiry issues.
  2. Add tests covering:
    • storing search results
    • subsequent filtering
    • pagination
    • TTL behavior
    • max-size eviction behavior
  3. Add (optional) diagnostic support for identifying server instance continuity.

E. Documentation and specs

  1. Update README.md for the new workflow and new tools.
  2. Update docs/troubleshooting.md with Cherry Studio guidance for multi-step flows (including ID copy hygiene and cache eviction notes).
  3. Update OpenSpec design and architecture notes for guided capture, confirmation, and categorical-first scoring (numeric scores remain visible).

Open Questions / Clarifications Needed

  • (resolved) infer_roles output should be Rank | Role | Rationale only (no numeric score).
  • (resolved) Guided time ranges should allow open-ended ranges (either start or end may be omitted).