--- inclusion: fileMatch fileMatchPattern: "bahn/**" --- # Bahn-Kontext When working on any file under `bahn/`, always consider the **Wissensdatenbank** (`bahn/wissensdatenbank/`) as relevant context. ## Wissensdatenbank (ETL-Pipeline für LLM/RAG) Location: `bahn/wissensdatenbank/` Remote: `https://git.tech.rz.db.de/einfachbahn-lab/tools/wissensdatenbank_group/wissensdatenbank.git` This is the central knowledge base ETL pipeline for DB InfraGO 1st-level support and chatbots. Key points: - **Purpose:** Collects knowledge from Confluence, web pages, PDFs, GitLab repos → processes into Markdown → serves as Single Source of Truth for vector DBs/chatbots - **Architecture:** ETL pipeline with strategies (confluence_page, confluence_tree, crawler, sitemap, gitlab_md, file, pdf) - **Output:** `data/processed///` (intern, extern, allgemein) - **Config:** `config/tools.yaml`, `config/general.yaml`, `config/filter_rules.json` - **Key files:** `src/` (pipeline code), `config/` (tool/source definitions), `data/` (processed output) ## When to reference it - Any bahn project that involves knowledge, documentation, or AI/chatbot integration - When discussing scope classification (intern/extern/allgemein) - When building MCP servers or tools that consume DB knowledge - When working on O2C-Harness, aisupport, or Confluence_Bot (they feed into or consume from the Wissensdatenbank) ## Sync ```powershell git subtree pull --prefix=bahn/wissensdatenbank wissensdatenbank main --squash git subtree push --prefix=bahn/wissensdatenbank wissensdatenbank main ```