1.6 KiB
1.6 KiB
inclusion, fileMatchPattern
| inclusion | fileMatchPattern |
|---|---|
| fileMatch | bahn/** |
Bahn-Kontext
When working on any file under bahn/, always consider the Wissensdatenbank (bahn/wissensdatenbank/) as relevant context.
Wissensdatenbank (ETL-Pipeline für LLM/RAG)
Location: bahn/wissensdatenbank/
Remote: https://git.tech.rz.db.de/einfachbahn-lab/tools/wissensdatenbank_group/wissensdatenbank.git
This is the central knowledge base ETL pipeline for DB InfraGO 1st-level support and chatbots. Key points:
- Purpose: Collects knowledge from Confluence, web pages, PDFs, GitLab repos → processes into Markdown → serves as Single Source of Truth for vector DBs/chatbots
- Architecture: ETL pipeline with strategies (confluence_page, confluence_tree, crawler, sitemap, gitlab_md, file, pdf)
- Output:
data/processed/<scope>/<domain>/(intern, extern, allgemein) - Config:
config/tools.yaml,config/general.yaml,config/filter_rules.json - Key files:
src/(pipeline code),config/(tool/source definitions),data/(processed output)
When to reference it
- Any bahn project that involves knowledge, documentation, or AI/chatbot integration
- When discussing scope classification (intern/extern/allgemein)
- When building MCP servers or tools that consume DB knowledge
- When working on O2C-Harness, aisupport, or Confluence_Bot (they feed into or consume from the Wissensdatenbank)
Sync
git subtree pull --prefix=bahn/wissensdatenbank wissensdatenbank main --squash
git subtree push --prefix=bahn/wissensdatenbank wissensdatenbank main