# Implementation Plan: NoteGraph Ingestion & Auto-Reference ## Overview This plan implements the NoteGraph ingestion pipeline as a Python package at `NoteGraph/ingestion/`. Tasks are ordered so each builds on the previous: scaffolding → extraction → LLM → enrichment → linking → output → CLI → web upload → Nextcloud sync → OrgMyLife → git → Docker/deploy → tests. Property-based tests use Hypothesis and are placed close to the code they validate. ## Tasks - [x] 1. Project scaffolding and configuration - [x] 1.1 Create `NoteGraph/ingestion/` package structure with `__init__.py`, `__main__.py`, and all subpackages (`extraction/`, `enrichment/`, `linking/`, `output/`, `integrations/`, `tests/`) - Create directory tree matching the design's package structure - Add empty `__init__.py` files in each subpackage - Create `__main__.py` with `from ingestion.cli import cli; cli()` entry point - _Requirements: 13.1_ - [x] 1.2 Create `NoteGraph/ingestion/pyproject.toml` with all dependencies - Include: click, watchdog, pytesseract, PyPDF2, pdfplumber, python-docx, litellm, httpx, pydantic, pydantic-settings, fastapi, uvicorn, python-multipart - Add dev dependencies: pytest, pytest-asyncio, hypothesis, pytest-mock, respx - Configure package metadata and entry points - _Requirements: 7.1, 8.7_ - [x] 1.3 Implement `NoteGraph/ingestion/config.py` with `IngestionConfig` (pydantic-settings) - Define all env vars: `LLM_PROVIDER`, `LLM_MODEL`, provider API keys, `ORGMYLIFE_API_URL`, `ORGMYLIFE_API_KEY`, `NEXTCLOUD_INBOX_URL`, `NEXTCLOUD_INBOX_USER`, `NEXTCLOUD_INBOX_PASSWORD`, `SB_USER`, paths, OCR settings, confidence threshold - Implement `llm_model_string` property for LiteLLM routing - Implement `validate_api_key()` that exits with clear error naming the missing variable - _Requirements: 8.1, 8.2, 8.3, 8.4, 8.5, 8.6_ - [x] 1.4 Create `NoteGraph/.env.example` with all supported environment variables and descriptions - Document every variable with comments explaining purpose and valid values - Include placeholders for: LLM keys, OrgMyLife, Nextcloud inbox, SB_USER - Ensure `.env` is in `.gitignore` - _Requirements: 8.5, 8.7_ - [x] 2. Text extraction layer - [x] 2.1 Implement `extraction/base.py` with `TextExtractor` protocol and `ExtractionResult` dataclass - Define `can_handle(file_path: Path) -> bool` and `extract(file_path: Path) -> ExtractionResult` - `ExtractionResult` fields: text, source_file, extraction_method, metadata - _Requirements: 2.1, 3.1_ - [x] 2.2 Implement `extraction/text.py` for plain text and markdown files - Handle `.md` and `.txt` extensions - Read as UTF-8, attempt encoding detection for non-UTF-8 files (chardet/charset-normalizer) - Preserve existing formatting, links, and Obsidian-style wiki-links - _Requirements: 3.1, 3.2, 3.4, 3.5, 3.6_ - [x] 2.3 Implement `extraction/pdf.py` for PDF text extraction - Use pdfplumber for text-based PDFs - Detect scanned pages (pages with no extractable text) and delegate to OCR - Handle mixed PDFs (text + scanned pages) by concatenating results - _Requirements: 2.2, 2.3, 2.4_ - [x] 2.4 Implement `extraction/ocr.py` for image OCR via pytesseract - Handle `.jpg`, `.jpeg`, `.png` files - Configure language as `deu+eng` (German + English) - Preserve paragraph structure by detecting line breaks - Return empty text (not error) when OCR cannot extract readable content - _Requirements: 2.1, 2.2, 2.5, 2.6_ - [x] 2.5 Implement `extraction/docx.py` for DOCX files - Use python-docx to extract text preserving heading structure and paragraph breaks - Handle `.docx` extension - _Requirements: 3.3_ - [ ]* 2.6 Write property test for text extraction identity (Property 3) - **Property 3: Text extraction preserves content for text-based formats** - For any valid UTF-8 string written to a .md or .txt file, extraction returns identical content - **Validates: Requirements 3.1, 3.2, 3.4** - [x] 3. File discovery - [x] 3.1 Implement `discovery.py` with recursive file discovery and extension filtering - Support extensions: `.pdf`, `.jpg`, `.jpeg`, `.png`, `.md`, `.txt`, `.docx` - Accept both single file path and directory path - Recursive directory traversal - _Requirements: 1.1, 1.2, 1.6_ - [ ]* 3.2 Write property test for file discovery filtering (Property 1) - **Property 1: File discovery returns only supported extensions** - For any directory tree with arbitrary extensions, discovery returns exactly those with supported extensions - **Validates: Requirements 1.1, 1.6** - [ ] 4. LLM provider layer - [x] 4.1 Implement `integrations/llm.py` with LiteLLM wrapper - `LLMClient` class with `complete(messages, response_format)` method - Route to provider based on `config.llm_model_string` (e.g., `openai/gpt-4o`, `anthropic/claude-sonnet-4-20250514`) - Implement retry logic: 3 retries with exponential backoff for rate limits, timeouts, 5xx errors - Raise descriptive exception on failure including provider name and error details - _Requirements: 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7_ - [ ]* 4.2 Write property test for provider routing (Property 9) - **Property 9: Provider routing from configuration** - For any valid provider name, configuring `LLM_PROVIDER` causes the LLM client to route to the corresponding provider - **Validates: Requirements 7.3, 8.1** - [ ]* 4.3 Write property test for missing API key error (Property 10) - **Property 10: Missing API key produces descriptive error** - For any provider requiring an API key, when the env var is unset, validation raises an error naming the missing variable - **Validates: Requirements 7.4, 8.6** - [x] 5. Enrichment agent - [x] 5.1 Implement `enrichment/models.py` with Pydantic models - `Entity` model: type (person/project/date/action_item), value, confidence (0.0-1.0), position - `ActionItem` model: description, assignee, deadline (ISO 8601) - `EnrichmentResult` model: title, category, tags, entities, action_items, summary - _Requirements: 4.1, 4.2, 4.3, 4.4, 4.5_ - [x] 5.2 Implement `enrichment/prompts.py` with structured LLM prompts - System prompt requesting structured JSON output for entity detection - Include context about known projects and people (read from existing NoteGraph pages) - Request confidence scores for each entity - _Requirements: 4.1, 4.2, 4.3, 4.4, 4.5_ - [x] 5.3 Implement `enrichment/agent.py` with enrichment orchestration - Call LLM with extracted text and structured prompt - Parse response into `EnrichmentResult` - Filter entities by confidence threshold (>= 0.7) - Handle LLM parse errors: retry once with stricter prompt, then use defaults - _Requirements: 4.1, 4.2, 4.3, 4.4, 4.5, 4.6_ - [ ]* 5.4 Write property tests for entity parsing (Properties 4, 5) - **Property 4: Entity parsing extracts all entities with valid structure** - For any valid enrichment response JSON, each entity has valid type, non-empty value, confidence in [0.0, 1.0] - **Property 5: Confidence threshold filtering** - For any list of entities, filtered output contains exactly those with confidence >= 0.7 - **Validates: Requirements 4.1, 4.2, 4.3, 4.4, 4.5, 4.6** - [x] 6. Checkpoint - Ensure all tests pass - Ensure all tests pass, ask the user if questions arise. - [x] 7. Reference linker - [x] 7.1 Implement `linking/linker.py` with wiki-link generation - Person entities → `[[people/firstname-lastname]]` (slugified) - Project entities → `[[projects/project-slug]]` (slugified) - Insert wiki-link inline at entity position (first occurrence only) - Leave subsequent mentions of the same entity as plain text - Add dates to frontmatter `date` field - _Requirements: 5.1, 5.2, 5.3, 5.5, 5.6_ - [x] 7.2 Implement `linking/stubs.py` for stub page creation - Check if target page exists in NoteGraph notes directory - Create stub page with entity name as title if it doesn't exist - Stub pages go in `notes/people/` or `notes/projects/` as appropriate - _Requirements: 5.4_ - [ ]* 7.3 Write property tests for wiki-link generation (Properties 6, 7) - **Property 6: Wiki-link generation follows correct format per entity type** - For any person name → `[[people/slugified-name]]`; for any project name → `[[projects/slugified-name]]` - **Property 7: Only first entity occurrence is linked** - For any text with N >= 2 occurrences of the same entity, only the first is converted to a wiki-link - **Validates: Requirements 5.1, 5.2, 5.6** - [x] 8. Output renderer and folder routing - [x] 8.1 Implement `output/renderer.py` for markdown + YAML frontmatter rendering - Render frontmatter with: title, date, tags, people, projects, source (file + timestamp) - Render body with wiki-links inserted - Render `## Action Items` section at the end with extracted action items - _Requirements: 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 9.2_ - [x] 8.2 Implement `output/router.py` for category-based folder routing - "meeting" → `notes/meetings/` - "project" → `notes/projects/` - "decision" → `notes/decisions/` - Default/inbox → `notes/inbox/` - Support `--output-dir` override - _Requirements: 10.1, 10.2, 10.3, 10.4_ - [x] 8.3 Implement `output/naming.py` for filename generation with collision avoidance - Pattern: `YYYY-MM-DD-slugified-title.md` - Slug: lowercase alphanumeric + hyphens only - Append numeric suffix (`-2`, `-3`) if file already exists - _Requirements: 10.5, 10.6_ - [ ]* 8.4 Write property tests for output (Properties 8, 11, 13, 14, 15) - **Property 8: Frontmatter round-trip validity** - For any valid EnrichmentResult, render to markdown then parse frontmatter → valid YAML with all required fields - **Property 11: Action items appear in output markdown section** - For any list of action items, rendered markdown contains `## Action Items` section listing every item - **Property 13: Category-based folder routing** - For any category, output path is within the correct folder - **Property 14: Output filename follows date-slug pattern** - For any date and title, filename matches `YYYY-MM-DD-slugified-title.md` - **Property 15: Filename collision avoidance** - When a file with the same name exists, a numeric suffix is appended - **Validates: Requirements 6.1, 6.3, 6.4, 6.6, 6.7, 9.2, 10.1, 10.2, 10.3, 10.5, 10.6** - [x] 9. Pipeline orchestration - [x] 9.1 Implement `pipeline.py` with main processing pipeline - Orchestrate: discovery → extraction → enrichment → linking → rendering → file write → git commit - Track progress: current file number, total count, file name - Handle errors per file: log + skip + continue - Output summary: total processed, success count, failed count - Support `--dry-run` mode (display what would be created without writing) - _Requirements: 1.3, 1.4, 1.5, 13.6_ - [ ]* 9.2 Write property tests for pipeline resilience and dry-run (Properties 2, 17) - **Property 2: Bulk import resilience — failed files do not block remaining processing** - For any batch where a subset fails, remaining files are processed successfully - **Property 17: Dry-run produces no file writes** - For any input with `--dry-run`, zero files are created on disk - **Validates: Requirements 1.4, 1.5, 13.6** - [x] 10. CLI interface - [x] 10.1 Implement `cli.py` with Click-based CLI - `import` subcommand: accepts path argument (file or directory) - `watch` subcommand: starts drop-folder watcher - Global flags: `--verbose`, `--dry-run`, `--no-commit` - Import flags: `--output-dir`, `--create-tasks`, `--provider` - Display usage help with examples when invoked without arguments - _Requirements: 13.1, 13.2, 13.3, 13.4, 13.5, 13.6, 13.7_ - [x] 11. Checkpoint - Ensure all tests pass - Ensure all tests pass, ask the user if questions arise. - [ ] 12. Web upload endpoint (FastAPI) - [x] 12.1 Implement `web/upload.py` with FastAPI app for file upload - HTML form with file upload accepting PDF, JPG, PNG - Mobile-friendly responsive HTML page (inline CSS, no external dependencies) - Basic auth protection using `SB_USER` env var (same credentials as SilverBullet) - On upload: save file to inbox folder, trigger processing pipeline - Return success/error response to the user - _Requirements: 11.2 (extends with web input)_ - [x] 12.2 Create `web/templates/upload.html` mobile-friendly upload page - Simple, clean form with drag-and-drop support - Responsive design for Android "Share to URL" workflow - Show upload progress and result feedback - _Requirements: (additional requirement: web upload for mobile)_ - [ ] 13. Nextcloud WebDAV inbox sync - [x] 13.1 Implement `integrations/nextcloud.py` with WebDAV folder polling - Poll `NEXTCLOUD_INBOX_URL` for new files using httpx WebDAV (PROPFIND/GET) - Download new files to local inbox for processing - After successful processing, move files to a `processed/` subfolder on Nextcloud (MOVE request) - Auth via `NEXTCLOUD_INBOX_USER` and `NEXTCLOUD_INBOX_PASSWORD` env vars - _Requirements: (additional requirement: Nextcloud WebDAV sync)_ - [x] 13.2 Add `sync` subcommand to CLI for Nextcloud polling - One-shot sync: poll once, process new files, exit - Can be called from cron or systemd timer - _Requirements: (additional requirement: Nextcloud WebDAV sync)_ - [x] 14. Drop-folder watcher - [x] 14.1 Implement `watcher.py` with watchdog-based folder monitoring - Monitor inbox folder for new files - Trigger pipeline on new file detection - Move processed files to `archive/` subdirectory - Move failed files to `failed/` subdirectory with error log - Ignore temporary files (names starting with `.` or ending with `.tmp`) - Process existing files in inbox on startup before entering watch mode - _Requirements: 11.1, 11.2, 11.3, 11.4, 11.5, 11.6_ - [ ]* 14.2 Write property test for temporary file filtering (Property 16) - **Property 16: Temporary file filtering** - For any filename, ignore logic returns true iff name starts with `.` or ends with `.tmp` - **Validates: Requirements 11.5** - [x] 15. OrgMyLife integration - [x] 15.1 Implement `integrations/orgmylife.py` with HTTP client for task creation - POST to `/api/tasks` with title, source_url, description - Auth via `ORGMYLIFE_API_KEY` header - Handle unreachable API gracefully: log warning, continue without creating tasks - _Requirements: 9.3, 9.4, 9.5, 9.6_ - [ ]* 15.2 Write property test for OrgMyLife task payload (Property 12) - **Property 12: OrgMyLife task payload correctness** - For any action item description and note path, payload has correct title and source_url - **Validates: Requirements 9.4** - [x] 16. Git auto-commit - [x] 16.1 Implement `integrations/git.py` for git commit operations - Stage new markdown files and stub pages - Commit with message format: `ingestion: import N notes from [source-type]` - Support `--no-commit` flag to skip - Handle git failures gracefully: log warning, don't fail the import - _Requirements: 12.1, 12.2, 12.3, 12.4, 12.5_ - [ ]* 16.2 Write property test for git commit message format (Property 18) - **Property 18: Git commit message format** - For any count N and source type, message matches `ingestion: import N notes from [source-type]` - **Validates: Requirements 12.2** - [x] 17. Checkpoint - Ensure all tests pass - Ensure all tests pass, ask the user if questions arise. - [x] 18. Docker integration and deployment - [x] 18.1 Create `NoteGraph/ingestion/Dockerfile` for the ingestion service - Python 3.11+ base image - Install system dependencies: tesseract-ocr, tesseract-ocr-deu - Install Python package with all dependencies - Run FastAPI upload server (uvicorn) as default command - Also support CLI entry point for import/watch/sync commands - _Requirements: 13.1_ - [x] 18.2 Update `NoteGraph/docker-compose.yml` to add ingestion service - Add `ingestion` service with build context `./ingestion` - Mount `./notes` volume for output - Mount `./ingestion/inbox` for drop-folder input - Expose upload endpoint port (8001) - Pass all env vars from `.env` - Add watcher mode as default or separate service - _Requirements: 11.1_ - [x] 18.3 Update `NoteGraph/.github/workflows/deploy.yml` with all secrets - Write ALL secrets from GitHub secrets into `.env` on VPS: LLM keys (OpenAI, Anthropic, Google, Mistral), Nextcloud inbox creds, OrgMyLife API key, SB_USER - Follow same pattern as OrgMyLife deploy workflow (ENV_CONTENT with printf) - Add `docker compose up -d --build` to rebuild with ingestion service - _Requirements: 8.5 (additional requirement: secrets via GitHub Actions)_ - [x] 19. Final checkpoint - Ensure all tests pass - Ensure all tests pass, ask the user if questions arise. ## Notes - Tasks marked with `*` are optional property-based tests and can be skipped for faster MVP - Each task references specific requirements for traceability - Checkpoints ensure incremental validation - Property tests use Hypothesis with `@settings(max_examples=100)` - The web upload endpoint and Nextcloud sync are additional input methods beyond the original CLI/watcher - All secrets are managed via GitHub Actions → `.env` on VPS, never committed to git