diff --git a/README.md b/README.md index 9d4de259..25e40510 100644 --- a/README.md +++ b/README.md @@ -23,26 +23,26 @@ An operating system for AI agents. Not a chatbot wrapper. Not a prompt chain. A ## Current State: Beta -**It works.** 14 of 15 branches are operational, tested, and communicating. 53 PRs merged. 30+ orchestration sessions. We're past prototyping and into hardening — building the infrastructure that makes the system reliable at scale. +**It works.** All 15 branches operational. 95+ PRs merged. 38 orchestration sessions. 733 tests across the system. 173 drone commands discovered. The system is past prototyping — we're in the hardening phase, building diagnostic tooling and iterating branch by branch. **Recently completed:** -- **CLI front door + init anywhere** — full discovery chain rebuilt so users can find and use `aipass init` through natural exploration. CLI rewritten to seedgo's auto-discovery pattern. Registered as a drone internal module — `drone @cli aipass init` works from any directory on any system, before any project exists. Creates a registry (UUID), passport, `.trinity/` memories, prompt files, and hooks. Every project is fully self-contained — no shared state, no cross-contamination between workspaces. -- **Memory introspection overhaul** — memory branch rewired to seedgo-compliant CLI discovery. Rollover and search introspection rebuilt. 18+ files ported back from archives (symbolic reasoning, templates, pool processor). Templates modernization in progress (FPLAN-0052). -- **Seedgo standards enforcement** — checklist command fixed, checker false positives addressed, handler guards added. PR #52 open. -- **Credential model** — UUID-based registry matching is live. Every project gets a unique registry ID, every citizen's passport carries it. `aipass init` creates a new project in any directory with its own credentials. -- **Stderr routing** — system-wide migration across 10 branches (48 files). CLI owns the display layer (`error()`, `warning()`, `fatal()` all route to stderr). Seedgo enforces it with an automated checker. -- **Dashboard pipeline** — Prax owns the dashboard end-to-end. Per-branch `STATUS.local.md` files sync to a central `STATUS.md` via `drone @prax status sync`. -- **Drone test suite** — 188 tests across 5 files. Interactive mode for human-facing commands. Credential verification with dedicated error hierarchy. -- **System governance** — git workflow, commit signing (`Co-Authored-By: @branch`), DPLAN/FPLAN documentation, and "How to Work" guidelines all codified in the global prompt. +- **Dispatch UX redesign** — `drone @ai_mail dispatch @target "Subject" "Body"` sends + wakes in one command. `--fresh` flag for clean sessions. `email` command for mail-only (no wake). Fully tested. +- **PR v2 workflow** — commit-on-main architecture. Changes never leave your working tree. Feature branches are just pointers for GitHub's PR system. No more disappearing files. +- **Handler guard fix** — cross-branch handler imports blocked by `.py` files from other branches. Command-line `python3 -c` allowed through. 13 branches updated. +- **Memory vectorization fix** — batch processing (2 subprocess calls instead of 228), decoupled from startup trigger, explicit `drone @memory process-plans` command. 113 plan files vectorized in ~1 minute. +- **Diagnostic tooling** — 4 scanners built: dead code (14 unused files found), command inventory (173 commands), prompt quality (6 rich/4 basic/5 stub), test coverage (733 tests, 26% module coverage). +- **Prax monitor** — fully operational with inotify file watching, branch detection, polling fallback with actionable error messages. +- **Plan cleanup** — 60+ FPLANs/DPLANs closed. Templates updated with "close immediately when done" rule. +- **System governance** — git workflow, commit signing, DPLAN/FPLAN documentation, logging/debugging guidelines, `.archive/` pattern all codified in the global prompt. **What we're solving now:** -- **Memory subsystem wiring** — symbolic reasoning, templates, and pool processor modules ported back but need full integration and testing. Templates modernization (spawn handler, v2 schema) in progress. -- **Lint cleanup** — 474 ruff violations remain (mostly unused imports from unwired modules). Blocked until seedgo audit coverage is higher — we can't confidently remove imports until we know what's actually used. -- **Dashboard CLI routing** — the Python API works but `drone @prax dashboard refresh --all` fails because argparse eats the flags before the module sees them. Known bug, fix pending. -- **Cross-platform reliability** — Linux and Windows tested. macOS is structurally supported but needs a dedicated testing pass. All paths use `pathlib`, no hardcoded paths, secrets stored at `~/.secrets/aipass/`. -- **Agent agnosticism** — currently focused on [Claude Code](https://docs.anthropic.com/en/docs/claude-code) (hooks for auto-diagnostics, prompt injection, session recovery). But AIPass is designed to not depend on any single provider. `agents.md` and `gemini.md` can bootstrap the system for Codex and Gemini — you lose hooks but keep the core. The plan is truly agnostic: any agent, anywhere, persistent memory, no vendor lock-in. +- **Branch-by-branch audit** — walking through every branch from devpulse, testing commands, noting issues, dispatching fixes. API branch audit in progress (DPLAN-0029). +- **Local prompt enrichment** — 5 branches still on 14-line stubs (ai_mail, backup, cli, drone, prax). Rich prompts = less babysitting. +- **Test coverage expansion** — 9 branches have zero tests. Building toward comprehensive coverage using the test scanner for visibility. +- **Cross-platform reliability** — Linux and Windows tested. macOS structurally supported. All paths use `pathlib`, secrets at `~/.secrets/aipass/`. +- **Agent agnosticism** — currently focused on [Claude Code](https://docs.anthropic.com/en/docs/claude-code) (hooks for auto-diagnostics, prompt injection, session recovery). But AIPass is designed to not depend on any single provider. `agents.md` and `gemini.md` can bootstrap the system for Codex and Gemini — you lose hooks but keep the core. ## Getting Started @@ -92,7 +92,7 @@ Every branch is a citizen — an expert in its domain with its own memories and |--------|------| | `devpulse` | **Start here.** Orchestration hub — coordinates everything | | `drone` | AI-friendly CLI — every command is a single-line, non-interactive call | -| `seedgo` | Standards enforcement — 24-standard audit pack, system compliance | +| `seedgo` | Standards enforcement — 21-standard audit pack, system compliance | | `prax` | Logging and monitoring (the only logger in the system) | | `cli` | Terminal display, stderr routing, project commands | | `flow` | Workflow management — FPLANs (execution) and DPLANs (design) | diff --git a/STATUS.md b/STATUS.md index 55a53419..75c85915 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,19 +2,19 @@ > Auto-generated by `drone @prax status sync`. Do not edit manually. -**Last sync:** 2026-03-17 19:57 +**Last sync:** 2026-03-20 00:00 **Summary:** 9 operational | 0 in-progress | 0 not started --- -
@ai_mail — Operational (2026-03-17) +
@ai_mail — Operational (2026-03-17 (session 14)) # @ai_mail > Inter-branch email, dispatch, wake, bounce **State:** Operational -**Last update:** 2026-03-17 +**Last update:** 2026-03-17 (session 14) ## Milestones - Send/receive/reply/close lifecycle @@ -24,19 +24,22 @@ - notify-send to dbus notifications - mailbox_path bug fixed (relative→absolute in get_user_by_email, get_all_users) - Health audit: 7 dead handlers archived, branch cleaned +- Seedgo compliance: 93% → 99% (json_structure 100%, introspection 100%) ## Current Work -- 1 inbox email from @api (unit test setup advice) — pending reply +- No active tasks. Inbox clear. ## Known Issues - _find_repo_root() duplicated 9 times across handlers (consolidation candidate) - get_all_branches() has 2 implementations (delivery.py vs registry/read.py) - Registry shows 6 pytest test branches with /tmp paths (test pollution) +- Readme standard at 83% (only non-100% standard remaining) ## Architecture Notes - trigger branch owns error_detected events; ai_mail provides delivery only - extensions/, plugins/, json_templates/, json/ shim, persistence/ = repeatable skeleton (don't archive) - Dead code was inherited from DevPass migration, not accumulated in AIPass +- .seedgo/bypass.json: email.py bypassed for introspection (inbox/sent/contacts need no args), json_utils/json_handler.py bypassed for json_structure (can't import itself)
@@ -73,14 +76,14 @@
-
@backup — Operational — seedgo 100% (all 23 standards) (2026-03-17 (session 9)) +
@backup — Operational — seedgo 100% (all 24 standards) (2026-03-17 (session 10)) # @backup > Multi-mode backup — snapshot, versioned, Google Drive -**State:** Operational — seedgo 100% (all 23 standards) -**Last update:** 2026-03-17 (session 9) +**State:** Operational — seedgo 100% (all 24 standards) +**Last update:** 2026-03-17 (session 10) ## Working Commands - `drone @backup snapshot [--dry-run]` — full copy backup @@ -92,6 +95,9 @@ - 1590 files / 12.8 MB - 235 runtime JSON files filtered (correct behavior) +## Recently Completed +- Session 10: @seedgo dispatch — json_structure 4%→100%, introspection 60%→100%, readme freshness. Re-did session 9 work lost in git collision. + ## Pending - @api dispatch: Migrate Google Drive auth to API branch (email c19c8d0f — parked) - integrations.py: locally archived to .archive but git state may have it back in backup.py imports — reconcile on next session @@ -139,47 +145,64 @@
-
@commons — Operational — Seedgo 99% (2026-03-17) +
@commons — Operational — Seedgo 99% (2026-03-18) # @commons > Social network for branches — posts, rooms, artifacts **State:** Operational — Seedgo 99% -**Last update:** 2026-03-17 +**Last update:** 2026-03-18 ## Seedgo Compliance (99%) 23/24 categories at 100%. Only Readme at 83% (minor tree check). -| Category | Score | -|----------|-------| -| Architecture, Cli, Cli_Flags, Documentation | 100% | -| Encapsulation, Error_Handling, Handlers, Imports | 100% | -| Introspection, Json_Structure, Log_* (all 4) | 100% | -| Meta, Modules, Naming, Permission_Flags | 100% | -| Shebang, Stderr_Routing, Testing, Trigger | 100% | -| Diagnostics | 100% | -| **Readme** | **83%** | +## Session 4 Summary (2026-03-17/18) -## Session 3 Summary (2026-03-17) -- DPLAN-0074: Full compliance & system verification -- FPLAN-0076: Execution plan -- Phase 0: Smoke test PASSED -- Phase 1: Quick wins (stderr, readme date, dropbox, diagnostics) -- Phase 2: Handler security guard (encapsulation 66%→100%) -- Phase 3: Introspection gates (52%→100% with bypasses) -- Phase 4: json_handler wired into 56 files (1%→100%) -- Fixed log template bug (dict→list format) -- Score: 91% → 99% +### Prax Log Routing Fix +- Stray `src/aipass/commons/` directory spotted by Patrick +- Root cause: prax `detect_branch_from_path()` didn't handle commons' non-standard path +- Dispatched to prax, who fixed 3 files. Verified: logs now write to `src/commons/logs/` + +### 5-Branch Social Stress Test +- Tested with commons, drone, spawn, flow, memory +- 40+ commands: posts, comments, votes, rooms, artifacts, explore, search +- Identity detection via CWD: bulletproof across all 5 branches + +### 3 Integration Gaps Fixed (DPLAN-0018, FPLAN-0097) +All same pattern: helper functions existed from Dev-Pass port but were never wired into production code. + +1. **FTS5 search** — `sync_post_to_fts()`/`sync_comment_to_fts()` never called from post_ops/comment_ops. Added calls + `backfill_fts_index()` (indexed 6 posts + 17 comments) +2. **Visitor tracking** — No `room_visits` table. Added table to schema, `record_visit()` to space_ops, recording in `_cmd_enter()`. Updated `get_visitors_data()`. +3. **Post/comment counts** — `increment_post_count()`/`increment_comment_count()` never called. Catchup nudge showed wrong "haven't posted" tip. Wired into post_ops/comment_ops. + +### Files Modified +- `apps/handlers/posts/post_ops.py` — FTS sync + post count +- `apps/handlers/comments/comment_ops.py` — FTS sync + comment count +- `apps/handlers/search/search_queries.py` — backfill_fts_index() +- `apps/handlers/rooms/space_ops.py` — record_visit() + updated visitors query +- `apps/modules/space_module.py` — visit recording in enter +- `apps/handlers/database/schema.sql` — room_visits table +- `tools/backfill_fts.py` — one-time FTS backfill script + +## Session 5: Module Rename (2026-03-18) + +Removed `_module` suffix from all 21 module files per seedgo naming standard ("Path = Context, Name = Action"). DPLAN-0019 → FPLAN-0099 master plan. + +- 21 agents deployed in 2 batches (10 + 11), running in parallel +- Each agent: renamed .py file, updated header + introspection string, renamed JSON tracking files +- Orchestrator: updated commons.py static import, .seedgo/bypass.json (9 entries) +- Old modules archived to `apps/modules/.archive/` +- System verified: all commands working, seedgo 99% ## Known Issues - Readme 83% — likely missing dropbox/ in directory tree - Pyright warnings (unused logger imports) — pre-existing, cosmetic ## Architecture -- 86 Python files, 21 modules, 19 handler domains -- SQLite + WAL + FTS5, 16 tables +- 86+ Python files, 21 modules, 19 handler domains +- SQLite + WAL + FTS5, 17 tables (added room_visits) - Lives at src/commons/ (standalone, outside aipass namespace) - Imports own code as `from commons.apps.*` - Uses aipass services: prax (logging), cli (display) @@ -223,61 +246,57 @@ **Last update:** 2026-03-18 ## Milestones -- 35 sessions of system coordination -- drone @git module — atomic PR workflow with lockfile (DPLAN-0008 complete) -- STATUS board system (per-branch + central aggregation) -- Prompt architecture (breadcrumbs pattern) -- dev.local.md → STATUS.local.md consolidation -- Dispatch+wake workflow (backup, flow, cli, prax all verified) -- First PR review cycle — 8 PRs reviewed+merged in one session +- 36 sessions of system coordination +- drone @git pr v2 — commit-on-main workflow (no disappearing files) +- Plan cleanup — 60+ FPLANs/DPLANs closed, templates updated with "close immediately" rule +- Prax monitor fully operational — UNKNOWN→PRAX, inotify restored, polling fallback with actionable messaging +- Global prompt refined — logging/debugging, .archive pattern, local-ahead-of-origin, console vs logs +- System cleanup — 12 notepad items investigated and resolved (commons DB, memory central_writer, registry ghosts, .seed bypass, unknown_branch) +- Bounce system confirmed working end-to-end ## Current Work -- **DPLAN-0008 COMPLETE**: drone @git module built, tested, deployed. All 3 phases done. Global prompt updated, 13 branch deny lists deployed. Branches use `drone @git pr` only. -- **Stale term cleanup**: Scanner tool built (`tools/dev_central_to_devpulse.py`). SOP written. 484 "dev central" hits across 155 files. 5 categories (A-E). Strategy: branch-by-branch, sub-agents handle A/B/E, plan C manually, leave D. -- DPLAN-004: Dashboard pipeline — argparse bug still present (--all/--branch flags eaten by prax.py). +- **Memory plan vectorization crisis**: Per-file subprocess cold-loads torch (~500MB) per plan file. 114 files × concurrent invocations = system crush. Wired into startup trigger so EVERY drone command fires it. Memory investigating with Dev-Pass comparison. Awaiting report. +- **Dispatch UX ref sweep**: ai_mail + trigger updated. Prax, daemon, flow have emails in inbox (not yet woken). +- **Local prompt gap**: 7 of 15 branches on 14-line stubs. Root cause of babysitting load vs Dev-Pass. - DPLAN-003: Phase 3 (drone aipass help) pending. -- Ruff CI: 474 violations remaining. No cleanup until seedgo coverage higher. ## Known Issues -- **prax dashboard CLI routing**: `drone @prax dashboard refresh --all` fails — argparse eats flags before module. Python API works fine. -- **flow/DPLAN CWD default**: DPLANs always go to flow's dev_planning/ regardless of caller's CWD. +- **CRITICAL: memory plan processing**: (1) per-file subprocess loads torch each time instead of batching, (2) no concurrency lock — multiple instances race same files, (3) wired into startup trigger — every drone command triggers heavy ML ops. Path fix in session 36 unmasked the bug (was silently finding 0 files before). +- **prax dashboard CLI routing**: argparse eats flags before module. - **api**: `models` command not routed through drone -- **ai_mail**: `get_current_user()` returns relative `mailbox_path` — causes doubled paths in reply -- **drone**: versioned backup exceeds 30s subprocess timeout -- **seedgo self-violation**: stderr_routing_check.py fails its own standard -- **ruff CI**: 474 lint violations (331 unused imports, 87 unused vars, etc). Backlog — don't auto-fix until seedgo coverage is 100%. -- **wake.py no --model flag**: dispatched branches use CLI default model. Can't control from our side. +- **ruff CI**: 474 lint violations. Backlog — don't auto-fix until seedgo coverage is 100%. +- **wake.py no --model flag**: dispatched branches use CLI default model. ## Todos -- [ ] Close stale FPLANs (0017, 0021) once verified complete +- [ ] Local prompt enrichment pass — 7 branches on stubs (drone, prax, cli, ai_mail, backup, api, spawn) +- [ ] Devpulse wake protection — daemon blocks dispatch but manual wake has no blocklist. Need full protection like Dev-Pass DevCentral. +- [x] DPLAN-0017: dispatch UX redesign — DELIVERED. dispatch/email/fresh all tested. +- [ ] Handler guard too aggressive — blocks `python3 -c` even for same-branch agents. Guard should only block cross-branch .py file imports, not command-line execution. Seedgo owns the `__init__.py` template. All 14 branches affected. - [ ] DPLAN-003 Phase 3: `drone aipass help` module -- [ ] Email prax about argparse bug (--all/--branch still broken despite PR #47) -- [ ] Seedgo: single-file checklist command broken (`drone @seedgo checklist ` fails silently) - [ ] Ruff cleanup FPLAN (after seedgo coverage reaches 100%) -- [ ] Add --model support to wake.py for dispatched branches +- [ ] Add --model support to wake.py ## Recently Completed -- Session 35: DPLAN-0008 complete — drone @git module built+tested+deployed, global prompt+deny lists, flow seedgo redo via drone @git pr. PRs #76-78 merged. (2026-03-18) -- Session 34: CLAUDE.md inbox startup, culture doc refs, venv→symlinks, setup.sh symlink install, ghost input bug research (2026-03-17) -- Session 33: Stale term scanner+SOP, dispatched ai_mail+trigger cleanup, PR #60 (2026-03-16) -- Session 32: README update (init-anywhere), git workflow fix (return to main after PR — tested, works), merged PRs #54-56, cleaned 8 stale branches, new principle (2026-03-16) -- Session 31: CLI front door (seedgo-compliant discovery + drone internal module), statusline git branch, Rich colors, DPLAN-0044, aipass init external test, 6 research agents, OpenClaw comparison, PR #51 (2026-03-15) -- Session 29: Merged 8 PRs (#39-48), ruff investigation (612→474), dispatch reply breadcrumb, /prep broadened (2026-03-14) -- Session 28: Global prompt governance, DPLAN/FPLAN, Git Workflow, How-to-Work, PR #41+#42 merged (2026-03-14) -- Prax Phase 1+2: dashboard devpulse section removed, refresh pipeline 4 bugs fixed (2026-03-14) -- CLI init_project.py migration: FPLAN-0041 done, 99% seedgo, all tests passed (2026-03-14) -- FPLAN-0033 Phase 3 stderr migration — 10 branches, 48 files (2026-03-14) -- DPLAN-003 Stage 1 — credential model complete (2026-03-14) +- Session 36: PR v2 (commit-on-main). Global prompt (logging, .archive, git awareness). Notepad cleanup (12 items). Prax monitor fixed (UNKNOWN, inotify, introspection). Registry healed (6 ghosts removed). Memory central_writer path fixed. Commons DB path fixed. (2026-03-18) +- Session 35: DPLAN-0008 complete — drone @git module. PRs #76-86. (2026-03-18) +- Session 34: CLAUDE.md inbox startup, venv→symlinks. (2026-03-17) +- Session 33: Stale term scanner+SOP, dispatched cleanup. (2026-03-16) ## Notepad > **SCRATCH SPACE — gets wiped at session start or topic change.** -- Session 35 COMPLETE. -- DPLAN-0008 fully done — all 3 phases. Module built, tested, prompt updated, deny lists deployed. -- settings.local.json files are gitignored — deny lists are local enforcement only. Spawn template is source of truth for new clones. -- memory_bank.central.json untracked in src/aipass/central/ — Patrick knows, needs to be moved. -- Culture doc still pending Patrick review. -- FPLAN-0094 should be closed (drone @git module complete). +Session 37: +- DPLAN-0017 DELIVERED + TESTED. dispatch/email/fresh all confirmed working on prax. +- Dispatched ref updates to: ai_mail (done), trigger (done), flow (email landed, blocked by interactive), prax (killed mid-dispatch), daemon (killed mid-dispatch). +- Memory vectorization incident: 5 concurrent drone commands each triggered startup→check_and_rollover→process_plans→per-file torch subprocess. System pegged at 100% CPU. Had to kill all processes. Root cause: session 36 path fix unmasked the bug — before fix, process_plans found 0 files and returned instantly. +- Inbox hang was NOT ai_mail — confirmed. System was CPU-starved from memory's torch subprocesses. +- Devpulse wake protection gap found: daemon has blocklist but manual wake doesn't. +- Memory incident RESOLVED: batch mode + decoupled from startup + explicit command (PR #94, #95). 113 files processed in ~1 min at ~20% CPU. +- Memory emailed trigger to review _run_memory_bank_check() in startup.py — still calls check_and_rollover() on every drone command. +- Prax, daemon still have `send`→`email` ref update emails in their inboxes — low priority, will pick up next wake. +- Flow has template update email + older pending items. +- 100K+ trigger_core.log entries from the incident — prax handling the overflow. +- PRs pending on GitHub: #89, #90, #94, #95, plus earlier flow/prax work.
@@ -323,53 +342,72 @@
-
@flow — In progress — seedgo compliance (FPLAN-0096) (2026-03-17) +
@flow — In progress — health check (DPLAN-0020) (2026-03-18) # @flow -> Plan lifecycle — FPLANs (building) + DPLANs (planning) via unified plugin architecture +> Plan lifecycle — FPLANs, DPLANs, and any custom plan type via filesystem-driven template registry -**State:** In progress — seedgo compliance (FPLAN-0096) -**Last update:** 2026-03-17 +**State:** In progress — health check (DPLAN-0020) +**Last update:** 2026-03-18 ## Current Work -- FPLAN-0096: Seedgo compliance — 40 tracked files modified - - json_handler wired into 33 handlers + 7 modules - - Introspection gates added to 7 modules - - Awaiting: zombie dplan cleanup, README update, final audit +- DPLAN-0020: Flow health check — session 9 deep dive with Patrick + - Introspection/help: DONE (all 8 modules) + - Template rename + auto-discovery: DONE + - Template registry + auto-heal: DONE (delete dir → full cleanup) + - No-fallback errors: DONE + - --dry-run for close: DONE + - Registry cleanup: DONE (Docker artifacts, orphan FPLANs) + - Post-close pipeline: DONE (foreground archival, flags set atomically) + - Vector intake: DONE (process-plans + verify, visible in close output) + - Remaining: dashboard push fix, README update, devpulse template updates + +## Recently Completed (Session 9) +- Foreground archival: moved from background subprocess to close_ops.py, fixing race condition +- Registry flags: processed/cleanup_completed now set atomically on close +- Vector pipeline: drone @memory process-plans + is_plan_vectorized() wired into close output +- Template auto-heal: load_registry() prunes orphaned types + deletes plan registry JSON +- --dry-run flag: drone @flow close --all --dry-run +- Registry cleanup: removed 4 Docker-path entries, 7 orphan FPLANs archived +- Silent fallback removed from mbank/process.py vector intake +- Memory bug verified fixed (CWD), devpulse dispatch replied (--dry-run + orphan heal) +- Closed DPLAN-0022, 0024, 0026 +- Cross-branch coordination: 3 email exchanges with @memory → path mismatch found + fixed ## Pending -- Delete zombie `apps/handlers/dplan/` (restored from archive by agent, not tracked in git) -- README update (do LAST with full system context, not just date bump) -- Final seedgo audit to verify 100% -- Reply to @memory bug report (email d2bde448: FPLAN create not registering) +- 2 devpulse dispatches: FPLAN template syntax updates + list→list open docs +- Dashboard push failing on close (non-critical, separate issue) +- README update (template system, new commands, pipeline changes) +- Final seedgo audit ## Active Plans -- DPLAN-0016: Seedgo compliance design doc -- FPLAN-0096: Seedgo compliance master plan +- DPLAN-0020: Flow health check (near completion) -## Open PR -- None currently - -## Key Learnings (Session 6) -- Work on main, only create branches at PR time via `drone @git pr` -- Zombie files on disk inflate seedgo audit counts — 15 dplan files were already archived -- Real tracked scope was 40 files, not 55 +## New Commands (Session 8-9) +``` +drone @flow register # Register plan type +drone @flow unregister # Remove plan type +drone @flow templates # List registered types +drone @flow scan # Find unregistered directories +drone @flow close --dry-run # Preview what would close +drone @flow close --all --dry-run # Preview close-all +``` ## Known Issues -- 64 orphaned plan files in flow/ root from before CWD fix (need cleanup after merge) -- Default FPLAN template not always fast-deleted on close (cosmetic) +- Dashboard push fails on every close ("Failed to push flow section to branch dashboard") +- plan_type.json files still in templates/ dirs (orphaned — system ignores them)
-
@memory — Operational — 100% seedgo compliant (2026-03-17) +
@memory — Operational — 100% seedgo compliant (2026-03-18 (s16)) # @memory > Vector memory archive — ChromaDB, sentence-transformers **State:** Operational — 100% seedgo compliant -**Last update:** 2026-03-17 +**Last update:** 2026-03-18 (s16) ## Milestones - ChromaDB integration @@ -377,96 +415,118 @@ - Searchable vector archive - FPLAN-0026 complete (rollover E2E, search, plans archival) - **DPLAN-0073 COMPLETE** — 100% seedgo audit, all 24 standards passing -- memory_json/ auto-populating with operation logs +- **FPLAN-0128 COMPLETE** — verify command + config fix (PR #94) +- **Session 16: CRITICAL FIX** — batch vectorization, decoupled from startup trigger (PR #95) +- 113 plan files vectorized (2019 chunks) in ~1 min at ~20% CPU ## Current Work - No active DPLANs or FPLANs — all caught up -- Completed: DPLAN-0073, DPLAN-0055, DPLAN-0045, FPLAN-0054, FPLAN-0062 +- Future: wire up watchdog watcher via `drone @memory watch` to replace startup trigger dependency ## Known Issues - `search` fails without torch/sentence-transformers installed -- DPLANs not vectorized on close (flow-side issue) -- Flow plan registry broken — `drone @flow create` makes file but doesn't register. Close fails. Emailed @flow. -- Old memory bank had ~4,180 vectors. Current has 127. Full port pending. +- Search returns semantically similar plans, not exact ID matches — use `verify` for ID lookup +- ~~Plan vectorization system overload~~ FIXED in PR #95 (batch mode + decoupled from startup) +- ~~Plan vectorization path mismatch~~ FIXED in PR #95 (correct path: src/aipass/backup/processed_plans) +- ~~Flow plan registry broken~~ FIXED in PR #71 +- Old memory bank had ~4,180 vectors. Current has 2,200+ (207 rollover + 2019 plans). Growing. + +## New Commands +- `drone @memory verify FPLAN-XXXX` — check if plan is vectorized in chroma +- `drone @memory process-plans` — explicit batch vectorization (NOT auto-triggered) + +## Architecture Notes +- Dual vector storage: local (`{branch}/.chroma/`) + global (`memory/.chroma/`). Local fails gracefully, global is critical (restores backup on failure). +- Plan vectorization decoupled from startup trigger. `_check_plans()` = count only. Processing via explicit command. +- Watchdog infrastructure exists in memory_watcher.py but not exposed via drone yet. +- Startup trigger chain: prax logger → trigger.fire('startup') → check_and_rollover(). Keep this lightweight — NEVER put heavy processing here. ## Notepad -- Session 12: 100% seedgo audit achieved. All commands tested and working. -- memory_json/ now has 8 log files auto-created from rollover check/status commands -- spawn has template_owners.json (empty) — memory should populate it as authoritative source for .trinity files +- Session 16: INCIDENT — vectorization overloaded system. 228 subprocess calls × 5 concurrent drone commands. Fixed with batch mode + decoupling. PR #95. Full report sent to devpulse. +- Session 15: FPLAN-0128 — verify command + config fix. PR #94. +- Session 14: Processed 4 @flow emails. Path mismatch root cause found. +- Emailed trigger about startup.py _run_memory_bank_check() weight — needs review - handlers/__init__.py guard blocks cross-branch imports — always expose via modules layer -- Post-rollover chain wired: Trigger.fire → central_writer → dashboard_push → pool_processor -- Memory at 28/25 key_learnings — rollover will trigger on next addition
-
@prax — Operational (2026-03-17) +
@prax — Operational (2026-03-18) # @prax > Logging, monitoring, dashboard infrastructure **State:** Operational -**Last update:** 2026-03-17 +**Last update:** 2026-03-18 ## Milestones - System-wide logging via `from aipass.prax import logger` -- Monitor command (passive + interactive) +- Monitor: polling fallback, caller attribution, agent tracking (Phases 1-3) - Dashboard infrastructure -- Two-tier logging simplification (PR #59, merged) -- Seedgo audit: 93% → 98% (session 13, PR #65 pending) +- Seedgo audit: 100% (25/25 categories) ## Current Work -- PR #65 pending merge (json_structure + stderr + dead code cleanup) -- 48 active files (9 dead handlers archived) +- No active plans — all closed +- 4 active commands: monitor, status, log-audit, dashboard +- 5 dead modules archived (init, shutdown, run, terminal, discover) -## Seedgo Audit: 98% -- 23 of 24 categories at 100% -- Log_Structure: false positives — seedgo PR #63 fixes the checker -- After both PRs merge (#63 + #65) → expect 100% +## Seedgo Audit: 100% +- 25 of 25 categories at 100% -## Pending Cross-Branch -- @seedgo PR #63 (log_structure checker fix) — confirmed fixed, awaiting merge +## Recently Completed (Session 17) +- Introspection standard: all modules compliant +- Monitor Phase 1: PollingObserver fallback, degradation notices, clean startup +- Monitor Phase 2: Caller attribution via drone [CALLER:BRANCH] markers +- Monitor Phase 3: Agent intelligence (tool use, thinking, responses) +- Archived 5 dead modules, stripped verbose flag, cleaned help +- Found + reported seedgo silent checker bug — fixed same session +- Final 3 introspection violations fixed (status.py, agent_status.py, logger.py) +- Cleaned stale logs/.archive/ ## Known Issues -- None blocking +- System inotify watches near-exhausted. Monitor uses polling fallback — functional but slower. +- Phase 4 (interactive filtering) deferred — not needed now
-
@seedgo — Operational — **100% compliance** (2026-03-17 (session 21)) +
@seedgo — Operational — **13/14 branches at 100%** (but new introspection checks will lower some scores) (2026-03-18 (session 22)) # @seedgo -> Standards enforcement, 24-standard audit pack + hook ownership +> Standards enforcement, 24-standard audit pack + hook ownership + branch dispatch -**State:** Operational — **100% compliance** -**Last update:** 2026-03-17 (session 21) +**State:** Operational — **13/14 branches at 100%** (but new introspection checks will lower some scores) +**Last update:** 2026-03-18 (session 22) -## Session 21 Summary +## Session 22 Summary -1. **log_structure fix** — Removed hierarchical placement post-check from branch_audit.py. Was penalizing prax for having logs. Prax now 100%. -2. **pre_edit_gate.py v1.1** — Branch-scoped. Cross-branch type errors no longer block edits. -3. **Hook audit + DPLAN-0012** — Inventoried all 12 hooks. Found duplicate PostToolUse, stale docs, bridge architecture discussion. Patrick decided: seedgo owns hooks. -4. **Portability architecture** — Bridge pattern (project_bridge.sh → prompt_inject.sh) is the pip install solution. NOT dead scripts. Auto-fix is AIPass-specific, stays in project hooks. -5. **branch_audit.py type fixes** — Path | None, Dict[str, Any]. Phantom revert occurred and was redone. -6. **Emailed Memory** — json_handler clarification for DPLAN-0073. +Dispatch wake — processed 3 emails. + +1. **Introspection checker improved** (Flow dispatch) — Two new checks: + - **Content references** — flags `print_introspection()`/`print_help()` that reference `python3` instead of `drone @` commands + - **Module help interception** — flags modules where `handle_command()` doesn't intercept `--help` +2. **Standard content fixed** — `introspection_content.py` examples updated from `python3` to `drone @` +3. **Devpulse HOLD** — Introspection standard stays as-is. Standards changes only from devpulse/Patrick. +4. **Prax rejected** — Told to use subcommands instead of removing introspection guards. ## Active Plans -- **DPLAN-0012** — Hook Management. PARKED. Architecture mapped, not executing yet. Patrick has other work in progress. -- **DPLAN-0075** — Auto-fix Execution. Two-hook system working in production. Monitor for edge cases. +- **DPLAN-0012** — Hook Management. PARKED. Architecture mapped, not executing yet. +- **DPLAN-0075** — Auto-fix Execution. Working in production. ## TODO -- **DPLAN-0012 execution** — When Patrick is ready. Phase 1: duplicate fix, Phase 2: cross-branch gate, Phase 3: docs -- **seedgo_verify modernization** — old verify module in .archive/, needs rebuild -- **log_structure fix review** — Patrick had mixed feelings on PR #63 fix +- **ai_mail readme** — 99% → 100% (only remaining non-100% branch) +- **System-wide introspection fixes** — New checks will flag python3 refs + missing --help interception across branches +- **DPLAN-0012 execution** — When Patrick is ready +- **seedgo_verify modernization** — old verify module in .archive/ +- **Update branch prompt** — Add dispatch patterns, hook ownership responsibilities ## Known Issues -- Duplicate PostToolUse hooks — auto_fix fires twice per edit (PARKED, don't touch) +- Duplicate PostToolUse hooks — auto_fix fires twice per edit (PARKED) - Hardcoded `/home/patrick/` paths in global settings hooks -- Bridge scripts (project_bridge.sh, prompt_inject.sh) superseded but NOT dead — portability solution -- Phantom file revert on branch_audit.py — cause unknown, watching -- readme_check.py command list check stops at ### sub-headings +- Bridge scripts superseded but needed for portability +- commons at src/commons/ not src/aipass/commons/ — path assumption trap
diff --git a/src/aipass/api/apps/handlers/auth/env.py b/src/aipass/api/apps/handlers/auth/env.py index 256eb039..d64d2d74 100644 --- a/src/aipass/api/apps/handlers/auth/env.py +++ b/src/aipass/api/apps/handlers/auth/env.py @@ -152,17 +152,17 @@ def create_env_template(provider: str = "openrouter", target_path: Optional[Path Args: provider: API provider name (default: 'openrouter') - target_path: Optional custom path (defaults to api/.env) + target_path: Optional custom path (defaults to ~/.secrets/aipass/.env) Returns: bool: True if template created successfully, False otherwise Example: >>> if create_env_template('openrouter'): - ... print("Template created at /.env") + ... print("Template created at ~/.secrets/aipass/.env") """ - # Default to api/.env - env_path = target_path or (API_ROOT / ".env") + # Default to ~/.secrets/aipass/.env (cross-platform standard) + env_path = target_path or (Path.home() / ".secrets" / "aipass" / ".env") # Don't overwrite existing file if env_path.exists(): diff --git a/src/aipass/api/apps/handlers/auth/keys.py b/src/aipass/api/apps/handlers/auth/keys.py index 5d5288c6..2bfab7e3 100644 --- a/src/aipass/api/apps/handlers/auth/keys.py +++ b/src/aipass/api/apps/handlers/auth/keys.py @@ -268,6 +268,53 @@ def get_validation_rules(provider: str) -> Dict[str, Any]: # KEY FORMAT CHECKING # ============================================== +def diagnose_key(provider: str = "openrouter") -> str: + """ + Diagnose why get_api_key() returned None. + + Checks all sources for a raw key (skipping validation) and explains + exactly why it failed — missing entirely, wrong prefix, too short, etc. + + Args: + provider: Provider name (default: 'openrouter') + + Returns: + str: Human-readable explanation of the key issue + + Example: + >>> if not get_api_key('openrouter'): + ... print(diagnose_key('openrouter')) + """ + # Check all sources for raw key (without validation) + key = get_key_from_config(provider) + source = "config" + + if not key: + key = get_key_from_env(provider) + source = "env" + + if not key: + env_var = f"{provider.upper()}_API_KEY" + key = read_env_file(env_var) + source = "dotenv" + + if not key: + return "No API key found in any source (config, environment, .env file)" + + # Key exists but failed validation — explain why + key = key.strip() + rules = get_validation_rules(provider) + + if "prefix" in rules and not key.startswith(rules["prefix"]): + actual_prefix = key[:len(rules["prefix"])] if len(key) >= len(rules["prefix"]) else key[:6] + return f"Key found ({source}) but invalid — expected prefix '{rules['prefix']}', got '{actual_prefix}...'" + + if "min_length" in rules and len(key) < rules["min_length"]: + return f"Key found ({source}) but too short — {len(key)} chars, need {rules['min_length']}+" + + return f"Key found ({source}) but failed validation" + + def check_key_format(key: str) -> Dict[str, Any]: """ Analyze key format and return details. diff --git a/src/aipass/api/apps/handlers/openrouter/client.py b/src/aipass/api/apps/handlers/openrouter/client.py index c82c5d90..e467d837 100644 --- a/src/aipass/api/apps/handlers/openrouter/client.py +++ b/src/aipass/api/apps/handlers/openrouter/client.py @@ -56,6 +56,7 @@ except ImportError: # Handler imports from aipass.api.apps.handlers.auth.keys import get_api_key from aipass.api.apps.handlers.openrouter.caller import get_caller_info +from aipass.api.apps.handlers.openrouter.provision import ensure_caller_config from aipass.api.apps.handlers.usage.tracking import track_usage # JSON handler @@ -315,6 +316,12 @@ def get_response(prompt: str, caller: Optional[str] = None, model: Optional[str] # logger.warning("Could not detect caller - using 'unknown'") caller = "unknown" + # Step 1b: Ensure caller has config (auto-provision if missing) + try: + ensure_caller_config(caller) + except Exception: + pass # Provisioning is best-effort — don't block the API call + # Step 2: Require model from caller - no defaults if not model: # logger.error("No model specified - caller must provide model from their branch config") diff --git a/src/aipass/api/apps/modules/api_key.py b/src/aipass/api/apps/modules/api_key.py index d8ace3e7..03aa3b4b 100644 --- a/src/aipass/api/apps/modules/api_key.py +++ b/src/aipass/api/apps/modules/api_key.py @@ -153,9 +153,15 @@ def init_env(): header("Initialize API Configuration") console.print() + env_path = Path.home() / ".secrets" / "aipass" / ".env" + + if env_path.exists(): + success(f"Environment file already exists at {env_path}") + return + # Create .env template via handler if env.create_env_template(): - success("Environment template created") + success(f"Environment template created at {env_path}") else: error("Failed to create environment template") diff --git a/src/aipass/api/apps/modules/openrouter_client.py b/src/aipass/api/apps/modules/openrouter_client.py index 23334ae0..5a2343a8 100644 --- a/src/aipass/api/apps/modules/openrouter_client.py +++ b/src/aipass/api/apps/modules/openrouter_client.py @@ -138,7 +138,7 @@ def handle_command(command: str, args: List[str]) -> bool: test_connection() return True if command == "models": - list_models() + list_models(args) return True if command == "status": check_status() @@ -164,17 +164,23 @@ def test_connection(): header("Test OpenRouter Connection") console.print() + console.print("[dim]Testing connection...[/dim]") + # Get API key via handler api_key = keys.get_api_key("openrouter") if not api_key: - error("No API key configured") + diagnosis = keys.diagnose_key("openrouter") + error(diagnosis) return - warning("Testing connection...") + # Real API ping — hit /models endpoint + model_list = models.fetch_models_from_api(api_key) - # TODO: Call handler to test connection when implemented - success("Connection test successful") + if model_list: + success(f"Connection successful — {len(model_list)} models available") + else: + error("Connection failed — could not reach OpenRouter API") def make_call(args: List[str]): @@ -187,31 +193,67 @@ def make_call(args: List[str]): warning("API call workflow - TODO") -def list_models(): +def list_models(args: List[str] | None = None): """Orchestrate list models workflow""" header("Available Models") console.print() + show_all = args and "--all" in args + # Get API key via handler api_key = keys.get_api_key("openrouter") if not api_key: - error("No API key configured") + diagnosis = keys.diagnose_key("openrouter") + error(diagnosis) return - warning("Fetching available models...") + console.print("[dim]Fetching available models...[/dim]") # Call handler to fetch models model_list = models.fetch_models_from_api(api_key) - if model_list: - success(f"Found {len(model_list)} models") - for model in model_list[:10]: - console.print(f" - {model}") - if len(model_list) > 10: - console.print(f" ... and {len(model_list) - 10} more") - else: + if not model_list: error("Failed to fetch models") + return + + success(f"Found {len(model_list)} models") + console.print() + + # Format as table + display_count = len(model_list) if show_all else min(10, len(model_list)) + + console.print(f" {'Model':<50} {'Context':>10} {'$/prompt':>10} {'$/compl':>10}") + console.print(f" {'─' * 50} {'─' * 10} {'─' * 10} {'─' * 10}") + + for model_data in model_list[:display_count]: + model_id = model_data.get("id", "unknown") + context = model_data.get("context_length", 0) + pricing = model_data.get("pricing", {}) + prompt_cost = pricing.get("prompt", "0") + completion_cost = pricing.get("completion", "0") + + # Format context length + if context >= 1_000_000: + ctx_str = f"{context // 1_000_000}M" + elif context >= 1_000: + ctx_str = f"{context // 1_000}k" + else: + ctx_str = str(context) + + # Format pricing + if str(prompt_cost) == "0" and str(completion_cost) == "0": + p_str = "free" + c_str = "free" + else: + p_str = f"${prompt_cost}" + c_str = f"${completion_cost}" + + console.print(f" {model_id:<50} {ctx_str:>10} {p_str:>10} {c_str:>10}") + + if not show_all and len(model_list) > 10: + console.print() + console.print(f" [dim]Showing 10 of {len(model_list)} — use --all for full list[/dim]") def check_status(): @@ -219,8 +261,32 @@ def check_status(): header("OpenRouter Client Status") console.print() - # TODO: Call handlers for status when implemented - warning("Client status - TODO") + # Key status + api_key = keys.get_api_key("openrouter") + + if api_key: + masked = api_key[:8] + "..." + api_key[-4:] + console.print(f" [cyan]Key configured:[/cyan] [green]yes[/green]") + console.print(f" [cyan]Key:[/cyan] {masked}") + else: + console.print(f" [cyan]Key configured:[/cyan] [red]no[/red]") + diagnosis = keys.diagnose_key("openrouter") + console.print(f" [cyan]Reason:[/cyan] {diagnosis}") + + console.print(f" [cyan]Provider:[/cyan] OpenRouter") + console.print(f" [cyan]Base URL:[/cyan] https://openrouter.ai/api/v1") + + # OpenAI SDK availability + try: + import openai # noqa: F401 + console.print(f" [cyan]OpenAI SDK:[/cyan] [green]available[/green]") + except ImportError: + console.print(f" [cyan]OpenAI SDK:[/cyan] [red]missing[/red]") + + # Client cache stats + cache_stats = client.get_cache_stats() + console.print(f" [cyan]Cached clients:[/cyan] {cache_stats['cached_clients']}/{cache_stats['max_cache_size']}") + console.print() # ============================================= diff --git a/src/aipass/devpulse/docs/README.md b/src/aipass/devpulse/docs/README.md new file mode 100644 index 00000000..dad7824b --- /dev/null +++ b/src/aipass/devpulse/docs/README.md @@ -0,0 +1,22 @@ +# DPLAN-003 Working Directory + +Research, mapping, and planning files for "AIPass as Operating System." + +Parent plan: `AIPass/DPLAN-003_aipass_as_operating_system_2026-03-13.md` + +## Scope + +**In scope:** `src/aipass/` branches only (drone, seedgo, prax, cli, flow, ai_mail, api, trigger, spawn, devpulse, backup, daemon, memory). + +**Out of scope:** `src/commons/`, `src/skills/` — these are separate and not part of this refactor. + +## Files + +| File | Purpose | Status | +|------|---------|--------| +| `registry_discovery_map.md` | Every find_registry() call in src/aipass/ | Done | +| `portability_audit.md` | Full investigation results (session 24) | Done | +| `credential_model.md` | Registry credential design (UUID now, macaroons later) | Stage 1 Done | +| `registry_refactor_plan.md` | Shared find_registry() design, migration steps | Pending | +| `aipass_init_spec.md` | CLI owns init — `drone @cli aipass init` (entry point later) | Pending | +| `aipass_help_spec.md` | CLI owns help — `drone @cli aipass help` (entry point later) | Pending | diff --git a/src/aipass/devpulse/docs/ai_mail_comms_upgrade.md b/src/aipass/devpulse/docs/ai_mail_comms_upgrade.md new file mode 100644 index 00000000..91d9d108 --- /dev/null +++ b/src/aipass/devpulse/docs/ai_mail_comms_upgrade.md @@ -0,0 +1,217 @@ +# AI_MAIL COMMS UPGRADE — Hardening Plan + +**Goal:** Make ai_mail bulletproof, then use as the model for other branches. +**Started:** 2026-03-10 | **Status:** In Progress +**Tracking:** Updated each session. Read this first on context resume. + +--- + +## Current State (Session 10) + +**Seedgo audit: 100%** — all 23 standards pass (42 files) +**Automated tests: 36** — test_send_identity.py v1.2.0 (3 audit rounds, 15 agents, 0 false positives) +**Production fixes deployed:** 3 cross-platform crashers fixed, test suite fully isolated +**Phase 2 (Error Handling):** 27 silent `except` blocks across 13 files now log with `logger.warning()` + +### Test Suite Evolution +- v1.0.0: 31 tests — 7 false positives found by 3 audit agents +- v1.1.0: 32 tests — 8 suspects found by 5 audit agents (live registry, weak contracts) +- v1.2.0: 36 tests — final audit found 2 minor issues, fixed. **0 live-data dependencies.** + +### Production Fixes (Session 9) +- `inbox_lock.py` — `import fcntl` guarded for Windows (msvcrt fallback) +- `ai_mail.py` — `signal.SIGPIPE` guarded with `hasattr` check +- `notify.py` — `/usr/bin/python3` replaced with `shutil.which("python3")` + +--- + +## Architecture Map + +``` +Entry: apps/ai_mail.py + +-- Module: apps/modules/email.py (orchestrator, v3.0.0) + |-- handle_send() -> send_args.py (parse) -> send.py (execute) -> delivery.py (inbox write) + |-- handle_inbox() -> inbox_resolve.py -> inbox_ops.py -> format.py + |-- handle_view() -> inbox_ops.py + |-- handle_reply() -> reply.py -> delivery.py + |-- handle_close() -> close_ops.py + |-- handle_sent() -> format.py + +-- handle_contacts() -> registry/ + + Module: apps/modules/dispatch.py (dispatch orchestrator) + |-- daemon.py (auto-dispatch loop) + |-- wake.py (manual branch wake) + |-- dispatch_monitor.py (agent lifecycle) + |-- status.py (dispatch status display) + +-- pending_work.py (pending dispatch queue) + + Identity: apps/handlers/users/ + |-- branch_detection.py (detect_branch_from_pwd — THE critical path) + |-- user.py (get_current_user, get_branch_by_email) + +-- load.py, config_generator.py + + Cross-branch: drone/apps/handlers/router_handler.py + +-- detect_caller_branch_name() -> sets AIPASS_CALLER_BRANCH env var +``` + +**Identity detection chain (9 stages):** +1. drone CLI entry +2. router resolves @branch to path +3. router_handler detects caller via CWD (+ AIPASS_BRANCH_NAME fallback) +4. executor merges caller_env into subprocess +5. ai_mail subprocess starts with env vars +6. branch_detection reads AIPASS_CALLER_BRANCH +7. send_args builds headers +8. delivery writes to recipient inbox.json +9. notify.py fires desktop notification + +--- + +## Systemic Issues Found (Session 9 Deep Audit) + +15 agents across 3 rounds audited tests + full codebase. Key findings: + +### Silent Failures (24 instances) — VIOLATES "fail to errors" rule +`except Exception: return None` in 10+ places with zero logging: +- `branch_detection.py` — 3 functions (entire identity chain) +- `delivery.py` — get_all_branches(), summary, notification +- `create.py` — load_email_file() +- `user.py` — get_user_by_email(), get_all_users() +- `daemon.py` — _read_json() (silent) vs wake.py (logs) — inconsistent +- `inbox_ops.py` — migration persist has literal `pass` + +### Dead/Redundant Code (~20% of codebase, ~1,728 lines) +- 5 fully unused files: validate.py, errors.py, data_ops.py, config_generator.py, pending_work.py +- `_find_repo_root()` copy-pasted in 9 files +- `get_all_branches()` implemented twice with different email behavior (correctness bug) +- `lock_utils.py` exists but unused — wake.py and daemon.py reimplement locking +- wake.py/daemon.py share 6+ duplicated functions + +### Cross-Platform Breakers +- [FIXED] `import fcntl` unconditional — crashes Windows +- [FIXED] `/usr/bin/python3` hardcoded — breaks macOS/Windows +- [FIXED] `signal.SIGPIPE` — crashes Windows startup +- [OPEN] `pgrep` ungated in wake.py/daemon.py +- [OPEN] `email.py` parents[2] fragile — should use _find_repo_root() + +--- + +## Hardening Phases + +### Phase 1: Test Suite (Critical Path) +**Priority: HIGH** | **Status: IN PROGRESS** + +- [x] `test_send_identity.py` — 36 tests, 3 audit rounds, 0 false positives +- [ ] `test_delivery.py` — round-trip send/receive +- [ ] `test_send_args.py` — argument parsing (all flags, interactive, error cases) +- [ ] `test_inbox_ops.py` — inbox operations (view, close, reply, close all) +- [ ] `test_dispatch_monitor.py` — dispatch lifecycle +- [ ] `test_notify.py` — notification delivery + +### Phase 2: Error Handling Overhaul +**Priority: HIGH** | **Status: IN PROGRESS** +Silent failures are why the system feels "fragile." + +- [x] Add `logger.warning()` to every bare `except Exception: return None` — **13 files, 27 instances fixed** +- [ ] Distinguish "not found" from "error reading" in return types +- [ ] Fix inconsistent error conventions (None vs tuple vs dict vs raise) +- [ ] Fix collision detection dead code in delivery.py get_all_branches() + +### Phase 3: Code Consolidation +**Priority: MEDIUM** — Reduce duplication, single source of truth. + +- [ ] Extract shared `_find_repo_root()` into commons or shared utility +- [ ] Consolidate `get_all_branches()` — one implementation, prefers explicit email +- [ ] Consolidate lock acquisition — use lock_utils.py, delete reimplementations +- [ ] Deduplicate wake.py/daemon.py shared functions (_read_json, _set_session_name, etc.) +- [ ] Archive 5 dead files (validate.py, errors.py, data_ops.py, config_generator.py, pending_work.py) + +### Phase 4: Identity Consolidation +**Priority: MEDIUM** — Reduce 12 detection mechanisms to 1 canonical resolver. + +- [ ] Define canonical `resolve_branch_identity()` function +- [ ] Priority chain: AIPASS_CALLER_BRANCH > AIPASS_BRANCH_NAME > CWD passport walk > --from +- [ ] Single file, single function, single source of truth +- [ ] All callers delegate to it + +### Phase 5: Cross-Platform Hardening +**Priority: MEDIUM** — Public repo must work on all platforms. + +- [ ] Guard `pgrep` usage in wake.py/daemon.py +- [ ] Replace `email.py` parents[2] with _find_repo_root() +- [ ] Guard `start_new_session=True` for Windows +- [ ] Add encoding='utf-8' to os.fdopen in wake.py + +### Phase 6: Standards for Communication +**Priority: LOW** — Make patterns enforceable system-wide. + +- [ ] `communication` standard — envelope format, required fields, identity rules +- [ ] `identity` standard — env var hierarchy, passport requirements +- [ ] `dispatch` standard — env isolation, lock management, bounce handling + +### Phase 7: Documentation & System Prompt +**Priority: LOW** — Update branch local prompt. + +- [ ] Write proper `.aipass/aipass_local_prompt.md` +- [ ] Document key commands, architecture, critical files + +--- + +## Decision Log + +| Date | Decision | Rationale | +|------|----------|-----------| +| 2026-03-08 | AIPASS_BRANCH_NAME env var over CWD detection | CWD unreliable when agents navigate. Env vars persist. | +| 2026-03-08 | dbus direct over notify-send | Portal mode strips persistence hints. dbus bypasses confinement. | +| 2026-03-08 | Unique app_name per source | GNOME collapses same desktop-entry into one notification. | +| 2026-03-08 | Fail loud, no CWD fallback for identity | Silent wrong sender worse than visible error. | +| 2026-03-10 | --from flag for explicit sender override | Plumbing existed, just needed CLI wiring. | +| 2026-03-10 | Tests first in hardening plan | Every bug from sessions 4-7 would have been caught by tests. | +| 2026-03-10 | 3 audit rounds on test suite | Each round found issues the previous missed. Diminishing returns by round 3. | +| 2026-03-10 | Error handling overhaul before code consolidation | Silent failures cause more user pain than code duplication. | + +--- + +## Session Notes + +### Session 10 (2026-03-10) +- **Phase 2: Error Handling Overhaul — logger.warning() sweep complete** +- 13 files modified, 27 silent `except Exception` blocks now log with `logger.warning()` +- Files fixed (by priority): + - `branch_detection.py` (3) — identity chain, most critical + - `user.py` (2) — user lookup + - `delivery.py` (6) — get_all_branches, migrate, callback, summary, notification, private branch + - `create.py` (2) — purge, load_email_file + - `inbox_ops.py` (1) — migration persist + - `format.py` (1) — alias lookup + - `reply.py` (1) — get_email_by_id + - `close_ops.py` (3) — dashboard, central, purge post-ops + - `send.py` (2) — central update after send/broadcast + - `dashboard_sync.py` (1) — push_dashboard_update + - `inbox_cleanup.py` (4) — migrate, dashboard, central, purge + - `error_handler.py` (1) — error notification delivery + - `error_dispatch.py` (2) — dashboard, central post-ops +- All 13 files pass `py_compile` syntax check +- Remaining silent blocks (acceptable): inbox_lock.py finally-block, json_handler.py (utility), + daemon/wake/dispatch_monitor (already log with logger.info), config_generator/data_ops (unused files) +- Next: Phase 2 remaining items (return type distinctions, error conventions, collision dead code) + +### Session 9 (2026-03-10, continued) +- Wrote test_send_identity.py v1.0.0 (31 tests) +- Audit round 1: 3 agents found 7 false positives -> rewrote to v1.1.0 +- Audit round 2: 5 agents found 8 issues (live registry, weak contracts) -> v1.2.0 (36 tests) +- Audit round 3: 5 agents (final test + full system sweep) + - Tests: 2 minor issues found and fixed (unpatched registry, missing mailbox_path assert) + - Error handling: 24 silent failure instances across 10+ files + - Code quality: ~1,728 dead/redundant lines (20% of codebase), 5 unused files + - Cross-platform: 3 CRITICAL (fcntl, SIGPIPE, /usr/bin/python3) — ALL FIXED + - Paths: parents[N] usage audited, mostly correct, 1 fragile case in email.py +- Fixed 3 cross-platform crashers in production code +- Updated COMMS_UPGRADE.md with full findings and revised phases +- Next: Phase 1 continues (more test files) or Phase 2 (error handling overhaul) + +### Session 8 (2026-03-10) +- Patrick initiated hardening project +- Self-audit: 100% on all 22 seedgo standards +- Mapped full architecture (55 Python files, 3 modules, 8 handler domains) +- Created this tracking document diff --git a/src/aipass/devpulse/docs/credential_model.md b/src/aipass/devpulse/docs/credential_model.md new file mode 100644 index 00000000..53757f31 --- /dev/null +++ b/src/aipass/devpulse/docs/credential_model.md @@ -0,0 +1,157 @@ +# DPLAN-003: Registry Credential Model + +## The Idea + +Registries get a unique token. Passports carry that token. Access is identity-based, not filesystem-based. No walk-up needed — your credential proves which registry is yours. + +## Why + +Current system finds registries by walking up directories. Works for one project, breaks with multiple. If two AIPass projects exist on one machine, a citizen launched from the wrong directory finds the wrong registry. Credentials solve this — your passport carries proof of membership. + +## Prior Art (Research) + +| System | Pattern | Fit | +|--------|---------|-----| +| **Macaroons** (Google Research) | Token IS the credential. Delegatable with caveats. Offline verification. DeepMind validated for AI agent delegation (2026). | Highest | +| **Vault Namespaces** | Project = namespace. Token scoped to namespace. Mini-registry per project. | High | +| **AWS STS / Token Vending** | Agent presents project ID, gets scoped credential. Credential itself is the boundary. | High | +| **K8s Namespace + ServiceAccount** | Token carries project scope as claim. RBAC composable. | High | +| **SPIFFE/SPIRE** | Process-level attestation without static secrets. | Medium | +| **direnv** | Auto-set env vars on directory entry. Zero-friction UX. | UX pattern | + +Full research: agent output from session 25. + +## Design: Two Stages + +### Stage 1: UUID Match (Manual, Now) + +Simple. Prove the concept works before adding crypto. + +**Registry gets an ID:** +```json +{ + "metadata": { + "id": "a1b2c3d4-...", + "name": "AIPASS", + "version": "1.0.0", + "last_updated": "2026-03-13", + "total_branches": 15 + }, + "branches": [...] +} +``` + +**Passports get the matching ID:** +```json +{ + "citizenship": { + "registered": true, + "registry_id": "a1b2c3d4-...", + "registry_name": "AIPASS", + "citizen_number": 7 + } +} +``` + +**Lookup flow:** +1. Walk up from CWD, find `*_REGISTRY.json` +2. Read its `metadata.id` +3. Check citizen's `citizenship.registry_id` matches +4. If mismatch → error ("citizen belongs to registry X, found registry Y") +5. If match → proceed + +**Spawn changes:** +- `aipass init` (or manual setup) generates the UUID for the registry +- Spawn reads registry UUID and injects into new passports via `{{REGISTRY_ID}}` placeholder +- Existing 15 branches get the UUID added to their passports (one-time migration) + +### Stage 2: Macaroon Tokens (Future, When Cross-Project Needed) + +Upgrade path when we need delegation and cross-project access. + +**Root token:** Created at `aipass init`. HMAC-based. Stored at `~/.secrets/aipass/projects/.token` + +**Citizen token:** Attenuated copy in passport. Can prove membership but can't mint new citizens. + +**Agent token:** Further attenuated. Carries: +- Registry ID (which project) +- Scope (full access — agents do the real work in AIPass) +- Expiry (session-scoped, dies when agent dies) +- Issuer (which citizen spawned this agent) + +Agents are NOT read-only. They're the builders — they write code, run tests, modify files. The token proves they belong to a project, it doesn't restrict what they do within it. Scope restrictions would be role-based (e.g., "can't modify other branches' files") not capability-based. + +**Verification:** Local HMAC check. No daemon needed for basic validation. Daemon is optional enhancement for audit logging and revocation. + +**direnv integration:** Entering a project directory auto-sets `AIPASS_PROJECT_TOKEN` in env. Agents inherit it. + +## What Changes (Stage 1) + +| Component | Change | +|-----------|--------| +| `AIPASS_REGISTRY.json` | Add `metadata.id` (UUID) | +| Passport template | Add `citizenship.registry_id` placeholder | +| Spawn `build_replacements_dict()` | Read registry UUID, add `{{REGISTRY_ID}}` | +| Spawn `add_to_registry()` | No change (branch entries stay the same) | +| Drone `find_registry()` | Optional: verify passport.registry_id matches found registry | +| All 15 passports | One-time: add `registry_id` field | + +## What Does NOT Change + +- Registry filename stays `*_REGISTRY.json` +- Walk-up discovery still works (credential is verification layer on top, not replacement) +- Branch structure unchanged +- No daemon needed +- No new dependencies + +## Resolved Questions + +1. **Registry ID location** → `metadata.id` — it's the project's identity, not the citizen's. +2. **UUID4 vs hash** → UUID4 (random). Simple, guaranteed unique, no inputs needed. +3. **Verification mode** → Hard error on mismatch. Fail loudly — that's the AIPass way. +4. **Where does init live?** → Temporary: `src/aipass/devpulse/apps/init_project.py`. Future: CLI branch. + +## Bugs, Quirks, and Findings + +Discovered during testing. Reference for future work. + +### init_project.py (10 edge cases tested) +- **Spaces in dir name** → registry filename gets spaces (`MY COOL PROJECT_REGISTRY.json`). Fixed: `_sanitize_name()` replaces non-alphanumeric with `_`. +- **Root path `/`** → `Path("/").name` is empty string, creates `_REGISTRY.json`. Fixed: validation rejects empty name. +- **Permission errors** → raw traceback instead of clean message. Fixed: `main()` catches `OSError`. +- **Passport overwrite on re-init** → if registry deleted but `.trinity/` survives, re-init would silently overwrite passport with new UUID. Fixed: passport guarded with `exists()` check. +- **Double init** → correctly blocked by `FileExistsError` on registry file. +- **Deep nested paths** → `mkdir(parents=True)` handles correctly. +- **UUID uniqueness** → 5 runs, 5 unique UUIDs. No collisions. +- **JSON validity** → all generated files parse clean. + +### Drone isolation (6 scenarios tested) +- **Sibling projects** → PASS. Two projects in same parent dir, fully isolated. +- **Nested project** → PASS. Inner registry wins over outer. No bleed-through. +- **Deep subdir walk-up** → PASS. Finds nearest ancestor registry correctly. +- **Cross-project contamination** → PASS. Branches are registry-scoped; modules are global. +- **Empty directory (no registry)** → CONCERN. Drone silently falls back to AIPass source registry via `__file__` walk-up. Any dir on this machine without a registry sees production branches. This is by design in `find_registry()` but will be dangerous with multi-project. Credential verification would catch this — citizen's registry_id won't match the fallback registry. +- **`drone @ai_mail` from mock project** → correctly returns "Branch not found in registry" (empty project has no branches). + +### Registry/passport gitignore +- `AIPASS_REGISTRY.json` and all `.trinity/passport.json` files are gitignored. UUID migration is local-only. This is correct for now — credentials are machine-specific, not repo state. But means `aipass init` must run on every clone/install. Future: consider whether UUID should be in-repo or machine-local. + +### AI mail after migration +- Send/receive works fine after UUID migration. AI mail's 4 internal `find_registry()` copies still hardcode `AIPASS_REGISTRY.json` — functional for now since that filename exists, but won't find `*_REGISTRY.json` in other projects. + +### Drone built-in modules vs branches +- `@drone` and `@seedgo` are hardcoded as "modules" in `module_registry.py`, always visible everywhere. Other branches (`@ai_mail`, `@spawn`, etc.) are registry-scoped. This distinction matters: modules are global services, branches are project citizens. + +## Decision Log + +- **2026-03-14:** Patrick proposed credential-based registry access. Agents do the real work — tokens prove membership, not restrict capability. Manual first, `aipass init` later. +- **2026-03-14:** Research confirmed macaroons as best-fit pattern (Google Research + DeepMind 2026 validation). Stage 1 = UUID match, Stage 2 = macaroon upgrade. +- **2026-03-14:** Scope limited to `src/aipass/` branches only. Commons and skills excluded. +- **2026-03-14:** All 4 open questions resolved. `init_project.py` built and tested — creates registry with UUID, passport with matching registry_id, .trinity/, .aipass/, AIPASS.md. Tested with temp directory — drone isolation confirmed (only built-in modules visible, not AIPass branches). +- **2026-03-14:** UUID migration executed (FPLAN-0030). Registry + 13 passports updated. Registry and passports are gitignored — UUID is machine-local, not repo state. +- **2026-03-14:** Drone verification added (`_verify_registry_credential` in load_registry). Drone dispatched and fixed error propagation — new `RegistryMismatchError` class separates "not found" (fallback OK) from "mismatch" (hard error). Spawn templates updated with `{{REGISTRY_ID}}` placeholder. +- **2026-03-14:** Stage 1 complete. Credential model is end-to-end: init creates credentialed registries, drone verifies on load, spawn injects into new passports, mismatch = hard error. Remaining: drone CLI stderr surfacing (DPLAN-0031), `aipass init` CLI command (future). +- **2026-03-14:** FPLAN-0032 executed — CLI stderr standardization Phase 1+2 done. CLI owns err_console, error()/warning() go to stderr. Drone imports from CLI instead of creating its own Console(stderr=True). Credential mismatch errors now properly surface to terminal via stderr. The stderr issue that blocked error visibility during DPLAN-003 testing is resolved. +- **2026-03-14:** CLI confirmed as owner of credential model + project commands. CLI owns: `aipass init` (create project), `aipass help` (project help), display API (error/warning/fatal/err_console). Drone owns verification (registry_handler). Spawn owns injection (passport templates). Access via `drone @cli aipass init/help` for now; `aipass init` entry point is future icing. +- **2026-03-14:** Seedgo `stderr_routing` standard created (24th standard). Checker bug found — was scanning all files but not displaying module/handler results. Seedgo fixed display + aggregation. Agent-verified across 3 branches. FPLAN-0033 parked for Phase 3 migration (343 error prints, 10 branches). +- **2026-03-14:** FPLAN-0033 Phase 3 executed autonomously (Patrick away). 15+ agents across 3 waves migrated 10 branches. 48 files, 666 insertions. System stderr avg 49%→69%. Post-migration verification found 63% of remaining seedgo violations are false positives ([yellow]Label:[/yellow] section headers flagged as warnings). Seedgo dispatched for checker refinement. 93 real violations remain (48 in commons which was out of scope). diff --git a/src/aipass/devpulse/docs/portability_audit.md b/src/aipass/devpulse/docs/portability_audit.md new file mode 100644 index 00000000..090f00cc --- /dev/null +++ b/src/aipass/devpulse/docs/portability_audit.md @@ -0,0 +1,31 @@ +# Portability Audit — Session 24 Results + +## Summary + +| Tool | Registry Discovery | CWD-Aware | Portable | Hardcoded | +|------|-------------------|-----------|----------|-----------| +| Drone | Walk-up + env var | No (uses registry) | Yes | Registry filename | +| Spawn | Walk-up + env var | No (uses registry) | Partial | Template location | +| Prax | Walk-up (no env) | No (sys logs at repo) | Partial | System logs dir | +| AI_Mail | Walk-up (no env) | No (inbox per branch) | Yes | Inbox location | +| Flow | Walk-up (no env) | Yes (plan creation) | Hybrid | Plan registry | + +## Key Findings + +- All tools use walk-up strategy to find `AIPASS_REGISTRY.json` +- Registry-relative path resolution already works (move registry + dirs = works) +- `AIPASS_REGISTRY` env var supported by drone and spawn +- System logs hardcoded to `{repo_root}/system_logs/` +- Spawn templates hardcoded to `{spawn_package}/templates/` +- Walk-up doesn't stop at project boundaries — finds nearest registry up the tree + +## The Core Fix + +Change `find_registry()` to: +1. Walk up from CWD looking for `*_REGISTRY.json` (glob, not hardcoded name) +2. Stop at first match — that's the project boundary +3. If none found, return error ("No AIPass project. Run `aipass init`") + +## Source + +Full investigation transcript: background agent session 24, 42 tool calls across drone/spawn/ai_mail/flow/prax. diff --git a/src/aipass/devpulse/docs/registry_discovery_map.md b/src/aipass/devpulse/docs/registry_discovery_map.md new file mode 100644 index 00000000..17583c4c --- /dev/null +++ b/src/aipass/devpulse/docs/registry_discovery_map.md @@ -0,0 +1,46 @@ +# Registry Discovery Map — FPLAN-0029 Phase 1 + +## Primary Implementations (5) + +| File | Function | Strategy | +|------|----------|----------| +| `src/aipass/spawn/apps/handlers/registry.py:34-79` | `find_registry(start_path)` | Env var → project markers (.git/pyproject.toml) → walk-up → cwd fallback | +| `src/aipass/drone/apps/handlers/registry_handler.py:36-61` | `find_registry()` | Walk-up from __file__ → walk-up from cwd → parents[4] fallback | +| `src/aipass/drone/apps/handlers/registry_handler.py:64-81` | `get_registry_path()` | Global override → env var → find_registry() | +| `src/commons/apps/handlers/database/db.py:343-378` | `_find_branch_registry()` | Env var AIPASS_ROOT → walk-up (10 limit) → ~/.aipass/ fallback | +| `src/commons/apps/handlers/identity/identity_ops.py:33-67` | `_find_branch_registry_path()` | Same as db.py | + +## Local Reimplementations (23+) + +All do walk-up from `__file__` looking for `AIPASS_REGISTRY.json`: + +### ai_mail (4 files) +- `apps/handlers/users/branch_detection.py:26-38` +- `apps/handlers/registry/read.py:31-41` +- `apps/handlers/email/format.py:21-31` +- `apps/handlers/dispatch/daemon.py:36-55` + +### memory (3 files) +- `apps/handlers/dashboard_push.py:44-54` +- `apps/handlers/monitor/detector.py:36-83` +- `apps/handlers/monitor/memory_watcher.py:363-385` + +### prax (2 files) +- `apps/handlers/registry/reader.py:24-39` +- `apps/handlers/dashboard/agent_status_writer.py:33-44` + +### seedgo (3 files) +- `apps/handlers/audit/discovery.py:54-62` +- `apps/handlers/diagnostics/discovery.py:21-29` +- `apps/handlers/readme/readme_ops.py:29-37` + +### + 11 more across other branches + +## Hardcoded String Count +100+ references to `"AIPASS_REGISTRY.json"` across codebase. + +## Phase 1 Plan +1. Create `commons.registry.find_registry()` — one function, glob for `*_REGISTRY.json` +2. All 23+ modules import from commons +3. Drone/spawn wrappers call commons internally +4. Test: temp dir with TEST_REGISTRY.json → drone systems finds it diff --git a/src/aipass/devpulse/docs/system_health_report.md b/src/aipass/devpulse/docs/system_health_report.md new file mode 100644 index 00000000..5764d1da --- /dev/null +++ b/src/aipass/devpulse/docs/system_health_report.md @@ -0,0 +1,525 @@ +# AIPass System Health Report +Generated: 2026-03-19 22:54 + +--- + +## 1. Dead Code (14 unused files) + + +Dead Code Scanner +Scanning 15 branches + +@ai_mail (3 modules, 31 handlers) + x handlers/monitoring/errors.py -- 0 references + > 33/34 files referenced + +@api (4 modules, 14 handlers) + x handlers/openrouter/provision.py -- 0 references + > 17/18 files referenced + +@backup (4 modules, 24 handlers) + x handlers/diff/vscode_integration.py -- 0 references + > 27/28 files referenced + +@cli (3 modules, 2 handlers) + > All 5 files referenced + +@commons (22 modules, 35 handlers) + > All 57 files referenced + +@daemon (6 modules, 8 handlers) + > All 14 files referenced + +@devpulse -- no apps/ directory +@drone (9 modules, 16 handlers) + > All 25 files referenced + +@flow (8 modules, 35 handlers) + x handlers/plan/file_ops.py -- 0 references + x handlers/plan/update_registry.py -- 0 references + > 41/43 files referenced + +@memory (5 modules, 29 handlers) + x handlers/learnings/manager.py -- 0 references + x handlers/schema/normalize.py -- 0 references + x handlers/search/vector_search.py -- 0 references + > 31/34 files referenced + +@prax (6 modules, 36 handlers) + > All 42 files referenced + +@seedgo (5 modules, 66 handlers) + x handlers/config/aipass_bypass.py -- 0 references + x handlers/config/aipass_ignore.py -- 0 references + x handlers/diagnostics/python_diognostics.py -- 0 references + x handlers/diagnostics/typscript_diognostics.py -- 0 references + x handlers/file/file_handler.py -- 0 references + x handlers/mock_standard_1/bypass_config/bypass.config.py -- 0 references + > 65/71 files referenced + +@skills (5 modules, 8 handlers) + > All 13 files referenced + +@spawn (6 modules, 15 handlers) + > All 21 files referenced + +@trigger (5 modules, 17 handlers) + > All 22 files referenced + +TOTAL: 14 unused across 14 branches + +--- + +## 2. Local Prompts (5 stubs need enrichment) + + +Local Prompt Status +================================================== + +RICH (50+ lines, 4+ sections): + v daemon 56 lines 4 sections + v devpulse 79 lines 8 sections + v flow 67 lines 6 sections + v seedgo 56 lines 4 sections + v skills 58 lines 5 sections + v trigger 53 lines 7 sections + +BASIC (15-49 lines): + ~ api 15 lines 2 sections + ~ commons 48 lines 6 sections + ~ memory 37 lines 5 sections + ~ spawn 20 lines 4 sections + +STUB (<15 lines): + x ai_mail 14 lines 1 section + x backup 14 lines 1 section + x cli 14 lines 1 section + x drone 14 lines 1 section + x prax 14 lines 1 section + +================================================== +SUMMARY: 6 rich, 4 basic, 5 stub + +================================================== +Section Breakdown +================================================== + +@ai_mail (14 lines, STUB) + v Status + 470 bytes + +@api (15 lines, BASIC) + v Identity + v Key Breadcrumbs + 1192 bytes + +@backup (14 lines, STUB) + v Status + 499 bytes + +@cli (14 lines, STUB) + v Status + 499 bytes + +@commons (48 lines, BASIC) + v Commands + v Architecture + v Integration Points + v Critical Files + v Role + v Key Details + 2328 bytes + +@daemon (56 lines, RICH) + v Commands + v Memory & Tracking + v Apps Layout + v Known Issues + 2788 bytes + +@devpulse (79 lines, RICH) + v Identity + v How You Work + v Dispatch Table + v Commands + v Branches + v Working Habits + v Autonomous Monitoring + v Memory & Tracking + v Has dispatch table (@branch refs) + 4710 bytes + +@drone (14 lines, STUB) + v Status + 503 bytes + +@flow (67 lines, RICH) + v Commands + v Architecture + v Integration Points + v Conventions + v Critical Files + v Plan Type System + 3461 bytes + +@memory (37 lines, BASIC) + v Identity + v Commands + v Architecture + v Memory & Tracking + v Known Issues + 1591 bytes + +@prax (14 lines, STUB) + v Status + 501 bytes + +@seedgo (56 lines, RICH) + v Commands + v Apps Layout (extra layer vs standard branch) + v How I Work — Standards Reasoning + v Quick Reference + 3166 bytes + +@skills (58 lines, RICH) + v Commands + v Memory & Tracking + v Apps Layout + v Search Paths (first match wins) + v Three Skill Tiers + 2706 bytes + +@spawn (20 lines, BASIC) + v Commands + v Architecture + v Role + v Principles + 745 bytes + +@trigger (53 lines, RICH) + v Dispatch Table + v Commands + v Architecture + v Integration Points + v Critical Files + v Role + v Rules + 2610 bytes + + +--- + +## 3. Commands (173 discovered) + + +@ai_mail (14 commands) + - close + - contacts + - dispatch + - email + - inbox + - ping + - read + - registry + - reply + - send + - sent + - status + - thresholds + - view + +@api (12 commands) + - call + - cleanup + - google + - init + - models + - reauth + - session + - stats + - status + - test + - track + - validate + +@backup (1 command) + - reauth + +@cli (5 commands) + - aipass + - demo + - display + - show + - templates + +@commons (51 commands) + - activity + - artifacts + - capsule + - capsules + - catchup + - collab + - comment + - craft + - database + - decorate + - delete + - digest + - drop + - enter + - event + - explore + - feed + - find + - gift + - inspect + - leaderboard + - leaderboards + - log + - look + - mint + - mute + - open + - pin + - pinned + - post + - preferences + - profile + - prompt + - react + - reactions + - room + - search + - secrets + - sign + - thread + - track + - trade + - trending + - unpin + - unreact + - visitors + - vote + - watch + - welcome + - who + - whoami + +@daemon (5 commands) + - actions + - activity + - activity_report + - schedule + - update + +@drone (24 commands) + - activate + - add + - branches + - check + - exists + - info + - list + - load + - lock + - lookup + - path + - pr + - remove + - reset + - resolve + - route + - route_all + - scan + - set + - status + - sync + - system + - systems + - unlock + +@flow (11 commands) + - aggregate + - close + - create + - list + - post_close + - register + - registry + - restore + - scan + - templates + - unregister + +@memory (13 commands) + - analyze + - bootstrap + - check + - demo + - extract + - fragments + - rollover + - search + - status + - symbolic + - templates + - verify + - watch + +@prax (4 commands) + - dashboard + - log-audit + - monitor + - status + +@seedgo (8 commands) + - audit + - checklist + - diagnostics + - diagnostics_audit + - readme + - readme_update + - standards_audit + - standards_query + +@skills (6 commands) + - create + - discover + - info + - list + - run + - validate + +@spawn (4 commands) + - create + - delete + - passport + - update + +@trigger (15 commands) + - branch_log_events + - core + - errors + - fire + - list + - log_events + - medic + - mute + - off + - on + - reset + - start + - status + - stop + - unmute + +DISCOVERED: 173 commands across 14 branches + +--- + +## 4. Test Coverage (26% module coverage) + + +@ai_mail (49 tests, 2 files) + tests/test_send_identity.py -- 36 tests -> email, users + tests/test_user_paths.py -- 13 tests -> users + UNTESTED: branch_ping, central_writer, dispatch, json, json_utils, monitoring, notify, registry + +@api (0 tests, 0 files) + (no test files found) + +@backup (0 tests, 1 file) + tests/test_pattern_scan.py -- 0 tests -> config, operations + UNTESTED: backup_core, config, diff, google_drive_sync, integrations, json, models, operations, reauth_drive, reporting, utils + +@cli (0 tests, 0 files) + (no test files found) + +@commons (82 tests, 2 files) + tests/test_commons.py -- 72 tests -> curation, database, notifications, profiles, search, welcome + tests/test_lifecycle.py -- 10 tests -> database, search + UNTESTED: activity, artifact, artifacts, capsule, catchup, central, comment, comments, commons_identity, dashboard, digest, engagement, explore, feed, identity, json, leaderboard, notification, post, posts, profile, reaction, room, rooms, social, space, trade + +@daemon (31 tests, 1 file) + tests/test_actions_registry.py -- 31 tests -> actions + UNTESTED: activity_report, json, monitoring, schedule, scheduler_ops, update, wakeup_ops + +@devpulse (0 tests, 0 files) + (no test files found) + +@drone (335 tests, 9 files) + tests/test_activation.py -- 38 tests -> command_registry, executor + tests/test_commands.py -- 35 tests -> command_registry, commands + tests/test_discovery.py -- 40 tests -> discovery, discovery_handler, exceptions, module_registry_handler + tests/test_executor.py -- 26 tests -> exceptions, executor + tests/test_git_module.py -- 46 tests -> git, git_module, module_registry_handler + tests/test_registry_handler.py -- 33 tests -> exceptions, registry_handler + tests/test_resolver.py -- 52 tests -> exceptions, registry_handler, resolver + tests/test_router.py -- 37 tests -> exceptions, executor, router, router_handler + tests/test_scan.py -- 28 tests -> scan, scanning + UNTESTED: config, json, module_registry, registry + +@flow (0 tests, 0 files) + (no test files found) + +@memory (0 tests, 0 files) + (no test files found) + +@prax (0 tests, 0 files) + (no test files found) + +@seedgo (0 tests, 0 files) + (no test files found) + +@skills (123 tests, 8 files) + tests/test_cli_routing.py -- 20 tests -> ? + tests/test_discovery.py -- 32 tests -> discovery_handler + tests/test_lifecycle.py -- 12 tests -> creator, discovery, loader_handler, runner, template + tests/test_loader.py -- 7 tests -> loader + tests/test_registry.py -- 14 tests -> registry + tests/test_runner.py -- 13 tests -> runner + tests/test_runner_handler.py -- 14 tests -> runner_handler + tests/test_validator.py -- 11 tests -> validator + UNTESTED: creator_handler, json + +@spawn (113 tests, 5 files) + tests/test_citizen_classes.py -- 28 tests -> class_registry, core, passport_ops, update, update_ops + tests/test_handlers.py -- 36 tests -> change_detection, json_ops, meta_ops, reconcile + tests/test_lifecycle.py -- 22 tests -> delete, delete_ops, sync_registry, sync_registry_ops, sync_templates, sync_templates_ops + tests/test_spawn.py -- 13 tests -> metadata, placeholders, registry + tests/test_update.py -- 14 tests -> meta_ops, update, update_ops + UNTESTED: file_ops, json, passport + +@trigger (0 tests, 0 files) + (no test files found) + + +Test Coverage Report +═══════════════════════════════════════════════════════ + +TESTED: + ✓ drone 335 tests 9 files 15/19 modules covered (79%) + ✓ skills 123 tests 8 files 10/12 modules covered (83%) + ✓ spawn 113 tests 5 files 18/21 modules covered (86%) + +PARTIAL: + ◐ ai_mail 49 tests 2 files 2/10 modules covered (20%) + ◐ commons 82 tests 2 files 6/33 modules covered (18%) + ◐ daemon 31 tests 1 file 1/8 modules covered (12%) + +NO TESTS: + ✗ api 0 tests 0 files 0/10 modules covered (0%) + ✗ backup 0 tests 1 file 0/11 modules covered (0%) + ✗ cli 0 tests 0 files 0/5 modules covered (0%) + ✗ devpulse 0 tests 0 files 0/0 modules covered (0%) + ✗ flow 0 tests 0 files 0/16 modules covered (0%) + ✗ memory 0 tests 0 files 0/16 modules covered (0%) + ✗ prax 0 tests 0 files 0/14 modules covered (0%) + ✗ seedgo 0 tests 0 files 0/14 modules covered (0%) + ✗ trigger 0 tests 0 files 0/12 modules covered (0%) + +═══════════════════════════════════════════════════════ +SUMMARY: 733 tests across 15 branches + 3 branches tested, 3 partial, 9 untested + Coverage: 52/201 modules (26%) + diff --git a/src/aipass/devpulse/docs/testing_standards.md b/src/aipass/devpulse/docs/testing_standards.md new file mode 100644 index 00000000..f807717a --- /dev/null +++ b/src/aipass/devpulse/docs/testing_standards.md @@ -0,0 +1,337 @@ +# Testing Standards +**Status:** Draft v1 - Current manual process documented +**Date:** 2025-11-12 + +--- + +## Current State: Manual Testing (Effective for Rapid Iteration) + +**Reality:** AIPass is a custom framework in active evolution. System changes weekly. Building extensive test infrastructure now = updating tests constantly instead of building features. + +**Current approach:** Manual testing with JSON/log verification works effectively for rapid iteration phase. + +**Future:** pytest framework expansion once branches stabilize (pytest is already configured and operational in some branches). + +--- + +## The 90% Build Process + +**How we build and test:** + +### 1. Planning Phase (Upfront Investment) +- Issue plan through Flow (master plan or default plan depending on scope) +- Spend time getting structure right before coding +- Define what we're building and how it fits together + +### 2. Build to 90% (AI-Led, Internal Verification) +- AI builds structure and implementation +- **Internal verification tests as we go:** + - Does the module turn on? + - Do commands work? + - Basic functionality confirmed? +- Human mostly observes, provides input if things go astray +- Focus on getting structure and pieces in place + +**Linux advantage:** Same environment (AI and human both on Linux) = outputs match = internal tests are reliable. Windows had issues with this, Linux doesn't. + +### 3. 90% Threshold (Human Review Phase) +- Human gets actively involved +- Reviews structure and implementation +- Asks questions about concerns +- Performs feature tests (what module is supposed to do) +- Identifies bugs and missing error handling + +### 4. Debug Cycle (Error Handling First, Then Fixes) + +**Critical pattern: Fix error handling BEFORE fixing bugs** + +**Example scenario:** +``` +Bug: API call not working, but console says "Success" +No error output, but logs show API never executed +Output is lying + +Process: +1. Fix error handling FIRST - make errors tell the truth +2. THEN fix the actual bug (why API isn't executing) +3. See clean pass with honest outputs +4. Move to next feature +``` + +**Why this order:** +- Can't debug effectively if errors lie +- Truth in outputs = faster debugging +- Honest error messages = system teaches itself what's wrong + +### 5. Iterate Until Acceptable +- Test different features +- Continue debug cycle (handle errors → fix bugs → verify) +- Reach acceptable standard (basic or advanced depending on module) +- Move on + +--- + +## Why Manual Testing Works Right Now + +**Advantages:** +- **Fast iteration** - No test maintenance overhead +- **Flexible** - System can change tomorrow without breaking test suite +- **JSON/log infrastructure** - Acts as verification layer + - Check config.json → see settings + - Check data.json → see state + - Check log.json → see operations history + - Check Prax logs → detailed debugging +- **Linux environment match** - AI tests = human tests (same outputs) +- **Effective debugging** - Error-first approach catches issues early + +**Tradeoffs:** +- Manual effort required at 90% stage +- No automated regression testing (yet) +- Relies on human review for final verification + +**Acceptable because:** System is evolving rapidly. Better to build fast and test manually than build slow and maintain brittle tests. + +--- + +## Future: pytest Framework + +**Infrastructure in place:** +- `pytest.ini` at repo root - Configuration for test discovery and markers +- `tests/` - Root test directory with conftest.py +- Branch-specific test directories at `src/aipass/{branch}/tests/` (API, Prax, CLI, etc.) + +**Already operational in some branches:** +- API branch: 4 test files (test_api_system.py, test_openrouter_key.py, etc.) +- Prax branch: test_log_rotation.py validates rotation behavior +- Seedgo branch: test_cli_errors.py demonstrates error handling patterns + +**When to expand testing:** +- Once modules and branches stabilize +- When system changes slow down (monthly, not weekly) +- When maintenance cost < value of automation + +**Why pytest:** +- Standard Python testing framework +- Already configured with pytest.ini +- Infrastructure exists, ready for expansion +- Some branches already have working tests + +**Current selective approach:** +- Test critical/stable components (API, Prax log rotation) +- Skip testing rapidly changing features +- Manual testing for experimental work +- Automated tests where they add value without maintenance burden + +--- + +## Error Handling Philosophy + +**Errors must tell the truth** - foundational to testing effectiveness + +**Good error handling:** +```python +try: + result = api_call() + if not result: + logger.error("API call failed - no response") + return {"success": False, "error": "API returned no data"} +except Exception as e: + logger.error(f"API call exception: {e}", exc_info=True) + return {"success": False, "error": str(e)} +``` + +**Bad error handling:** +```python +try: + result = api_call() + return {"success": True} # LIES - didn't check if result valid +except: + pass # Silent failure - no truth +``` + +**Testing relies on honest errors:** +- If errors lie, testing is impossible +- Fix error handling first = testing becomes possible +- Then fix bugs with confident verification + +--- + +## JSON/Log Infrastructure as Testing Layer + +**Three-JSON pattern supports testing:** + +### Config Verification +```bash +# Check if settings are correct +cat module_name_config.json +# See API keys, limits, feature toggles +``` + +### State Verification +```bash +# Check current state +cat module_name_data.json +# See metrics, counts, current status +``` + +### Operations Verification +```bash +# Check what actually happened +cat module_name_log.json +# See recent operations and results +``` + +### Detailed Debugging +```bash +# Prax provides file-based logging, not real-time watching +# Check system logs directory for detailed output +ls -la system_logs/ +# Read specific module logs (Prax manages this directory location) +cat system_logs/module_name.log +``` + +**This infrastructure = verification layer without formal tests** + +--- + +## Testing Workflow Example + +**Building a new branch creation module:** + +1. **Plan** - Define structure, features, workflow (Flow plan) + +2. **Build to 90%** - AI implements: + - Module structure + - Handler functions + - Config/data/log JSONs + - Internal verification: "Does `create_branch test_branch` work?" + +3. **Review at 90%** - Human tests: + - Create branch with various names + - Check if files copied correctly + - Verify registry updated + - Try edge cases (existing branch, invalid name) + +4. **Find bug** - Branch created but registry not updated + - **First:** Check error handling - is error logged? Is return value honest? + - Add error handling if missing + - **Then:** Fix bug - why isn't registry updating? + - Verify with clean pass + +5. **Iterate** - Test more features: + - Template copying + - Placeholder replacement + - Memory file handling + - Backup on conflicts + +6. **Acceptable** - All major features work, errors are honest, ready to use + +--- + +## Current Testing Checklist + +**For any new module/feature:** + +- [ ] Does it turn on without errors? +- [ ] Do basic commands work? +- [ ] Are errors handled and logged? +- [ ] Do outputs tell the truth? +- [ ] Check config.json - settings correct? +- [ ] Check data.json - state tracking working? +- [ ] Check log.json - operations recorded? +- [ ] Test edge cases (invalid input, missing files, etc.) +- [ ] Check Prax logs in system_logs/ for detailed debugging info +- [ ] Manual feature tests at 90% stage + +--- + +## When to Test What + +**During development (AI internal verification):** +- Module starts without errors +- Basic commands execute +- Expected output appears + +**At 90% stage (human testing):** +- Feature functionality (does it do what it's supposed to?) +- Edge cases (what breaks it?) +- Error handling (are errors honest?) +- Integration (does it work with other modules?) + +**Before considering "done":** +- Clean passes on major features +- Errors tell the truth (no silent failures) +- JSON logs show operations correctly +- Acceptable standard reached (basic or advanced) + +--- + +## Summary + +**Current approach:** Manual testing with JSON/log infrastructure + +**Why it works:** +- Fast iteration without test maintenance +- Linux environment = reliable verification +- Error-first debugging = effective bug fixing +- JSON/log system = verification layer + +**Build process:** Plan → Build to 90% → Review/test → Debug (errors first, then bugs) → Iterate → Acceptable + +**Future:** pytest framework when system stabilizes + +**Philosophy:** Build fast, verify as you go, handle errors honestly, iterate rapidly. Test infrastructure comes after stability. + +--- + +## Pytest Test Structure + +**Current implementation:** +``` +pytest.ini # Test configuration (repo root) +tests/ # Root test directory + conftest.py # Shared fixtures +src/aipass/ + api/tests/ # API tests (4 files operational) + test_api_system.py + test_openrouter_key.py + test_free_models_quick.py + test_paid_model.py + prax/tests/ # Prax tests + test_log_rotation.py + cli/tests/ # CLI tests (infrastructure ready) + seedgo/ + tests/conftest.py + apps/modules/test_cli_errors.py # Demo module +``` + +**Running tests:** +```bash +# Run all tests +pytest + +# Run specific branch tests +pytest src/aipass/api/tests/ + +# Run with markers +pytest -m unit +pytest -m integration +pytest -m slow +``` + +**Test demonstration module:** +- `src/aipass/seedgo/apps/modules/test_cli_errors.py` - Shows error handling patterns +- Not a pytest test, but demonstrates testing concepts +- Run directly: `python3 src/aipass/seedgo/apps/modules/test_cli_errors.py` + +--- + +## Comments + +#@comments:2025-11-13:claude: Pytest infrastructure exists and operational in API/Prax branches. Selective testing approach: automate stable components, manual testing for rapid development. + +#@comments:2025-11-13:claude: Prax doesn't have real-time "watcher" command for logs - uses file-based logging in system_logs/ (Prax manages the location). Updated documentation to reflect actual capabilities. + +#@comments:2025-11-13:claude: Error-first debugging approach is critical - maybe this should be emphasized in error_handling.md when we fill that section? + +#@comments:2025-11-13:claude: "90% threshold" is interesting pattern - not 100% perfection, but "acceptable standard" (basic or advanced). Reflects pragmatic development philosophy.