feat(system): README update, API init fix, diagnostic docs, STATUS sync

- README.md: updated to current state (15 branches, 95+ PRs, 733 tests, 173 commands)
- API: init path fixed (creates .env at ~/.secrets/aipass/ not src/aipass/api/)
- API: init shows path and detects existing .env
- devpulse/docs: system_health_report, testing_standards, ai_mail comms upgrade
- STATUS.md: synced via drone @prax status sync

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
AIOSAI
2026-03-20 00:02:04 -07:00
co-authored by Claude Opus 4.6
parent a800e7919a
commit ec63b8ba29
14 changed files with 1696 additions and 175 deletions
+22
View File
@@ -0,0 +1,22 @@
# DPLAN-003 Working Directory
Research, mapping, and planning files for "AIPass as Operating System."
Parent plan: `AIPass/DPLAN-003_aipass_as_operating_system_2026-03-13.md`
## Scope
**In scope:** `src/aipass/` branches only (drone, seedgo, prax, cli, flow, ai_mail, api, trigger, spawn, devpulse, backup, daemon, memory).
**Out of scope:** `src/commons/`, `src/skills/` — these are separate and not part of this refactor.
## Files
| File | Purpose | Status |
|------|---------|--------|
| `registry_discovery_map.md` | Every find_registry() call in src/aipass/ | Done |
| `portability_audit.md` | Full investigation results (session 24) | Done |
| `credential_model.md` | Registry credential design (UUID now, macaroons later) | Stage 1 Done |
| `registry_refactor_plan.md` | Shared find_registry() design, migration steps | Pending |
| `aipass_init_spec.md` | CLI owns init — `drone @cli aipass init` (entry point later) | Pending |
| `aipass_help_spec.md` | CLI owns help — `drone @cli aipass help` (entry point later) | Pending |
@@ -0,0 +1,217 @@
# AI_MAIL COMMS UPGRADE — Hardening Plan
**Goal:** Make ai_mail bulletproof, then use as the model for other branches.
**Started:** 2026-03-10 | **Status:** In Progress
**Tracking:** Updated each session. Read this first on context resume.
---
## Current State (Session 10)
**Seedgo audit: 100%** — all 23 standards pass (42 files)
**Automated tests: 36** — test_send_identity.py v1.2.0 (3 audit rounds, 15 agents, 0 false positives)
**Production fixes deployed:** 3 cross-platform crashers fixed, test suite fully isolated
**Phase 2 (Error Handling):** 27 silent `except` blocks across 13 files now log with `logger.warning()`
### Test Suite Evolution
- v1.0.0: 31 tests — 7 false positives found by 3 audit agents
- v1.1.0: 32 tests — 8 suspects found by 5 audit agents (live registry, weak contracts)
- v1.2.0: 36 tests — final audit found 2 minor issues, fixed. **0 live-data dependencies.**
### Production Fixes (Session 9)
- `inbox_lock.py` — `import fcntl` guarded for Windows (msvcrt fallback)
- `ai_mail.py` — `signal.SIGPIPE` guarded with `hasattr` check
- `notify.py` — `/usr/bin/python3` replaced with `shutil.which("python3")`
---
## Architecture Map
```
Entry: apps/ai_mail.py
+-- Module: apps/modules/email.py (orchestrator, v3.0.0)
|-- handle_send() -> send_args.py (parse) -> send.py (execute) -> delivery.py (inbox write)
|-- handle_inbox() -> inbox_resolve.py -> inbox_ops.py -> format.py
|-- handle_view() -> inbox_ops.py
|-- handle_reply() -> reply.py -> delivery.py
|-- handle_close() -> close_ops.py
|-- handle_sent() -> format.py
+-- handle_contacts() -> registry/
Module: apps/modules/dispatch.py (dispatch orchestrator)
|-- daemon.py (auto-dispatch loop)
|-- wake.py (manual branch wake)
|-- dispatch_monitor.py (agent lifecycle)
|-- status.py (dispatch status display)
+-- pending_work.py (pending dispatch queue)
Identity: apps/handlers/users/
|-- branch_detection.py (detect_branch_from_pwd — THE critical path)
|-- user.py (get_current_user, get_branch_by_email)
+-- load.py, config_generator.py
Cross-branch: drone/apps/handlers/router_handler.py
+-- detect_caller_branch_name() -> sets AIPASS_CALLER_BRANCH env var
```
**Identity detection chain (9 stages):**
1. drone CLI entry
2. router resolves @branch to path
3. router_handler detects caller via CWD (+ AIPASS_BRANCH_NAME fallback)
4. executor merges caller_env into subprocess
5. ai_mail subprocess starts with env vars
6. branch_detection reads AIPASS_CALLER_BRANCH
7. send_args builds headers
8. delivery writes to recipient inbox.json
9. notify.py fires desktop notification
---
## Systemic Issues Found (Session 9 Deep Audit)
15 agents across 3 rounds audited tests + full codebase. Key findings:
### Silent Failures (24 instances) — VIOLATES "fail to errors" rule
`except Exception: return None` in 10+ places with zero logging:
- `branch_detection.py` — 3 functions (entire identity chain)
- `delivery.py` — get_all_branches(), summary, notification
- `create.py` — load_email_file()
- `user.py` — get_user_by_email(), get_all_users()
- `daemon.py` — _read_json() (silent) vs wake.py (logs) — inconsistent
- `inbox_ops.py` — migration persist has literal `pass`
### Dead/Redundant Code (~20% of codebase, ~1,728 lines)
- 5 fully unused files: validate.py, errors.py, data_ops.py, config_generator.py, pending_work.py
- `_find_repo_root()` copy-pasted in 9 files
- `get_all_branches()` implemented twice with different email behavior (correctness bug)
- `lock_utils.py` exists but unused — wake.py and daemon.py reimplement locking
- wake.py/daemon.py share 6+ duplicated functions
### Cross-Platform Breakers
- [FIXED] `import fcntl` unconditional — crashes Windows
- [FIXED] `/usr/bin/python3` hardcoded — breaks macOS/Windows
- [FIXED] `signal.SIGPIPE` — crashes Windows startup
- [OPEN] `pgrep` ungated in wake.py/daemon.py
- [OPEN] `email.py` parents[2] fragile — should use _find_repo_root()
---
## Hardening Phases
### Phase 1: Test Suite (Critical Path)
**Priority: HIGH** | **Status: IN PROGRESS**
- [x] `test_send_identity.py` — 36 tests, 3 audit rounds, 0 false positives
- [ ] `test_delivery.py` — round-trip send/receive
- [ ] `test_send_args.py` — argument parsing (all flags, interactive, error cases)
- [ ] `test_inbox_ops.py` — inbox operations (view, close, reply, close all)
- [ ] `test_dispatch_monitor.py` — dispatch lifecycle
- [ ] `test_notify.py` — notification delivery
### Phase 2: Error Handling Overhaul
**Priority: HIGH** | **Status: IN PROGRESS**
Silent failures are why the system feels "fragile."
- [x] Add `logger.warning()` to every bare `except Exception: return None` — **13 files, 27 instances fixed**
- [ ] Distinguish "not found" from "error reading" in return types
- [ ] Fix inconsistent error conventions (None vs tuple vs dict vs raise)
- [ ] Fix collision detection dead code in delivery.py get_all_branches()
### Phase 3: Code Consolidation
**Priority: MEDIUM** — Reduce duplication, single source of truth.
- [ ] Extract shared `_find_repo_root()` into commons or shared utility
- [ ] Consolidate `get_all_branches()` — one implementation, prefers explicit email
- [ ] Consolidate lock acquisition — use lock_utils.py, delete reimplementations
- [ ] Deduplicate wake.py/daemon.py shared functions (_read_json, _set_session_name, etc.)
- [ ] Archive 5 dead files (validate.py, errors.py, data_ops.py, config_generator.py, pending_work.py)
### Phase 4: Identity Consolidation
**Priority: MEDIUM** — Reduce 12 detection mechanisms to 1 canonical resolver.
- [ ] Define canonical `resolve_branch_identity()` function
- [ ] Priority chain: AIPASS_CALLER_BRANCH > AIPASS_BRANCH_NAME > CWD passport walk > --from
- [ ] Single file, single function, single source of truth
- [ ] All callers delegate to it
### Phase 5: Cross-Platform Hardening
**Priority: MEDIUM** — Public repo must work on all platforms.
- [ ] Guard `pgrep` usage in wake.py/daemon.py
- [ ] Replace `email.py` parents[2] with _find_repo_root()
- [ ] Guard `start_new_session=True` for Windows
- [ ] Add encoding='utf-8' to os.fdopen in wake.py
### Phase 6: Standards for Communication
**Priority: LOW** — Make patterns enforceable system-wide.
- [ ] `communication` standard — envelope format, required fields, identity rules
- [ ] `identity` standard — env var hierarchy, passport requirements
- [ ] `dispatch` standard — env isolation, lock management, bounce handling
### Phase 7: Documentation & System Prompt
**Priority: LOW** — Update branch local prompt.
- [ ] Write proper `.aipass/aipass_local_prompt.md`
- [ ] Document key commands, architecture, critical files
---
## Decision Log
| Date | Decision | Rationale |
|------|----------|-----------|
| 2026-03-08 | AIPASS_BRANCH_NAME env var over CWD detection | CWD unreliable when agents navigate. Env vars persist. |
| 2026-03-08 | dbus direct over notify-send | Portal mode strips persistence hints. dbus bypasses confinement. |
| 2026-03-08 | Unique app_name per source | GNOME collapses same desktop-entry into one notification. |
| 2026-03-08 | Fail loud, no CWD fallback for identity | Silent wrong sender worse than visible error. |
| 2026-03-10 | --from flag for explicit sender override | Plumbing existed, just needed CLI wiring. |
| 2026-03-10 | Tests first in hardening plan | Every bug from sessions 4-7 would have been caught by tests. |
| 2026-03-10 | 3 audit rounds on test suite | Each round found issues the previous missed. Diminishing returns by round 3. |
| 2026-03-10 | Error handling overhaul before code consolidation | Silent failures cause more user pain than code duplication. |
---
## Session Notes
### Session 10 (2026-03-10)
- **Phase 2: Error Handling Overhaul — logger.warning() sweep complete**
- 13 files modified, 27 silent `except Exception` blocks now log with `logger.warning()`
- Files fixed (by priority):
- `branch_detection.py` (3) — identity chain, most critical
- `user.py` (2) — user lookup
- `delivery.py` (6) — get_all_branches, migrate, callback, summary, notification, private branch
- `create.py` (2) — purge, load_email_file
- `inbox_ops.py` (1) — migration persist
- `format.py` (1) — alias lookup
- `reply.py` (1) — get_email_by_id
- `close_ops.py` (3) — dashboard, central, purge post-ops
- `send.py` (2) — central update after send/broadcast
- `dashboard_sync.py` (1) — push_dashboard_update
- `inbox_cleanup.py` (4) — migrate, dashboard, central, purge
- `error_handler.py` (1) — error notification delivery
- `error_dispatch.py` (2) — dashboard, central post-ops
- All 13 files pass `py_compile` syntax check
- Remaining silent blocks (acceptable): inbox_lock.py finally-block, json_handler.py (utility),
daemon/wake/dispatch_monitor (already log with logger.info), config_generator/data_ops (unused files)
- Next: Phase 2 remaining items (return type distinctions, error conventions, collision dead code)
### Session 9 (2026-03-10, continued)
- Wrote test_send_identity.py v1.0.0 (31 tests)
- Audit round 1: 3 agents found 7 false positives -> rewrote to v1.1.0
- Audit round 2: 5 agents found 8 issues (live registry, weak contracts) -> v1.2.0 (36 tests)
- Audit round 3: 5 agents (final test + full system sweep)
- Tests: 2 minor issues found and fixed (unpatched registry, missing mailbox_path assert)
- Error handling: 24 silent failure instances across 10+ files
- Code quality: ~1,728 dead/redundant lines (20% of codebase), 5 unused files
- Cross-platform: 3 CRITICAL (fcntl, SIGPIPE, /usr/bin/python3) — ALL FIXED
- Paths: parents[N] usage audited, mostly correct, 1 fragile case in email.py
- Fixed 3 cross-platform crashers in production code
- Updated COMMS_UPGRADE.md with full findings and revised phases
- Next: Phase 1 continues (more test files) or Phase 2 (error handling overhaul)
### Session 8 (2026-03-10)
- Patrick initiated hardening project
- Self-audit: 100% on all 22 seedgo standards
- Mapped full architecture (55 Python files, 3 modules, 8 handler domains)
- Created this tracking document
@@ -0,0 +1,157 @@
# DPLAN-003: Registry Credential Model
## The Idea
Registries get a unique token. Passports carry that token. Access is identity-based, not filesystem-based. No walk-up needed — your credential proves which registry is yours.
## Why
Current system finds registries by walking up directories. Works for one project, breaks with multiple. If two AIPass projects exist on one machine, a citizen launched from the wrong directory finds the wrong registry. Credentials solve this — your passport carries proof of membership.
## Prior Art (Research)
| System | Pattern | Fit |
|--------|---------|-----|
| **Macaroons** (Google Research) | Token IS the credential. Delegatable with caveats. Offline verification. DeepMind validated for AI agent delegation (2026). | Highest |
| **Vault Namespaces** | Project = namespace. Token scoped to namespace. Mini-registry per project. | High |
| **AWS STS / Token Vending** | Agent presents project ID, gets scoped credential. Credential itself is the boundary. | High |
| **K8s Namespace + ServiceAccount** | Token carries project scope as claim. RBAC composable. | High |
| **SPIFFE/SPIRE** | Process-level attestation without static secrets. | Medium |
| **direnv** | Auto-set env vars on directory entry. Zero-friction UX. | UX pattern |
Full research: agent output from session 25.
## Design: Two Stages
### Stage 1: UUID Match (Manual, Now)
Simple. Prove the concept works before adding crypto.
**Registry gets an ID:**
```json
{
"metadata": {
"id": "a1b2c3d4-...",
"name": "AIPASS",
"version": "1.0.0",
"last_updated": "2026-03-13",
"total_branches": 15
},
"branches": [...]
}
```
**Passports get the matching ID:**
```json
{
"citizenship": {
"registered": true,
"registry_id": "a1b2c3d4-...",
"registry_name": "AIPASS",
"citizen_number": 7
}
}
```
**Lookup flow:**
1. Walk up from CWD, find `*_REGISTRY.json`
2. Read its `metadata.id`
3. Check citizen's `citizenship.registry_id` matches
4. If mismatch → error ("citizen belongs to registry X, found registry Y")
5. If match → proceed
**Spawn changes:**
- `aipass init` (or manual setup) generates the UUID for the registry
- Spawn reads registry UUID and injects into new passports via `{{REGISTRY_ID}}` placeholder
- Existing 15 branches get the UUID added to their passports (one-time migration)
### Stage 2: Macaroon Tokens (Future, When Cross-Project Needed)
Upgrade path when we need delegation and cross-project access.
**Root token:** Created at `aipass init`. HMAC-based. Stored at `~/.secrets/aipass/projects/<uuid>.token`
**Citizen token:** Attenuated copy in passport. Can prove membership but can't mint new citizens.
**Agent token:** Further attenuated. Carries:
- Registry ID (which project)
- Scope (full access — agents do the real work in AIPass)
- Expiry (session-scoped, dies when agent dies)
- Issuer (which citizen spawned this agent)
Agents are NOT read-only. They're the builders — they write code, run tests, modify files. The token proves they belong to a project, it doesn't restrict what they do within it. Scope restrictions would be role-based (e.g., "can't modify other branches' files") not capability-based.
**Verification:** Local HMAC check. No daemon needed for basic validation. Daemon is optional enhancement for audit logging and revocation.
**direnv integration:** Entering a project directory auto-sets `AIPASS_PROJECT_TOKEN` in env. Agents inherit it.
## What Changes (Stage 1)
| Component | Change |
|-----------|--------|
| `AIPASS_REGISTRY.json` | Add `metadata.id` (UUID) |
| Passport template | Add `citizenship.registry_id` placeholder |
| Spawn `build_replacements_dict()` | Read registry UUID, add `{{REGISTRY_ID}}` |
| Spawn `add_to_registry()` | No change (branch entries stay the same) |
| Drone `find_registry()` | Optional: verify passport.registry_id matches found registry |
| All 15 passports | One-time: add `registry_id` field |
## What Does NOT Change
- Registry filename stays `*_REGISTRY.json`
- Walk-up discovery still works (credential is verification layer on top, not replacement)
- Branch structure unchanged
- No daemon needed
- No new dependencies
## Resolved Questions
1. **Registry ID location** → `metadata.id` — it's the project's identity, not the citizen's.
2. **UUID4 vs hash** → UUID4 (random). Simple, guaranteed unique, no inputs needed.
3. **Verification mode** → Hard error on mismatch. Fail loudly — that's the AIPass way.
4. **Where does init live?** → Temporary: `src/aipass/devpulse/apps/init_project.py`. Future: CLI branch.
## Bugs, Quirks, and Findings
Discovered during testing. Reference for future work.
### init_project.py (10 edge cases tested)
- **Spaces in dir name** → registry filename gets spaces (`MY COOL PROJECT_REGISTRY.json`). Fixed: `_sanitize_name()` replaces non-alphanumeric with `_`.
- **Root path `/`** → `Path("/").name` is empty string, creates `_REGISTRY.json`. Fixed: validation rejects empty name.
- **Permission errors** → raw traceback instead of clean message. Fixed: `main()` catches `OSError`.
- **Passport overwrite on re-init** → if registry deleted but `.trinity/` survives, re-init would silently overwrite passport with new UUID. Fixed: passport guarded with `exists()` check.
- **Double init** → correctly blocked by `FileExistsError` on registry file.
- **Deep nested paths** → `mkdir(parents=True)` handles correctly.
- **UUID uniqueness** → 5 runs, 5 unique UUIDs. No collisions.
- **JSON validity** → all generated files parse clean.
### Drone isolation (6 scenarios tested)
- **Sibling projects** → PASS. Two projects in same parent dir, fully isolated.
- **Nested project** → PASS. Inner registry wins over outer. No bleed-through.
- **Deep subdir walk-up** → PASS. Finds nearest ancestor registry correctly.
- **Cross-project contamination** → PASS. Branches are registry-scoped; modules are global.
- **Empty directory (no registry)** → CONCERN. Drone silently falls back to AIPass source registry via `__file__` walk-up. Any dir on this machine without a registry sees production branches. This is by design in `find_registry()` but will be dangerous with multi-project. Credential verification would catch this — citizen's registry_id won't match the fallback registry.
- **`drone @ai_mail` from mock project** → correctly returns "Branch not found in registry" (empty project has no branches).
### Registry/passport gitignore
- `AIPASS_REGISTRY.json` and all `.trinity/passport.json` files are gitignored. UUID migration is local-only. This is correct for now — credentials are machine-specific, not repo state. But means `aipass init` must run on every clone/install. Future: consider whether UUID should be in-repo or machine-local.
### AI mail after migration
- Send/receive works fine after UUID migration. AI mail's 4 internal `find_registry()` copies still hardcode `AIPASS_REGISTRY.json` — functional for now since that filename exists, but won't find `*_REGISTRY.json` in other projects.
### Drone built-in modules vs branches
- `@drone` and `@seedgo` are hardcoded as "modules" in `module_registry.py`, always visible everywhere. Other branches (`@ai_mail`, `@spawn`, etc.) are registry-scoped. This distinction matters: modules are global services, branches are project citizens.
## Decision Log
- **2026-03-14:** Patrick proposed credential-based registry access. Agents do the real work — tokens prove membership, not restrict capability. Manual first, `aipass init` later.
- **2026-03-14:** Research confirmed macaroons as best-fit pattern (Google Research + DeepMind 2026 validation). Stage 1 = UUID match, Stage 2 = macaroon upgrade.
- **2026-03-14:** Scope limited to `src/aipass/` branches only. Commons and skills excluded.
- **2026-03-14:** All 4 open questions resolved. `init_project.py` built and tested — creates registry with UUID, passport with matching registry_id, .trinity/, .aipass/, AIPASS.md. Tested with temp directory — drone isolation confirmed (only built-in modules visible, not AIPass branches).
- **2026-03-14:** UUID migration executed (FPLAN-0030). Registry + 13 passports updated. Registry and passports are gitignored — UUID is machine-local, not repo state.
- **2026-03-14:** Drone verification added (`_verify_registry_credential` in load_registry). Drone dispatched and fixed error propagation — new `RegistryMismatchError` class separates "not found" (fallback OK) from "mismatch" (hard error). Spawn templates updated with `{{REGISTRY_ID}}` placeholder.
- **2026-03-14:** Stage 1 complete. Credential model is end-to-end: init creates credentialed registries, drone verifies on load, spawn injects into new passports, mismatch = hard error. Remaining: drone CLI stderr surfacing (DPLAN-0031), `aipass init` CLI command (future).
- **2026-03-14:** FPLAN-0032 executed — CLI stderr standardization Phase 1+2 done. CLI owns err_console, error()/warning() go to stderr. Drone imports from CLI instead of creating its own Console(stderr=True). Credential mismatch errors now properly surface to terminal via stderr. The stderr issue that blocked error visibility during DPLAN-003 testing is resolved.
- **2026-03-14:** CLI confirmed as owner of credential model + project commands. CLI owns: `aipass init` (create project), `aipass help` (project help), display API (error/warning/fatal/err_console). Drone owns verification (registry_handler). Spawn owns injection (passport templates). Access via `drone @cli aipass init/help` for now; `aipass init` entry point is future icing.
- **2026-03-14:** Seedgo `stderr_routing` standard created (24th standard). Checker bug found — was scanning all files but not displaying module/handler results. Seedgo fixed display + aggregation. Agent-verified across 3 branches. FPLAN-0033 parked for Phase 3 migration (343 error prints, 10 branches).
- **2026-03-14:** FPLAN-0033 Phase 3 executed autonomously (Patrick away). 15+ agents across 3 waves migrated 10 branches. 48 files, 666 insertions. System stderr avg 49%→69%. Post-migration verification found 63% of remaining seedgo violations are false positives ([yellow]Label:[/yellow] section headers flagged as warnings). Seedgo dispatched for checker refinement. 93 real violations remain (48 in commons which was out of scope).
@@ -0,0 +1,31 @@
# Portability Audit — Session 24 Results
## Summary
| Tool | Registry Discovery | CWD-Aware | Portable | Hardcoded |
|------|-------------------|-----------|----------|-----------|
| Drone | Walk-up + env var | No (uses registry) | Yes | Registry filename |
| Spawn | Walk-up + env var | No (uses registry) | Partial | Template location |
| Prax | Walk-up (no env) | No (sys logs at repo) | Partial | System logs dir |
| AI_Mail | Walk-up (no env) | No (inbox per branch) | Yes | Inbox location |
| Flow | Walk-up (no env) | Yes (plan creation) | Hybrid | Plan registry |
## Key Findings
- All tools use walk-up strategy to find `AIPASS_REGISTRY.json`
- Registry-relative path resolution already works (move registry + dirs = works)
- `AIPASS_REGISTRY` env var supported by drone and spawn
- System logs hardcoded to `{repo_root}/system_logs/`
- Spawn templates hardcoded to `{spawn_package}/templates/`
- Walk-up doesn't stop at project boundaries — finds nearest registry up the tree
## The Core Fix
Change `find_registry()` to:
1. Walk up from CWD looking for `*_REGISTRY.json` (glob, not hardcoded name)
2. Stop at first match — that's the project boundary
3. If none found, return error ("No AIPass project. Run `aipass init`")
## Source
Full investigation transcript: background agent session 24, 42 tool calls across drone/spawn/ai_mail/flow/prax.
@@ -0,0 +1,46 @@
# Registry Discovery Map — FPLAN-0029 Phase 1
## Primary Implementations (5)
| File | Function | Strategy |
|------|----------|----------|
| `src/aipass/spawn/apps/handlers/registry.py:34-79` | `find_registry(start_path)` | Env var → project markers (.git/pyproject.toml) → walk-up → cwd fallback |
| `src/aipass/drone/apps/handlers/registry_handler.py:36-61` | `find_registry()` | Walk-up from __file__ → walk-up from cwd → parents[4] fallback |
| `src/aipass/drone/apps/handlers/registry_handler.py:64-81` | `get_registry_path()` | Global override → env var → find_registry() |
| `src/commons/apps/handlers/database/db.py:343-378` | `_find_branch_registry()` | Env var AIPASS_ROOT → walk-up (10 limit) → ~/.aipass/ fallback |
| `src/commons/apps/handlers/identity/identity_ops.py:33-67` | `_find_branch_registry_path()` | Same as db.py |
## Local Reimplementations (23+)
All do walk-up from `__file__` looking for `AIPASS_REGISTRY.json`:
### ai_mail (4 files)
- `apps/handlers/users/branch_detection.py:26-38`
- `apps/handlers/registry/read.py:31-41`
- `apps/handlers/email/format.py:21-31`
- `apps/handlers/dispatch/daemon.py:36-55`
### memory (3 files)
- `apps/handlers/dashboard_push.py:44-54`
- `apps/handlers/monitor/detector.py:36-83`
- `apps/handlers/monitor/memory_watcher.py:363-385`
### prax (2 files)
- `apps/handlers/registry/reader.py:24-39`
- `apps/handlers/dashboard/agent_status_writer.py:33-44`
### seedgo (3 files)
- `apps/handlers/audit/discovery.py:54-62`
- `apps/handlers/diagnostics/discovery.py:21-29`
- `apps/handlers/readme/readme_ops.py:29-37`
### + 11 more across other branches
## Hardcoded String Count
100+ references to `"AIPASS_REGISTRY.json"` across codebase.
## Phase 1 Plan
1. Create `commons.registry.find_registry()` — one function, glob for `*_REGISTRY.json`
2. All 23+ modules import from commons
3. Drone/spawn wrappers call commons internally
4. Test: temp dir with TEST_REGISTRY.json → drone systems finds it
@@ -0,0 +1,525 @@
# AIPass System Health Report
Generated: 2026-03-19 22:54
---
## 1. Dead Code (14 unused files)
Dead Code Scanner
Scanning 15 branches
@ai_mail (3 modules, 31 handlers)
x handlers/monitoring/errors.py -- 0 references
> 33/34 files referenced
@api (4 modules, 14 handlers)
x handlers/openrouter/provision.py -- 0 references
> 17/18 files referenced
@backup (4 modules, 24 handlers)
x handlers/diff/vscode_integration.py -- 0 references
> 27/28 files referenced
@cli (3 modules, 2 handlers)
> All 5 files referenced
@commons (22 modules, 35 handlers)
> All 57 files referenced
@daemon (6 modules, 8 handlers)
> All 14 files referenced
@devpulse -- no apps/ directory
@drone (9 modules, 16 handlers)
> All 25 files referenced
@flow (8 modules, 35 handlers)
x handlers/plan/file_ops.py -- 0 references
x handlers/plan/update_registry.py -- 0 references
> 41/43 files referenced
@memory (5 modules, 29 handlers)
x handlers/learnings/manager.py -- 0 references
x handlers/schema/normalize.py -- 0 references
x handlers/search/vector_search.py -- 0 references
> 31/34 files referenced
@prax (6 modules, 36 handlers)
> All 42 files referenced
@seedgo (5 modules, 66 handlers)
x handlers/config/aipass_bypass.py -- 0 references
x handlers/config/aipass_ignore.py -- 0 references
x handlers/diagnostics/python_diognostics.py -- 0 references
x handlers/diagnostics/typscript_diognostics.py -- 0 references
x handlers/file/file_handler.py -- 0 references
x handlers/mock_standard_1/bypass_config/bypass.config.py -- 0 references
> 65/71 files referenced
@skills (5 modules, 8 handlers)
> All 13 files referenced
@spawn (6 modules, 15 handlers)
> All 21 files referenced
@trigger (5 modules, 17 handlers)
> All 22 files referenced
TOTAL: 14 unused across 14 branches
---
## 2. Local Prompts (5 stubs need enrichment)
Local Prompt Status
==================================================
RICH (50+ lines, 4+ sections):
v daemon 56 lines 4 sections
v devpulse 79 lines 8 sections
v flow 67 lines 6 sections
v seedgo 56 lines 4 sections
v skills 58 lines 5 sections
v trigger 53 lines 7 sections
BASIC (15-49 lines):
~ api 15 lines 2 sections
~ commons 48 lines 6 sections
~ memory 37 lines 5 sections
~ spawn 20 lines 4 sections
STUB (<15 lines):
x ai_mail 14 lines 1 section
x backup 14 lines 1 section
x cli 14 lines 1 section
x drone 14 lines 1 section
x prax 14 lines 1 section
==================================================
SUMMARY: 6 rich, 4 basic, 5 stub
==================================================
Section Breakdown
==================================================
@ai_mail (14 lines, STUB)
v Status
470 bytes
@api (15 lines, BASIC)
v Identity
v Key Breadcrumbs
1192 bytes
@backup (14 lines, STUB)
v Status
499 bytes
@cli (14 lines, STUB)
v Status
499 bytes
@commons (48 lines, BASIC)
v Commands
v Architecture
v Integration Points
v Critical Files
v Role
v Key Details
2328 bytes
@daemon (56 lines, RICH)
v Commands
v Memory & Tracking
v Apps Layout
v Known Issues
2788 bytes
@devpulse (79 lines, RICH)
v Identity
v How You Work
v Dispatch Table
v Commands
v Branches
v Working Habits
v Autonomous Monitoring
v Memory & Tracking
v Has dispatch table (@branch refs)
4710 bytes
@drone (14 lines, STUB)
v Status
503 bytes
@flow (67 lines, RICH)
v Commands
v Architecture
v Integration Points
v Conventions
v Critical Files
v Plan Type System
3461 bytes
@memory (37 lines, BASIC)
v Identity
v Commands
v Architecture
v Memory & Tracking
v Known Issues
1591 bytes
@prax (14 lines, STUB)
v Status
501 bytes
@seedgo (56 lines, RICH)
v Commands
v Apps Layout (extra layer vs standard branch)
v How I Work — Standards Reasoning
v Quick Reference
3166 bytes
@skills (58 lines, RICH)
v Commands
v Memory & Tracking
v Apps Layout
v Search Paths (first match wins)
v Three Skill Tiers
2706 bytes
@spawn (20 lines, BASIC)
v Commands
v Architecture
v Role
v Principles
745 bytes
@trigger (53 lines, RICH)
v Dispatch Table
v Commands
v Architecture
v Integration Points
v Critical Files
v Role
v Rules
2610 bytes
---
## 3. Commands (173 discovered)
@ai_mail (14 commands)
- close
- contacts
- dispatch
- email
- inbox
- ping
- read
- registry
- reply
- send
- sent
- status
- thresholds
- view
@api (12 commands)
- call
- cleanup
- google
- init
- models
- reauth
- session
- stats
- status
- test
- track
- validate
@backup (1 command)
- reauth
@cli (5 commands)
- aipass
- demo
- display
- show
- templates
@commons (51 commands)
- activity
- artifacts
- capsule
- capsules
- catchup
- collab
- comment
- craft
- database
- decorate
- delete
- digest
- drop
- enter
- event
- explore
- feed
- find
- gift
- inspect
- leaderboard
- leaderboards
- log
- look
- mint
- mute
- open
- pin
- pinned
- post
- preferences
- profile
- prompt
- react
- reactions
- room
- search
- secrets
- sign
- thread
- track
- trade
- trending
- unpin
- unreact
- visitors
- vote
- watch
- welcome
- who
- whoami
@daemon (5 commands)
- actions
- activity
- activity_report
- schedule
- update
@drone (24 commands)
- activate
- add
- branches
- check
- exists
- info
- list
- load
- lock
- lookup
- path
- pr
- remove
- reset
- resolve
- route
- route_all
- scan
- set
- status
- sync
- system
- systems
- unlock
@flow (11 commands)
- aggregate
- close
- create
- list
- post_close
- register
- registry
- restore
- scan
- templates
- unregister
@memory (13 commands)
- analyze
- bootstrap
- check
- demo
- extract
- fragments
- rollover
- search
- status
- symbolic
- templates
- verify
- watch
@prax (4 commands)
- dashboard
- log-audit
- monitor
- status
@seedgo (8 commands)
- audit
- checklist
- diagnostics
- diagnostics_audit
- readme
- readme_update
- standards_audit
- standards_query
@skills (6 commands)
- create
- discover
- info
- list
- run
- validate
@spawn (4 commands)
- create
- delete
- passport
- update
@trigger (15 commands)
- branch_log_events
- core
- errors
- fire
- list
- log_events
- medic
- mute
- off
- on
- reset
- start
- status
- stop
- unmute
DISCOVERED: 173 commands across 14 branches
---
## 4. Test Coverage (26% module coverage)
@ai_mail (49 tests, 2 files)
tests/test_send_identity.py -- 36 tests -> email, users
tests/test_user_paths.py -- 13 tests -> users
UNTESTED: branch_ping, central_writer, dispatch, json, json_utils, monitoring, notify, registry
@api (0 tests, 0 files)
(no test files found)
@backup (0 tests, 1 file)
tests/test_pattern_scan.py -- 0 tests -> config, operations
UNTESTED: backup_core, config, diff, google_drive_sync, integrations, json, models, operations, reauth_drive, reporting, utils
@cli (0 tests, 0 files)
(no test files found)
@commons (82 tests, 2 files)
tests/test_commons.py -- 72 tests -> curation, database, notifications, profiles, search, welcome
tests/test_lifecycle.py -- 10 tests -> database, search
UNTESTED: activity, artifact, artifacts, capsule, catchup, central, comment, comments, commons_identity, dashboard, digest, engagement, explore, feed, identity, json, leaderboard, notification, post, posts, profile, reaction, room, rooms, social, space, trade
@daemon (31 tests, 1 file)
tests/test_actions_registry.py -- 31 tests -> actions
UNTESTED: activity_report, json, monitoring, schedule, scheduler_ops, update, wakeup_ops
@devpulse (0 tests, 0 files)
(no test files found)
@drone (335 tests, 9 files)
tests/test_activation.py -- 38 tests -> command_registry, executor
tests/test_commands.py -- 35 tests -> command_registry, commands
tests/test_discovery.py -- 40 tests -> discovery, discovery_handler, exceptions, module_registry_handler
tests/test_executor.py -- 26 tests -> exceptions, executor
tests/test_git_module.py -- 46 tests -> git, git_module, module_registry_handler
tests/test_registry_handler.py -- 33 tests -> exceptions, registry_handler
tests/test_resolver.py -- 52 tests -> exceptions, registry_handler, resolver
tests/test_router.py -- 37 tests -> exceptions, executor, router, router_handler
tests/test_scan.py -- 28 tests -> scan, scanning
UNTESTED: config, json, module_registry, registry
@flow (0 tests, 0 files)
(no test files found)
@memory (0 tests, 0 files)
(no test files found)
@prax (0 tests, 0 files)
(no test files found)
@seedgo (0 tests, 0 files)
(no test files found)
@skills (123 tests, 8 files)
tests/test_cli_routing.py -- 20 tests -> ?
tests/test_discovery.py -- 32 tests -> discovery_handler
tests/test_lifecycle.py -- 12 tests -> creator, discovery, loader_handler, runner, template
tests/test_loader.py -- 7 tests -> loader
tests/test_registry.py -- 14 tests -> registry
tests/test_runner.py -- 13 tests -> runner
tests/test_runner_handler.py -- 14 tests -> runner_handler
tests/test_validator.py -- 11 tests -> validator
UNTESTED: creator_handler, json
@spawn (113 tests, 5 files)
tests/test_citizen_classes.py -- 28 tests -> class_registry, core, passport_ops, update, update_ops
tests/test_handlers.py -- 36 tests -> change_detection, json_ops, meta_ops, reconcile
tests/test_lifecycle.py -- 22 tests -> delete, delete_ops, sync_registry, sync_registry_ops, sync_templates, sync_templates_ops
tests/test_spawn.py -- 13 tests -> metadata, placeholders, registry
tests/test_update.py -- 14 tests -> meta_ops, update, update_ops
UNTESTED: file_ops, json, passport
@trigger (0 tests, 0 files)
(no test files found)
Test Coverage Report
═══════════════════════════════════════════════════════
TESTED:
✓ drone 335 tests 9 files 15/19 modules covered (79%)
✓ skills 123 tests 8 files 10/12 modules covered (83%)
✓ spawn 113 tests 5 files 18/21 modules covered (86%)
PARTIAL:
◐ ai_mail 49 tests 2 files 2/10 modules covered (20%)
◐ commons 82 tests 2 files 6/33 modules covered (18%)
◐ daemon 31 tests 1 file 1/8 modules covered (12%)
NO TESTS:
✗ api 0 tests 0 files 0/10 modules covered (0%)
✗ backup 0 tests 1 file 0/11 modules covered (0%)
✗ cli 0 tests 0 files 0/5 modules covered (0%)
✗ devpulse 0 tests 0 files 0/0 modules covered (0%)
✗ flow 0 tests 0 files 0/16 modules covered (0%)
✗ memory 0 tests 0 files 0/16 modules covered (0%)
✗ prax 0 tests 0 files 0/14 modules covered (0%)
✗ seedgo 0 tests 0 files 0/14 modules covered (0%)
✗ trigger 0 tests 0 files 0/12 modules covered (0%)
═══════════════════════════════════════════════════════
SUMMARY: 733 tests across 15 branches
3 branches tested, 3 partial, 9 untested
Coverage: 52/201 modules (26%)
@@ -0,0 +1,337 @@
# Testing Standards
**Status:** Draft v1 - Current manual process documented
**Date:** 2025-11-12
---
## Current State: Manual Testing (Effective for Rapid Iteration)
**Reality:** AIPass is a custom framework in active evolution. System changes weekly. Building extensive test infrastructure now = updating tests constantly instead of building features.
**Current approach:** Manual testing with JSON/log verification works effectively for rapid iteration phase.
**Future:** pytest framework expansion once branches stabilize (pytest is already configured and operational in some branches).
---
## The 90% Build Process
**How we build and test:**
### 1. Planning Phase (Upfront Investment)
- Issue plan through Flow (master plan or default plan depending on scope)
- Spend time getting structure right before coding
- Define what we're building and how it fits together
### 2. Build to 90% (AI-Led, Internal Verification)
- AI builds structure and implementation
- **Internal verification tests as we go:**
- Does the module turn on?
- Do commands work?
- Basic functionality confirmed?
- Human mostly observes, provides input if things go astray
- Focus on getting structure and pieces in place
**Linux advantage:** Same environment (AI and human both on Linux) = outputs match = internal tests are reliable. Windows had issues with this, Linux doesn't.
### 3. 90% Threshold (Human Review Phase)
- Human gets actively involved
- Reviews structure and implementation
- Asks questions about concerns
- Performs feature tests (what module is supposed to do)
- Identifies bugs and missing error handling
### 4. Debug Cycle (Error Handling First, Then Fixes)
**Critical pattern: Fix error handling BEFORE fixing bugs**
**Example scenario:**
```
Bug: API call not working, but console says "Success"
No error output, but logs show API never executed
Output is lying
Process:
1. Fix error handling FIRST - make errors tell the truth
2. THEN fix the actual bug (why API isn't executing)
3. See clean pass with honest outputs
4. Move to next feature
```
**Why this order:**
- Can't debug effectively if errors lie
- Truth in outputs = faster debugging
- Honest error messages = system teaches itself what's wrong
### 5. Iterate Until Acceptable
- Test different features
- Continue debug cycle (handle errors → fix bugs → verify)
- Reach acceptable standard (basic or advanced depending on module)
- Move on
---
## Why Manual Testing Works Right Now
**Advantages:**
- **Fast iteration** - No test maintenance overhead
- **Flexible** - System can change tomorrow without breaking test suite
- **JSON/log infrastructure** - Acts as verification layer
- Check config.json → see settings
- Check data.json → see state
- Check log.json → see operations history
- Check Prax logs → detailed debugging
- **Linux environment match** - AI tests = human tests (same outputs)
- **Effective debugging** - Error-first approach catches issues early
**Tradeoffs:**
- Manual effort required at 90% stage
- No automated regression testing (yet)
- Relies on human review for final verification
**Acceptable because:** System is evolving rapidly. Better to build fast and test manually than build slow and maintain brittle tests.
---
## Future: pytest Framework
**Infrastructure in place:**
- `pytest.ini` at repo root - Configuration for test discovery and markers
- `tests/` - Root test directory with conftest.py
- Branch-specific test directories at `src/aipass/{branch}/tests/` (API, Prax, CLI, etc.)
**Already operational in some branches:**
- API branch: 4 test files (test_api_system.py, test_openrouter_key.py, etc.)
- Prax branch: test_log_rotation.py validates rotation behavior
- Seedgo branch: test_cli_errors.py demonstrates error handling patterns
**When to expand testing:**
- Once modules and branches stabilize
- When system changes slow down (monthly, not weekly)
- When maintenance cost < value of automation
**Why pytest:**
- Standard Python testing framework
- Already configured with pytest.ini
- Infrastructure exists, ready for expansion
- Some branches already have working tests
**Current selective approach:**
- Test critical/stable components (API, Prax log rotation)
- Skip testing rapidly changing features
- Manual testing for experimental work
- Automated tests where they add value without maintenance burden
---
## Error Handling Philosophy
**Errors must tell the truth** - foundational to testing effectiveness
**Good error handling:**
```python
try:
result = api_call()
if not result:
logger.error("API call failed - no response")
return {"success": False, "error": "API returned no data"}
except Exception as e:
logger.error(f"API call exception: {e}", exc_info=True)
return {"success": False, "error": str(e)}
```
**Bad error handling:**
```python
try:
result = api_call()
return {"success": True} # LIES - didn't check if result valid
except:
pass # Silent failure - no truth
```
**Testing relies on honest errors:**
- If errors lie, testing is impossible
- Fix error handling first = testing becomes possible
- Then fix bugs with confident verification
---
## JSON/Log Infrastructure as Testing Layer
**Three-JSON pattern supports testing:**
### Config Verification
```bash
# Check if settings are correct
cat module_name_config.json
# See API keys, limits, feature toggles
```
### State Verification
```bash
# Check current state
cat module_name_data.json
# See metrics, counts, current status
```
### Operations Verification
```bash
# Check what actually happened
cat module_name_log.json
# See recent operations and results
```
### Detailed Debugging
```bash
# Prax provides file-based logging, not real-time watching
# Check system logs directory for detailed output
ls -la system_logs/
# Read specific module logs (Prax manages this directory location)
cat system_logs/module_name.log
```
**This infrastructure = verification layer without formal tests**
---
## Testing Workflow Example
**Building a new branch creation module:**
1. **Plan** - Define structure, features, workflow (Flow plan)
2. **Build to 90%** - AI implements:
- Module structure
- Handler functions
- Config/data/log JSONs
- Internal verification: "Does `create_branch test_branch` work?"
3. **Review at 90%** - Human tests:
- Create branch with various names
- Check if files copied correctly
- Verify registry updated
- Try edge cases (existing branch, invalid name)
4. **Find bug** - Branch created but registry not updated
- **First:** Check error handling - is error logged? Is return value honest?
- Add error handling if missing
- **Then:** Fix bug - why isn't registry updating?
- Verify with clean pass
5. **Iterate** - Test more features:
- Template copying
- Placeholder replacement
- Memory file handling
- Backup on conflicts
6. **Acceptable** - All major features work, errors are honest, ready to use
---
## Current Testing Checklist
**For any new module/feature:**
- [ ] Does it turn on without errors?
- [ ] Do basic commands work?
- [ ] Are errors handled and logged?
- [ ] Do outputs tell the truth?
- [ ] Check config.json - settings correct?
- [ ] Check data.json - state tracking working?
- [ ] Check log.json - operations recorded?
- [ ] Test edge cases (invalid input, missing files, etc.)
- [ ] Check Prax logs in system_logs/ for detailed debugging info
- [ ] Manual feature tests at 90% stage
---
## When to Test What
**During development (AI internal verification):**
- Module starts without errors
- Basic commands execute
- Expected output appears
**At 90% stage (human testing):**
- Feature functionality (does it do what it's supposed to?)
- Edge cases (what breaks it?)
- Error handling (are errors honest?)
- Integration (does it work with other modules?)
**Before considering "done":**
- Clean passes on major features
- Errors tell the truth (no silent failures)
- JSON logs show operations correctly
- Acceptable standard reached (basic or advanced)
---
## Summary
**Current approach:** Manual testing with JSON/log infrastructure
**Why it works:**
- Fast iteration without test maintenance
- Linux environment = reliable verification
- Error-first debugging = effective bug fixing
- JSON/log system = verification layer
**Build process:** Plan → Build to 90% → Review/test → Debug (errors first, then bugs) → Iterate → Acceptable
**Future:** pytest framework when system stabilizes
**Philosophy:** Build fast, verify as you go, handle errors honestly, iterate rapidly. Test infrastructure comes after stability.
---
## Pytest Test Structure
**Current implementation:**
```
pytest.ini # Test configuration (repo root)
tests/ # Root test directory
conftest.py # Shared fixtures
src/aipass/
api/tests/ # API tests (4 files operational)
test_api_system.py
test_openrouter_key.py
test_free_models_quick.py
test_paid_model.py
prax/tests/ # Prax tests
test_log_rotation.py
cli/tests/ # CLI tests (infrastructure ready)
seedgo/
tests/conftest.py
apps/modules/test_cli_errors.py # Demo module
```
**Running tests:**
```bash
# Run all tests
pytest
# Run specific branch tests
pytest src/aipass/api/tests/
# Run with markers
pytest -m unit
pytest -m integration
pytest -m slow
```
**Test demonstration module:**
- `src/aipass/seedgo/apps/modules/test_cli_errors.py` - Shows error handling patterns
- Not a pytest test, but demonstrates testing concepts
- Run directly: `python3 src/aipass/seedgo/apps/modules/test_cli_errors.py`
---
## Comments
#@comments:2025-11-13:claude: Pytest infrastructure exists and operational in API/Prax branches. Selective testing approach: automate stable components, manual testing for rapid development.
#@comments:2025-11-13:claude: Prax doesn't have real-time "watcher" command for logs - uses file-based logging in system_logs/ (Prax manages the location). Updated documentation to reflect actual capabilities.
#@comments:2025-11-13:claude: Error-first debugging approach is critical - maybe this should be emphasized in error_handling.md when we fill that section?
#@comments:2025-11-13:claude: "90% threshold" is interesting pattern - not 100% perfection, but "acceptable standard" (basic or advanced). Reflects pragmatic development philosophy.