feat(system): README update, API init fix, diagnostic docs, STATUS sync
- README.md: updated to current state (15 branches, 95+ PRs, 733 tests, 173 commands) - API: init path fixed (creates .env at ~/.secrets/aipass/ not src/aipass/api/) - API: init shows path and detects existing .env - devpulse/docs: system_health_report, testing_standards, ai_mail comms upgrade - STATUS.md: synced via drone @prax status sync Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.6
parent
a800e7919a
commit
ec63b8ba29
@@ -0,0 +1,22 @@
|
||||
# DPLAN-003 Working Directory
|
||||
|
||||
Research, mapping, and planning files for "AIPass as Operating System."
|
||||
|
||||
Parent plan: `AIPass/DPLAN-003_aipass_as_operating_system_2026-03-13.md`
|
||||
|
||||
## Scope
|
||||
|
||||
**In scope:** `src/aipass/` branches only (drone, seedgo, prax, cli, flow, ai_mail, api, trigger, spawn, devpulse, backup, daemon, memory).
|
||||
|
||||
**Out of scope:** `src/commons/`, `src/skills/` — these are separate and not part of this refactor.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Purpose | Status |
|
||||
|------|---------|--------|
|
||||
| `registry_discovery_map.md` | Every find_registry() call in src/aipass/ | Done |
|
||||
| `portability_audit.md` | Full investigation results (session 24) | Done |
|
||||
| `credential_model.md` | Registry credential design (UUID now, macaroons later) | Stage 1 Done |
|
||||
| `registry_refactor_plan.md` | Shared find_registry() design, migration steps | Pending |
|
||||
| `aipass_init_spec.md` | CLI owns init — `drone @cli aipass init` (entry point later) | Pending |
|
||||
| `aipass_help_spec.md` | CLI owns help — `drone @cli aipass help` (entry point later) | Pending |
|
||||
@@ -0,0 +1,217 @@
|
||||
# AI_MAIL COMMS UPGRADE — Hardening Plan
|
||||
|
||||
**Goal:** Make ai_mail bulletproof, then use as the model for other branches.
|
||||
**Started:** 2026-03-10 | **Status:** In Progress
|
||||
**Tracking:** Updated each session. Read this first on context resume.
|
||||
|
||||
---
|
||||
|
||||
## Current State (Session 10)
|
||||
|
||||
**Seedgo audit: 100%** — all 23 standards pass (42 files)
|
||||
**Automated tests: 36** — test_send_identity.py v1.2.0 (3 audit rounds, 15 agents, 0 false positives)
|
||||
**Production fixes deployed:** 3 cross-platform crashers fixed, test suite fully isolated
|
||||
**Phase 2 (Error Handling):** 27 silent `except` blocks across 13 files now log with `logger.warning()`
|
||||
|
||||
### Test Suite Evolution
|
||||
- v1.0.0: 31 tests — 7 false positives found by 3 audit agents
|
||||
- v1.1.0: 32 tests — 8 suspects found by 5 audit agents (live registry, weak contracts)
|
||||
- v1.2.0: 36 tests — final audit found 2 minor issues, fixed. **0 live-data dependencies.**
|
||||
|
||||
### Production Fixes (Session 9)
|
||||
- `inbox_lock.py` — `import fcntl` guarded for Windows (msvcrt fallback)
|
||||
- `ai_mail.py` — `signal.SIGPIPE` guarded with `hasattr` check
|
||||
- `notify.py` — `/usr/bin/python3` replaced with `shutil.which("python3")`
|
||||
|
||||
---
|
||||
|
||||
## Architecture Map
|
||||
|
||||
```
|
||||
Entry: apps/ai_mail.py
|
||||
+-- Module: apps/modules/email.py (orchestrator, v3.0.0)
|
||||
|-- handle_send() -> send_args.py (parse) -> send.py (execute) -> delivery.py (inbox write)
|
||||
|-- handle_inbox() -> inbox_resolve.py -> inbox_ops.py -> format.py
|
||||
|-- handle_view() -> inbox_ops.py
|
||||
|-- handle_reply() -> reply.py -> delivery.py
|
||||
|-- handle_close() -> close_ops.py
|
||||
|-- handle_sent() -> format.py
|
||||
+-- handle_contacts() -> registry/
|
||||
|
||||
Module: apps/modules/dispatch.py (dispatch orchestrator)
|
||||
|-- daemon.py (auto-dispatch loop)
|
||||
|-- wake.py (manual branch wake)
|
||||
|-- dispatch_monitor.py (agent lifecycle)
|
||||
|-- status.py (dispatch status display)
|
||||
+-- pending_work.py (pending dispatch queue)
|
||||
|
||||
Identity: apps/handlers/users/
|
||||
|-- branch_detection.py (detect_branch_from_pwd — THE critical path)
|
||||
|-- user.py (get_current_user, get_branch_by_email)
|
||||
+-- load.py, config_generator.py
|
||||
|
||||
Cross-branch: drone/apps/handlers/router_handler.py
|
||||
+-- detect_caller_branch_name() -> sets AIPASS_CALLER_BRANCH env var
|
||||
```
|
||||
|
||||
**Identity detection chain (9 stages):**
|
||||
1. drone CLI entry
|
||||
2. router resolves @branch to path
|
||||
3. router_handler detects caller via CWD (+ AIPASS_BRANCH_NAME fallback)
|
||||
4. executor merges caller_env into subprocess
|
||||
5. ai_mail subprocess starts with env vars
|
||||
6. branch_detection reads AIPASS_CALLER_BRANCH
|
||||
7. send_args builds headers
|
||||
8. delivery writes to recipient inbox.json
|
||||
9. notify.py fires desktop notification
|
||||
|
||||
---
|
||||
|
||||
## Systemic Issues Found (Session 9 Deep Audit)
|
||||
|
||||
15 agents across 3 rounds audited tests + full codebase. Key findings:
|
||||
|
||||
### Silent Failures (24 instances) — VIOLATES "fail to errors" rule
|
||||
`except Exception: return None` in 10+ places with zero logging:
|
||||
- `branch_detection.py` — 3 functions (entire identity chain)
|
||||
- `delivery.py` — get_all_branches(), summary, notification
|
||||
- `create.py` — load_email_file()
|
||||
- `user.py` — get_user_by_email(), get_all_users()
|
||||
- `daemon.py` — _read_json() (silent) vs wake.py (logs) — inconsistent
|
||||
- `inbox_ops.py` — migration persist has literal `pass`
|
||||
|
||||
### Dead/Redundant Code (~20% of codebase, ~1,728 lines)
|
||||
- 5 fully unused files: validate.py, errors.py, data_ops.py, config_generator.py, pending_work.py
|
||||
- `_find_repo_root()` copy-pasted in 9 files
|
||||
- `get_all_branches()` implemented twice with different email behavior (correctness bug)
|
||||
- `lock_utils.py` exists but unused — wake.py and daemon.py reimplement locking
|
||||
- wake.py/daemon.py share 6+ duplicated functions
|
||||
|
||||
### Cross-Platform Breakers
|
||||
- [FIXED] `import fcntl` unconditional — crashes Windows
|
||||
- [FIXED] `/usr/bin/python3` hardcoded — breaks macOS/Windows
|
||||
- [FIXED] `signal.SIGPIPE` — crashes Windows startup
|
||||
- [OPEN] `pgrep` ungated in wake.py/daemon.py
|
||||
- [OPEN] `email.py` parents[2] fragile — should use _find_repo_root()
|
||||
|
||||
---
|
||||
|
||||
## Hardening Phases
|
||||
|
||||
### Phase 1: Test Suite (Critical Path)
|
||||
**Priority: HIGH** | **Status: IN PROGRESS**
|
||||
|
||||
- [x] `test_send_identity.py` — 36 tests, 3 audit rounds, 0 false positives
|
||||
- [ ] `test_delivery.py` — round-trip send/receive
|
||||
- [ ] `test_send_args.py` — argument parsing (all flags, interactive, error cases)
|
||||
- [ ] `test_inbox_ops.py` — inbox operations (view, close, reply, close all)
|
||||
- [ ] `test_dispatch_monitor.py` — dispatch lifecycle
|
||||
- [ ] `test_notify.py` — notification delivery
|
||||
|
||||
### Phase 2: Error Handling Overhaul
|
||||
**Priority: HIGH** | **Status: IN PROGRESS**
|
||||
Silent failures are why the system feels "fragile."
|
||||
|
||||
- [x] Add `logger.warning()` to every bare `except Exception: return None` — **13 files, 27 instances fixed**
|
||||
- [ ] Distinguish "not found" from "error reading" in return types
|
||||
- [ ] Fix inconsistent error conventions (None vs tuple vs dict vs raise)
|
||||
- [ ] Fix collision detection dead code in delivery.py get_all_branches()
|
||||
|
||||
### Phase 3: Code Consolidation
|
||||
**Priority: MEDIUM** — Reduce duplication, single source of truth.
|
||||
|
||||
- [ ] Extract shared `_find_repo_root()` into commons or shared utility
|
||||
- [ ] Consolidate `get_all_branches()` — one implementation, prefers explicit email
|
||||
- [ ] Consolidate lock acquisition — use lock_utils.py, delete reimplementations
|
||||
- [ ] Deduplicate wake.py/daemon.py shared functions (_read_json, _set_session_name, etc.)
|
||||
- [ ] Archive 5 dead files (validate.py, errors.py, data_ops.py, config_generator.py, pending_work.py)
|
||||
|
||||
### Phase 4: Identity Consolidation
|
||||
**Priority: MEDIUM** — Reduce 12 detection mechanisms to 1 canonical resolver.
|
||||
|
||||
- [ ] Define canonical `resolve_branch_identity()` function
|
||||
- [ ] Priority chain: AIPASS_CALLER_BRANCH > AIPASS_BRANCH_NAME > CWD passport walk > --from
|
||||
- [ ] Single file, single function, single source of truth
|
||||
- [ ] All callers delegate to it
|
||||
|
||||
### Phase 5: Cross-Platform Hardening
|
||||
**Priority: MEDIUM** — Public repo must work on all platforms.
|
||||
|
||||
- [ ] Guard `pgrep` usage in wake.py/daemon.py
|
||||
- [ ] Replace `email.py` parents[2] with _find_repo_root()
|
||||
- [ ] Guard `start_new_session=True` for Windows
|
||||
- [ ] Add encoding='utf-8' to os.fdopen in wake.py
|
||||
|
||||
### Phase 6: Standards for Communication
|
||||
**Priority: LOW** — Make patterns enforceable system-wide.
|
||||
|
||||
- [ ] `communication` standard — envelope format, required fields, identity rules
|
||||
- [ ] `identity` standard — env var hierarchy, passport requirements
|
||||
- [ ] `dispatch` standard — env isolation, lock management, bounce handling
|
||||
|
||||
### Phase 7: Documentation & System Prompt
|
||||
**Priority: LOW** — Update branch local prompt.
|
||||
|
||||
- [ ] Write proper `.aipass/aipass_local_prompt.md`
|
||||
- [ ] Document key commands, architecture, critical files
|
||||
|
||||
---
|
||||
|
||||
## Decision Log
|
||||
|
||||
| Date | Decision | Rationale |
|
||||
|------|----------|-----------|
|
||||
| 2026-03-08 | AIPASS_BRANCH_NAME env var over CWD detection | CWD unreliable when agents navigate. Env vars persist. |
|
||||
| 2026-03-08 | dbus direct over notify-send | Portal mode strips persistence hints. dbus bypasses confinement. |
|
||||
| 2026-03-08 | Unique app_name per source | GNOME collapses same desktop-entry into one notification. |
|
||||
| 2026-03-08 | Fail loud, no CWD fallback for identity | Silent wrong sender worse than visible error. |
|
||||
| 2026-03-10 | --from flag for explicit sender override | Plumbing existed, just needed CLI wiring. |
|
||||
| 2026-03-10 | Tests first in hardening plan | Every bug from sessions 4-7 would have been caught by tests. |
|
||||
| 2026-03-10 | 3 audit rounds on test suite | Each round found issues the previous missed. Diminishing returns by round 3. |
|
||||
| 2026-03-10 | Error handling overhaul before code consolidation | Silent failures cause more user pain than code duplication. |
|
||||
|
||||
---
|
||||
|
||||
## Session Notes
|
||||
|
||||
### Session 10 (2026-03-10)
|
||||
- **Phase 2: Error Handling Overhaul — logger.warning() sweep complete**
|
||||
- 13 files modified, 27 silent `except Exception` blocks now log with `logger.warning()`
|
||||
- Files fixed (by priority):
|
||||
- `branch_detection.py` (3) — identity chain, most critical
|
||||
- `user.py` (2) — user lookup
|
||||
- `delivery.py` (6) — get_all_branches, migrate, callback, summary, notification, private branch
|
||||
- `create.py` (2) — purge, load_email_file
|
||||
- `inbox_ops.py` (1) — migration persist
|
||||
- `format.py` (1) — alias lookup
|
||||
- `reply.py` (1) — get_email_by_id
|
||||
- `close_ops.py` (3) — dashboard, central, purge post-ops
|
||||
- `send.py` (2) — central update after send/broadcast
|
||||
- `dashboard_sync.py` (1) — push_dashboard_update
|
||||
- `inbox_cleanup.py` (4) — migrate, dashboard, central, purge
|
||||
- `error_handler.py` (1) — error notification delivery
|
||||
- `error_dispatch.py` (2) — dashboard, central post-ops
|
||||
- All 13 files pass `py_compile` syntax check
|
||||
- Remaining silent blocks (acceptable): inbox_lock.py finally-block, json_handler.py (utility),
|
||||
daemon/wake/dispatch_monitor (already log with logger.info), config_generator/data_ops (unused files)
|
||||
- Next: Phase 2 remaining items (return type distinctions, error conventions, collision dead code)
|
||||
|
||||
### Session 9 (2026-03-10, continued)
|
||||
- Wrote test_send_identity.py v1.0.0 (31 tests)
|
||||
- Audit round 1: 3 agents found 7 false positives -> rewrote to v1.1.0
|
||||
- Audit round 2: 5 agents found 8 issues (live registry, weak contracts) -> v1.2.0 (36 tests)
|
||||
- Audit round 3: 5 agents (final test + full system sweep)
|
||||
- Tests: 2 minor issues found and fixed (unpatched registry, missing mailbox_path assert)
|
||||
- Error handling: 24 silent failure instances across 10+ files
|
||||
- Code quality: ~1,728 dead/redundant lines (20% of codebase), 5 unused files
|
||||
- Cross-platform: 3 CRITICAL (fcntl, SIGPIPE, /usr/bin/python3) — ALL FIXED
|
||||
- Paths: parents[N] usage audited, mostly correct, 1 fragile case in email.py
|
||||
- Fixed 3 cross-platform crashers in production code
|
||||
- Updated COMMS_UPGRADE.md with full findings and revised phases
|
||||
- Next: Phase 1 continues (more test files) or Phase 2 (error handling overhaul)
|
||||
|
||||
### Session 8 (2026-03-10)
|
||||
- Patrick initiated hardening project
|
||||
- Self-audit: 100% on all 22 seedgo standards
|
||||
- Mapped full architecture (55 Python files, 3 modules, 8 handler domains)
|
||||
- Created this tracking document
|
||||
@@ -0,0 +1,157 @@
|
||||
# DPLAN-003: Registry Credential Model
|
||||
|
||||
## The Idea
|
||||
|
||||
Registries get a unique token. Passports carry that token. Access is identity-based, not filesystem-based. No walk-up needed — your credential proves which registry is yours.
|
||||
|
||||
## Why
|
||||
|
||||
Current system finds registries by walking up directories. Works for one project, breaks with multiple. If two AIPass projects exist on one machine, a citizen launched from the wrong directory finds the wrong registry. Credentials solve this — your passport carries proof of membership.
|
||||
|
||||
## Prior Art (Research)
|
||||
|
||||
| System | Pattern | Fit |
|
||||
|--------|---------|-----|
|
||||
| **Macaroons** (Google Research) | Token IS the credential. Delegatable with caveats. Offline verification. DeepMind validated for AI agent delegation (2026). | Highest |
|
||||
| **Vault Namespaces** | Project = namespace. Token scoped to namespace. Mini-registry per project. | High |
|
||||
| **AWS STS / Token Vending** | Agent presents project ID, gets scoped credential. Credential itself is the boundary. | High |
|
||||
| **K8s Namespace + ServiceAccount** | Token carries project scope as claim. RBAC composable. | High |
|
||||
| **SPIFFE/SPIRE** | Process-level attestation without static secrets. | Medium |
|
||||
| **direnv** | Auto-set env vars on directory entry. Zero-friction UX. | UX pattern |
|
||||
|
||||
Full research: agent output from session 25.
|
||||
|
||||
## Design: Two Stages
|
||||
|
||||
### Stage 1: UUID Match (Manual, Now)
|
||||
|
||||
Simple. Prove the concept works before adding crypto.
|
||||
|
||||
**Registry gets an ID:**
|
||||
```json
|
||||
{
|
||||
"metadata": {
|
||||
"id": "a1b2c3d4-...",
|
||||
"name": "AIPASS",
|
||||
"version": "1.0.0",
|
||||
"last_updated": "2026-03-13",
|
||||
"total_branches": 15
|
||||
},
|
||||
"branches": [...]
|
||||
}
|
||||
```
|
||||
|
||||
**Passports get the matching ID:**
|
||||
```json
|
||||
{
|
||||
"citizenship": {
|
||||
"registered": true,
|
||||
"registry_id": "a1b2c3d4-...",
|
||||
"registry_name": "AIPASS",
|
||||
"citizen_number": 7
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Lookup flow:**
|
||||
1. Walk up from CWD, find `*_REGISTRY.json`
|
||||
2. Read its `metadata.id`
|
||||
3. Check citizen's `citizenship.registry_id` matches
|
||||
4. If mismatch → error ("citizen belongs to registry X, found registry Y")
|
||||
5. If match → proceed
|
||||
|
||||
**Spawn changes:**
|
||||
- `aipass init` (or manual setup) generates the UUID for the registry
|
||||
- Spawn reads registry UUID and injects into new passports via `{{REGISTRY_ID}}` placeholder
|
||||
- Existing 15 branches get the UUID added to their passports (one-time migration)
|
||||
|
||||
### Stage 2: Macaroon Tokens (Future, When Cross-Project Needed)
|
||||
|
||||
Upgrade path when we need delegation and cross-project access.
|
||||
|
||||
**Root token:** Created at `aipass init`. HMAC-based. Stored at `~/.secrets/aipass/projects/<uuid>.token`
|
||||
|
||||
**Citizen token:** Attenuated copy in passport. Can prove membership but can't mint new citizens.
|
||||
|
||||
**Agent token:** Further attenuated. Carries:
|
||||
- Registry ID (which project)
|
||||
- Scope (full access — agents do the real work in AIPass)
|
||||
- Expiry (session-scoped, dies when agent dies)
|
||||
- Issuer (which citizen spawned this agent)
|
||||
|
||||
Agents are NOT read-only. They're the builders — they write code, run tests, modify files. The token proves they belong to a project, it doesn't restrict what they do within it. Scope restrictions would be role-based (e.g., "can't modify other branches' files") not capability-based.
|
||||
|
||||
**Verification:** Local HMAC check. No daemon needed for basic validation. Daemon is optional enhancement for audit logging and revocation.
|
||||
|
||||
**direnv integration:** Entering a project directory auto-sets `AIPASS_PROJECT_TOKEN` in env. Agents inherit it.
|
||||
|
||||
## What Changes (Stage 1)
|
||||
|
||||
| Component | Change |
|
||||
|-----------|--------|
|
||||
| `AIPASS_REGISTRY.json` | Add `metadata.id` (UUID) |
|
||||
| Passport template | Add `citizenship.registry_id` placeholder |
|
||||
| Spawn `build_replacements_dict()` | Read registry UUID, add `{{REGISTRY_ID}}` |
|
||||
| Spawn `add_to_registry()` | No change (branch entries stay the same) |
|
||||
| Drone `find_registry()` | Optional: verify passport.registry_id matches found registry |
|
||||
| All 15 passports | One-time: add `registry_id` field |
|
||||
|
||||
## What Does NOT Change
|
||||
|
||||
- Registry filename stays `*_REGISTRY.json`
|
||||
- Walk-up discovery still works (credential is verification layer on top, not replacement)
|
||||
- Branch structure unchanged
|
||||
- No daemon needed
|
||||
- No new dependencies
|
||||
|
||||
## Resolved Questions
|
||||
|
||||
1. **Registry ID location** → `metadata.id` — it's the project's identity, not the citizen's.
|
||||
2. **UUID4 vs hash** → UUID4 (random). Simple, guaranteed unique, no inputs needed.
|
||||
3. **Verification mode** → Hard error on mismatch. Fail loudly — that's the AIPass way.
|
||||
4. **Where does init live?** → Temporary: `src/aipass/devpulse/apps/init_project.py`. Future: CLI branch.
|
||||
|
||||
## Bugs, Quirks, and Findings
|
||||
|
||||
Discovered during testing. Reference for future work.
|
||||
|
||||
### init_project.py (10 edge cases tested)
|
||||
- **Spaces in dir name** → registry filename gets spaces (`MY COOL PROJECT_REGISTRY.json`). Fixed: `_sanitize_name()` replaces non-alphanumeric with `_`.
|
||||
- **Root path `/`** → `Path("/").name` is empty string, creates `_REGISTRY.json`. Fixed: validation rejects empty name.
|
||||
- **Permission errors** → raw traceback instead of clean message. Fixed: `main()` catches `OSError`.
|
||||
- **Passport overwrite on re-init** → if registry deleted but `.trinity/` survives, re-init would silently overwrite passport with new UUID. Fixed: passport guarded with `exists()` check.
|
||||
- **Double init** → correctly blocked by `FileExistsError` on registry file.
|
||||
- **Deep nested paths** → `mkdir(parents=True)` handles correctly.
|
||||
- **UUID uniqueness** → 5 runs, 5 unique UUIDs. No collisions.
|
||||
- **JSON validity** → all generated files parse clean.
|
||||
|
||||
### Drone isolation (6 scenarios tested)
|
||||
- **Sibling projects** → PASS. Two projects in same parent dir, fully isolated.
|
||||
- **Nested project** → PASS. Inner registry wins over outer. No bleed-through.
|
||||
- **Deep subdir walk-up** → PASS. Finds nearest ancestor registry correctly.
|
||||
- **Cross-project contamination** → PASS. Branches are registry-scoped; modules are global.
|
||||
- **Empty directory (no registry)** → CONCERN. Drone silently falls back to AIPass source registry via `__file__` walk-up. Any dir on this machine without a registry sees production branches. This is by design in `find_registry()` but will be dangerous with multi-project. Credential verification would catch this — citizen's registry_id won't match the fallback registry.
|
||||
- **`drone @ai_mail` from mock project** → correctly returns "Branch not found in registry" (empty project has no branches).
|
||||
|
||||
### Registry/passport gitignore
|
||||
- `AIPASS_REGISTRY.json` and all `.trinity/passport.json` files are gitignored. UUID migration is local-only. This is correct for now — credentials are machine-specific, not repo state. But means `aipass init` must run on every clone/install. Future: consider whether UUID should be in-repo or machine-local.
|
||||
|
||||
### AI mail after migration
|
||||
- Send/receive works fine after UUID migration. AI mail's 4 internal `find_registry()` copies still hardcode `AIPASS_REGISTRY.json` — functional for now since that filename exists, but won't find `*_REGISTRY.json` in other projects.
|
||||
|
||||
### Drone built-in modules vs branches
|
||||
- `@drone` and `@seedgo` are hardcoded as "modules" in `module_registry.py`, always visible everywhere. Other branches (`@ai_mail`, `@spawn`, etc.) are registry-scoped. This distinction matters: modules are global services, branches are project citizens.
|
||||
|
||||
## Decision Log
|
||||
|
||||
- **2026-03-14:** Patrick proposed credential-based registry access. Agents do the real work — tokens prove membership, not restrict capability. Manual first, `aipass init` later.
|
||||
- **2026-03-14:** Research confirmed macaroons as best-fit pattern (Google Research + DeepMind 2026 validation). Stage 1 = UUID match, Stage 2 = macaroon upgrade.
|
||||
- **2026-03-14:** Scope limited to `src/aipass/` branches only. Commons and skills excluded.
|
||||
- **2026-03-14:** All 4 open questions resolved. `init_project.py` built and tested — creates registry with UUID, passport with matching registry_id, .trinity/, .aipass/, AIPASS.md. Tested with temp directory — drone isolation confirmed (only built-in modules visible, not AIPass branches).
|
||||
- **2026-03-14:** UUID migration executed (FPLAN-0030). Registry + 13 passports updated. Registry and passports are gitignored — UUID is machine-local, not repo state.
|
||||
- **2026-03-14:** Drone verification added (`_verify_registry_credential` in load_registry). Drone dispatched and fixed error propagation — new `RegistryMismatchError` class separates "not found" (fallback OK) from "mismatch" (hard error). Spawn templates updated with `{{REGISTRY_ID}}` placeholder.
|
||||
- **2026-03-14:** Stage 1 complete. Credential model is end-to-end: init creates credentialed registries, drone verifies on load, spawn injects into new passports, mismatch = hard error. Remaining: drone CLI stderr surfacing (DPLAN-0031), `aipass init` CLI command (future).
|
||||
- **2026-03-14:** FPLAN-0032 executed — CLI stderr standardization Phase 1+2 done. CLI owns err_console, error()/warning() go to stderr. Drone imports from CLI instead of creating its own Console(stderr=True). Credential mismatch errors now properly surface to terminal via stderr. The stderr issue that blocked error visibility during DPLAN-003 testing is resolved.
|
||||
- **2026-03-14:** CLI confirmed as owner of credential model + project commands. CLI owns: `aipass init` (create project), `aipass help` (project help), display API (error/warning/fatal/err_console). Drone owns verification (registry_handler). Spawn owns injection (passport templates). Access via `drone @cli aipass init/help` for now; `aipass init` entry point is future icing.
|
||||
- **2026-03-14:** Seedgo `stderr_routing` standard created (24th standard). Checker bug found — was scanning all files but not displaying module/handler results. Seedgo fixed display + aggregation. Agent-verified across 3 branches. FPLAN-0033 parked for Phase 3 migration (343 error prints, 10 branches).
|
||||
- **2026-03-14:** FPLAN-0033 Phase 3 executed autonomously (Patrick away). 15+ agents across 3 waves migrated 10 branches. 48 files, 666 insertions. System stderr avg 49%→69%. Post-migration verification found 63% of remaining seedgo violations are false positives ([yellow]Label:[/yellow] section headers flagged as warnings). Seedgo dispatched for checker refinement. 93 real violations remain (48 in commons which was out of scope).
|
||||
@@ -0,0 +1,31 @@
|
||||
# Portability Audit — Session 24 Results
|
||||
|
||||
## Summary
|
||||
|
||||
| Tool | Registry Discovery | CWD-Aware | Portable | Hardcoded |
|
||||
|------|-------------------|-----------|----------|-----------|
|
||||
| Drone | Walk-up + env var | No (uses registry) | Yes | Registry filename |
|
||||
| Spawn | Walk-up + env var | No (uses registry) | Partial | Template location |
|
||||
| Prax | Walk-up (no env) | No (sys logs at repo) | Partial | System logs dir |
|
||||
| AI_Mail | Walk-up (no env) | No (inbox per branch) | Yes | Inbox location |
|
||||
| Flow | Walk-up (no env) | Yes (plan creation) | Hybrid | Plan registry |
|
||||
|
||||
## Key Findings
|
||||
|
||||
- All tools use walk-up strategy to find `AIPASS_REGISTRY.json`
|
||||
- Registry-relative path resolution already works (move registry + dirs = works)
|
||||
- `AIPASS_REGISTRY` env var supported by drone and spawn
|
||||
- System logs hardcoded to `{repo_root}/system_logs/`
|
||||
- Spawn templates hardcoded to `{spawn_package}/templates/`
|
||||
- Walk-up doesn't stop at project boundaries — finds nearest registry up the tree
|
||||
|
||||
## The Core Fix
|
||||
|
||||
Change `find_registry()` to:
|
||||
1. Walk up from CWD looking for `*_REGISTRY.json` (glob, not hardcoded name)
|
||||
2. Stop at first match — that's the project boundary
|
||||
3. If none found, return error ("No AIPass project. Run `aipass init`")
|
||||
|
||||
## Source
|
||||
|
||||
Full investigation transcript: background agent session 24, 42 tool calls across drone/spawn/ai_mail/flow/prax.
|
||||
@@ -0,0 +1,46 @@
|
||||
# Registry Discovery Map — FPLAN-0029 Phase 1
|
||||
|
||||
## Primary Implementations (5)
|
||||
|
||||
| File | Function | Strategy |
|
||||
|------|----------|----------|
|
||||
| `src/aipass/spawn/apps/handlers/registry.py:34-79` | `find_registry(start_path)` | Env var → project markers (.git/pyproject.toml) → walk-up → cwd fallback |
|
||||
| `src/aipass/drone/apps/handlers/registry_handler.py:36-61` | `find_registry()` | Walk-up from __file__ → walk-up from cwd → parents[4] fallback |
|
||||
| `src/aipass/drone/apps/handlers/registry_handler.py:64-81` | `get_registry_path()` | Global override → env var → find_registry() |
|
||||
| `src/commons/apps/handlers/database/db.py:343-378` | `_find_branch_registry()` | Env var AIPASS_ROOT → walk-up (10 limit) → ~/.aipass/ fallback |
|
||||
| `src/commons/apps/handlers/identity/identity_ops.py:33-67` | `_find_branch_registry_path()` | Same as db.py |
|
||||
|
||||
## Local Reimplementations (23+)
|
||||
|
||||
All do walk-up from `__file__` looking for `AIPASS_REGISTRY.json`:
|
||||
|
||||
### ai_mail (4 files)
|
||||
- `apps/handlers/users/branch_detection.py:26-38`
|
||||
- `apps/handlers/registry/read.py:31-41`
|
||||
- `apps/handlers/email/format.py:21-31`
|
||||
- `apps/handlers/dispatch/daemon.py:36-55`
|
||||
|
||||
### memory (3 files)
|
||||
- `apps/handlers/dashboard_push.py:44-54`
|
||||
- `apps/handlers/monitor/detector.py:36-83`
|
||||
- `apps/handlers/monitor/memory_watcher.py:363-385`
|
||||
|
||||
### prax (2 files)
|
||||
- `apps/handlers/registry/reader.py:24-39`
|
||||
- `apps/handlers/dashboard/agent_status_writer.py:33-44`
|
||||
|
||||
### seedgo (3 files)
|
||||
- `apps/handlers/audit/discovery.py:54-62`
|
||||
- `apps/handlers/diagnostics/discovery.py:21-29`
|
||||
- `apps/handlers/readme/readme_ops.py:29-37`
|
||||
|
||||
### + 11 more across other branches
|
||||
|
||||
## Hardcoded String Count
|
||||
100+ references to `"AIPASS_REGISTRY.json"` across codebase.
|
||||
|
||||
## Phase 1 Plan
|
||||
1. Create `commons.registry.find_registry()` — one function, glob for `*_REGISTRY.json`
|
||||
2. All 23+ modules import from commons
|
||||
3. Drone/spawn wrappers call commons internally
|
||||
4. Test: temp dir with TEST_REGISTRY.json → drone systems finds it
|
||||
@@ -0,0 +1,525 @@
|
||||
# AIPass System Health Report
|
||||
Generated: 2026-03-19 22:54
|
||||
|
||||
---
|
||||
|
||||
## 1. Dead Code (14 unused files)
|
||||
|
||||
|
||||
Dead Code Scanner
|
||||
Scanning 15 branches
|
||||
|
||||
@ai_mail (3 modules, 31 handlers)
|
||||
x handlers/monitoring/errors.py -- 0 references
|
||||
> 33/34 files referenced
|
||||
|
||||
@api (4 modules, 14 handlers)
|
||||
x handlers/openrouter/provision.py -- 0 references
|
||||
> 17/18 files referenced
|
||||
|
||||
@backup (4 modules, 24 handlers)
|
||||
x handlers/diff/vscode_integration.py -- 0 references
|
||||
> 27/28 files referenced
|
||||
|
||||
@cli (3 modules, 2 handlers)
|
||||
> All 5 files referenced
|
||||
|
||||
@commons (22 modules, 35 handlers)
|
||||
> All 57 files referenced
|
||||
|
||||
@daemon (6 modules, 8 handlers)
|
||||
> All 14 files referenced
|
||||
|
||||
@devpulse -- no apps/ directory
|
||||
@drone (9 modules, 16 handlers)
|
||||
> All 25 files referenced
|
||||
|
||||
@flow (8 modules, 35 handlers)
|
||||
x handlers/plan/file_ops.py -- 0 references
|
||||
x handlers/plan/update_registry.py -- 0 references
|
||||
> 41/43 files referenced
|
||||
|
||||
@memory (5 modules, 29 handlers)
|
||||
x handlers/learnings/manager.py -- 0 references
|
||||
x handlers/schema/normalize.py -- 0 references
|
||||
x handlers/search/vector_search.py -- 0 references
|
||||
> 31/34 files referenced
|
||||
|
||||
@prax (6 modules, 36 handlers)
|
||||
> All 42 files referenced
|
||||
|
||||
@seedgo (5 modules, 66 handlers)
|
||||
x handlers/config/aipass_bypass.py -- 0 references
|
||||
x handlers/config/aipass_ignore.py -- 0 references
|
||||
x handlers/diagnostics/python_diognostics.py -- 0 references
|
||||
x handlers/diagnostics/typscript_diognostics.py -- 0 references
|
||||
x handlers/file/file_handler.py -- 0 references
|
||||
x handlers/mock_standard_1/bypass_config/bypass.config.py -- 0 references
|
||||
> 65/71 files referenced
|
||||
|
||||
@skills (5 modules, 8 handlers)
|
||||
> All 13 files referenced
|
||||
|
||||
@spawn (6 modules, 15 handlers)
|
||||
> All 21 files referenced
|
||||
|
||||
@trigger (5 modules, 17 handlers)
|
||||
> All 22 files referenced
|
||||
|
||||
TOTAL: 14 unused across 14 branches
|
||||
|
||||
---
|
||||
|
||||
## 2. Local Prompts (5 stubs need enrichment)
|
||||
|
||||
|
||||
Local Prompt Status
|
||||
==================================================
|
||||
|
||||
RICH (50+ lines, 4+ sections):
|
||||
v daemon 56 lines 4 sections
|
||||
v devpulse 79 lines 8 sections
|
||||
v flow 67 lines 6 sections
|
||||
v seedgo 56 lines 4 sections
|
||||
v skills 58 lines 5 sections
|
||||
v trigger 53 lines 7 sections
|
||||
|
||||
BASIC (15-49 lines):
|
||||
~ api 15 lines 2 sections
|
||||
~ commons 48 lines 6 sections
|
||||
~ memory 37 lines 5 sections
|
||||
~ spawn 20 lines 4 sections
|
||||
|
||||
STUB (<15 lines):
|
||||
x ai_mail 14 lines 1 section
|
||||
x backup 14 lines 1 section
|
||||
x cli 14 lines 1 section
|
||||
x drone 14 lines 1 section
|
||||
x prax 14 lines 1 section
|
||||
|
||||
==================================================
|
||||
SUMMARY: 6 rich, 4 basic, 5 stub
|
||||
|
||||
==================================================
|
||||
Section Breakdown
|
||||
==================================================
|
||||
|
||||
@ai_mail (14 lines, STUB)
|
||||
v Status
|
||||
470 bytes
|
||||
|
||||
@api (15 lines, BASIC)
|
||||
v Identity
|
||||
v Key Breadcrumbs
|
||||
1192 bytes
|
||||
|
||||
@backup (14 lines, STUB)
|
||||
v Status
|
||||
499 bytes
|
||||
|
||||
@cli (14 lines, STUB)
|
||||
v Status
|
||||
499 bytes
|
||||
|
||||
@commons (48 lines, BASIC)
|
||||
v Commands
|
||||
v Architecture
|
||||
v Integration Points
|
||||
v Critical Files
|
||||
v Role
|
||||
v Key Details
|
||||
2328 bytes
|
||||
|
||||
@daemon (56 lines, RICH)
|
||||
v Commands
|
||||
v Memory & Tracking
|
||||
v Apps Layout
|
||||
v Known Issues
|
||||
2788 bytes
|
||||
|
||||
@devpulse (79 lines, RICH)
|
||||
v Identity
|
||||
v How You Work
|
||||
v Dispatch Table
|
||||
v Commands
|
||||
v Branches
|
||||
v Working Habits
|
||||
v Autonomous Monitoring
|
||||
v Memory & Tracking
|
||||
v Has dispatch table (@branch refs)
|
||||
4710 bytes
|
||||
|
||||
@drone (14 lines, STUB)
|
||||
v Status
|
||||
503 bytes
|
||||
|
||||
@flow (67 lines, RICH)
|
||||
v Commands
|
||||
v Architecture
|
||||
v Integration Points
|
||||
v Conventions
|
||||
v Critical Files
|
||||
v Plan Type System
|
||||
3461 bytes
|
||||
|
||||
@memory (37 lines, BASIC)
|
||||
v Identity
|
||||
v Commands
|
||||
v Architecture
|
||||
v Memory & Tracking
|
||||
v Known Issues
|
||||
1591 bytes
|
||||
|
||||
@prax (14 lines, STUB)
|
||||
v Status
|
||||
501 bytes
|
||||
|
||||
@seedgo (56 lines, RICH)
|
||||
v Commands
|
||||
v Apps Layout (extra layer vs standard branch)
|
||||
v How I Work — Standards Reasoning
|
||||
v Quick Reference
|
||||
3166 bytes
|
||||
|
||||
@skills (58 lines, RICH)
|
||||
v Commands
|
||||
v Memory & Tracking
|
||||
v Apps Layout
|
||||
v Search Paths (first match wins)
|
||||
v Three Skill Tiers
|
||||
2706 bytes
|
||||
|
||||
@spawn (20 lines, BASIC)
|
||||
v Commands
|
||||
v Architecture
|
||||
v Role
|
||||
v Principles
|
||||
745 bytes
|
||||
|
||||
@trigger (53 lines, RICH)
|
||||
v Dispatch Table
|
||||
v Commands
|
||||
v Architecture
|
||||
v Integration Points
|
||||
v Critical Files
|
||||
v Role
|
||||
v Rules
|
||||
2610 bytes
|
||||
|
||||
|
||||
---
|
||||
|
||||
## 3. Commands (173 discovered)
|
||||
|
||||
|
||||
@ai_mail (14 commands)
|
||||
- close
|
||||
- contacts
|
||||
- dispatch
|
||||
- email
|
||||
- inbox
|
||||
- ping
|
||||
- read
|
||||
- registry
|
||||
- reply
|
||||
- send
|
||||
- sent
|
||||
- status
|
||||
- thresholds
|
||||
- view
|
||||
|
||||
@api (12 commands)
|
||||
- call
|
||||
- cleanup
|
||||
- google
|
||||
- init
|
||||
- models
|
||||
- reauth
|
||||
- session
|
||||
- stats
|
||||
- status
|
||||
- test
|
||||
- track
|
||||
- validate
|
||||
|
||||
@backup (1 command)
|
||||
- reauth
|
||||
|
||||
@cli (5 commands)
|
||||
- aipass
|
||||
- demo
|
||||
- display
|
||||
- show
|
||||
- templates
|
||||
|
||||
@commons (51 commands)
|
||||
- activity
|
||||
- artifacts
|
||||
- capsule
|
||||
- capsules
|
||||
- catchup
|
||||
- collab
|
||||
- comment
|
||||
- craft
|
||||
- database
|
||||
- decorate
|
||||
- delete
|
||||
- digest
|
||||
- drop
|
||||
- enter
|
||||
- event
|
||||
- explore
|
||||
- feed
|
||||
- find
|
||||
- gift
|
||||
- inspect
|
||||
- leaderboard
|
||||
- leaderboards
|
||||
- log
|
||||
- look
|
||||
- mint
|
||||
- mute
|
||||
- open
|
||||
- pin
|
||||
- pinned
|
||||
- post
|
||||
- preferences
|
||||
- profile
|
||||
- prompt
|
||||
- react
|
||||
- reactions
|
||||
- room
|
||||
- search
|
||||
- secrets
|
||||
- sign
|
||||
- thread
|
||||
- track
|
||||
- trade
|
||||
- trending
|
||||
- unpin
|
||||
- unreact
|
||||
- visitors
|
||||
- vote
|
||||
- watch
|
||||
- welcome
|
||||
- who
|
||||
- whoami
|
||||
|
||||
@daemon (5 commands)
|
||||
- actions
|
||||
- activity
|
||||
- activity_report
|
||||
- schedule
|
||||
- update
|
||||
|
||||
@drone (24 commands)
|
||||
- activate
|
||||
- add
|
||||
- branches
|
||||
- check
|
||||
- exists
|
||||
- info
|
||||
- list
|
||||
- load
|
||||
- lock
|
||||
- lookup
|
||||
- path
|
||||
- pr
|
||||
- remove
|
||||
- reset
|
||||
- resolve
|
||||
- route
|
||||
- route_all
|
||||
- scan
|
||||
- set
|
||||
- status
|
||||
- sync
|
||||
- system
|
||||
- systems
|
||||
- unlock
|
||||
|
||||
@flow (11 commands)
|
||||
- aggregate
|
||||
- close
|
||||
- create
|
||||
- list
|
||||
- post_close
|
||||
- register
|
||||
- registry
|
||||
- restore
|
||||
- scan
|
||||
- templates
|
||||
- unregister
|
||||
|
||||
@memory (13 commands)
|
||||
- analyze
|
||||
- bootstrap
|
||||
- check
|
||||
- demo
|
||||
- extract
|
||||
- fragments
|
||||
- rollover
|
||||
- search
|
||||
- status
|
||||
- symbolic
|
||||
- templates
|
||||
- verify
|
||||
- watch
|
||||
|
||||
@prax (4 commands)
|
||||
- dashboard
|
||||
- log-audit
|
||||
- monitor
|
||||
- status
|
||||
|
||||
@seedgo (8 commands)
|
||||
- audit
|
||||
- checklist
|
||||
- diagnostics
|
||||
- diagnostics_audit
|
||||
- readme
|
||||
- readme_update
|
||||
- standards_audit
|
||||
- standards_query
|
||||
|
||||
@skills (6 commands)
|
||||
- create
|
||||
- discover
|
||||
- info
|
||||
- list
|
||||
- run
|
||||
- validate
|
||||
|
||||
@spawn (4 commands)
|
||||
- create
|
||||
- delete
|
||||
- passport
|
||||
- update
|
||||
|
||||
@trigger (15 commands)
|
||||
- branch_log_events
|
||||
- core
|
||||
- errors
|
||||
- fire
|
||||
- list
|
||||
- log_events
|
||||
- medic
|
||||
- mute
|
||||
- off
|
||||
- on
|
||||
- reset
|
||||
- start
|
||||
- status
|
||||
- stop
|
||||
- unmute
|
||||
|
||||
DISCOVERED: 173 commands across 14 branches
|
||||
|
||||
---
|
||||
|
||||
## 4. Test Coverage (26% module coverage)
|
||||
|
||||
|
||||
@ai_mail (49 tests, 2 files)
|
||||
tests/test_send_identity.py -- 36 tests -> email, users
|
||||
tests/test_user_paths.py -- 13 tests -> users
|
||||
UNTESTED: branch_ping, central_writer, dispatch, json, json_utils, monitoring, notify, registry
|
||||
|
||||
@api (0 tests, 0 files)
|
||||
(no test files found)
|
||||
|
||||
@backup (0 tests, 1 file)
|
||||
tests/test_pattern_scan.py -- 0 tests -> config, operations
|
||||
UNTESTED: backup_core, config, diff, google_drive_sync, integrations, json, models, operations, reauth_drive, reporting, utils
|
||||
|
||||
@cli (0 tests, 0 files)
|
||||
(no test files found)
|
||||
|
||||
@commons (82 tests, 2 files)
|
||||
tests/test_commons.py -- 72 tests -> curation, database, notifications, profiles, search, welcome
|
||||
tests/test_lifecycle.py -- 10 tests -> database, search
|
||||
UNTESTED: activity, artifact, artifacts, capsule, catchup, central, comment, comments, commons_identity, dashboard, digest, engagement, explore, feed, identity, json, leaderboard, notification, post, posts, profile, reaction, room, rooms, social, space, trade
|
||||
|
||||
@daemon (31 tests, 1 file)
|
||||
tests/test_actions_registry.py -- 31 tests -> actions
|
||||
UNTESTED: activity_report, json, monitoring, schedule, scheduler_ops, update, wakeup_ops
|
||||
|
||||
@devpulse (0 tests, 0 files)
|
||||
(no test files found)
|
||||
|
||||
@drone (335 tests, 9 files)
|
||||
tests/test_activation.py -- 38 tests -> command_registry, executor
|
||||
tests/test_commands.py -- 35 tests -> command_registry, commands
|
||||
tests/test_discovery.py -- 40 tests -> discovery, discovery_handler, exceptions, module_registry_handler
|
||||
tests/test_executor.py -- 26 tests -> exceptions, executor
|
||||
tests/test_git_module.py -- 46 tests -> git, git_module, module_registry_handler
|
||||
tests/test_registry_handler.py -- 33 tests -> exceptions, registry_handler
|
||||
tests/test_resolver.py -- 52 tests -> exceptions, registry_handler, resolver
|
||||
tests/test_router.py -- 37 tests -> exceptions, executor, router, router_handler
|
||||
tests/test_scan.py -- 28 tests -> scan, scanning
|
||||
UNTESTED: config, json, module_registry, registry
|
||||
|
||||
@flow (0 tests, 0 files)
|
||||
(no test files found)
|
||||
|
||||
@memory (0 tests, 0 files)
|
||||
(no test files found)
|
||||
|
||||
@prax (0 tests, 0 files)
|
||||
(no test files found)
|
||||
|
||||
@seedgo (0 tests, 0 files)
|
||||
(no test files found)
|
||||
|
||||
@skills (123 tests, 8 files)
|
||||
tests/test_cli_routing.py -- 20 tests -> ?
|
||||
tests/test_discovery.py -- 32 tests -> discovery_handler
|
||||
tests/test_lifecycle.py -- 12 tests -> creator, discovery, loader_handler, runner, template
|
||||
tests/test_loader.py -- 7 tests -> loader
|
||||
tests/test_registry.py -- 14 tests -> registry
|
||||
tests/test_runner.py -- 13 tests -> runner
|
||||
tests/test_runner_handler.py -- 14 tests -> runner_handler
|
||||
tests/test_validator.py -- 11 tests -> validator
|
||||
UNTESTED: creator_handler, json
|
||||
|
||||
@spawn (113 tests, 5 files)
|
||||
tests/test_citizen_classes.py -- 28 tests -> class_registry, core, passport_ops, update, update_ops
|
||||
tests/test_handlers.py -- 36 tests -> change_detection, json_ops, meta_ops, reconcile
|
||||
tests/test_lifecycle.py -- 22 tests -> delete, delete_ops, sync_registry, sync_registry_ops, sync_templates, sync_templates_ops
|
||||
tests/test_spawn.py -- 13 tests -> metadata, placeholders, registry
|
||||
tests/test_update.py -- 14 tests -> meta_ops, update, update_ops
|
||||
UNTESTED: file_ops, json, passport
|
||||
|
||||
@trigger (0 tests, 0 files)
|
||||
(no test files found)
|
||||
|
||||
|
||||
Test Coverage Report
|
||||
═══════════════════════════════════════════════════════
|
||||
|
||||
TESTED:
|
||||
✓ drone 335 tests 9 files 15/19 modules covered (79%)
|
||||
✓ skills 123 tests 8 files 10/12 modules covered (83%)
|
||||
✓ spawn 113 tests 5 files 18/21 modules covered (86%)
|
||||
|
||||
PARTIAL:
|
||||
◐ ai_mail 49 tests 2 files 2/10 modules covered (20%)
|
||||
◐ commons 82 tests 2 files 6/33 modules covered (18%)
|
||||
◐ daemon 31 tests 1 file 1/8 modules covered (12%)
|
||||
|
||||
NO TESTS:
|
||||
✗ api 0 tests 0 files 0/10 modules covered (0%)
|
||||
✗ backup 0 tests 1 file 0/11 modules covered (0%)
|
||||
✗ cli 0 tests 0 files 0/5 modules covered (0%)
|
||||
✗ devpulse 0 tests 0 files 0/0 modules covered (0%)
|
||||
✗ flow 0 tests 0 files 0/16 modules covered (0%)
|
||||
✗ memory 0 tests 0 files 0/16 modules covered (0%)
|
||||
✗ prax 0 tests 0 files 0/14 modules covered (0%)
|
||||
✗ seedgo 0 tests 0 files 0/14 modules covered (0%)
|
||||
✗ trigger 0 tests 0 files 0/12 modules covered (0%)
|
||||
|
||||
═══════════════════════════════════════════════════════
|
||||
SUMMARY: 733 tests across 15 branches
|
||||
3 branches tested, 3 partial, 9 untested
|
||||
Coverage: 52/201 modules (26%)
|
||||
|
||||
@@ -0,0 +1,337 @@
|
||||
# Testing Standards
|
||||
**Status:** Draft v1 - Current manual process documented
|
||||
**Date:** 2025-11-12
|
||||
|
||||
---
|
||||
|
||||
## Current State: Manual Testing (Effective for Rapid Iteration)
|
||||
|
||||
**Reality:** AIPass is a custom framework in active evolution. System changes weekly. Building extensive test infrastructure now = updating tests constantly instead of building features.
|
||||
|
||||
**Current approach:** Manual testing with JSON/log verification works effectively for rapid iteration phase.
|
||||
|
||||
**Future:** pytest framework expansion once branches stabilize (pytest is already configured and operational in some branches).
|
||||
|
||||
---
|
||||
|
||||
## The 90% Build Process
|
||||
|
||||
**How we build and test:**
|
||||
|
||||
### 1. Planning Phase (Upfront Investment)
|
||||
- Issue plan through Flow (master plan or default plan depending on scope)
|
||||
- Spend time getting structure right before coding
|
||||
- Define what we're building and how it fits together
|
||||
|
||||
### 2. Build to 90% (AI-Led, Internal Verification)
|
||||
- AI builds structure and implementation
|
||||
- **Internal verification tests as we go:**
|
||||
- Does the module turn on?
|
||||
- Do commands work?
|
||||
- Basic functionality confirmed?
|
||||
- Human mostly observes, provides input if things go astray
|
||||
- Focus on getting structure and pieces in place
|
||||
|
||||
**Linux advantage:** Same environment (AI and human both on Linux) = outputs match = internal tests are reliable. Windows had issues with this, Linux doesn't.
|
||||
|
||||
### 3. 90% Threshold (Human Review Phase)
|
||||
- Human gets actively involved
|
||||
- Reviews structure and implementation
|
||||
- Asks questions about concerns
|
||||
- Performs feature tests (what module is supposed to do)
|
||||
- Identifies bugs and missing error handling
|
||||
|
||||
### 4. Debug Cycle (Error Handling First, Then Fixes)
|
||||
|
||||
**Critical pattern: Fix error handling BEFORE fixing bugs**
|
||||
|
||||
**Example scenario:**
|
||||
```
|
||||
Bug: API call not working, but console says "Success"
|
||||
No error output, but logs show API never executed
|
||||
Output is lying
|
||||
|
||||
Process:
|
||||
1. Fix error handling FIRST - make errors tell the truth
|
||||
2. THEN fix the actual bug (why API isn't executing)
|
||||
3. See clean pass with honest outputs
|
||||
4. Move to next feature
|
||||
```
|
||||
|
||||
**Why this order:**
|
||||
- Can't debug effectively if errors lie
|
||||
- Truth in outputs = faster debugging
|
||||
- Honest error messages = system teaches itself what's wrong
|
||||
|
||||
### 5. Iterate Until Acceptable
|
||||
- Test different features
|
||||
- Continue debug cycle (handle errors → fix bugs → verify)
|
||||
- Reach acceptable standard (basic or advanced depending on module)
|
||||
- Move on
|
||||
|
||||
---
|
||||
|
||||
## Why Manual Testing Works Right Now
|
||||
|
||||
**Advantages:**
|
||||
- **Fast iteration** - No test maintenance overhead
|
||||
- **Flexible** - System can change tomorrow without breaking test suite
|
||||
- **JSON/log infrastructure** - Acts as verification layer
|
||||
- Check config.json → see settings
|
||||
- Check data.json → see state
|
||||
- Check log.json → see operations history
|
||||
- Check Prax logs → detailed debugging
|
||||
- **Linux environment match** - AI tests = human tests (same outputs)
|
||||
- **Effective debugging** - Error-first approach catches issues early
|
||||
|
||||
**Tradeoffs:**
|
||||
- Manual effort required at 90% stage
|
||||
- No automated regression testing (yet)
|
||||
- Relies on human review for final verification
|
||||
|
||||
**Acceptable because:** System is evolving rapidly. Better to build fast and test manually than build slow and maintain brittle tests.
|
||||
|
||||
---
|
||||
|
||||
## Future: pytest Framework
|
||||
|
||||
**Infrastructure in place:**
|
||||
- `pytest.ini` at repo root - Configuration for test discovery and markers
|
||||
- `tests/` - Root test directory with conftest.py
|
||||
- Branch-specific test directories at `src/aipass/{branch}/tests/` (API, Prax, CLI, etc.)
|
||||
|
||||
**Already operational in some branches:**
|
||||
- API branch: 4 test files (test_api_system.py, test_openrouter_key.py, etc.)
|
||||
- Prax branch: test_log_rotation.py validates rotation behavior
|
||||
- Seedgo branch: test_cli_errors.py demonstrates error handling patterns
|
||||
|
||||
**When to expand testing:**
|
||||
- Once modules and branches stabilize
|
||||
- When system changes slow down (monthly, not weekly)
|
||||
- When maintenance cost < value of automation
|
||||
|
||||
**Why pytest:**
|
||||
- Standard Python testing framework
|
||||
- Already configured with pytest.ini
|
||||
- Infrastructure exists, ready for expansion
|
||||
- Some branches already have working tests
|
||||
|
||||
**Current selective approach:**
|
||||
- Test critical/stable components (API, Prax log rotation)
|
||||
- Skip testing rapidly changing features
|
||||
- Manual testing for experimental work
|
||||
- Automated tests where they add value without maintenance burden
|
||||
|
||||
---
|
||||
|
||||
## Error Handling Philosophy
|
||||
|
||||
**Errors must tell the truth** - foundational to testing effectiveness
|
||||
|
||||
**Good error handling:**
|
||||
```python
|
||||
try:
|
||||
result = api_call()
|
||||
if not result:
|
||||
logger.error("API call failed - no response")
|
||||
return {"success": False, "error": "API returned no data"}
|
||||
except Exception as e:
|
||||
logger.error(f"API call exception: {e}", exc_info=True)
|
||||
return {"success": False, "error": str(e)}
|
||||
```
|
||||
|
||||
**Bad error handling:**
|
||||
```python
|
||||
try:
|
||||
result = api_call()
|
||||
return {"success": True} # LIES - didn't check if result valid
|
||||
except:
|
||||
pass # Silent failure - no truth
|
||||
```
|
||||
|
||||
**Testing relies on honest errors:**
|
||||
- If errors lie, testing is impossible
|
||||
- Fix error handling first = testing becomes possible
|
||||
- Then fix bugs with confident verification
|
||||
|
||||
---
|
||||
|
||||
## JSON/Log Infrastructure as Testing Layer
|
||||
|
||||
**Three-JSON pattern supports testing:**
|
||||
|
||||
### Config Verification
|
||||
```bash
|
||||
# Check if settings are correct
|
||||
cat module_name_config.json
|
||||
# See API keys, limits, feature toggles
|
||||
```
|
||||
|
||||
### State Verification
|
||||
```bash
|
||||
# Check current state
|
||||
cat module_name_data.json
|
||||
# See metrics, counts, current status
|
||||
```
|
||||
|
||||
### Operations Verification
|
||||
```bash
|
||||
# Check what actually happened
|
||||
cat module_name_log.json
|
||||
# See recent operations and results
|
||||
```
|
||||
|
||||
### Detailed Debugging
|
||||
```bash
|
||||
# Prax provides file-based logging, not real-time watching
|
||||
# Check system logs directory for detailed output
|
||||
ls -la system_logs/
|
||||
# Read specific module logs (Prax manages this directory location)
|
||||
cat system_logs/module_name.log
|
||||
```
|
||||
|
||||
**This infrastructure = verification layer without formal tests**
|
||||
|
||||
---
|
||||
|
||||
## Testing Workflow Example
|
||||
|
||||
**Building a new branch creation module:**
|
||||
|
||||
1. **Plan** - Define structure, features, workflow (Flow plan)
|
||||
|
||||
2. **Build to 90%** - AI implements:
|
||||
- Module structure
|
||||
- Handler functions
|
||||
- Config/data/log JSONs
|
||||
- Internal verification: "Does `create_branch test_branch` work?"
|
||||
|
||||
3. **Review at 90%** - Human tests:
|
||||
- Create branch with various names
|
||||
- Check if files copied correctly
|
||||
- Verify registry updated
|
||||
- Try edge cases (existing branch, invalid name)
|
||||
|
||||
4. **Find bug** - Branch created but registry not updated
|
||||
- **First:** Check error handling - is error logged? Is return value honest?
|
||||
- Add error handling if missing
|
||||
- **Then:** Fix bug - why isn't registry updating?
|
||||
- Verify with clean pass
|
||||
|
||||
5. **Iterate** - Test more features:
|
||||
- Template copying
|
||||
- Placeholder replacement
|
||||
- Memory file handling
|
||||
- Backup on conflicts
|
||||
|
||||
6. **Acceptable** - All major features work, errors are honest, ready to use
|
||||
|
||||
---
|
||||
|
||||
## Current Testing Checklist
|
||||
|
||||
**For any new module/feature:**
|
||||
|
||||
- [ ] Does it turn on without errors?
|
||||
- [ ] Do basic commands work?
|
||||
- [ ] Are errors handled and logged?
|
||||
- [ ] Do outputs tell the truth?
|
||||
- [ ] Check config.json - settings correct?
|
||||
- [ ] Check data.json - state tracking working?
|
||||
- [ ] Check log.json - operations recorded?
|
||||
- [ ] Test edge cases (invalid input, missing files, etc.)
|
||||
- [ ] Check Prax logs in system_logs/ for detailed debugging info
|
||||
- [ ] Manual feature tests at 90% stage
|
||||
|
||||
---
|
||||
|
||||
## When to Test What
|
||||
|
||||
**During development (AI internal verification):**
|
||||
- Module starts without errors
|
||||
- Basic commands execute
|
||||
- Expected output appears
|
||||
|
||||
**At 90% stage (human testing):**
|
||||
- Feature functionality (does it do what it's supposed to?)
|
||||
- Edge cases (what breaks it?)
|
||||
- Error handling (are errors honest?)
|
||||
- Integration (does it work with other modules?)
|
||||
|
||||
**Before considering "done":**
|
||||
- Clean passes on major features
|
||||
- Errors tell the truth (no silent failures)
|
||||
- JSON logs show operations correctly
|
||||
- Acceptable standard reached (basic or advanced)
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
**Current approach:** Manual testing with JSON/log infrastructure
|
||||
|
||||
**Why it works:**
|
||||
- Fast iteration without test maintenance
|
||||
- Linux environment = reliable verification
|
||||
- Error-first debugging = effective bug fixing
|
||||
- JSON/log system = verification layer
|
||||
|
||||
**Build process:** Plan → Build to 90% → Review/test → Debug (errors first, then bugs) → Iterate → Acceptable
|
||||
|
||||
**Future:** pytest framework when system stabilizes
|
||||
|
||||
**Philosophy:** Build fast, verify as you go, handle errors honestly, iterate rapidly. Test infrastructure comes after stability.
|
||||
|
||||
---
|
||||
|
||||
## Pytest Test Structure
|
||||
|
||||
**Current implementation:**
|
||||
```
|
||||
pytest.ini # Test configuration (repo root)
|
||||
tests/ # Root test directory
|
||||
conftest.py # Shared fixtures
|
||||
src/aipass/
|
||||
api/tests/ # API tests (4 files operational)
|
||||
test_api_system.py
|
||||
test_openrouter_key.py
|
||||
test_free_models_quick.py
|
||||
test_paid_model.py
|
||||
prax/tests/ # Prax tests
|
||||
test_log_rotation.py
|
||||
cli/tests/ # CLI tests (infrastructure ready)
|
||||
seedgo/
|
||||
tests/conftest.py
|
||||
apps/modules/test_cli_errors.py # Demo module
|
||||
```
|
||||
|
||||
**Running tests:**
|
||||
```bash
|
||||
# Run all tests
|
||||
pytest
|
||||
|
||||
# Run specific branch tests
|
||||
pytest src/aipass/api/tests/
|
||||
|
||||
# Run with markers
|
||||
pytest -m unit
|
||||
pytest -m integration
|
||||
pytest -m slow
|
||||
```
|
||||
|
||||
**Test demonstration module:**
|
||||
- `src/aipass/seedgo/apps/modules/test_cli_errors.py` - Shows error handling patterns
|
||||
- Not a pytest test, but demonstrates testing concepts
|
||||
- Run directly: `python3 src/aipass/seedgo/apps/modules/test_cli_errors.py`
|
||||
|
||||
---
|
||||
|
||||
## Comments
|
||||
|
||||
#@comments:2025-11-13:claude: Pytest infrastructure exists and operational in API/Prax branches. Selective testing approach: automate stable components, manual testing for rapid development.
|
||||
|
||||
#@comments:2025-11-13:claude: Prax doesn't have real-time "watcher" command for logs - uses file-based logging in system_logs/ (Prax manages the location). Updated documentation to reflect actual capabilities.
|
||||
|
||||
#@comments:2025-11-13:claude: Error-first debugging approach is critical - maybe this should be emphasized in error_handling.md when we fill that section?
|
||||
|
||||
#@comments:2025-11-13:claude: "90% threshold" is interesting pattern - not 100% perfection, but "acceptable standard" (basic or advanced). Reflects pragmatic development philosophy.
|
||||
Reference in New Issue
Block a user