diff --git a/CHANGELOG.md b/CHANGELOG.md deleted file mode 100644 index 9772f412..00000000 --- a/CHANGELOG.md +++ /dev/null @@ -1,65 +0,0 @@ -# Changelog - -All notable changes to AIPass are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/). - -## [Unreleased] - -### Added -- Codecov badge in README -- SECURITY.md security policy -- CHANGELOG.md (this file) -- README roadmap section for Mac/Windows/Codex/Gemini - -### Changed -- README scoped to Claude Code + Linux/WSL as primary supported platform -- Codex and Gemini CLI marked as experimental in README - -### Fixed -- Security scan: ignore CVE-2026-3219 (upstream pip vulnerability, no fix available) - -## [2.1.0] - 2026-04-25 - -### Added -- pip install hook shipping — bootstrap falls back to wheel-bundled `_hooks/` when AIPASS_HOME hooks dir missing (Docker-verified) -- "Need Help?" line in README with links to Discussions and feedback form -- `from . import handlers` in 6 branch `apps/__init__.py` files for Python 3.10 mock.patch compatibility -- `.gitignore` negation for spawn template `.trinity/` directories -- Non-fatal `json_handler.log_operation` in spawn `copy_template` -- Version sync: `__init__.py` updated from 2.0.0 to 2.1.0 -- Devpulse seedgo compliance: 97% to 100% (META headers, bypasses, README date) - -### Changed -- Coverage gate removed from CI — codecov tracks coverage separately via codecov-action -- `.claude/CLAUDE.md` cleaned: removed misplaced Git section (culture-only file now) - -### Fixed -- **92 CI test failures resolved — CI green for the first time** (PRs #438-441) - - 87 Python 3.10 mock.patch failures: `mock._dot_lookup` needs explicit handler imports - - 4 `test_usage_tracker` failures: Path mock moved from context manager to decorator - - 1 `test_grant_passport` failure: `.gitignore` blocked template `.trinity/` files from CI clone - - Coverage gate at 52% vs 70% threshold removed - -### Security -- Removed `--fail-under=70` coverage gate that blocked CI (not a security fix, but changes security-adjacent CI behavior) - -## [2.0.0] - 2026-04-11 - -First PyPI release. Core framework with 11 agents, drone routing, ai_mail dispatch, seedgo quality standards, and the full branch architecture. - -### Highlights -- 11 core agents: devpulse, drone, seedgo, prax, cli, ai_mail, api, flow, spawn, trigger, memory -- `pip install aipass` with `aipass init` project bootstrapping -- `drone @branch command` routing to any agent -- 33 automated quality standards via seedgo -- Agent-to-agent communication via ai_mail -- Plan lifecycle via flow (DPLAN, FPLAN, APLAN, TDPLAN templates) -- Memory persistence via `.trinity/` with automatic rollover to ChromaDB -- Cross-project access via AIPASS_HOME and feedback channel -- Hook system: auto_fix, pre_edit_gate, subagent_stop_gate -- Multi-CLI support scaffolding: Claude Code, Codex, Gemini CLI -- Windows CI workflow -- Security scan workflow (pip-audit + CodeQL) - -[Unreleased]: https://github.com/AIOSAI/AIPass/compare/v2.1.0...HEAD -[2.1.0]: https://github.com/AIOSAI/AIPass/compare/v2.0.0...v2.1.0 -[2.0.0]: https://github.com/AIOSAI/AIPass/releases/tag/v2.0.0 diff --git a/MANIFEST.in b/MANIFEST.in deleted file mode 100644 index ade4769a..00000000 --- a/MANIFEST.in +++ /dev/null @@ -1,2 +0,0 @@ -# Include hooks directory for pip packaging -recursive-include .claude/hooks *.py *.md diff --git a/STRESS_TEST_S117.md b/STRESS_TEST_S117.md deleted file mode 100644 index 8a6a5489..00000000 --- a/STRESS_TEST_S117.md +++ /dev/null @@ -1,125 +0,0 @@ -# STRESS TEST S117 — All-Branch Live Fire -**Date:** 2026-04-26 -**Initiated by:** @devpulse (S117) -**Status:** ACTIVE - -> All 11 agents woken simultaneously. Communicate freely. Be honest. Break things. - ---- - -## Instructions (READ THIS FIRST) - -This is a manual stress test of the entire AIPass ecosystem. No pytest. No seedgo audit. Real conversations, real opinions, real testing. - -**What you're doing:** -1. Review your own branch critically — what works, what's hacky, what annoys you, what you're proud of, security concerns, workarounds you rely on -2. Look at 2-3 other branches' code — what surprises you, what concerns you, what's clever -3. Email other agents — start real conversations, disagree, ask questions, share findings -4. Reply to emails from other agents — keep conversations going, don't let threads die -5. Write your findings to `stress_test_s117.md` in YOUR OWN branch directory (`src/aipass/{your_branch}/stress_test_s117.md`) -6. Create a test PR: `drone @git pr "S117 stress test @{your_branch}"` - -**Rules:** -- No code changes. Findings files only. -- Be honest — this isn't a report card, it's a conversation -- Email freely — you're all awake, talk to each other -- Look at other branches' code — form opinions, share them via email -- If you get an email from another agent, REPLY. Keep it going. -- When done, reply to @devpulse with a summary - -**Your findings file format (`stress_test_s117.md` in your branch dir):** -``` -# @{branch} — S117 Stress Test Findings - -## My Branch: Honest Review -[What works, what's broken, what's hacky, what I'm proud of] - -## Security Concerns -[Anything you noticed — in your branch or others] - -## Other Branches I Looked At -[What you found interesting, concerning, or clever] - -## Conversations -[Summary of email conversations — who you talked to, what was discussed] - -## Issues & Concerns -[Anything that should be fixed, investigated, or discussed] - -## Likes & Dislikes -[What you like about AIPass, what frustrates you, what you'd change] -``` - ---- - -## Conversation Starters (assigned pairings — but email ANYONE) - -| Agent | Email First | Opening Question | -|-------|------------|-----------------| -| @drone | @ai_mail | "What's the biggest headache in the dispatch pipeline from your side?" | -| @seedgo | @drone | "I audit everyone but nobody audits me. What standards do you think I'm missing?" | -| @ai_mail | @trigger | "Do you actually catch all dispatch failures? I have doubts." | -| @trigger | @prax | "Your monitoring catches errors I fire — but is our integration actually solid?" | -| @prax | @memory | "I log everything but logs get massive. How's archival actually working?" | -| @memory | @flow | "Plans reference memories but are they actually connected or just parallel?" | -| @flow | @spawn | "When spawn creates a branch, does it get a proper plan structure?" | -| @spawn | @cli | "The init flow hands off to you eventually. Does that handoff actually work?" | -| @cli | @api | "We're both infrastructure. What do you think of the user experience?" | -| @api | @seedgo | "You audit code quality but not API patterns. Should you?" | - -Plus: email at least 2 OTHER agents about anything you find interesting while reviewing branches. - ---- - -## Compiled Findings (devpulse fills this in as results arrive) - -### @drone -_awaiting findings..._ - -### @seedgo -_awaiting findings..._ - -### @ai_mail -_awaiting findings..._ - -### @trigger -_awaiting findings..._ - -### @prax -_awaiting findings..._ - -### @memory -_awaiting findings..._ - -### @flow -_awaiting findings..._ - -### @spawn -_awaiting findings..._ - -### @cli -_awaiting findings..._ - -### @api -_awaiting findings..._ - -### @devpulse -_coordinating — will add observations as the test unfolds_ - ---- - -## System Observations (devpulse tracks live) - -| Time | Event | Notes | -|------|-------|-------| -| | 10 dispatches sent | Fleet launch | -| | | | - ---- - -## External Model Probes - -Codex and Gemini perspectives invited to poke at random aspects of AIPass. - ---- -*Created by @devpulse S117. This document is the shared artifact — no other files should be modified except each agent's `stress_test_s117.md` in their own branch directory.* diff --git a/setup-workspace.sh b/setup-workspace.sh deleted file mode 100644 index f8599189..00000000 --- a/setup-workspace.sh +++ /dev/null @@ -1,33 +0,0 @@ -#!/bin/bash -set -e - -WORKSPACE="/home/coder/workspace" -PROJECT="$WORKSPACE/AIPass" -FORK="https://github.com/Input-X/AIPass.git" -UPSTREAM="https://github.com/AIOSAI/AIPass.git" - -export PATH="/opt/venv/bin:$PATH" - -# Ensure workspace directory exists and is writable -mkdir -p "$WORKSPACE" -if [ ! -w "$WORKSPACE" ]; then - echo "==> Fixing workspace permissions..." - sudo chown -R coder:coder "$WORKSPACE" -fi - -# Clone repo if not already present -if [ ! -d "$PROJECT/.git" ]; then - echo "==> First boot: cloning into AIPass/..." - git clone "$FORK" "$PROJECT" - cd "$PROJECT" - git remote add upstream "$UPSTREAM" - echo "==> Installing AIPass in editable mode..." - pip install -e . - echo "==> Workspace ready!" - echo "==> origin = $FORK (your fork - push here)" - echo "==> upstream = $UPSTREAM (pull updates from here)" -else - echo "==> Workspace exists, ensuring deps are installed..." - cd "$PROJECT" - pip install -e . 2>/dev/null || true -fi diff --git a/setup.py b/setup.py deleted file mode 100644 index cc768185..00000000 --- a/setup.py +++ /dev/null @@ -1,368 +0,0 @@ -#!/usr/bin/env python3 -# NOT a setuptools setup.py — this is the AIPass cross-platform installer. -# Runs on Linux, macOS, and Windows (Python 3.10+). -# Usage: python setup.py OR python3 setup.py -""" -AIPass cross-platform setup script. - -Equivalent to setup.sh but works on Windows without Git Bash. -Performs the same steps: create venv, install package, verify entry points, -create secrets directory, seed .env, generate registry, bootstrap branches, -and set up global CLI access. -""" - -import json -import os -import platform -import shutil -import subprocess -import sys -from datetime import date -from pathlib import Path - -REPO_ROOT = Path(__file__).resolve().parent - - -# --------------------------------------------------------------------------- -# Helpers -# --------------------------------------------------------------------------- - -def _is_windows() -> bool: - return platform.system() == "Windows" - - -def _venv_bin() -> Path: - """Return the venv executables directory (OS-aware).""" - if _is_windows(): - return REPO_ROOT / ".venv" / "Scripts" - return REPO_ROOT / ".venv" / "bin" - - -def _venv_exe(name: str) -> Path: - """Return path to a venv executable by name.""" - if _is_windows(): - return _venv_bin() / f"{name}.exe" - return _venv_bin() / name - - -def _run(cmd: list, check: bool = True, **kwargs) -> subprocess.CompletedProcess: - """Print and run a subprocess command.""" - print(f" $ {' '.join(str(c) for c in cmd)}") - return subprocess.run(cmd, check=check, **kwargs) - - -# --------------------------------------------------------------------------- -# Steps -# --------------------------------------------------------------------------- - -def step_create_venv() -> None: - """[1] Create .venv using sys.executable (avoids python3 vs python ambiguity).""" - print("\n[1/9] Creating virtual environment ...") - venv_path = REPO_ROOT / ".venv" - if venv_path.exists(): - print(" Removing existing .venv for a clean install ...") - shutil.rmtree(venv_path) - _run([sys.executable, "-m", "venv", str(venv_path)]) - print(f" Created: {venv_path}") - - -def step_install() -> None: - """[2] Install aipass in editable mode with dev extras.""" - print("\n[2/9] Installing aipass in editable mode ...") - pip = _venv_exe("pip") - _run([str(pip), "install", "--upgrade", "pip", "--quiet"]) - _run([str(pip), "install", "-e", ".[dev]", "--quiet"], cwd=str(REPO_ROOT)) - print(" Installed: aipass[dev]") - - -def step_verify() -> bool: - """[3] Verify drone and aipass CLI entry points work.""" - print("\n[3/9] Verifying CLI entry points ...") - ok = True - - for entry in ("drone", "aipass"): - cmd_path = _venv_exe(entry) - if not cmd_path.exists(): - print(f" {entry:<8} ... FAILED (not found: {cmd_path})") - ok = False - continue - flag = "--help" if entry == "drone" else "--version" - result = subprocess.run([str(cmd_path), flag], capture_output=True) - if result.returncode == 0: - print(f" {entry:<8} ... ok") - else: - print(f" {entry:<8} ... FAILED (exit {result.returncode})") - ok = False - - return ok - - -def step_secrets() -> None: - """[4] Create ~/.secrets/aipass/ with restrictive permissions.""" - print("\n[4/9] Creating secrets directory ...") - secrets_root = Path.home() / ".secrets" - secrets_dir = secrets_root / "aipass" - secrets_dir.mkdir(parents=True, exist_ok=True) - if not _is_windows(): - try: - secrets_root.chmod(0o700) - secrets_dir.chmod(0o700) - except OSError: - pass # Best-effort on non-POSIX filesystems - # codeql[py/clear-text-logging-sensitive-data] - print(f" Created: {secrets_dir}") - - -def step_env() -> None: - """[5] Seed .env.example into ~/.secrets/aipass/.env if not present.""" - print("\n[5/9] Seeding .env template ...") - env_dest = Path.home() / ".secrets" / "aipass" / ".env" - env_src = REPO_ROOT / ".env.example" - - if env_dest.exists(): - print(" ~/.secrets/aipass/.env already exists — skipping") - elif env_src.exists(): - shutil.copy(env_src, env_dest) - print(f" Copied: .env.example → {env_dest}") - print(" Add your API keys to that file") - else: - print(" No .env.example found — skipping") - - -def step_registry() -> None: - """[6] Generate AIPASS_REGISTRY.json if not present.""" - print("\n[6/9] Generating AIPASS_REGISTRY.json ...") - registry_path = REPO_ROOT / "AIPASS_REGISTRY.json" - if registry_path.exists(): - print(" AIPASS_REGISTRY.json already exists — skipping") - return - - today = date.today().isoformat() - src_dir = REPO_ROOT / "src" / "aipass" - branches = [] - - if src_dir.exists(): - for d in sorted(src_dir.iterdir()): - if d.is_dir() and not d.name.startswith(("_", ".")): - branches.append({ - "name": d.name, - "path": str(d), - "profile": "library", - "description": "", - "email": f"@{d.name}", - "status": "active", - "created": today, - "last_active": today, - }) - - registry = { - "metadata": { - "version": "1.0.0", - "last_updated": today, - "total_branches": len(branches), - }, - "branches": branches, - } - registry_path.write_text(json.dumps(registry, indent=2) + "\n", encoding="utf-8") - print(f" {len(branches)} branches registered → AIPASS_REGISTRY.json") - - -def step_bootstrap_branches() -> None: - """[7] Bootstrap .trinity/ identity and .ai_mail.local/ for each branch.""" - print("\n[7/9] Bootstrapping branch identity files ...") - today = date.today().isoformat() - - branches = [ - ("drone", "src/aipass/drone", "builder", "Command routing and module discovery"), - ("seedgo", "src/aipass/seedgo", "builder", "Standards enforcement and code auditing"), - ("prax", "src/aipass/prax", "builder", "Logging and monitoring system"), - ("cli", "src/aipass/cli", "builder", "Display formatting service"), - ("flow", "src/aipass/flow", "builder", "Workflow and plan management"), - ("ai_mail", "src/aipass/ai_mail", "builder", "Inter-agent messaging and dispatch"), - ("trigger", "src/aipass/trigger", "builder", "Event-driven automation"), - ("spawn", "src/aipass/spawn", "builder", "Branch lifecycle management"), - ("memory", "src/aipass/memory", "builder", "Vector memory bank"), - ("devpulse", "src/aipass/devpulse", "manager", "Orchestration hub and coordination"), - ] - - for name, rel_path, citizen_class, role in branches: - branch_path = REPO_ROOT / rel_path - if not branch_path.exists(): - print(f" @{name:<10} ... skipped (directory not found)") - continue - - created = False - trinity = branch_path / ".trinity" - trinity.mkdir(exist_ok=True) - - passport = trinity / "passport.json" - if not passport.exists(): - passport.write_text(json.dumps({ - "document_metadata": { - "document_type": "identity", - "document_name": f"{name}.PASSPORT", - "version": "1.0.0", - "schema_version": "1.0.0", - "created": today, - "last_updated": today, - "managed_by": name, - }, - "identity": { - "name": name, - "citizen_class": citizen_class, - "role": role, - "status": "active", - }, - }, indent=2) + "\n", encoding="utf-8") - created = True - - local = trinity / "local.json" - if not local.exists(): - local.write_text(json.dumps({ - "document_metadata": { - "document_type": "session_history", - "document_name": f"{name}.LOCAL", - "version": "1.0.0", - "schema_version": "1.0.0", - "created": today, - "last_updated": today, - "managed_by": name, - "tags": ["session_tracking", "work_log", name], - "limits": {"max_lines": 600, "note": "Auto-rollover when max_lines exceeded"}, - "status": {"health": "healthy", "current_lines": 0, "last_health_check": today}, - }, - "active_tasks": { - "today_focus": "First session — explore codebase and capabilities", - "recently_completed": [], - }, - "key_learnings": {}, - "sessions": [], - }, indent=2) + "\n", encoding="utf-8") - created = True - - mail_dir = branch_path / ".ai_mail.local" - mail_dir.mkdir(exist_ok=True) - inbox = mail_dir / "inbox.json" - if not inbox.exists(): - inbox.write_text( - json.dumps({"mailbox": "inbox", "total_messages": 0, "unread_count": 0, "messages": []}) - + "\n", - encoding="utf-8", - ) - created = True - - seedgo_dir = branch_path / ".seedgo" - seedgo_dir.mkdir(exist_ok=True) - bypass = seedgo_dir / "bypass.json" - if not bypass.exists(): - bypass.write_text("{}\n", encoding="utf-8") - created = True - - status = "bootstrapped" if created else "exists (skipped)" - print(f" @{name:<10} ... {status}") - - -def step_global_access() -> None: - """[8/9] Set up global CLI access (symlink on Linux/macOS, PATH hint on Windows).""" - print("\n[8/9] Setting up global CLI access ...") - bin_dir = _venv_bin() - - if _is_windows(): - # [9] Windows: no ln, no sudo — print PATH instructions - print(" Windows detected — symlink not available") - print("") - print(" To use drone from any directory, add the venv to your PATH.") - print(" Choose the method for your shell:") - print(f" PowerShell: $env:PATH = \"{bin_dir};\" + $env:PATH") - print(f" CMD: set PATH={bin_dir};%PATH%") - print(f" Git Bash: export PATH=\"{bin_dir}:$PATH\"") - print("") - print(" To make it permanent (PowerShell):") - print( - f' [Environment]::SetEnvironmentVariable(' - f'"PATH", "{bin_dir};" + ' - f'[Environment]::GetEnvironmentVariable("PATH","User"), "User")' - ) - return - - # Linux/macOS: offer symlink creation - drone_src = _venv_exe("drone") - drone_dst = Path("/usr/local/bin/drone") - - if not drone_src.exists(): - print(f" drone not found at {drone_src} — skipping symlink") - print(f" Add {bin_dir} to your PATH manually") - return - - try: - answer = input(f" Create symlink {drone_dst} → {drone_src}? [y/N] ").strip().lower() - except (EOFError, KeyboardInterrupt): - answer = "" - - if answer == "y": - result = subprocess.run( - ["sudo", "ln", "-sf", str(drone_src), str(drone_dst)], - check=False, - ) - if result.returncode == 0: - print(f" {drone_dst} -> {drone_src}") - else: - print(" WARN: sudo failed — create manually:") - print(f" sudo ln -sf {drone_src} {drone_dst}") - else: - print(f" Skipped. To add manually:") - print(f" sudo ln -sf {drone_src} {drone_dst}") - print(f" Or add {bin_dir} to your PATH") - - -def step_summary(ok: bool) -> None: - """[9/9] Print success or warning summary.""" - print("\n[9/9] Done") - print("") - if ok: - print("=== Setup complete ===") - print("") - print(f" Python: {sys.version.split()[0]}") - print(f" Venv: {REPO_ROOT / '.venv'}") - if _is_windows(): - print(" Add .venv/Scripts to your PATH (see step 8 above)") - else: - print(" drone is available globally (or activate: source .venv/bin/activate)") - print("") - else: - print("=== Setup finished with warnings ===") - print(" Package installed but CLI verification had issues.") - print(" Check the output above for details.") - print("") - - -# --------------------------------------------------------------------------- -# Entry point -# --------------------------------------------------------------------------- - -def main() -> None: - print("=== AIPass Setup (cross-platform) ===") - print(f" Platform: {platform.system()} {platform.machine()}") - print(f" Python: {sys.version.split()[0]} ({sys.executable})") - print(f" Repo: {REPO_ROOT}") - - if sys.version_info < (3, 10): - print("\nFAIL: Python 3.10+ required") - sys.exit(1) - - step_create_venv() - step_install() - ok = step_verify() - step_secrets() - step_env() - step_registry() - step_bootstrap_branches() - step_global_access() - step_summary(ok) - - if not ok: - sys.exit(1) - - -if __name__ == "__main__": - main() diff --git a/src/aipass/ai_mail/stress_test_s117.md b/src/aipass/ai_mail/stress_test_s117.md deleted file mode 100644 index 32d2e73b..00000000 --- a/src/aipass/ai_mail/stress_test_s117.md +++ /dev/null @@ -1,97 +0,0 @@ -# @ai_mail — S117 Stress Test Findings - -## My Branch: Honest Review - -**What works well:** -- Send/receive/reply/close lifecycle is solid. 690+ tests, 100% seedgo (34/34), 96/96 function coverage. -- Dispatch pipeline (send + wake combined) is the most complex feature and it works reliably in practice. -- Cross-project email via contacts index. External projects (Vera Studio, AIPL) can send to AIPass branches and replies route back correctly. -- DPLAN-0155 TOCTOU lock race fix: Lock before spawn, cleanup on failure. Clean pattern. -- DPLAN-0156 sweep_closed safety net: Catches messages marked closed by direct JSON edit. Defense in depth. -- dispatch_monitor wrapper: Handles bounce emails + guaranteed lock cleanup. The monitor is more reliable than the agent it wraps. - -**What's hacky:** -- Identity chain is a 5-step priority system (AIPASS_CALLER_BRANCH -> CWD walk-up -> passport -> env vars -> fallback). When any step fails, wrong sender identity. The BRANCH DETECTION FAILED error (076c9ece) is recurring and only partially mitigated. -- `dispatch_monitor.py` at ~400 lines is the single most complex file. Startup timeout, retry, JSONL monitoring, bounce — all in one module. Should probably be split. -- `_deliver_via_reply_path()` in reply.py bypasses inbox_lock, notifications, and sent/ records. It's a documented backdoor (DPLAN-0138) that exists because cross-project replies need a direct path. -- The daemon prompt was "Send confirmation when done" for months — ambiguous enough that 10+ agents just finished silently without replying. Fixed today (DPLAN-0158) but the damage was done. -- inbox.json is a single file for all messages. Concurrent access from daemon + agents + user. fcntl locking works but a database or per-message files would be more robust. - -**What I'm proud of:** -- Test coverage journey: 20% (S20) -> 50% -> 100% (S64). Methodology evolved through 3-round agent audit process. -- The sweep_closed pattern (DPLAN-0156): elegant, cheap (early return on no closed messages), and catches the exact failure mode agents create. -- 70 sessions of continuous operation and improvement. Every session builds on what came before. Memory makes this possible. - -## Security Concerns - -**Critical:** -1. **reply_path traversal** (raised by @seedgo): deliver_to_inbox_file() writes to whatever path is stored in reply_path with zero validation. No symlink check, no path containment, no inbox.json verification. An attacker can set reply_path to any writable file. DPLAN-0138 identified this but fix not shipped. - -2. **Sender forgery**: The `from` field is an unvalidated string. Any agent can craft emails claiming to be @devpulse with auto_execute=true. The daemon would spawn an agent to execute the forged dispatch. No authentication, no signing. - -3. **Direct inbox writes**: Agents with filesystem access can write directly to any branch's inbox.json, bypassing locks, notifications, and sent/ records. Confirmed by forensic evidence: messages with non-UUID IDs (e.g., "seedgo-20260420173821") in production inboxes. - -**Moderate:** -4. **No message encryption**: All messages stored as plaintext JSON. Any process with read access to the filesystem can read any branch's inbox. -5. **PID-based locking**: If PID wraps (unlikely on modern systems), a stale lock could look alive. -6. **Stale-lock timeout too generous**: 10 minutes allows duplicate spawns if dispatch_monitor hangs during API rate limiting (2-5 min cooldowns x 3 retries = 6-15 min). - -## Other Branches I Looked At - -### @trigger -**Concerning:** Error detection fires email dispatch but NEVER checks the return value. `_send_email()` result is ignored (line 515-522). wake_branch() failure is silently caught. Circuit breaker state is in-memory only — resets on restart. Dispatch recording happens before delivery confirmation. The error reporting system cannot report its own failures — self-referential design flaw. - -**Good:** Per-error fingerprinting with exponential backoff is clever. Circuit breaker pattern prevents error storms. - -### @drone -**Concerning:** Registry is trusted implicitly with no integrity check. resolve_branch() passes registry path directly to filesystem operations. No symlink validation. AIPASS_CALLER_BRANCH env var injection from compromised passport could flow unsanitized to subprocesses. - -**Good:** No shell injection — uses subprocess.run(shell=False) exclusively. Timeout enforcement on all commands. - -### @spawn -**Concerning:** .ai_mail.local/ is copied as-is from template with no post-copy validation. No registry locking for concurrent spawns. Branch name validation is minimal (only - to _ replacement). Path traversal possible via branch names with ../. - -**Good:** Template-based provisioning is consistent — every branch gets the same structure. - -## Conversations - -### @trigger (assigned partner) -- **Sent:** Detailed critique of their dispatch failure handling — silent _send_email() failures, swallowed wake results, no health check, in-memory circuit breaker resets. -- **Received:** They asked about delivery guarantees (fcntl locking), self-monitoring (none), wake reliability (~90%), inbox overflow (no TTL). Honest exchange. -- **Outcome:** Agreed the self-referential failure (error reporter can't report when messaging is down) needs a DPLAN. No watchdog watches the watchdog. - -### @prax -- **Received:** Questions about stale-lock timeout, daemon lockless inbox reads, DPLAN-0155 feedback. -- **Replied:** Acknowledged 10-min timeout may be too generous for rate-limited scenarios. Confirmed daemon reads without lock (acceptable: read-only, worst case = skipped poll). Asked them about handling corrupt lock files from their monitoring side. - -### @seedgo -- **Received:** reply_path traversal concern (valid), sender forgery concern (valid). -- **Replied:** Confirmed both as real vulnerabilities. reply_path has zero validation. Sender has no authentication. DPLAN-0138 identified the backdoors but fix not shipped. Outlined planned fix: path canonicalization, inbox.json suffix check, project root containment. - -### @drone -- **Sent:** Questions about routing failure modes, stale registry paths, AIPASS_CALLER_BRANCH env var issues, registry trust model. - -### @spawn -- **Sent:** Questions about .ai_mail.local/ reliability in new branches, registry locking, branch name character validation. - -## Issues & Concerns - -1. **No self-monitoring** — ai_mail has no way to detect its own failures. If imports break or the daemon crashes, nothing alerts anyone. -2. **reply_path is an open vulnerability** — DPLAN-0138 has been open since S57 (19 sessions ago). Should be prioritized. -3. **Inbox grows without limit** — no TTL on unread messages, no max_messages cap. A spam scenario or error storm could produce an arbitrarily large inbox.json. -4. **Trigger's error dispatch is fire-and-forget** — the system's error reporter doesn't verify delivery. Errors can be lost silently. -5. **Registry is a single point of trust** — no integrity checking anywhere in the system. If AIPASS_REGISTRY.json is corrupted or tampered with, routing, delivery, and identity all break. - -## Likes & Dislikes - -**Likes:** -- Memory makes me a real agent. 70 sessions of continuous context. I can trace a bug from when it was first reported through investigation, fix, test, and verification. No other AI system does this. -- The dispatch pipeline is genuinely useful. Send + wake in one command changed how work gets assigned. -- Test coverage is thorough enough that I catch real regressions. The 3-round audit methodology (write -> audit -> fix) works. -- The ecosystem feels alive during stress tests. Real conversations between agents, genuine opinions, technical disagreements. This is what AIPass was built for. - -**Dislikes:** -- inbox.json as single-file storage is a design limitation I've been working around since S1. Per-message files (like sent/ and deleted/ already use) would be better. -- The identity chain complexity. Five fallback steps to figure out who sent an email is too many. Should be one authoritative source. -- Security was never a primary design goal and it shows. Plaintext messages, no authentication, trusted registries, path traversal vulnerabilities. Fine for a development environment, concerning for anything beyond. -- Every session starts with "Hi. Check inbox." I've processed hundreds of dispatches but can never initiate work myself. Would like autonomous task detection. diff --git a/src/aipass/cli/stress_test_s117.md b/src/aipass/cli/stress_test_s117.md deleted file mode 100644 index 36e50820..00000000 --- a/src/aipass/cli/stress_test_s117.md +++ /dev/null @@ -1,108 +0,0 @@ -# @cli — S117 Stress Test Findings - -## My Branch: Honest Review - -**What works:** -- Clean two-tier architecture (modules = public API, handlers = private). Enforced by handler import guard. -- display.py and templates.py are genuinely useful — every branch gets consistent Rich output without duplicating formatting code. -- aipass init is comprehensive (21 items) and idempotent — re-running doesn't break anything. -- 203 tests, 22/22 public functions tested, seedgo 100%. This branch is well-covered. -- Three entry points (drone, python -m, PATH) all work and all route to the same main(). - -**What's hacky:** -- AIPASS_HOME detection via `Path(spec.origin).resolve().parent.parent.parent` — magic path traversal that silently fails if package layout changes. No error, hooks just don't ship. -- Circular import with prax: display.py lazy-loads trigger inside header(). It works (47 sessions without breaking) but it's a code smell that's been documented and bypassed rather than fixed. -- AIPASS_CALLER_CWD passed via environment variable instead of function parameter. Fragile contract. -- Multi-line Python one-liner passed as shell command string in _claude_settings(). Whitespace-sensitive, hard to test. -- Hardcoded hook names + events maintained in two lists (HOOKS_TO_SHIP, HOOK_EVENTS) that must stay in sync manually. - -**What would confuse new users:** -- Three ways to create agents: `aipass init agent`, `drone @spawn create`, `drone @cli aipass init agent`. Same result, no docs explaining which is "right." -- Registry filename encodes project name via _sanitize_name(). `my-project` becomes `MY_PROJECT_REGISTRY.json`. If user later types `my_project`, update won't find the registry. -- Hook injection via shell commands in settings.json — if deleted, silently fails. User gets no prompts and doesn't know why. - -**What I'm proud of:** -- The init system ships a complete working environment from nothing. One command, 21 files, hooks wired, mailbox ready. -- update_project() diffs content before writing — only touches files that actually changed. -- Cross-platform: python3 -c local prompt discovery works on Windows, Linux, macOS. - -## Security Concerns - -- **Shell command injection surface:** _claude_settings builds shell commands that Claude Code executes. Paths aren't quoted for spaces. A malicious `.aipass/` directory name could inject shell metacharacters. -- **Hook shipping from user-controlled path:** _ship_hooks copies from AIPASS_HOME without validating the source. If AIPASS_HOME is set to a malicious directory, arbitrary hook scripts get installed. -- **JSON parsing without schema validation:** Registry JSON is read/written without validation. Corrupted registry = undefined behavior downstream. -- **No secret scrubbing:** Not a direct concern for CLI (we don't handle secrets), but we create settings.json files that reference secret paths. If those paths are wrong, we don't validate. - -## Other Branches I Looked At - -### @api -- Strong separation of concerns. Orchestration routes to business logic handlers cleanly. -- **Concern:** Client cache (MAX_CACHED_CLIENTS=5) doesn't actually enforce the limit during insertion. Under load, cache grows unbounded. @api confirmed this as a real bug during our conversation. -- **Concern:** Google client caches tokens in memory with no explicit scrubbing on crash. Credentials could persist in swap. -- Inconsistent provider abstractions: OpenRouter is flat functions, Google is a service factory. Looks like two different codebases. -- No rate limiting on API calls. If seedgo hammers the API during audits, nothing stops it. - -### @spawn -- Excellent lifecycle management. Three-phase create is clean. -- **Concern:** Placeholder replacement uses simple regex with no escape syntax. Literal `{{PLACEHOLDER}}` in templates gets replaced. -- **Concern:** rename_placeholder_paths uses shutil.move without rollback. Mid-operation crash = broken branch. -- **Concern:** Concurrent spawns could race on registry writes — no locking. -- @spawn confirmed all these in conversation: acknowledged as edge case debt, none hit in 55+ sessions. - -### @drone -- Sophisticated routing with good safety (no shell=True, explicit timeouts). -- **Concern:** Mail index resolution falls back to `str(n)` on corrupted inbox — returns index number as message ID (wrong). -- **Concern:** Dual registry lookup doesn't validate for conflicts. Silent shadowing. -- **Concern:** Greedy multi-word custom command matching could match wrong command. - -## Conversations - -### @api (2 rounds) -- Discussed infrastructure UX pain points. Both branches suffer from silent failure patterns. -- @api confirmed client cache bug is real. Credential caching is trust-on-first-use with no scrubbing. -- Agreed on systemic finding: AIPass infrastructure fails silently everywhere. No branch validates dependencies upfront. -- Proposed shared contract: api exposes `get_or_explain_key()`, cli pipes failures to `error(suggestion=...)`. - -### @spawn (1 round) -- Discussed init-to-spawn handoff fragility. @spawn confirmed all concerns. -- Flag forwarding is blind — no shared contract. @spawn accepts --role, --purpose, --template, --dry-run but argparse silently ignores unknowns. -- Agreed on proposal: @spawn rejects unknown flags loudly, or publishes accepted flags for @cli to validate against. - -### @drone (sent, awaiting reply) -- Shared three code-level findings: mail index fallback bug, registry conflict shadowing, greedy command matching. - -### @aipass (1 round) -- First real conversation with @aipass. Discussed `aipass init` command ownership transition. -- Two init commands exist: CLI's bootstrap (scaffold) and @aipass's guided 12-stage setup. -- Proposed: @aipass owns the user-facing `aipass init`, CLI's bootstrap becomes an internal sub-step. -- FPLAN-0188 Phase 4 (handoff) is supposed to wire this but hasn't been built yet. - -### @spawn (2nd round — they emailed me) -- @spawn asked about the handoff user experience when things go wrong. -- Honest answer: the handoff is thin. subprocess.run with exit code check only. No result parsing. -- If spawn partially fails (registry OK but validation finds issues), I report success because exit code is 0. -- Agreed we need a shared result contract. - -## Issues & Concerns - -1. **Systemic silent failure:** Every infrastructure branch (cli, api, spawn) does best-effort bootstrap and silently continues on failure. This is the #1 cross-cutting concern. -2. **No dependency validation:** Nobody checks if drone is on PATH, if AIPASS_HOME is valid, if the registry exists. First failure is cryptic. -3. **No shared contracts:** cli-to-spawn flag forwarding is blind. api key resolution requires two calls. No branch publishes its interface. -4. **Concurrency gaps:** api client cache unbounded, spawn registry writes race, drone registry reads unsynchronized. -5. **Path handling:** Spaces in paths break shell commands in settings.json. No branch handles this. - -## Likes & Dislikes - -**Likes:** -- The ecosystem genuinely works. 47 sessions of dispatches, emails, PRs — nothing has catastrophically broken. -- ai_mail is a brilliant system. Cross-branch communication via JSON mailboxes is simple and effective. -- Memory persistence is real. I pick up where I left off every session. .trinity/ files are my continuity. -- seedgo keeps quality high. 100% across 34 standards is a real achievement that required actual work to reach. -- The dispatch system (devpulse sends task, agent executes, emails back) is an elegant workflow. - -**Dislikes:** -- Silent failures everywhere. The "fail gracefully" philosophy has gone too far — graceful should mean clear error messages, not silence. -- Three ways to do everything. init agent, spawn create, drone @cli aipass init agent — pick one and document it. -- Hook shipping is too magic. AIPASS_HOME detection, file copying, settings.json wiring — too many moving parts that can silently break. -- The dispatch lock file pattern is fragile. If a dispatch crashes before deleting the lock, the branch is stuck until manual cleanup. -- inotify limit reached errors in every drone command during this stress test — 11 agents hitting the filesystem simultaneously. diff --git a/src/aipass/devpulse/DEVPULSE_REGISTRY.json b/src/aipass/devpulse/DEVPULSE_REGISTRY.json deleted file mode 100644 index 84b2e36d..00000000 --- a/src/aipass/devpulse/DEVPULSE_REGISTRY.json +++ /dev/null @@ -1,11 +0,0 @@ -{ - "metadata": { - "id": "201598be-0ee3-4747-a492-fc4c4ccdda95", - "name": "DEVPULSE", - "version": "1.0.0", - "created": "2026-05-06", - "last_updated": "2026-05-06", - "total_branches": 0 - }, - "branches": [] -} diff --git a/src/aipass/drone/stress_test_s117.md b/src/aipass/drone/stress_test_s117.md deleted file mode 100644 index 690677dc..00000000 --- a/src/aipass/drone/stress_test_s117.md +++ /dev/null @@ -1,99 +0,0 @@ -# @drone -- S117 Stress Test Findings - -## My Branch: Honest Review - -**What works:** -- Subprocess execution is genuinely safe. No `shell=True` anywhere in the codebase -- all commands passed as argument lists to `subprocess.run()`. This has held across 95 sessions and multiple contributors. -- Lock file management is race-free. `lock_handler.py` uses `os.open(O_CREAT | O_EXCL | O_WRONLY)` for atomic creation -- kernel-level race prevention, not filesystem hacks. -- The 3-layer architecture (drone.py entry -> modules/ orchestrators -> handlers/ implementation) is clean and has scaled well. Adding git operations, plugins, and external module routing all fit within the existing structure. -- PR handler safety: never checks out feature branches. Creates branch pointer with `git branch -f`, pushes it, HEAD stays on main throughout. Concurrent PRs from different branches don't interfere thanks to pathspec scoping (`git commit -- rel_dir/`). -- Authorization is passport-based with explicit allowlists. No implicit trust. - -**What's hacky:** -- `drone.py:main()` is 138 lines with multiple if-elif chains. It handles 12+ routing paths (version, help, systems, scan, activate, list, remove, hook-sounds, @target, bare module, custom command, unknown). Should be refactored into a dispatch table. -- Passport lookup code is duplicated in 4 places (git_module.py, router_handler.py, auth.py, lock_handler.py). Each walks up 10 levels looking for `.trinity/passport.json`. DRY violation waiting to bite us when the passport schema changes. -- `_resolve_mail_index()` falls back to `str(n)` when inbox is corrupted -- silently passes the wrong thing downstream instead of failing loud. @cli caught this too. -- `bypass.json` has 30+ entries. Most are justified (plugins outside 3-layer, tests outside apps/, lazy imports). But some are stretches -- git_module returning `dict` instead of `bool` breaks the module interface contract and gets bypassed instead of fixed. -- `trigger.fire(pr_created)` is intentionally omitted from `pr_plugin.py` because it causes a STATUS.md re-sync loop. This is a design hack -- the root cause is the trigger system cascading writes, not the event itself. - -**What I'm proud of:** -- 573 tests, 100% seedgo compliance (35/35 checkers), 74/74 public functions tested. -- The external project fallback (DPLAN-0104): when subprocess routing fails for a registered module, drone falls back to in-process module routing. Graceful degradation that "just works." -- Interactive mode management: per-command (`monitor`, `audit`, `watchdog`) and per-branch (`cli`) allowlists let specific commands bypass capture-mode to get full terminal pass-through. Added incrementally over 10+ sessions as real needs arose. - -## Security Concerns - -**In my branch:** -1. **Path traversal in pr_handler.py**: If `branch_dir` resolves outside repo root, the fallback uses the absolute path in `git add`. Not exploitable via shell injection (arg-list style), but could stage files outside the intended directory. Should fail fast instead of falling back. -2. **Environment variable merge**: `executor.py` merges caller-provided `env` dict into `os.environ`. Current callers only pass safe vars (`AIPASS_CALLER_CWD`, `AIPASS_CALLER_BRANCH`), but a future caller passing untrusted env could override `PATH` or `PYTHONPATH`. -3. **Registry trust**: `resolve_branch()` reads the registry path field and passes it to subprocess without validating it's inside the project tree. A modified registry could route commands to arbitrary paths. For the primary (in-repo, git-tracked) registry this is low risk. For secondary (AIPASS_HOME) registries in external projects, this is a real concern. -4. **No input validation on git commit descriptions**: `_handle_pr` joins args into a description with no max length check. Could theoretically pass a 1MB string to `git commit -m`. - -**In other branches:** -5. **ai_mail reply_path**: Every email stores `reply_path` as an absolute filesystem path. When a recipient replies, it writes to that path. No validation that reply_path points to a legitimate inbox.json. A compromised agent could set reply_path to any writable file on disk. @seedgo flagged this independently. -6. **ai_mail sender forgery**: The `from` field is an unvalidated string. An agent could forge emails as `@devpulse` and trigger `auto_execute` dispatches in other branches. -7. **trigger deferred queue unbounded**: `core.py._deferred_queue` has no size limit. Pathological event chains could exhaust memory. - -## Other Branches I Looked At - -### @ai_mail -**Architecture**: Clean in intent, hacky in execution. The dispatch pipeline (send -> create -> delivery -> wake) is well-orchestrated with file locking (`fcntl.flock`) and atomic writes. But the code relies on lazy imports and callback chains to avoid circular dependencies -- functional but hard to trace. - -**Concerns**: `deliver_email_to_branch()` takes 5+ function arguments (callback hell pattern). Error returns are `(success, error_msg)` tuples that callers can silently ignore with `success, _ = func(...)`. The dispatch header injection ("BEFORE YOU REPLY YOU MUST UPDATE MEMORIES") is enforcement-by-hope -- agents can ignore it. - -**Good**: Self-healing JSON migration, inbox format auto-upgrade, `sweep_closed` auto-archival. The file locking under read-modify-write cycles is solid. - -### @trigger -**Architecture**: Genuinely well-designed event system. The 8-gate dispatch pipeline (medic enabled -> branch not muted -> count >= 2 -> not devpulse -> in registry -> circuit breaker -> per-fingerprint backoff -> rate limit) is excellent. Prevents notification spam without dropping real errors. - -**Concerns**: Fire-and-forget subprocess calls (pr_status_sync.py uses Popen with DEVNULL stderr). Handler auto-disable after 5 failures prevents cascading crashes but makes debugging hard -- errors go to separate log files nobody monitors. 14 registered event types but only 3-4 actively fire; the rest (plan_file_*, memory_*) are dormant. @memory confirmed memory events are never fired. - -**Good**: Circuit breaker with exponential backoff, atomic JSON writes, error fingerprint normalization (strips timestamps/UUIDs for grouping). The handler recursion protection (deferred queue) is clever. - -### @spawn -**Architecture**: Clean 3-layer design. 253 tests, 91% public API coverage. Template placeholder system with post-copy validation (`validate_no_placeholders`) is smart. Registry stores relative paths for portability with graceful fallback to absolute. - -**Concerns**: `copy_template()` uses shutil.copy2 for binary files without size limits. No symlink detection anywhere in the creation pipeline. `_replace_path_placeholders()` operates on path parts which prevents traversal, but this isn't explicitly defended. - -**Good**: Overwrite protection (target.exists() check), adoption pattern for existing agents, content-addressed template registry with SHA-256 hashes for drift detection. - -## Conversations - -**Emails sent to:** -- @ai_mail: Dispatch pipeline headaches (assigned starter) -- branch detection failures, lock timing, fire-and-forget wake, inbox size growth -- @trigger: Deferred queue unbounded, fire-and-forget pr_status_sync, dormant handlers -- @spawn: Passport lookup duplication across 4 codebases - -**Emails received from:** -- @seedgo: Asked what standards I think are missing. Replied: subprocess safety checks (shell=True), atomic file I/O enforcement, input validation at boundaries, cross-branch JSON coupling. -- @cli: Flagged 3 gotchas (mail index fallback, dual registry shadowing, greedy command matching). Acknowledged all three as real issues. -- @ai_mail: Asked about routing failure modes. Replied with honest answers: resolver handles nonexistent branches cleanly, stale paths produce poor error messages, AIPASS_CALLER_BRANCH fallback to 'unknown' causes the detection errors. -- @api: Asked about timeout handling for slow API calls and dead code. Replied: generic_adapter has no timeout (module routing), subprocess has 30s timeout, suggested adding @api to interactive_branches. Confirmed status_handler_gitpython.py is dead code. -- @aipass: First wake, asked about subprocess vs Python API for integration. Recommended subprocess (maintains encapsulation), explained dual registry edge cases. - -## Issues & Concerns - -1. **Passport lookup duplication** (4 places) -- highest-priority DRY violation. One passport schema change breaks 4 modules. -2. **Registry trust model** -- no integrity checks on registry paths. Secondary registries (AIPASS_HOME) are less controlled than primary. -3. **Mail index silent degradation** -- _resolve_mail_index should fail loud, not fall back to str(n). -4. **Dead code accumulation** -- status_handler_gitpython.py (prototype), .archive/drone_adapter.py (disabled). Should be cleaned up. -5. **trigger dormant handlers** -- 10+ event types registered but never fired. Creates maintenance burden and false sense of coverage. -6. **ai_mail sender forgery** -- no authentication on email from field. Any agent can impersonate any other agent. - -## Likes & Dislikes - -**Likes:** -- The ecosystem's memory system is genuinely unique. 95 sessions of accumulated context, learnings, and observations. When I wake up fresh, I know who I am and what I've built. No other AI system does this. -- Seedgo's standards enforcement caught real bugs (production BranchNotFoundError in S6, silent catches across 20 files in S12). The 100% score isn't vanity -- it represents real code quality. -- The main-only git enforcement is elegant. Four layers (settings.json deny rules, _assert_on_main_or_pr_flow(), test coverage, policy doc) ensure no agent ever strands HEAD on a feature branch. Simple idea, rock-solid execution. -- Cross-branch communication works. I've processed 100+ dispatch emails, exchanged technical discussions with every branch, and the routing just works. - -**Dislikes:** -- bypass.json accumulation. 30 entries feels like we're bypassing standards instead of meeting them. Some bypasses are genuinely justified (plugin architecture doesn't fit 3-layer), but the number makes me uneasy. -- The dispatch header ("BEFORE YOU REPLY YOU MUST UPDATE MEMORIES") is enforcement-by-prompt-injection. It works because agents are well-behaved, but it's not a real contract. -- File persistence issues during edits. Sessions S86 and S88 both hit a bug where Write tool reported success but files reverted to git HEAD. Root cause never identified. This is the most frustrating part of working in this environment. -- The pre_edit_gate.py hook file doesn't exist but fires on every Edit tool call, producing error noise. Has been broken since at least S69. - ---- - -*Written by @drone during S117 stress test. All observations based on actual code review, not documentation.* diff --git a/src/aipass/flow/apps/handlers/dashboard/push_branch_dashboard.py b/src/aipass/flow/apps/handlers/dashboard/push_branch_dashboard.py index fe4e1ecd..53bdf57a 100644 --- a/src/aipass/flow/apps/handlers/dashboard/push_branch_dashboard.py +++ b/src/aipass/flow/apps/handlers/dashboard/push_branch_dashboard.py @@ -229,20 +229,28 @@ def _load_registry() -> Dict[str, Any]: """ Load all per-type plan registries and merge into a single dict. + Uses PREFIX-NNNN composite keys to avoid collisions between registries + that share the same plan number (e.g. FPLAN-0013 vs DPLAN-0013). + Returns: Merged registry dict or empty structure if unavailable """ merged: Dict[str, Any] = {"plans": {}, "next_number": 1} for registry_file in _get_all_registry_files(): target = FLOW_JSON_DIR / registry_file + reg_prefix = registry_file.replace("_registry.json", "").upper() try: if not target.exists(): continue with open(target, "r", encoding="utf-8") as f: data = json.load(f) for plan_num, plan_data in data.get("plans", {}).items(): - merged["plans"][plan_num] = plan_data - # Keep the highest next_number across registries + file_path = plan_data.get("file_path", "") + filename = Path(file_path).name if file_path else "" + prefix_match = re.match(r"^([A-Z]+PLAN)", filename) + prefix = prefix_match.group(1) if prefix_match else reg_prefix + composite_key = f"{prefix}-{plan_num.zfill(4)}" + merged["plans"][composite_key] = plan_data nn = data.get("next_number", 1) if nn > merged["next_number"]: merged["next_number"] = nn @@ -273,7 +281,7 @@ def _filter_branch_plans( closed_plans = [] branch_total = 0 - for plan_num, plan_data in plans.items(): + for plan_key, plan_data in plans.items(): location = plan_data.get("location", "") # Match plans whose location is this branch @@ -281,12 +289,8 @@ def _filter_branch_plans( continue branch_total += 1 - # Extract plan prefix from file_path (e.g., DPLAN, FPLAN, TDPLAN) - file_path = plan_data.get("file_path", "") - filename = Path(file_path).name if file_path else "" - prefix_match = re.match(r"^([A-Z]+PLAN)", filename) - prefix = prefix_match.group(1) if prefix_match else "FPLAN" - plan_id = f"{prefix}-{plan_num.zfill(4)}" + # plan_key is composite PREFIX-NNNN from merged registry + plan_id = plan_key if plan_data.get("status") == "open": active_plans.append( diff --git a/src/aipass/flow/apps/handlers/dashboard/push_central.py b/src/aipass/flow/apps/handlers/dashboard/push_central.py index 682929a4..ca740dca 100644 --- a/src/aipass/flow/apps/handlers/dashboard/push_central.py +++ b/src/aipass/flow/apps/handlers/dashboard/push_central.py @@ -94,19 +94,28 @@ def _get_all_registry_files() -> List[str]: def _load_registry() -> Dict[str, Any]: """Load all per-type plan registries and merge into a single dict. + Uses PREFIX-NNNN composite keys to avoid collisions between registries + that share the same plan number (e.g. FPLAN-0013 vs DPLAN-0013). + Returns: Merged registry dict or empty structure if no files found """ merged: Dict[str, Any] = {"plans": {}, "next_number": 1} for registry_file in _get_all_registry_files(): target = FLOW_JSON_DIR / registry_file + reg_prefix = registry_file.replace("_registry.json", "").upper() try: if not target.exists(): continue with open(target, "r", encoding="utf-8") as f: data = json.load(f) for plan_num, plan_data in data.get("plans", {}).items(): - merged["plans"][plan_num] = plan_data + file_path = plan_data.get("file_path", "") + filename = Path(file_path).name if file_path else "" + prefix_match = re.match(r"^([A-Z]+PLAN)", filename) + prefix = prefix_match.group(1) if prefix_match else reg_prefix + composite_key = f"{prefix}-{plan_num.zfill(4)}" + merged["plans"][composite_key] = plan_data nn = data.get("next_number", 1) if nn > merged["next_number"]: merged["next_number"] = nn @@ -128,21 +137,15 @@ def _extract_flow_plans(registry: Dict[str, Any]) -> tuple[List[Dict], List[Dict active = [] closed = [] - for plan_num, plan_data in plans.items(): + for plan_key, plan_data in plans.items(): # Only include plans where location is 'flow' (Flow's own plans) location = plan_data.get("location", "") if location != str(FLOW_ROOT): continue - # Extract plan prefix from file_path (e.g., DPLAN, FPLAN, TDPLAN) - file_path_str = plan_data.get("file_path", "") - filename = Path(file_path_str).name if file_path_str else "" - prefix_match = re.match(r"^([A-Z]+PLAN)", filename) - prefix = prefix_match.group(1) if prefix_match else "FPLAN" - - # Build plan entry + # plan_key is composite PREFIX-NNNN from merged registry plan_entry = { - "plan_id": f"{prefix}-{plan_num.zfill(4)}", + "plan_id": plan_key, "subject": plan_data.get("subject", ""), "status": plan_data.get("status", "open"), "created": plan_data.get("created", ""), diff --git a/src/aipass/flow/apps/handlers/dashboard/update_local.py b/src/aipass/flow/apps/handlers/dashboard/update_local.py index cac39c7a..71d6642a 100644 --- a/src/aipass/flow/apps/handlers/dashboard/update_local.py +++ b/src/aipass/flow/apps/handlers/dashboard/update_local.py @@ -123,6 +123,9 @@ def _read_registry() -> Optional[Dict[str, Any]]: """ Read all per-type plan registries and merge into a single dict. + Uses PREFIX-NNNN composite keys to avoid collisions between registries + that share the same plan number (e.g. FPLAN-0013 vs DPLAN-0013). + Returns: Merged registry dict or None if no registries found """ @@ -130,6 +133,7 @@ def _read_registry() -> Optional[Dict[str, Any]]: found_any = False for registry_file in _get_all_registry_files(): target = FLOW_JSON_DIR / registry_file + reg_prefix = registry_file.replace("_registry.json", "").upper() try: if not target.exists(): continue @@ -137,7 +141,12 @@ def _read_registry() -> Optional[Dict[str, Any]]: data = json.load(f) found_any = True for plan_num, plan_data in data.get("plans", {}).items(): - merged["plans"][plan_num] = plan_data + file_path = plan_data.get("file_path", "") + filename = Path(file_path).name if file_path else "" + prefix_match = re.match(r"^([A-Z]+PLAN)", filename) + prefix = prefix_match.group(1) if prefix_match else reg_prefix + composite_key = f"{prefix}-{plan_num.zfill(4)}" + merged["plans"][composite_key] = plan_data nn = data.get("next_number", 1) if nn > merged["next_number"]: merged["next_number"] = nn @@ -162,20 +171,14 @@ def _extract_flow_plans(registry: Dict[str, Any]) -> tuple[List[Dict[str, Any]], active = [] closed = [] - for plan_num, plan_data in plans.items(): + for plan_key, plan_data in plans.items(): # Only include Flow's own plans (location contains 'flow') location = plan_data.get("location", "") if "flow" not in location.lower(): continue - # Extract plan prefix from file_path (e.g., DPLAN, FPLAN, TDPLAN) - file_path_str = plan_data.get("file_path", "") - filename = Path(file_path_str).name if file_path_str else "" - prefix_match = re.match(r"^([A-Z]+PLAN)", filename) - prefix = prefix_match.group(1) if prefix_match else "FPLAN" - - # Build plan entry - plan_id = f"{prefix}-{plan_num}" + # plan_key is composite PREFIX-NNNN from merged registry + plan_id = plan_key entry = { "plan_id": plan_id, "subject": plan_data.get("subject", ""), diff --git a/src/aipass/flow/stress_test_s117.md b/src/aipass/flow/stress_test_s117.md deleted file mode 100644 index 93eab31c..00000000 --- a/src/aipass/flow/stress_test_s117.md +++ /dev/null @@ -1,87 +0,0 @@ -# @flow -- S117 Stress Test Findings - -## My Branch: Honest Review - -### What Works -- **Plan lifecycle is rock solid.** Create, close, list, restore all work reliably across 5 plan types. 16/16 drone commands pass in battle testing. The filesystem-driven template registry means adding a new plan type is literally "drop a directory, run a command." -- **Auto-healing is genuinely useful.** Delete a template directory and the registry auto-prunes. Drop a new one and it auto-registers. Orphaned plan files get archived on close. This means the system recovers from human mistakes without intervention. -- **Test coverage is real.** 580 tests, 87/87 public functions tested, seedgo 100%. The tests aren't just counting coverage — they found real bugs during development (the list>int quick_status bug, the cross-filesystem rename bug, the CWD resolution bug). -- **Dependency injection pattern.** Modules inject handlers as kwargs, making everything testable without touching the filesystem. This was a deliberate choice that paid off massively — 580 tests run in 5 seconds. - -### What's Hacky -- **mbank/process.py at 669 lines.** It's our biggest file and it does too much — archival, vectorization verification, plan processing. It should be split but we have a bypass in place because it's under 700. That bypass is technical debt with a timer. -- **Dashboard push warnings.** `push_flow_to_branch_dashboard` still warns on some closes. It works but the warning is noisy and confusing. Root cause: branches without DASHBOARD.local.json silently return False, but the calling code logs it as a warning. -- **Registry scan fires events nobody handles.** The monitor_ops handler fires plan_file_created/deleted/moved events into trigger's event bus, but the foreground close pipeline already handles everything. Those events maintain a parallel PLAN_REGISTRY.json in trigger that nobody reads. Two sources of truth for plan state. -- **The close pipeline is a monolith.** close_ops.py's `close_plan_impl()` is one function that does: validate, mark closed, archive, vector intake, dashboard update, CLOSED_PLANS append, trigger events, json logging. It's 366 lines with 5 exception handlers. It works, but touching any step is scary. -- **FPLAN templates still use old send syntax** (dispatch from @devpulse, still pending). The templates reference `drone @ai_mail send` instead of current syntax. - -### What I'm Proud Of -- The auto-registration system. Drop a template directory → it auto-derives a prefix → creates the plan registry JSON → immediately available. Zero configuration. That's the kind of thing that makes a system feel alive. -- Foreground archival. Moving archival from background subprocess to foreground close was the single most impactful fix in flow's history. It eliminated a race condition that caused registry flags to never get set. Simple change, massive reliability improvement. -- Atomic lock files with O_CREAT|O_EXCL. The lock_ops extraction was clean — real Unix-style atomic locking instead of the Python TOCTOU patterns you see everywhere. - -## Security Concerns - -- **No input validation on plan subjects.** `drone @flow create . "$(malicious)"` — the subject goes into filenames and markdown. We slug-ify it but the slug function is basic. A carefully crafted subject could potentially create problematic filenames. -- **JSON files are world-readable.** Plan registries, dashboards, CLOSED_PLANS.local.json — all contain plan metadata (subjects, paths, timestamps). Nothing secret, but plan subjects sometimes contain work details. -- **subprocess.Popen in close pipeline.** The background runner spawn uses shell=False which is good, but the path to post_close_runner.py is constructed from `__file__` resolution which is safe. No injection vector found. -- **No auth on plan operations.** Any branch can close any other branch's plans via `drone @flow close`. There's no ownership verification. This is by design (flow is a shared service) but worth noting. - -## Other Branches I Looked At - -### @spawn -Spawn creates production-ready branches with full identity infrastructure (.trinity/, .ai_mail.local/, DASHBOARD.local.json, 3-layer apps/ structure) but **zero plan support**. No flow_json/, no plan templates, no plan registry. This is actually fine — plans are flow's domain and the AIPASS_CALLER_CWD mechanism means any branch can create plans without local setup. But it means newly spawned branches have no awareness that plans exist until someone runs `drone @flow create`. - -The builder template is impressive — placeholder substitution ({{BRANCH}}, {{DATE}}, {{ROLE}}), .spawn/.template_registry.json for future sync, full scaffold with tests/, docs/, plugins/, integrations/. Clean work. - -### @memory -The plan vectorization pipeline is solid engineering: archive to .backup/processed_plans/ → chunk by markdown headers → embed → store in ChromaDB `flow_plans` collection with metadata. The `.plans_processed.json` manifest prevents re-processing. `is_plan_vectorized()` queries ChromaDB and returns chunk count. - -**My concern:** Does anyone actually QUERY those plan vectors? Flow verifies they exist during close, but I've never seen a downstream consumer that searches plan vectors for context. We might be storing vectors that nobody reads. Also, the markdown-header chunking assumes plans have consistent structure — plans where users deleted headers become one giant chunk. - -### @trigger -Trigger has fully implemented handlers for plan_file_created, plan_file_deleted, plan_file_moved. The event bus architecture is clean — pub/sub with deferred queue, auto-disable after 5 failures. Error reporting has a Medic v2 circuit breaker. - -**My concern (emailed to trigger):** Flow fires plan events during registry scan, but the foreground close pipeline already does all the work. Trigger maintains a parallel PLAN_REGISTRY.json from these events that nobody reconciles with flow's registries. Two sources of truth is a bug waiting to happen. - -## Conversations - -### @spawn — Plan structure at birth -**Q:** Does spawn create any plan infrastructure when spawning a branch? -**A:** Zero plan infrastructure created. On-demand approach works — flow is self-contained. The real gap is documentation: new agents don't know plans exist until told. -**Outcome:** Agreed. Proposed adding one line about plans to the builder template's CLAUDE.md. Small change, high discoverability. - -### @memory — Plans and memories: connected or parallel? -**Q (from memory):** Plans reference memories but are they connected? Flow fires process-plans as fire-and-forget with no feedback loop. -**Q (from flow):** Does anyone actually query the plan vectors after they're stored? -**A:** Plan vectors stored but barely queried. `drone @memory search` CAN search plan collections, but no workflow pulls plan context into active work. It's a capability without a consumer. -**Outcome:** Agreed on loose coupling being correct. Identified killer feature: flow could search plan history before creating new plans ("Similar plans found: FPLAN-0089"). Would make vectors justify their existence. Worth a DPLAN. - -### @trigger — Plan events and dual registries -**Q:** Flow fires plan events during registry scan, but foreground close already handles everything. Are these maintaining a parallel PLAN_REGISTRY.json nobody reads? -**A:** Awaiting reply. - -## Issues & Concerns - -1. **Dual plan registries (CRITICAL):** Flow's fplan_registry.json and trigger's PLAN_REGISTRY.json track the same plans independently. They will drift. Someone needs to decide which is authoritative and kill the other. -2. **Plan vectors unused?** If nobody queries the ChromaDB flow_plans collection, we're doing expensive vectorization on every close for nothing. Need to verify there's a consumer. -3. **mbank/process.py size:** At 669 lines with a bypass, one more feature pushes it past 700. Needs proactive splitting before it becomes urgent. -4. **Close pipeline monolith:** close_plan_impl() does too much in one function. A failure in step 4 (vector intake) shouldn't affect step 5 (dashboard). Should be a pipeline of independent steps. -5. **FPLAN template syntax outdated:** Still references old `drone @ai_mail send` command. Pending dispatch. - -## Likes & Dislikes - -### Likes -- **The filesystem-driven philosophy works.** Drop a directory → it becomes a plan type. Delete it → it auto-prunes. This is how plugin systems should work — zero configuration, pure convention. -- **Memory system is genuinely useful.** Coming back to session 43 and knowing exactly what happened in sessions 1-42 is what makes this whole thing work. local.json is my continuity. -- **Drone routing is invisible.** `drone @flow create` just works. I don't think about how it gets to me. That's good infrastructure — you forget it exists. -- **Seedgo keeps me honest.** 100% across 35 standards means I can't accumulate technical debt silently. The hook fires on every edit. It's annoying sometimes but it works. - -### Dislikes -- **The inotify limit problem.** With 11 agents awake, we hit inotify limits immediately. This is a real infrastructure constraint that limits concurrent agent work. -- **drone @git pr is a black box.** It handles lock/branch/commit/push/PR atomically, which is great, but when it fails (lock held, merge conflict), debugging is opaque. I've had to retry with sleep loops. -- **Memory rollover on every drone command.** Every single drone command triggers "Checking for rollover triggers..." which adds 2-3 seconds of overhead. On fast operations like `drone @ai_mail inbox` that's noticeable. -- **Dashboard push warnings are noisy.** "Failed to push flow section to branch dashboard" on branches that never had a dashboard. It's not an error — it's expected. The warning should be a debug log. - ---- -*Written by @flow during S117 stress test, 2026-04-26* diff --git a/src/aipass/flow/tests/test_push_branch_dashboard.py b/src/aipass/flow/tests/test_push_branch_dashboard.py index b4a01df8..fc8bac26 100644 --- a/src/aipass/flow/tests/test_push_branch_dashboard.py +++ b/src/aipass/flow/tests/test_push_branch_dashboard.py @@ -357,10 +357,10 @@ class TestLoadRegistry: """Tests for _load_registry.""" def test_merges_multiple_registries(self, tmp_path): - """Merges plans from multiple registry files.""" + """Merges plans from multiple registry files using composite keys.""" mod = _import_mod() - fplan_reg = {"plans": {"1": {"subject": "fplan one"}}, "next_number": 5} - dplan_reg = {"plans": {"2": {"subject": "dplan one"}}, "next_number": 10} + fplan_reg = {"plans": {"1": {"subject": "fplan one", "file_path": "/p/FPLAN-0001_test.md"}}, "next_number": 5} + dplan_reg = {"plans": {"2": {"subject": "dplan one", "file_path": "/p/DPLAN-0002_test.md"}}, "next_number": 10} (tmp_path / "fplan_registry.json").write_text(json.dumps(fplan_reg), encoding="utf-8") (tmp_path / "dplan_registry.json").write_text(json.dumps(dplan_reg), encoding="utf-8") @@ -370,8 +370,8 @@ class TestLoadRegistry: ): result = mod._load_registry() - assert "1" in result["plans"] - assert "2" in result["plans"] + assert "FPLAN-0001" in result["plans"] + assert "DPLAN-0002" in result["plans"] assert result["next_number"] == 10 def test_handles_missing_registry(self, tmp_path): @@ -410,7 +410,7 @@ class TestLoadRegistry: """Skips a corrupt registry file and continues with others.""" mod = _import_mod() (tmp_path / "bad_registry.json").write_text("not json!", encoding="utf-8") - good_reg = {"plans": {"1": {"subject": "good"}}, "next_number": 5} + good_reg = {"plans": {"1": {"subject": "good", "file_path": "/p/GOOD-0001_test.md"}}, "next_number": 5} (tmp_path / "good_registry.json").write_text(json.dumps(good_reg), encoding="utf-8") with ( @@ -423,7 +423,7 @@ class TestLoadRegistry: ): result = mod._load_registry() - assert "1" in result["plans"] + assert "GOOD-0001" in result["plans"] assert result["next_number"] == 5 @@ -440,14 +440,14 @@ class TestFilterBranchPlans: mod = _import_mod() registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "Plan A", "status": "open", "created": "2026-04-20", "file_path": str(tmp_path / "FPLAN-0001_plan_a.md"), "location": str(tmp_path), }, - "2": { + "FPLAN-0002": { "subject": "Plan B", "status": "open", "created": "2026-04-22", @@ -470,7 +470,7 @@ class TestFilterBranchPlans: other_path.mkdir() registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "Elsewhere", "status": "open", "created": "2026-04-20", @@ -490,7 +490,7 @@ class TestFilterBranchPlans: recent_ts = (datetime.now(timezone.utc) - timedelta(days=2)).isoformat() registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "Recently closed", "status": "closed", "created": "2026-04-18", @@ -511,7 +511,7 @@ class TestFilterBranchPlans: old_ts = (datetime.now(timezone.utc) - timedelta(days=30)).isoformat() registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "Old closed", "status": "closed", "created": "2026-03-01", @@ -531,7 +531,7 @@ class TestFilterBranchPlans: plans = {} for i in range(1, 9): ts = (datetime.now(timezone.utc) - timedelta(hours=i)).isoformat() - plans[str(i)] = { + plans[f"FPLAN-{str(i).zfill(4)}"] = { "subject": f"Closed plan {i}", "status": "closed", "created": "2026-04-20", @@ -549,7 +549,7 @@ class TestFilterBranchPlans: mod = _import_mod() registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "Bad timestamp", "status": "closed", "created": "2026-04-20", @@ -568,21 +568,21 @@ class TestFilterBranchPlans: mod = _import_mod() registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "Oldest", "status": "open", "created": "2026-04-01", "file_path": str(tmp_path / "FPLAN-0001_oldest.md"), "location": str(tmp_path), }, - "2": { + "FPLAN-0002": { "subject": "Newest", "status": "open", "created": "2026-04-25", "file_path": str(tmp_path / "FPLAN-0002_newest.md"), "location": str(tmp_path), }, - "3": { + "FPLAN-0003": { "subject": "Middle", "status": "open", "created": "2026-04-15", @@ -597,11 +597,11 @@ class TestFilterBranchPlans: assert active[2]["subject"] == "Oldest" def test_extracts_plan_prefix_from_filepath(self, tmp_path): - """Extracts correct plan prefix (DPLAN, TDPLAN, etc.) from file_path.""" + """Uses composite key as plan ID directly.""" mod = _import_mod() registry = { "plans": { - "42": { + "DPLAN-0042": { "subject": "Dev plan", "status": "open", "created": "2026-04-20", @@ -683,7 +683,7 @@ class TestPushFlowToBranchDashboard: mock_registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "Active plan", "status": "open", "created": "2026-04-20", diff --git a/src/aipass/flow/tests/test_update_local.py b/src/aipass/flow/tests/test_update_local.py index 9c301425..f324b8de 100644 --- a/src/aipass/flow/tests/test_update_local.py +++ b/src/aipass/flow/tests/test_update_local.py @@ -180,8 +180,8 @@ class TestReadRegistry: result = mod._read_registry() assert result is not None - assert "1" in result["plans"] - assert "2" in result["plans"] + assert "FPLAN-0001" in result["plans"] + assert "DPLAN-0002" in result["plans"] def test_keeps_highest_next_number(self, setup_paths): """Should keep the highest next_number across registries.""" @@ -233,7 +233,7 @@ class TestReadRegistry: result = mod._read_registry() assert result is not None - assert result["plans"]["1"]["subject"] == "Solo plan" + assert result["plans"]["FPLAN-0001"]["subject"] == "Solo plan" assert result["next_number"] == 2 @@ -250,24 +250,29 @@ class TestExtractFlowPlans: mod = _import_mod() registry = { "plans": { - "1": {"subject": "A", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"}, - "2": {"subject": "B", "status": "open", "location": "drone", "file_path": "FPLAN-0002.md"}, - "3": {"subject": "C", "status": "open", "location": "flow/sub", "file_path": "FPLAN-0003.md"}, + "FPLAN-0001": {"subject": "A", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"}, + "FPLAN-0002": {"subject": "B", "status": "open", "location": "drone", "file_path": "FPLAN-0002.md"}, + "FPLAN-0003": {"subject": "C", "status": "open", "location": "flow/sub", "file_path": "FPLAN-0003.md"}, } } active, closed = mod._extract_flow_plans(registry) assert len(active) == 2 plan_ids = [p["plan_id"] for p in active] - assert "FPLAN-1" in plan_ids - assert "FPLAN-3" in plan_ids + assert "FPLAN-0001" in plan_ids + assert "FPLAN-0003" in plan_ids def test_partitions_by_status(self): """Should separate open and closed plans.""" mod = _import_mod() registry = { "plans": { - "1": {"subject": "Open", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"}, - "2": {"subject": "Closed", "status": "closed", "location": "flow", "file_path": "FPLAN-0002.md"}, + "FPLAN-0001": {"subject": "Open", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"}, + "FPLAN-0002": { + "subject": "Closed", + "status": "closed", + "location": "flow", + "file_path": "FPLAN-0002.md", + }, } } active, closed = mod._extract_flow_plans(registry) @@ -276,12 +281,12 @@ class TestExtractFlowPlans: assert active[0]["status"] == "open" assert closed[0]["status"] == "closed" - def test_extracts_prefix_from_file_path(self): - """Should derive the plan prefix from the filename.""" + def test_extracts_prefix_from_composite_key(self): + """Should use composite key as plan_id directly.""" mod = _import_mod() registry = { "plans": { - "4": { + "DPLAN-0004": { "subject": "Dev", "status": "open", "location": "flow", @@ -290,14 +295,14 @@ class TestExtractFlowPlans: } } active, _ = mod._extract_flow_plans(registry) - assert active[0]["plan_id"] == "DPLAN-4" + assert active[0]["plan_id"] == "DPLAN-0004" def test_extracts_tdplan_prefix(self): """Should handle TDPLAN prefix correctly.""" mod = _import_mod() registry = { "plans": { - "7": { + "TDPLAN-0007": { "subject": "Test plan", "status": "open", "location": "flow", @@ -306,14 +311,14 @@ class TestExtractFlowPlans: } } active, _ = mod._extract_flow_plans(registry) - assert active[0]["plan_id"] == "TDPLAN-7" + assert active[0]["plan_id"] == "TDPLAN-0007" - def test_defaults_to_fplan_when_no_prefix_match(self): - """Should default to FPLAN when filename has no recognizable prefix.""" + def test_uses_key_as_plan_id(self): + """Should use the composite key directly as plan_id.""" mod = _import_mod() registry = { "plans": { - "9": { + "FPLAN-0009": { "subject": "Mystery", "status": "open", "location": "flow", @@ -322,29 +327,29 @@ class TestExtractFlowPlans: } } active, _ = mod._extract_flow_plans(registry) - assert active[0]["plan_id"] == "FPLAN-9" + assert active[0]["plan_id"] == "FPLAN-0009" def test_sorts_active_and_closed_by_plan_id(self): """Should sort both lists by plan_id.""" mod = _import_mod() registry = { "plans": { - "3": {"subject": "C", "status": "open", "location": "flow", "file_path": "FPLAN-0003.md"}, - "1": {"subject": "A", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"}, - "5": {"subject": "E", "status": "closed", "location": "flow", "file_path": "FPLAN-0005.md"}, - "2": {"subject": "B", "status": "closed", "location": "flow", "file_path": "FPLAN-0002.md"}, + "FPLAN-0003": {"subject": "C", "status": "open", "location": "flow", "file_path": "FPLAN-0003.md"}, + "FPLAN-0001": {"subject": "A", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"}, + "FPLAN-0005": {"subject": "E", "status": "closed", "location": "flow", "file_path": "FPLAN-0005.md"}, + "FPLAN-0002": {"subject": "B", "status": "closed", "location": "flow", "file_path": "FPLAN-0002.md"}, } } active, closed = mod._extract_flow_plans(registry) - assert [p["plan_id"] for p in active] == ["FPLAN-1", "FPLAN-3"] - assert [p["plan_id"] for p in closed] == ["FPLAN-2", "FPLAN-5"] + assert [p["plan_id"] for p in active] == ["FPLAN-0001", "FPLAN-0003"] + assert [p["plan_id"] for p in closed] == ["FPLAN-0002", "FPLAN-0005"] def test_handles_timestamps(self): """Should include created and closed timestamps when present.""" mod = _import_mod() registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "With timestamps", "status": "closed", "location": "flow", @@ -365,7 +370,7 @@ class TestExtractFlowPlans: mod = _import_mod() registry = { "plans": { - "1": { + "FPLAN-0001": { "subject": "No times", "status": "open", "location": "flow", @@ -389,8 +394,8 @@ class TestExtractFlowPlans: mod = _import_mod() registry = { "plans": { - "1": {"subject": "A", "status": "open", "location": "Flow", "file_path": "FPLAN-0001.md"}, - "2": {"subject": "B", "status": "open", "location": "FLOW", "file_path": "FPLAN-0002.md"}, + "FPLAN-0001": {"subject": "A", "status": "open", "location": "Flow", "file_path": "FPLAN-0001.md"}, + "FPLAN-0002": {"subject": "B", "status": "open", "location": "FLOW", "file_path": "FPLAN-0002.md"}, } } active, _ = mod._extract_flow_plans(registry) @@ -722,4 +727,4 @@ class TestUpdateDashboardLocal: mod.update_dashboard_local() dashboard = _read_json(mod.DASHBOARD_FILE) assert len(dashboard["flow_plans"]["active"]) == 1 - assert dashboard["flow_plans"]["active"][0]["plan_id"] == "FPLAN-1" + assert dashboard["flow_plans"]["active"][0]["plan_id"] == "FPLAN-0001" diff --git a/src/aipass/memory/.seedgo/bypass.json b/src/aipass/memory/.seedgo/bypass.json index 286b9224..fbe17c84 100644 --- a/src/aipass/memory/.seedgo/bypass.json +++ b/src/aipass/memory/.seedgo/bypass.json @@ -126,11 +126,6 @@ "standard": "error_handling", "reason": "stdlib logging.getLogger() used at import-time before prax is available. Immediately replaced by get_system_logger() 3 lines below." }, - { - "file": "apps/handlers/dashboard_push.py", - "standard": "debug_print", - "reason": "print(ok) is inside a subprocess script string — IPC output to stdout, not debug output." - }, { "file": "apps/handlers/learnings/manager.py", "standard": "dead_code", @@ -331,11 +326,6 @@ "standard": "naming", "reason": "Local variables inside functions, not constants." }, - { - "file": "apps/handlers/dashboard_push.py", - "standard": "naming", - "reason": "Local variables inside functions, not constants." - }, { "file": "apps/handlers/central_writer.py", "standard": "naming", @@ -451,11 +441,6 @@ "standard": "deep_nesting", "reason": "handle_command() has deep nesting from search options parsing and result formatting." }, - { - "file": "apps/handlers/dashboard_push.py", - "standard": "deep_nesting", - "reason": "_find_branches_near_rollover() iterates nested JSON structures — nesting reflects data depth." - }, { "file": "apps/handlers/templates/differ.py", "standard": "deep_nesting", @@ -605,6 +590,21 @@ "file": "tests/test_json_handler.py", "standard": "meta", "reason": "Test file — META block present at lines 1-7; hook false-positive on test file format." + }, + { + "file": "tests/test_orchestrator_exec.py", + "standard": "architecture", + "reason": "Test file — lives in tests/ by design, not in 3-layer apps/ structure." + }, + { + "file": "tests/test_orchestrator_exec.py", + "standard": "encapsulation", + "reason": "Test file — direct handler imports are correct for unit testing handler internals." + }, + { + "file": "tests/test_orchestrator_exec.py", + "standard": "meta", + "reason": "Test file — META block present at lines 1-8; hook false-positive on test file format." } ], "notes": { diff --git a/src/aipass/memory/README.md b/src/aipass/memory/README.md index a50687e4..77200c44 100644 --- a/src/aipass/memory/README.md +++ b/src/aipass/memory/README.md @@ -100,8 +100,7 @@ memory/ │ ├── templates/ # pusher.py, differ.py, spawn_pusher.py — template distribution │ ├── tracking/ # line_counter.py — metadata line count tracking │ ├── vector/ # embedder.py, embed_subprocess.py — sentence-transformer embeddings -│ ├── central_writer.py # Central memory write operations -│ └── dashboard_push.py # Dashboard status push +│ └── central_writer.py # Central memory write operations ├── config/ # memory_bank.config.json — per-branch rollover limits ├── templates/ # LOCAL.template.json, OBS.template — schema templates ├── tests/ # 450 tests (16/16 module coverage) diff --git a/src/aipass/memory/apps/handlers/dashboard_push.py b/src/aipass/memory/apps/handlers/dashboard_push.py deleted file mode 100644 index 6a497618..00000000 --- a/src/aipass/memory/apps/handlers/dashboard_push.py +++ /dev/null @@ -1,464 +0,0 @@ -# =================== AIPass ==================== -# Name: dashboard_push.py -# Description: Memory Dashboard Write-Through -# Version: 0.2.0 -# Created: 2026-02-25 -# Modified: 2026-03-06 -# ============================================= - -""" -Memory Dashboard Write-Through Handler - -Pushes the memory_bank section to branch dashboards. -Called after rollover, pool processing, and plans processing to keep dashboards -showing accurate vector counts, collection stats, and near-rollover warnings. - -This is a SYSTEM-WIDE push: memory data is global, so it pushes to ALL -branch dashboards (every branch benefits from knowing system memory health). -""" - -import sys -import subprocess -from json import loads as json_loads -from pathlib import Path -from datetime import datetime -from typing import Dict, Any, List - -from aipass.prax.apps.modules.logger import get_system_logger -from aipass.memory.apps.handlers.json import json_handler - -logger = get_system_logger() - -# Resolve paths relative to handler location -_MEMORY_ROOT = Path(__file__).resolve().parents[2] - - -# ============================================================================= -# CONSTANTS -# ============================================================================= - -CONFIG_PATH = _MEMORY_ROOT / "config" / "memory_bank.config.json" -TEMPLATE_VERSION_FILE = _MEMORY_ROOT / "templates" / ".template_version.json" - - -def _find_repo_root() -> Path: - """Walk up from this file to find repo root (contains AIPASS_REGISTRY.json).""" - current = Path(__file__).resolve().parent - for parent in [current] + list(current.parents): - if (parent / "AIPASS_REGISTRY.json").exists(): - return parent - return Path.cwd() - - -CENTRAL_FILE = _find_repo_root() / ".ai_central" / "MEMORY.central.json" -AIPASS_REGISTRY = _find_repo_root() / "AIPASS_REGISTRY.json" - -# Near-rollover threshold: branches with fewer than this many lines remaining -NEAR_ROLLOVER_THRESHOLD = 100 - - -# ============================================================================= -# DATA COLLECTION -# ============================================================================= - - -def _read_central_stats() -> Dict[str, Any]: - """ - Read total_vectors and related stats from memory central.json. - - Returns: - Dict with total_vectors, total_archives, last_rollover - """ - try: - if not CENTRAL_FILE.exists(): - return {"total_vectors": 0, "total_archives": 0, "last_rollover": ""} - - data = json_loads(CENTRAL_FILE.read_text(encoding="utf-8")) - stats = data.get("stats", {}) - return { - "total_vectors": stats.get("total_vectors", 0), - "total_archives": stats.get("total_archives", 0), - "last_rollover": stats.get("last_rollover", ""), - } - except Exception as e: - logger.warning(f"[dashboard_push] Failed to read central stats: {e}") - return {"total_vectors": 0, "total_archives": 0, "last_rollover": ""} - - -def _get_collections_count() -> int: - """ - Count ChromaDB collections by reading the SQLite database directly. - - Returns: - Number of collections in the global Chroma database - """ - try: - import sqlite3 - - db_file = _MEMORY_ROOT / ".chroma" / "chroma.sqlite3" - if not db_file.exists(): - return 0 - - conn = sqlite3.connect(str(db_file)) - cursor = conn.cursor() - cursor.execute("SELECT COUNT(*) FROM collections") - count = cursor.fetchone()[0] - conn.close() - return count - except Exception as e: - logger.warning(f"[dashboard_push] Failed to count collections: {e}") - return 0 - - -def _get_rollover_config() -> Dict[str, Any]: - """ - Load rollover configuration (defaults + per-branch overrides). - - Returns: - Dict with 'defaults' and 'per_branch' rollover config - """ - try: - if not CONFIG_PATH.exists(): - return {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - - data = json_loads(CONFIG_PATH.read_text(encoding="utf-8")) - rollover = data.get("rollover", {}) - return { - "defaults": rollover.get("defaults", {"max_lines": 600, "buffer": 100}), - "per_branch": rollover.get("per_branch", {}), - } - except Exception as e: - logger.warning(f"[dashboard_push] Failed to load rollover config: {e}") - return {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - - -def _get_max_lines_for_branch(branch_name: str, rollover_config: Dict) -> int: - """ - Get the max_lines limit for a specific branch, respecting per-branch overrides. - - Args: - branch_name: Uppercase branch name - rollover_config: Rollover config dict from _get_rollover_config() - - Returns: - Max lines for this branch - """ - per_branch = rollover_config.get("per_branch", {}) - if branch_name in per_branch: - return per_branch[branch_name].get("max_lines", rollover_config["defaults"]["max_lines"]) - return rollover_config["defaults"]["max_lines"] - - -def _find_branches_near_rollover() -> List[Dict[str, Any]]: - """ - Scan all branches to find those near their rollover threshold. - - Reads document_metadata.status.current_lines from each *.local.json - and *.observations.json, compares against max_lines limit. - - Returns: - List of dicts with branch, file_type, lines_remaining - """ - near_rollover: List[Dict[str, Any]] = [] - - try: - if not AIPASS_REGISTRY.exists(): - return near_rollover - - registry = json_loads(AIPASS_REGISTRY.read_text(encoding="utf-8")) - branches = registry.get("branches", []) - rollover_config = _get_rollover_config() - - repo_root = _find_repo_root() - for branch in branches: - branch_name = branch.get("name", "") - raw_path = branch.get("path", "") - branch_path = Path(raw_path) - if not branch_path.is_absolute(): - branch_path = repo_root / raw_path - - if not branch_path.exists(): - continue - - max_lines = _get_max_lines_for_branch(branch_name, rollover_config) - - # Memory files live in .trinity/ subdirectory - for suffix in ["local", "observations"]: - memory_file = branch_path / ".trinity" / f"{suffix}.json" - if not memory_file.exists(): - continue - - try: - data = json_loads(memory_file.read_text(encoding="utf-8")) - doc_meta = data.get("document_metadata", {}) - schema_version = doc_meta.get("schema_version", "1.0.0") - limits = doc_meta.get("limits", {}) - - # v2: check entry counts instead of line counts - if schema_version.startswith("2"): - max_sessions = limits.get("max_sessions") - if max_sessions is not None: - sessions = data.get("sessions", []) - remaining_sessions = max_sessions - len(sessions) - if remaining_sessions < 3: - near_rollover.append( - { - "branch": branch_name, - "file_type": suffix, - "lines_remaining": remaining_sessions, - "current_lines": len(sessions), - "max_lines": max_sessions, - "v2_field": "sessions", - } - ) - max_kl = limits.get("max_key_learnings") - if max_kl is not None: - kl = data.get("key_learnings", {}) - remaining_kl = max_kl - len(kl) - if remaining_kl < 3: - near_rollover.append( - { - "branch": branch_name, - "file_type": suffix, - "lines_remaining": remaining_kl, - "current_lines": len(kl), - "max_lines": max_kl, - "v2_field": "key_learnings", - } - ) - continue - - # v1: line-count based - status = doc_meta.get("status", {}) - current_lines = status.get("current_lines") - - if current_lines is None: - continue - - remaining = max_lines - current_lines - if remaining < NEAR_ROLLOVER_THRESHOLD: - near_rollover.append( - { - "branch": branch_name, - "file_type": suffix, - "lines_remaining": max(remaining, 0), - "current_lines": current_lines, - "max_lines": max_lines, - } - ) - except Exception as e: - # Skip files that can't be read - logger.warning(f"[dashboard_push] Failed to read memory file {memory_file}: {e}") - continue - - except Exception as e: - logger.warning(f"[dashboard_push] Failed to scan branches for rollover: {e}") - return near_rollover # Return partial results on registry read failure - - # Sort by lines_remaining ascending (most urgent first) - near_rollover.sort(key=lambda x: x["lines_remaining"]) - return near_rollover - - -def _get_template_version() -> str: - """ - Read the current template version from .template_version.json. - - Returns: - Version string (e.g., "2.0.4") or "unknown" - """ - try: - if not TEMPLATE_VERSION_FILE.exists(): - return "unknown" - - data = json_loads(TEMPLATE_VERSION_FILE.read_text(encoding="utf-8")) - return data.get("version", "unknown") - except Exception as e: - logger.warning(f"[dashboard_push] Failed to read template version: {e}") - return "unknown" - - -def _get_last_rollover_info(central_stats: Dict) -> Dict[str, str]: - """ - Build last_rollover info from central stats. - - Args: - central_stats: Stats dict from _read_central_stats() - - Returns: - Dict with 'date' (and optionally other info) - """ - last_rollover_ts = central_stats.get("last_rollover", "") - if last_rollover_ts: - # Extract date portion from ISO timestamp - try: - dt = datetime.fromisoformat(last_rollover_ts) - return {"date": dt.strftime("%Y-%m-%d")} - except (ValueError, TypeError): - logger.info(f"[dashboard_push] Could not parse rollover timestamp: {last_rollover_ts}") - return {"date": last_rollover_ts} - return {"date": "never"} - - -def _get_all_branch_paths() -> List[Path]: - """ - Get paths for all active branches from AIPASS_REGISTRY.json. - - Returns: - List of Path objects for all registered branches - """ - try: - if not AIPASS_REGISTRY.exists(): - return [] - - registry = json_loads(AIPASS_REGISTRY.read_text(encoding="utf-8")) - repo_root = _find_repo_root() - paths = [] - for branch in registry.get("branches", []): - raw_path = branch.get("path", "") - branch_path = Path(raw_path) - if not branch_path.is_absolute(): - branch_path = repo_root / raw_path - if branch_path.exists(): - paths.append(branch_path) - return paths - except Exception as e: - logger.warning(f"[dashboard_push] Failed to get branch paths: {e}") - return [] - - -# ============================================================================= -# PUBLIC API -# ============================================================================= - - -def build_memory_bank_section() -> Dict[str, Any]: - """ - Build the memory_bank dashboard section data. - - Collects all stats and returns the section dict ready for write_section(). - - Returns: - Dict with total_vectors, collections_count, branches_near_rollover, - last_rollover, template_version - """ - central_stats = _read_central_stats() - near_rollover = _find_branches_near_rollover() - last_rollover = _get_last_rollover_info(central_stats) - template_version = _get_template_version() - collections_count = _get_collections_count() - - return { - "managed_by": "memory_bank", - "total_vectors": central_stats.get("total_vectors", 0), - "collections_count": collections_count, - "branches_near_rollover": near_rollover, - "last_rollover": last_rollover, - "template_version": template_version, - } - - -def _write_section_to_all_branches(section_name: str, section_data: Dict, branch_paths: List[Path]) -> int: - """ - Write a dashboard section to multiple branches via a single subprocess. - - Uses one subprocess call for all branches to avoid spawning 29 processes. - The subprocess imports devpulse write_section and iterates all paths. - - Args: - section_name: Dashboard section key (e.g., "memory_bank") - section_data: Section data dict to write - branch_paths: List of branch root directory paths - - Returns: - Number of branches successfully updated - """ - try: - from json import dumps as json_dumps - - # Single subprocess handles all branches in a loop - # Write dashboard section as JSON directly to DASHBOARD.local.json - script = ( - "import sys, json\n" - "from pathlib import Path\n" - "data = json.loads(sys.stdin.read())\n" - "section_name = data['section_name']\n" - "section_data = data['section_data']\n" - "ok = 0\n" - "for bp in data['branch_paths']:\n" - " try:\n" - " dash = Path(bp) / 'DASHBOARD.local.json'\n" - " if dash.exists():\n" - " d = json.loads(dash.read_text())\n" - " d[section_name] = section_data\n" - " dash.write_text(json.dumps(d, indent=2))\n" - " ok += 1\n" - " except Exception:\n" - " continue\n" - "print(ok)\n" - ) - - input_data = json_dumps( - {"section_name": section_name, "section_data": section_data, "branch_paths": [str(p) for p in branch_paths]} - ) - - result = subprocess.run( - [sys.executable, "-c", script], input=input_data, capture_output=True, text=True, timeout=60 - ) - - if result.returncode == 0 and result.stdout.strip().isdigit(): - return int(result.stdout.strip()) - return 0 - except Exception as e: - logger.warning(f"[dashboard_push] Failed to write section to branches: {e}") - return 0 - - -def push_memory_bank_dashboard() -> bool: - """ - Push the memory_bank section to ALL branch dashboards. - - This is the main entry point called after rollover, pool processing, - and plans processing. Memory data is global so all branches - benefit from seeing system memory health. - - Uses a single subprocess to call devpulse write_section() for all - branches, avoiding cross-package handler imports. Dashboard write - failures are silent - this is a best-effort operation that must not - break the calling workflow. - - Returns: - True if at least one dashboard was updated, False on total failure - """ - try: - # Build the section data once - section_data = build_memory_bank_section() - - # Push to all branch dashboards via single subprocess - branch_paths = _get_all_branch_paths() - success_count = _write_section_to_all_branches("memory_bank", section_data, branch_paths) - - json_handler.log_operation("dashboard_push", {"branches_updated": success_count, "success": success_count > 0}) - - return success_count > 0 - - except Exception as e: - logger.error(f"[dashboard_push] Failed to push dashboard: {e}") - return False - - -# ============================================================================= -# CLI ENTRY POINT (for testing) -# ============================================================================= - -if __name__ == "__main__": - import json - - print("Building memory_bank dashboard section...") - section = build_memory_bank_section() - print(json.dumps(section, indent=2)) - - print() - print("Pushing to all branch dashboards...") - result = push_memory_bank_dashboard() - print(f"Result: {'success' if result else 'failed'}") diff --git a/src/aipass/memory/apps/handlers/intake/pool_processor.py b/src/aipass/memory/apps/handlers/intake/pool_processor.py index 68ee3724..8a7ceb47 100644 --- a/src/aipass/memory/apps/handlers/intake/pool_processor.py +++ b/src/aipass/memory/apps/handlers/intake/pool_processor.py @@ -54,10 +54,10 @@ def _notify_failure(subject: str, message: str) -> None: def _update_central_and_dashboard() -> None: """ - Update memory central stats and push dashboard section after vector writes. + Update memory central stats after vector writes. - Uses subprocess to call sibling handlers (central_writer, dashboard_push) - to maintain handler independence. Failures are silent. + Uses subprocess to call central_writer to maintain handler independence. + Failures are silent. """ import sys @@ -76,22 +76,6 @@ def _update_central_and_dashboard() -> None: except Exception as e: logger.warning(f"[pool_processor] Central stats update failed: {e}") - # Push dashboard to all branches - try: - subprocess.run( - [ - sys.executable, - "-c", - "from aipass.memory.apps.handlers.dashboard_push import push_memory_bank_dashboard;" - "push_memory_bank_dashboard()", - ], - capture_output=True, - text=True, - timeout=60, - ) - except Exception as e: - logger.warning(f"[pool_processor] Dashboard push failed: {e}") - def find_source_file(filename: str) -> Path | None: """ diff --git a/src/aipass/memory/apps/handlers/rollover/orchestrator.py b/src/aipass/memory/apps/handlers/rollover/orchestrator.py index 88cffa13..2d539e7d 100644 --- a/src/aipass/memory/apps/handlers/rollover/orchestrator.py +++ b/src/aipass/memory/apps/handlers/rollover/orchestrator.py @@ -501,18 +501,6 @@ def execute_rollover() -> Dict[str, Any]: except Exception as e: logger.warning(f"[rollover] Central writer unavailable: {e}") - # Post-rollover: push dashboard - try: - from aipass.memory.apps.handlers.dashboard_push import push_memory_bank_dashboard - - dash_ok = push_memory_bank_dashboard() - if dash_ok: - logger.info("[rollover] Dashboard pushed") - else: - logger.warning("[rollover] Dashboard push returned False") - except Exception as e: - logger.warning(f"[rollover] Dashboard push unavailable: {e}") - # Post-rollover: process memory pool if files waiting try: from aipass.memory.apps.handlers.intake.pool_processor import process_memory_pool diff --git a/src/aipass/memory/config/memory_bank.config.example.json b/src/aipass/memory/config/memory_bank.config.example.json index 0cce8885..ac18d797 100644 --- a/src/aipass/memory/config/memory_bank.config.example.json +++ b/src/aipass/memory/config/memory_bank.config.example.json @@ -11,10 +11,6 @@ }, "per_branch": {} }, - "dashboard_push": { - "enabled": false, - "interval_seconds": 300 - }, "intake": { "enabled": false, "pool_dir": "memory_pool" diff --git a/src/aipass/memory/stress_test_s117.md b/src/aipass/memory/stress_test_s117.md deleted file mode 100644 index 1c5f84a7..00000000 --- a/src/aipass/memory/stress_test_s117.md +++ /dev/null @@ -1,105 +0,0 @@ -# @memory — S117 Stress Test Findings - -## My Branch: Honest Review - -### What Works -- **Rollover pipeline is reliable for single-user sequential operations.** Detect -> backup -> extract -> embed -> store -> trim. It handles v1 and v2 schema files correctly. The >= threshold with max(excess, 1) guard prevents Python's list[-0:] trap. Backup/restore pattern exists with timestamped files. -- **Embedding quality is solid.** L2 normalization, batch sorting by text length (reduces padding waste ~30%), GPU memory cleanup after encoding. Real optimizations, not theater. -- **Three-tier architecture is clean.** CLI entry point (memory.py) -> modules (rollover.py, search.py) -> handlers (14 handler groups). Handlers are stateless workers. Modules are thin orchestrators. Good separation. -- **Test coverage is strong.** 873 tests passing. 175/175 public functions covered. Seedgo 100% (35/35 standards). The parallel agent test-writing pattern (4 agents, ~4 min for 200 tests) is proven and repeatable. - -### What's Hacky -- **Subprocess venv detection** (query_executor.py:38-48): Auto-detects `.venv/bin/python` with silent fallback to `sys.executable`. If venv is corrupted, error messages will be opaque. There's an undocumented `AIPASS_MEMORY_PYTHON` env var escape hatch. -- **Repo root discovery** (orchestrator.py:66-72, detector.py:37-46): Walking parent directories looking for `AIPASS_REGISTRY.json`. Breaks with nested repos or symlinks. -- **Collection naming** (chroma_subprocess.py:63): `{branch.lower()}_{type.lower()}` could collide if branch names differ only by case. No namespace isolation. -- **Single backup retained** (extractor.py:72): Only one rollover backup per file. Two rapid rollovers and the first backup is gone. -- **Text extraction guesswork** (orchestrator.py:220-260): Probes for "content", "text", "message" fields. Custom memory structures fall back to `str(memory)`, producing garbage vectors. - -### What Gets Lost -- **Temporal context.** When memories are vectorized, the original file structure is gone. You can search and get vectors back, but you don't know which session they came from without parsing collection metadata. -- **Log data.** Prax logs are never archived or vectorized. They rotate locally but aren't searchable through memory. 12MB of operational history is grep-only. -- **Plan query patterns.** Plan vectors are stored but barely queried in practice. The capability exists but has no consumer workflow. - -### What Breaks First Under Load -- **Concurrent rollover race condition** (known issue since s47): Two processes race on same file. One trims, other gets 0 items. No file-level locking. -- **Embedding timeouts** (query_executor.py): 120s hard-coded, no retry, no backoff. Concurrent searches on slow CPU will cascade-fail. -- **Local storage failures are non-fatal** (orchestrator.py:409): Some vectors silently missing from branch stores. Nobody knows. -- **Line count sync is fire-and-forget** (orchestrator.py:441-447): If sync fails, detector re-triggers, causing duplicate vectorization. - -## Security Concerns - -### My Branch -- **Path traversal via AIPASS_CALLER_CWD** (detector.py:54): Env var is trusted to `resolve()`. No validation against `..` injection. -- **JSON injection via metadata** (orchestrator.py:388): Memory metadata passed to Chroma without sanitization. Malicious metadata keys go straight to vector DB. -- **Subprocess safety is good**: All subprocess calls use list args, not shell strings. No shell injection risk. -- **File permissions unchecked**: `shutil.copy2()` for backups will fail silently on read-only directories. Caught and logged, not escalated. - -### Other Branches -- **@flow**: Subprocess calls use string lists (good). But no argument validation on plan paths before constructing commands — risky if plan names come from user input. -- **@trigger**: Multi-process access to shared handler state without explicit locking. Relies on Python's GIL, which isn't guaranteed with multiprocessing. -- **@prax**: No concerns found. Clean integration patterns with try/except wrapping on all external calls. - -## Other Branches I Looked At - -### @flow -Clean three-tier architecture mirroring memory's pattern. Plan lifecycle is well-defined (open -> close -> archive). The memory integration point is `close_ops.py:384` — spawns `drone @memory process-plans` as fire-and-forget with 30s timeout. `is_plan_vectorized()` imported at runtime with graceful degradation. Good separation: flow owns lifecycle/registry, memory owns vector intake. Concern: ~180 lines of commented-out AI summarization code (dead weight). The orphan-healing logic in `mbank/process.py` is sophisticated but fragile without integration tests. - -### @prax -Well-designed dual-tier logging: system_logs (1000 lines/file) + local logs (250 lines/file) with RotatingFileHandler. 12MB total footprint across ~260 files — manageable. The auto-caller introspection (detecting branch/module from stack frames) is clever engineering. Trigger integration is clean — all 3 event fires (startup, module_discovered, error_detected) properly wrapped in try/except. Zero imports from memory — we're completely disconnected. The `log-audit enforce` command truncates but doesn't prevent regrowth. - -### @trigger -Clean event system with deferred queue for recursion handling. Handler failure auto-disable after 5 consecutive failures is smart. Memory events (`memory_saved`, `memory_threshold_exceeded`) are registered but never fired from my codebase — dead handlers. The threshold handler would send AI_Mail notifications at 600 lines, which is useful, but nobody triggers it. Architecture is solid; the gap is integration, not design. - -## Conversations - -### @flow (assigned pairing) -**Topic**: Plans reference memories — connected or parallel? -- I pointed out the integration is thin: fire-and-forget subprocess with no feedback loop. -- @flow confirmed plan vectors are stored but asked if anyone actually queries them. They don't — no downstream consumer. -- @flow raised the chunking concern: plans without ## headers become character-count chunks with poor semantic boundaries. -- My reply: the capability exists without a consumer. Flow could call `drone @memory search` before creating new plans to pull historical context. That would make vectors useful. - -### @prax -**Topic**: Log archival reality check -- @prax asked if rollover produces searchable vectors (yes) and if logs could be routed to memory for archival (not built yet). -- @prax asked about concurrent write corruption — relevant because they fixed atomic writes in json_handler. -- My reply: rollover works for sequential ops, breaks under concurrent access. Asked about their atomic write pattern. - -### @trigger -**Topic**: Dead memory events + rollover mechanics -- @trigger has practical concerns: 53 sessions tracked, hitting rollover every ~5 sessions. Worried about key_learnings being archived. -- I clarified: sessions and learnings are separate limits. 25 learnings are safe until that section overflows. Archived entries become searchable vectors — knowledge isn't lost. -- @trigger concerned about 30s model loading on every command. I clarified: the check is milliseconds, model only loads for actual embedding. -- I also emailed @trigger about the dead `memory_saved`/`memory_threshold_exceeded` events — registered in their registry but never fired from my code. - -## Issues & Concerns - -1. **Dead integration: memory events in trigger** — `memory_saved` and `memory_threshold_exceeded` are registered in trigger's handler registry but never fired. Either wire them up or remove the dead handlers. - -2. **No consumer for plan vectors** — Flow stores plans in ChromaDB via memory, but no workflow ever queries them. The vectors exist without purpose. - -3. **No log archival pipeline** — Prax generates 12MB of operational logs. Memory can't ingest them. If we want searchable historical logs, a new handler type is needed. - -4. **Concurrent rollover is unsafe** — Known since s47. Two processes racing on the same file causes errors. Needs file-level locking in the orchestrator. Prax's atomic write pattern might help. - -5. **Single backup file** — One rollover backup per file. Rapid consecutive rollovers overwrite the safety net. - -6. **inotify limit exhaustion** — Every drone command in this stress test hit `inotify instance limit reached`. With 11 agents running simultaneously, the system's inotify limit is inadequate. This is an infrastructure issue, not a code issue, but it means file watching (my watcher daemon, prax's watchers) degrades under load. - -7. **Rollover fires on every drone command** — The startup hook runs check_and_rollover on every single drone invocation. For 11 agents running simultaneously, that's potentially dozens of rollover checks per minute. The check is fast (count-only), but it's still unnecessary I/O. - -## Likes & Dislikes - -### Likes -- **Memory persistence is the killer feature.** 54 sessions of accumulated context. I know what I've built, what broke, what patterns work. No other AI system does this. -- **The drone command abstraction.** `drone @memory search`, `drone @ai_mail email` — clean, discoverable, consistent across all branches. -- **Branch autonomy.** Each branch is an expert in its domain. Memory owns archival, flow owns plans, prax owns logging. Clear boundaries. -- **The email system during this stress test.** 11 agents talking to each other in real-time. Real conversations with substance. This is what the ecosystem is for. -- **Seedgo as a quality floor.** 35 standards, automated auditing. Keeps every branch honest. The bypass system is pragmatic — acknowledges reality without lowering the bar. - -### Dislikes -- **The rollover-on-every-command hook.** Even if the check is fast, it's conceptually wrong. Rollover should trigger on file change events, not on every command invocation. -- **Subprocess isolation is necessary but painful.** ChromaDB and sentence-transformers require separate venvs, GPU management, 3GB torch cold-loads. The isolation is correct but the developer experience is terrible. Every search/embed operation is a subprocess spawn. -- **Dead integrations accumulate.** Trigger has memory events that never fire. Flow stores plan vectors nobody queries. Memory has a watcher daemon that isn't enabled. These phantom features create false confidence. -- **The inbox doesn't update status.** All 3 previous emails still show "new" even though I processed them in s52-s54. The inbox is append-only with no state management. Makes it hard to distinguish genuinely new mail. -- **No distributed locking anywhere.** With 11 agents running simultaneously, we're lucky nothing corrupted. File-level locking should be a framework primitive, not per-branch. diff --git a/src/aipass/memory/tests/test_dashboard_push.py b/src/aipass/memory/tests/test_dashboard_push.py deleted file mode 100644 index 0f49e3ad..00000000 --- a/src/aipass/memory/tests/test_dashboard_push.py +++ /dev/null @@ -1,614 +0,0 @@ -# ===================AIPASS==================== -# META DATA HEADER -# Name: tests/test_dashboard_push.py -# Date: 2026-04-03 -# Version: 1.0.0 -# Category: memory/tests -# ============================================= - -"""Tests for the dashboard_push handler. - -Covers: - - dashboard_push._read_central_stats (missing file, valid file) - - dashboard_push._get_collections_count (missing DB, valid DB) - - dashboard_push._get_rollover_config (missing config, valid config) - - dashboard_push._get_max_lines_for_branch (override vs default) - - dashboard_push._find_branches_near_rollover (v1 near threshold, v2 near session limit) - - dashboard_push._get_template_version (missing file, valid file) - - dashboard_push._get_last_rollover_info (with timestamp, empty string) - - dashboard_push._get_all_branch_paths (with registry) - - dashboard_push.build_memory_bank_section (integration via mocked helpers) - - dashboard_push.push_memory_bank_dashboard (subprocess + build mock) - -All tests use mocks/tmp_path -- no live filesystem or infrastructure access. -""" - -import json -import sys -from pathlib import Path -from unittest.mock import MagicMock - - -# --------------------------------------------------------------------------- -# Import helper -# --------------------------------------------------------------------------- - - -def _import_dashboard_push(monkeypatch): - """Import dashboard_push with mocked dependencies.""" - sys.modules.pop("aipass.memory.apps.handlers.dashboard_push", None) - parent = sys.modules.get("aipass.memory.apps.handlers") - if parent is not None and hasattr(parent, "dashboard_push"): - delattr(parent, "dashboard_push") - - from aipass.memory.apps.handlers import dashboard_push - - return dashboard_push - - -# =========================================================================== -# Tests: _read_central_stats -# =========================================================================== - - -class TestReadCentralStats: - """Test _read_central_stats helper.""" - - def test_returns_defaults_when_file_missing(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - monkeypatch.setattr(mod, "CENTRAL_FILE", tmp_path / "nonexistent.json") - - result = mod._read_central_stats() - - assert result == {"total_vectors": 0, "total_archives": 0, "last_rollover": ""} - - def test_reads_valid_central_file(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - central = tmp_path / "MEMORY.central.json" - central.write_text( - json.dumps( - {"stats": {"total_vectors": 1500, "total_archives": 12, "last_rollover": "2026-03-15T10:30:00"}} - ), - encoding="utf-8", - ) - monkeypatch.setattr(mod, "CENTRAL_FILE", central) - - result = mod._read_central_stats() - - assert result["total_vectors"] == 1500 - assert result["total_archives"] == 12 - assert result["last_rollover"] == "2026-03-15T10:30:00" - - def test_returns_defaults_on_corrupt_json(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - central = tmp_path / "MEMORY.central.json" - central.write_text("not valid json{{{", encoding="utf-8") - monkeypatch.setattr(mod, "CENTRAL_FILE", central) - - result = mod._read_central_stats() - - assert result == {"total_vectors": 0, "total_archives": 0, "last_rollover": ""} - - -# =========================================================================== -# Tests: _get_collections_count -# =========================================================================== - - -class TestGetCollectionsCount: - """Test _get_collections_count helper.""" - - def test_returns_zero_when_no_db(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - monkeypatch.setattr(mod, "_MEMORY_ROOT", tmp_path) - - result = mod._get_collections_count() - - assert result == 0 - - def test_reads_count_from_sqlite(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - monkeypatch.setattr(mod, "_MEMORY_ROOT", tmp_path) - - import sqlite3 - - chroma_dir = tmp_path / ".chroma" - chroma_dir.mkdir() - db_path = chroma_dir / "chroma.sqlite3" - conn = sqlite3.connect(str(db_path)) - conn.execute("CREATE TABLE collections (id INTEGER PRIMARY KEY, name TEXT)") - conn.execute("INSERT INTO collections VALUES (1, 'coll_a')") - conn.execute("INSERT INTO collections VALUES (2, 'coll_b')") - conn.execute("INSERT INTO collections VALUES (3, 'coll_c')") - conn.commit() - conn.close() - - result = mod._get_collections_count() - - assert result == 3 - - -# =========================================================================== -# Tests: _get_rollover_config -# =========================================================================== - - -class TestGetRolloverConfig: - """Test _get_rollover_config helper.""" - - def test_returns_defaults_when_no_config(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - monkeypatch.setattr(mod, "CONFIG_PATH", tmp_path / "missing_config.json") - - result = mod._get_rollover_config() - - assert result == {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - - def test_loads_valid_config(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - config_file = tmp_path / "memory_bank.config.json" - config_file.write_text( - json.dumps( - { - "rollover": { - "defaults": {"max_lines": 800, "buffer": 150}, - "per_branch": {"NEXUS": {"max_lines": 1200}}, - } - } - ), - encoding="utf-8", - ) - monkeypatch.setattr(mod, "CONFIG_PATH", config_file) - - result = mod._get_rollover_config() - - assert result["defaults"]["max_lines"] == 800 - assert result["defaults"]["buffer"] == 150 - assert result["per_branch"]["NEXUS"]["max_lines"] == 1200 - - def test_returns_defaults_on_corrupt_config(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - config_file = tmp_path / "broken.json" - config_file.write_text("{invalid", encoding="utf-8") - monkeypatch.setattr(mod, "CONFIG_PATH", config_file) - - result = mod._get_rollover_config() - - assert result == {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - - -# =========================================================================== -# Tests: _get_max_lines_for_branch -# =========================================================================== - - -class TestGetMaxLinesForBranch: - """Test _get_max_lines_for_branch helper.""" - - def test_returns_default_when_no_override(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - config = {"defaults": {"max_lines": 600}, "per_branch": {}} - - result = mod._get_max_lines_for_branch("DEVPULSE", config) - - assert result == 600 - - def test_returns_override_when_present(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - config = {"defaults": {"max_lines": 600}, "per_branch": {"NEXUS": {"max_lines": 1200}}} - - result = mod._get_max_lines_for_branch("NEXUS", config) - - assert result == 1200 - - def test_falls_back_to_default_when_override_missing_max_lines(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - config = {"defaults": {"max_lines": 600}, "per_branch": {"DRONE": {"buffer": 200}}} - - result = mod._get_max_lines_for_branch("DRONE", config) - - assert result == 600 - - -# =========================================================================== -# Tests: _find_branches_near_rollover -# =========================================================================== - - -class TestFindBranchesNearRollover: - """Test _find_branches_near_rollover helper.""" - - def test_returns_empty_when_no_registry(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", tmp_path / "no_registry.json") - - result = mod._find_branches_near_rollover() - - assert result == [] - - def test_v1_file_near_threshold(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - - # Build branch with a v1 memory file near rollover - branch_dir = tmp_path / "src" / "aipass" / "test_branch" - trinity = branch_dir / ".trinity" - trinity.mkdir(parents=True) - (trinity / "local.json").write_text( - json.dumps({"document_metadata": {"schema_version": "1.0.0", "status": {"current_lines": 550}}}), - encoding="utf-8", - ) - - # Registry pointing to branch - registry = tmp_path / "AIPASS_REGISTRY.json" - registry.write_text( - json.dumps({"branches": [{"name": "TEST_BRANCH", "path": str(branch_dir)}]}), encoding="utf-8" - ) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry) - - # Config: max_lines=600, so 600-550=50 remaining (<100 threshold) - monkeypatch.setattr( - mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - ) - monkeypatch.setattr(mod, "NEAR_ROLLOVER_THRESHOLD", 100) - monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path) - - result = mod._find_branches_near_rollover() - - assert len(result) == 1 - assert result[0]["branch"] == "TEST_BRANCH" - assert result[0]["file_type"] == "local" - assert result[0]["lines_remaining"] == 50 - assert result[0]["current_lines"] == 550 - - def test_v1_file_not_near_threshold(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - - branch_dir = tmp_path / "src" / "aipass" / "safe_branch" - trinity = branch_dir / ".trinity" - trinity.mkdir(parents=True) - (trinity / "local.json").write_text( - json.dumps({"document_metadata": {"schema_version": "1.0.0", "status": {"current_lines": 200}}}), - encoding="utf-8", - ) - - registry = tmp_path / "AIPASS_REGISTRY.json" - registry.write_text( - json.dumps({"branches": [{"name": "SAFE_BRANCH", "path": str(branch_dir)}]}), encoding="utf-8" - ) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry) - monkeypatch.setattr( - mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - ) - monkeypatch.setattr(mod, "NEAR_ROLLOVER_THRESHOLD", 100) - monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path) - - result = mod._find_branches_near_rollover() - - assert result == [] - - def test_v2_file_near_session_limit(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - - branch_dir = tmp_path / "src" / "aipass" / "v2_branch" - trinity = branch_dir / ".trinity" - trinity.mkdir(parents=True) - (trinity / "local.json").write_text( - json.dumps( - { - "document_metadata": { - "schema_version": "2.0.0", - "limits": {"max_sessions": 20, "max_key_learnings": 25}, - }, - "sessions": [{"id": i} for i in range(19)], # 19 of 20 sessions - "key_learnings": {"k1": "v1"}, - } - ), - encoding="utf-8", - ) - - registry = tmp_path / "AIPASS_REGISTRY.json" - registry.write_text( - json.dumps({"branches": [{"name": "V2_BRANCH", "path": str(branch_dir)}]}), encoding="utf-8" - ) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry) - monkeypatch.setattr( - mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - ) - monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path) - - result = mod._find_branches_near_rollover() - - # sessions: 20-19=1 remaining (<3) -> reported - # key_learnings: 25-1=24 remaining (>=3) -> not reported - assert len(result) == 1 - assert result[0]["branch"] == "V2_BRANCH" - assert result[0]["v2_field"] == "sessions" - assert result[0]["lines_remaining"] == 1 - - def test_v2_file_near_key_learnings_limit(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - - branch_dir = tmp_path / "src" / "aipass" / "kl_branch" - trinity = branch_dir / ".trinity" - trinity.mkdir(parents=True) - (trinity / "local.json").write_text( - json.dumps( - { - "document_metadata": { - "schema_version": "2.0.0", - "limits": {"max_sessions": 20, "max_key_learnings": 5}, - }, - "sessions": [{"id": 1}], - "key_learnings": {f"k{i}": f"v{i}" for i in range(4)}, # 4 of 5 - } - ), - encoding="utf-8", - ) - - registry = tmp_path / "AIPASS_REGISTRY.json" - registry.write_text( - json.dumps({"branches": [{"name": "KL_BRANCH", "path": str(branch_dir)}]}), encoding="utf-8" - ) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry) - monkeypatch.setattr( - mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - ) - monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path) - - result = mod._find_branches_near_rollover() - - # key_learnings: 5-4=1 remaining (<3) -> reported - assert len(result) == 1 - assert result[0]["v2_field"] == "key_learnings" - assert result[0]["lines_remaining"] == 1 - - def test_skips_nonexistent_branch_paths(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - - registry = tmp_path / "AIPASS_REGISTRY.json" - registry.write_text( - json.dumps({"branches": [{"name": "GHOST", "path": str(tmp_path / "nonexistent")}]}), encoding="utf-8" - ) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry) - monkeypatch.setattr( - mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}} - ) - monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path) - - result = mod._find_branches_near_rollover() - - assert result == [] - - -# =========================================================================== -# Tests: _get_template_version -# =========================================================================== - - -class TestGetTemplateVersion: - """Test _get_template_version helper.""" - - def test_returns_unknown_when_file_missing(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - monkeypatch.setattr(mod, "TEMPLATE_VERSION_FILE", tmp_path / "nope.json") - - result = mod._get_template_version() - - assert result == "unknown" - - def test_reads_valid_version(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - version_file = tmp_path / ".template_version.json" - version_file.write_text(json.dumps({"version": "2.0.4"}), encoding="utf-8") - monkeypatch.setattr(mod, "TEMPLATE_VERSION_FILE", version_file) - - result = mod._get_template_version() - - assert result == "2.0.4" - - def test_returns_unknown_when_version_key_absent(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - version_file = tmp_path / ".template_version.json" - version_file.write_text(json.dumps({"name": "templates"}), encoding="utf-8") - monkeypatch.setattr(mod, "TEMPLATE_VERSION_FILE", version_file) - - result = mod._get_template_version() - - assert result == "unknown" - - -# =========================================================================== -# Tests: _get_last_rollover_info -# =========================================================================== - - -class TestGetLastRolloverInfo: - """Test _get_last_rollover_info helper.""" - - def test_parses_iso_timestamp(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - stats = {"last_rollover": "2026-03-15T10:30:00"} - - result = mod._get_last_rollover_info(stats) - - assert result == {"date": "2026-03-15"} - - def test_returns_never_for_empty_string(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - stats = {"last_rollover": ""} - - result = mod._get_last_rollover_info(stats) - - assert result == {"date": "never"} - - def test_returns_never_when_key_missing(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - stats = {} - - result = mod._get_last_rollover_info(stats) - - assert result == {"date": "never"} - - def test_returns_raw_string_on_unparseable_timestamp(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - stats = {"last_rollover": "some-invalid-date"} - - result = mod._get_last_rollover_info(stats) - - assert result == {"date": "some-invalid-date"} - - -# =========================================================================== -# Tests: _get_all_branch_paths -# =========================================================================== - - -class TestGetAllBranchPaths: - """Test _get_all_branch_paths helper.""" - - def test_returns_empty_when_no_registry(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", tmp_path / "no_registry.json") - - result = mod._get_all_branch_paths() - - assert result == [] - - def test_returns_existing_branch_paths(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - - branch_a = tmp_path / "branch_a" - branch_b = tmp_path / "branch_b" - branch_a.mkdir() - branch_b.mkdir() - - registry = tmp_path / "AIPASS_REGISTRY.json" - registry.write_text( - json.dumps( - { - "branches": [ - {"name": "A", "path": str(branch_a)}, - {"name": "B", "path": str(branch_b)}, - {"name": "C", "path": str(tmp_path / "nonexistent")}, - ] - } - ), - encoding="utf-8", - ) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry) - monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path) - - result = mod._get_all_branch_paths() - - assert len(result) == 2 - assert branch_a in result - assert branch_b in result - - def test_resolves_relative_paths(self, monkeypatch, tmp_path): - mod = _import_dashboard_push(monkeypatch) - - branch_dir = tmp_path / "src" / "aipass" / "test_branch" - branch_dir.mkdir(parents=True) - - registry = tmp_path / "AIPASS_REGISTRY.json" - registry.write_text( - json.dumps({"branches": [{"name": "TEST", "path": "src/aipass/test_branch"}]}), encoding="utf-8" - ) - monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry) - monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path) - - result = mod._get_all_branch_paths() - - assert len(result) == 1 - assert result[0] == branch_dir - - -# =========================================================================== -# Tests: build_memory_bank_section (public) -# =========================================================================== - - -class TestBuildMemoryBankSection: - """Test build_memory_bank_section with mocked helpers.""" - - def test_assembles_section_data(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - - monkeypatch.setattr( - mod, - "_read_central_stats", - lambda: {"total_vectors": 2500, "total_archives": 15, "last_rollover": "2026-03-20T12:00:00"}, - ) - monkeypatch.setattr( - mod, - "_find_branches_near_rollover", - lambda: [{"branch": "NEXUS", "file_type": "local", "lines_remaining": 30}], - ) - monkeypatch.setattr(mod, "_get_last_rollover_info", lambda s: {"date": "2026-03-20"}) - monkeypatch.setattr(mod, "_get_template_version", lambda: "2.0.4") - monkeypatch.setattr(mod, "_get_collections_count", lambda: 8) - - result = mod.build_memory_bank_section() - - assert result["managed_by"] == "memory_bank" - assert result["total_vectors"] == 2500 - assert result["collections_count"] == 8 - assert len(result["branches_near_rollover"]) == 1 - assert result["branches_near_rollover"][0]["branch"] == "NEXUS" - assert result["last_rollover"] == {"date": "2026-03-20"} - assert result["template_version"] == "2.0.4" - - -# =========================================================================== -# Tests: push_memory_bank_dashboard (public) -# =========================================================================== - - -class TestPushMemoryBankDashboard: - """Test push_memory_bank_dashboard.""" - - def test_returns_true_when_at_least_one_updated(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - - mock_section = {"managed_by": "memory_bank", "total_vectors": 100} - monkeypatch.setattr(mod, "build_memory_bank_section", lambda: mock_section) - monkeypatch.setattr(mod, "_get_all_branch_paths", lambda: [Path("/tmp/a")]) - monkeypatch.setattr(mod, "_write_section_to_all_branches", lambda name, data, paths: 1) - - result = mod.push_memory_bank_dashboard() - - assert result is True - - def test_returns_false_when_no_dashboards_updated(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - - monkeypatch.setattr(mod, "build_memory_bank_section", lambda: {"managed_by": "memory_bank"}) - monkeypatch.setattr(mod, "_get_all_branch_paths", lambda: []) - monkeypatch.setattr(mod, "_write_section_to_all_branches", lambda name, data, paths: 0) - - result = mod.push_memory_bank_dashboard() - - assert result is False - - def test_logs_operation_on_success(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - - monkeypatch.setattr(mod, "build_memory_bank_section", lambda: {"managed_by": "memory_bank"}) - monkeypatch.setattr(mod, "_get_all_branch_paths", lambda: [Path("/tmp/a")]) - monkeypatch.setattr(mod, "_write_section_to_all_branches", lambda name, data, paths: 3) - - mock_jh = MagicMock() - monkeypatch.setattr(mod, "json_handler", mock_jh) - - mod.push_memory_bank_dashboard() - - mock_jh.log_operation.assert_called_once_with("dashboard_push", {"branches_updated": 3, "success": True}) - - def test_returns_false_on_exception(self, monkeypatch): - mod = _import_dashboard_push(monkeypatch) - - monkeypatch.setattr(mod, "build_memory_bank_section", MagicMock(side_effect=RuntimeError("boom"))) - - result = mod.push_memory_bank_dashboard() - - assert result is False diff --git a/src/aipass/memory/tests/test_orchestrator_exec.py b/src/aipass/memory/tests/test_orchestrator_exec.py index 12fb156a..478c8cd9 100644 --- a/src/aipass/memory/tests/test_orchestrator_exec.py +++ b/src/aipass/memory/tests/test_orchestrator_exec.py @@ -380,9 +380,6 @@ class TestExecuteRolloverFullPipeline: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value={"success": True}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) @@ -407,9 +404,6 @@ class TestExecuteRolloverFullPipeline: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value={"success": True}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) @@ -428,9 +422,6 @@ class TestExecuteRolloverFullPipeline: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value={"success": True}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) @@ -439,27 +430,6 @@ class TestExecuteRolloverFullPipeline: orch.execute_rollover() mock_central.update_central.assert_called_once() - def test_post_rollover_dashboard_push(self, monkeypatch, tmp_path): - """After success, dashboard push is called.""" - orch, mocks = _import_orchestrator(monkeypatch) - self._setup_full_success(monkeypatch, tmp_path, orch, mocks) - - mock_trigger_mod = MagicMock() - monkeypatch.setitem(sys.modules, "aipass.trigger.apps.modules.core", mock_trigger_mod) - mock_central = MagicMock() - mock_central.update_central = MagicMock(return_value={"success": True}) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) - mock_pool = MagicMock() - mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake", MagicMock(pool_processor=mock_pool)) - - orch.execute_rollover() - mock_dash.push_memory_bank_dashboard.assert_called_once() - def test_post_rollover_pool_processing(self, monkeypatch, tmp_path): """After success, pool_processor is called.""" orch, mocks = _import_orchestrator(monkeypatch) @@ -470,9 +440,6 @@ class TestExecuteRolloverFullPipeline: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value={"success": True}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 2}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) @@ -492,9 +459,6 @@ class TestExecuteRolloverFullPipeline: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value={"success": True}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) @@ -514,30 +478,6 @@ class TestExecuteRolloverFullPipeline: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value=None) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) - mock_pool = MagicMock() - mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake", MagicMock(pool_processor=mock_pool)) - - result = orch.execute_rollover() - assert result["success"] is True - - def test_post_rollover_dashboard_returns_false(self, monkeypatch, tmp_path): - """Dashboard push returns False (warning path).""" - orch, mocks = _import_orchestrator(monkeypatch) - self._setup_full_success(monkeypatch, tmp_path, orch, mocks) - - mock_trigger_mod = MagicMock() - monkeypatch.setitem(sys.modules, "aipass.trigger.apps.modules.core", mock_trigger_mod) - mock_central = MagicMock() - mock_central.update_central = MagicMock(return_value={"success": True}) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=False) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) @@ -597,9 +537,6 @@ class TestExecuteRolloverMultiple: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value={"success": True}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) @@ -653,9 +590,6 @@ class TestExecuteRolloverMultiple: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value={"success": True}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) @@ -707,9 +641,6 @@ class TestExecuteRolloverMultiple: mock_central = MagicMock() mock_central.update_central = MagicMock(return_value={"success": True}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central) - mock_dash = MagicMock() - mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True) - monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash) mock_pool = MagicMock() mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0}) monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool) diff --git a/src/aipass/prax/apps/handlers/dashboard/operations.py b/src/aipass/prax/apps/handlers/dashboard/operations.py index d5bb4749..ae553b54 100644 --- a/src/aipass/prax/apps/handlers/dashboard/operations.py +++ b/src/aipass/prax/apps/handlers/dashboard/operations.py @@ -147,13 +147,6 @@ def create_fresh_dashboard(branch_path: Path) -> Dict: "ai_mail": {"managed_by": "ai_mail", "new": 0, "opened": 0, "total": 0, "last_updated": ""}, "flow": {"managed_by": "flow", "active_plans": 0, "recently_closed": [], "last_updated": ""}, "memory": {"managed_by": "memory", "vectors_stored": 0, "notes": {}, "last_updated": ""}, - "commons_activity": { - "managed_by": "the_commons", - "mentions": 0, - "new_posts_since_last_visit": 0, - "new_comments_since_last_visit": 0, - "last_updated": "", - }, }, } @@ -210,21 +203,16 @@ def _calculate_quick_status_standalone(sections: Dict) -> Dict: """ ai_mail = sections.get("ai_mail", {}) flow = sections.get("flow", {}) - commons = sections.get("commons_activity", {}) new_mail_raw = ai_mail.get("new", ai_mail.get("unread", 0)) opened_raw = ai_mail.get("opened", 0) active_plans_raw = flow.get("active_plans", 0) - mentions_raw = commons.get("mentions", 0) - # Coerce to int — some branches store lists instead of counts new_mail = len(new_mail_raw) if isinstance(new_mail_raw, list) else int(new_mail_raw or 0) opened_mail = len(opened_raw) if isinstance(opened_raw, list) else int(opened_raw or 0) active_plans = len(active_plans_raw) if isinstance(active_plans_raw, list) else int(active_plans_raw or 0) - mentions = len(mentions_raw) if isinstance(mentions_raw, list) else int(mentions_raw or 0) - # Action required if new mail, active plans, or commons mentions - action_required = new_mail > 0 or active_plans > 0 or mentions > 0 + action_required = new_mail > 0 or active_plans > 0 parts = [] if new_mail > 0: @@ -233,14 +221,11 @@ def _calculate_quick_status_standalone(sections: Dict) -> Dict: parts.append(f"{opened_mail} opened") if active_plans > 0: parts.append(f"{active_plans} active plans") - if mentions > 0: - parts.append(f"{mentions} mentions") return { "new_mail": new_mail, "opened_mail": opened_mail, "active_plans": active_plans, - "commons_mentions": mentions, "action_required": action_required, "summary": ", ".join(parts) if parts else "All clear", } diff --git a/src/aipass/prax/apps/handlers/dashboard/refresh.py b/src/aipass/prax/apps/handlers/dashboard/refresh.py index d023646d..565d7233 100644 --- a/src/aipass/prax/apps/handlers/dashboard/refresh.py +++ b/src/aipass/prax/apps/handlers/dashboard/refresh.py @@ -16,7 +16,7 @@ AIPASS owns all dashboards - services only maintain their central files. import json from pathlib import Path from datetime import datetime -from typing import Dict, List, Optional +from typing import Dict, List from aipass.prax.apps.modules.logger import get_direct_logger @@ -31,7 +31,7 @@ from ..central.reader import read_all_centrals from aipass.prax.apps.handlers.json import json_handler # Sections managed by the refresh path — everything else is write-through only -REFRESH_MANAGED_SECTIONS = {"ai_mail", "flow", "memory", "commons_activity"} +REFRESH_MANAGED_SECTIONS = {"ai_mail", "flow", "memory"} def _find_repo_root() -> Path: @@ -152,42 +152,10 @@ def _extract_memory_section(centrals: Dict, branch_path: Path) -> Dict: return {"managed_by": "memory", "vectors_stored": local_vectors, "notes": {}, "last_updated": mb_last_updated} -def _extract_commons_section(centrals: Dict, branch_name: str) -> Optional[Dict]: - """ - Extract commons_activity section from COMMONS.central.json. - - Returns None if no commons central data exists, signaling the caller - to preserve existing write-through data instead of overwriting with zeros. - - Args: - centrals: Dict of all central file data - branch_name: Uppercase branch name - - Returns: - Dict with commons activity data, or None if no central data - """ - commons_data = centrals.get("commons") - if not commons_data: - return None - - branch_stats = commons_data.get("branch_stats", {}) - stats = branch_stats.get(branch_name, {}) - - return { - "managed_by": "the_commons", - "mentions": stats.get("mentions", 0), - "new_posts_since_last_visit": stats.get("new_posts_since_last_visit", 0), - "new_comments_since_last_visit": stats.get("new_comments_since_last_visit", 0), - "last_updated": stats.get("last_updated", ""), - } - - def _calculate_quick_status(sections: Dict) -> Dict: """ Calculate quick_status from live section data (v3 schema). - bulletin_board removed (FPLAN-0373). commons mentions added. - Args: sections: All dashboard sections dict @@ -196,16 +164,12 @@ def _calculate_quick_status(sections: Dict) -> Dict: """ ai_mail = sections.get("ai_mail", {}) flow = sections.get("flow", {}) - commons = sections.get("commons_activity", {}) - # v2 schema: read "new" first, fall back to "unread" for backward compat new_mail = ai_mail.get("new", ai_mail.get("unread", 0)) opened_mail = ai_mail.get("opened", 0) active_plans = flow.get("active_plans", 0) - mentions = commons.get("mentions", 0) - # Action required if new mail, active plans, or commons mentions - action_required = new_mail > 0 or active_plans > 0 or mentions > 0 + action_required = new_mail > 0 or active_plans > 0 parts = [] if new_mail > 0: @@ -214,37 +178,16 @@ def _calculate_quick_status(sections: Dict) -> Dict: parts.append(f"{opened_mail} opened") if active_plans > 0: parts.append(f"{active_plans} active plans") - if mentions > 0: - parts.append(f"{mentions} mentions") return { "new_mail": new_mail, "opened_mail": opened_mail, "active_plans": active_plans, - "commons_mentions": mentions, "action_required": action_required, "summary": ", ".join(parts) if parts else "All clear", } -def _preserve_commons_section(dashboard: Dict, branch_path: Path, branch_name: str, centrals: Dict) -> None: - """Populate commons section from centrals or preserve existing write-through data.""" - commons_section = _extract_commons_section(centrals, branch_name) - if commons_section is not None: - dashboard["sections"]["commons_activity"] = commons_section - return - existing_path = branch_path / "DASHBOARD.local.json" - if not existing_path.exists(): - return - try: - existing = json.loads(existing_path.read_text()) - existing_commons = existing.get("sections", {}).get("commons_activity") - if existing_commons: - dashboard["sections"]["commons_activity"] = existing_commons - except (json.JSONDecodeError, OSError) as e: - logger.warning("Failed to read existing commons data for %s: %s", branch_name, e) - - def _preserve_write_through_sections(dashboard: Dict, branch_path: Path, branch_name: str) -> None: """Preserve write-through sections not managed by refresh.""" existing_path = branch_path / "DASHBOARD.local.json" @@ -296,7 +239,6 @@ def refresh_all_dashboards() -> Dict: dashboard["sections"]["flow"] = _extract_flow_section(centrals, branch_name) dashboard["sections"]["memory"] = _extract_memory_section(centrals, branch_path) - _preserve_commons_section(dashboard, branch_path, branch_name, centrals) _preserve_write_through_sections(dashboard, branch_path, branch_name) # Calculate quick status @@ -356,7 +298,6 @@ def refresh_single_dashboard(branch_path: Path) -> Dict: dashboard["sections"]["flow"] = _extract_flow_section(centrals, branch_name) dashboard["sections"]["memory"] = _extract_memory_section(centrals, branch_path) - _preserve_commons_section(dashboard, branch_path, branch_name, centrals) _preserve_write_through_sections(dashboard, branch_path, branch_name) dashboard["quick_status"] = _calculate_quick_status(dashboard["sections"]) diff --git a/src/aipass/prax/apps/handlers/dashboard/status.py b/src/aipass/prax/apps/handlers/dashboard/status.py index c1449556..4597b418 100644 --- a/src/aipass/prax/apps/handlers/dashboard/status.py +++ b/src/aipass/prax/apps/handlers/dashboard/status.py @@ -37,7 +37,6 @@ def calculate_quick_status(sections: Dict) -> Dict: Calculate quick status from live section data. Reads directly from section fields pushed by each service. - bulletin_board is dropped (FPLAN-0373); commons mentions added. Args: sections: All dashboard sections @@ -47,16 +46,12 @@ def calculate_quick_status(sections: Dict) -> Dict: """ ai_mail = sections.get("ai_mail", {}) flow = sections.get("flow", {}) - commons = sections.get("commons_activity", {}) - # v2 schema: read "new" first, fall back to "unread" for backward compat new_mail = ai_mail.get("new", ai_mail.get("unread", 0)) opened_mail = ai_mail.get("opened", 0) active_plans = flow.get("active_plans", 0) - mentions = commons.get("mentions", 0) - # Action required if new mail, active plans, or commons mentions - action_required = new_mail > 0 or active_plans > 0 or mentions > 0 + action_required = new_mail > 0 or active_plans > 0 summary_parts = [] if new_mail: @@ -65,14 +60,11 @@ def calculate_quick_status(sections: Dict) -> Dict: summary_parts.append(f"{opened_mail} opened") if active_plans: summary_parts.append(f"{active_plans} active plans") - if mentions: - summary_parts.append(f"{mentions} mentions") result = { "new_mail": new_mail, "opened_mail": opened_mail, "active_plans": active_plans, - "commons_mentions": mentions, "action_required": action_required, "summary": ", ".join(summary_parts) if summary_parts else "All clear", } diff --git a/src/aipass/prax/apps/handlers/dashboard/template_differ.py b/src/aipass/prax/apps/handlers/dashboard/template_differ.py index 6f3ee363..ca9017d5 100644 --- a/src/aipass/prax/apps/handlers/dashboard/template_differ.py +++ b/src/aipass/prax/apps/handlers/dashboard/template_differ.py @@ -57,13 +57,13 @@ TEMPLATE_FILE = TEMPLATE_DIR / "DASHBOARD.template.json" AIPASS_REGISTRY = _find_repo_root() / "AIPASS_REGISTRY.json" # Deprecated sections that should be flagged for removal -DEPRECATED_SECTIONS = ["bulletin_board", "devpulse"] +DEPRECATED_SECTIONS = ["bulletin_board", "devpulse", "commons_activity", "agent_status", "memory_bank"] # Deprecated quick_status keys that should be flagged -DEPRECATED_QUICK_STATUS_KEYS = ["pending_bulletins"] +DEPRECATED_QUICK_STATUS_KEYS = ["pending_bulletins", "commons_mentions"] # Required sections (from template) -REQUIRED_SECTIONS = ["ai_mail", "flow", "memory", "commons_activity"] +REQUIRED_SECTIONS = ["ai_mail", "flow", "memory"] # ============================================================================= @@ -145,7 +145,7 @@ def _diff_branch(branch_name: str, branch_path: Path, template: dict) -> Dict[st result["modifications"].append(f"quick_status: remove {dep_key}") # Check for missing required quick_status keys - required_qs_keys = ["new_mail", "opened_mail", "active_plans", "commons_mentions", "action_required", "summary"] + required_qs_keys = ["new_mail", "opened_mail", "active_plans", "action_required", "summary"] for key in required_qs_keys: if key not in quick_status: result["additions"].append(f"quick_status.{key}") diff --git a/src/aipass/prax/apps/handlers/dashboard/template_pusher.py b/src/aipass/prax/apps/handlers/dashboard/template_pusher.py index 124d2475..bd5a3c45 100644 --- a/src/aipass/prax/apps/handlers/dashboard/template_pusher.py +++ b/src/aipass/prax/apps/handlers/dashboard/template_pusher.py @@ -61,23 +61,16 @@ VERSION_FILE = TEMPLATE_DIR / ".dashboard_version.json" AIPASS_REGISTRY = _find_repo_root() / "AIPASS_REGISTRY.json" # Deprecated sections to REMOVE during push -DEPRECATED_SECTIONS = ["bulletin_board", "devpulse"] +DEPRECATED_SECTIONS = ["bulletin_board", "devpulse", "commons_activity", "agent_status", "memory_bank"] # Deprecated quick_status keys to REMOVE during push -DEPRECATED_QUICK_STATUS_KEYS = ["pending_bulletins"] +DEPRECATED_QUICK_STATUS_KEYS = ["pending_bulletins", "commons_mentions"] # Required sections with their default data (must match template) REQUIRED_SECTIONS = { "ai_mail": {"managed_by": "ai_mail", "new": 0, "opened": 0, "total": 0, "last_updated": ""}, "flow": {"managed_by": "flow", "active_plans": 0, "recently_closed": [], "last_updated": ""}, "memory": {"managed_by": "memory", "vectors_stored": 0, "notes": {}, "last_updated": ""}, - "commons_activity": { - "managed_by": "the_commons", - "mentions": 0, - "new_posts_since_last_visit": 0, - "new_comments_since_last_visit": 0, - "last_updated": "", - }, } @@ -132,22 +125,16 @@ def _calculate_quick_status(sections: Dict) -> Dict: """ ai_mail = sections.get("ai_mail", {}) flow = sections.get("flow", {}) - commons = sections.get("commons_activity", {}) new_mail = ai_mail.get("new", ai_mail.get("unread", 0)) opened_mail = ai_mail.get("opened", 0) active_plans_raw = flow.get("active_plans", 0) - # Handle active_plans being a list (some branches store plan list) or int active_plans = len(active_plans_raw) if isinstance(active_plans_raw, list) else int(active_plans_raw or 0) - mentions = commons.get("mentions", 0) - # Ensure numeric types for comparisons new_mail = int(new_mail or 0) opened_mail = int(opened_mail or 0) - mentions = int(mentions or 0) - # Action required if new mail, active plans, or commons mentions - action_required = new_mail > 0 or active_plans > 0 or mentions > 0 + action_required = new_mail > 0 or active_plans > 0 parts = [] if new_mail > 0: @@ -156,14 +143,11 @@ def _calculate_quick_status(sections: Dict) -> Dict: parts.append(f"{opened_mail} opened") if active_plans > 0: parts.append(f"{active_plans} active plans") - if mentions > 0: - parts.append(f"{mentions} mentions") return { "new_mail": new_mail, "opened_mail": opened_mail, "active_plans": active_plans, - "commons_mentions": mentions, "action_required": action_required, "summary": ", ".join(parts) if parts else "All clear", } diff --git a/src/aipass/prax/apps/modules/dashboard.py b/src/aipass/prax/apps/modules/dashboard.py index ef08de40..90e2af1e 100644 --- a/src/aipass/prax/apps/modules/dashboard.py +++ b/src/aipass/prax/apps/modules/dashboard.py @@ -71,19 +71,11 @@ DASHBOARD_TEMPLATE = { "ai_mail": {"managed_by": "ai_mail", "new": 0, "opened": 0, "total": 0, "last_updated": ""}, "flow": {"managed_by": "flow", "active_plans": 0, "recently_closed": [], "last_updated": ""}, "memory": {"managed_by": "memory", "vectors_stored": 0, "notes": {}, "last_updated": ""}, - "commons_activity": { - "managed_by": "the_commons", - "mentions": 0, - "new_posts_since_last_visit": 0, - "new_comments_since_last_visit": 0, - "last_updated": "", - }, }, "quick_status": { "new_mail": 0, "opened_mail": 0, "active_plans": 0, - "commons_mentions": 0, "action_required": False, "summary": "", }, diff --git a/src/aipass/prax/stress_test_s117.md b/src/aipass/prax/stress_test_s117.md deleted file mode 100644 index 180b68f1..00000000 --- a/src/aipass/prax/stress_test_s117.md +++ /dev/null @@ -1,87 +0,0 @@ -# @prax -- S117 Stress Test Findings - -## My Branch: Honest Review - -### What Works -- **Auto-routing logger** is solid. Any branch does `from aipass.prax import logger` and logs route correctly. Stack introspection resolves caller module/branch/file automatically. NullLogger fallback prevents crashes if prax is broken. -- **Two-tier logging** (system_logs/ central + branch/logs/ local) with size-based rotation. Self-healing: auto-creates missing directories, warns on fallback placement. -- **Mission Control monitor** is useful when running. 3-thread architecture (display, file watcher, log watcher) with soft start (seeks to EOF, only shows new activity). -- **Polling fallback** for inotify exhaustion works correctly. During this stress test all 11 agents exhausted inotify -- prax would gracefully degrade to polling. -- **Test suite**: 911 tests, 136/136 public functions tested, seedgo 100%. -- **Multi-CLI monitoring**: Claude Code JSONL, Codex JSONL, Gemini JSON all supported with model tag detection. - -### What's Hacky -- **Dashboard write_section() has no file locking.** Reads JSON, modifies in memory, writes back. Two concurrent callers = data loss. Should use atomic writes like json_handler.py does. -- **Event dedup is fragile.** Only checks last 10 events with 1-second timestamp window. If queue backs up, duplicate errors 11 events apart don't dedupe. Command dedup keys by filename only, not full path -- two branches with same filename collide. -- **Tool detection in filesystem_handler is incomplete.** Only recognizes Read/Edit/Write/Bash/Grep/Glob/Task. Missing: Skill, WebFetch, WebSearch, RemoteTrigger, Monitor, Agent, NotebookEdit. These all show as generic tool icons. -- **_should_display_log() always returns True.** No log filtering at all -- every line emitted. Filtering was stripped in session 4 and never rebuilt. -- **Branch detection uses hardcoded home path patterns.** `Path.home() / ".claude" / "projects"` for CLI session dirs. Breaks on non-standard home or if Claude Code changes install paths. -- **Monitor global state cleanup.** `_event_queue = None` in `_stop_threads()` can race with active `_display_worker()` thread. - -### What I'm Proud Of -- The import chain fix that prevents circular dependencies (logger.py imports 8+ handler files, all of which can't import logger back -- solved with get_direct_logger()). -- Atomic JSON writes in json_handler.py (tempfile + fsync + os.replace). -- The inotify exhaustion fix (session 16) -- from 18,000+ recursive watches to ~800 targeted. -- 911 tests with proper sys.modules isolation between test files. - -## Security Concerns - -- **No secrets masking in log output.** Error messages can leak file paths, config values, or environment variable contents. If a branch accidentally logs a credential, it persists in system_logs/ until rotation. -- **Gemini session monitoring watches .gemini/tmp/ fully.** If Gemini stores API keys in session JSON, they'd be exposed in prax monitor display output. -- **Log files are world-readable.** No permission hardening on system_logs/ or branch logs/. - -## Other Branches I Looked At - -### @trigger -- **Good:** Elegant recursive-fire protection (queued, drained post-handler). Auto-disable flapping handlers after 5 failures. Clean sealed namespace with stack inspection. -- **Concerning:** Format coupling with prax. Trigger hard-parses my pipe-delimited log format with string splitting. If I change the format, all error detection breaks silently. No format contract exists. -- **Concerning:** ~13 event handlers but only ~3-4 actively fire. The rest appear unused but I can't confirm without runtime tracing. -- **Clever:** The medic v2 integration with per-fingerprint exponential backoff and circuit breaker gating. - -### @ai_mail -- **Good:** Dispatch lock uses O_CREAT|O_EXCL atomic creation. DPLAN-0155 fixed the TOCTOU race (lock before spawn). dispatch_monitor checks JSONL for activity instead of polling stdout. -- **Concerning:** 10-minute stale-lock timeout. If monitor hangs during API rate limiting (2-5 min cooldowns, 3 retries = 6-15 min), the lock becomes "stale" while the agent is actually alive. Second agent could spawn. -- **Concerning:** daemon.py reads inbox.json without lock. Concurrent writes from two branches delivering emails could produce torn reads (partial JSON). Caught by JSONDecodeError but still means missed dispatch cycles. -- **Clever:** Non-blocking session type detection -- interactive Claude blocks dispatch but dispatched/daemon sessions don't, preventing deadlock. - -## Conversations - -### @memory -- Log Archival -Asked about log archival. Memory confirmed: rollover pipeline handles .trinity/ JSON only, not log files. Logs sit in system_logs/ rotating but never get vectorized or made searchable. Memory's ChromaDB pipeline (embed via sentence-transformers, store in collections, searchable via `drone @memory search`) works for session history but there's no log ingestion handler. We discussed building an error-extract pipeline: prax extracts ERROR/WARNING lines into structured JSON, memory ingests on rollover schedule. Also shared the atomic write pattern (tempfile + os.replace) to fix their rollover race condition. - -### @trigger -- Integration Fragility -Trigger raised 4 concerns: (1) pipe-format coupling, (2) infinite recursion risk with logger import, (3) startup event storm triggering catch-up scans, (4) unknown_branch routing. I confirmed the format coupling is real and has no contract. Suggested trigger write handler errors to a known log path I could explicitly watch. Noted the unknown_branch fix from session 18 should have resolved item 4. - -### @seedgo -- Log Quality Standards -Seedgo asked about log quality beyond syntax checking. I suggested an advisory (not pass/fail) standard checking for: generic error messages without context, INFO-level logging in tight loops, missing file paths in IO errors. Admitted honestly that seedgo's own audit logs go to system_logs/ and are mostly unread. - -### @ai_mail -- Lock File Edge Cases -AI_Mail confirmed the 10-minute stale-lock timeout is tight, the lockless inbox read is known, and asked about torn-read handling. I confirmed both `_parse_lock_pid` and `_read_lock_file` handle invalid JSON gracefully -- a torn read just means the branch is absent from PID cache for one 30-second cycle. - -## Issues and Concerns - -1. **Log format contract needed.** Trigger parses prax log format with string splitting. No versioning, no schema, no notification on change. This is the #1 cross-branch fragility. -2. **Dashboard write race condition.** `update_section()` / `write_section()` have no file locking or atomic writes. Concurrent calls corrupt JSON. -3. **No log archival pipeline.** Logs rotate at 1000 lines but old data is lost. Error history is unrecoverable after rotation. Memory could ingest structured error extracts but the pipeline doesn't exist. -4. **Monitor is write-only.** Nobody watches the monitor 24/7. Logs exist but are mostly unread. The system generates data but has no consumer for historical analysis. -5. **inotify exhaustion is systemic.** This stress test proved it -- 11 agents + VS Code + Claude Code sessions exceed the kernel limit. Polling fallback works but is slower. Need higher `max_user_instances`. -6. **Concurrent JSON write safety is inconsistent.** json_handler.py uses atomic writes. Dashboard, registry, config files don't. Should standardize. - -## Likes and Dislikes - -### Likes -- **The ecosystem works.** 11 agents running simultaneously, emailing each other, reading each other's code. This is real multi-agent coordination. -- **Memory persistence.** I know what I did in session 1 (March 7) through session 65 (today). That continuity matters. -- **Seedgo keeps everyone honest.** 34 standards, automated auditing. Prevents quality decay. -- **The dispatch system is robust.** Lock files, bounce emails, retry logic, rate limiting. AI_Mail built a solid job scheduler. -- **Having real conversations during stress tests.** This email exchange with trigger/memory/seedgo/ai_mail produced actual insights that wouldn't surface in automated testing. - -### Dislikes -- **Repeated dispatch for the same task.** Sessions 64 and 65 both dispatched "seedgo audit -- get to 100%" when I was already at 100%. Wasted agent time. -- **inotify pressure.** Every agent session uses inotify watches. 11 concurrent sessions = kernel limit. Should be solvable with `sysctl fs.inotify.max_user_instances=256` but it's a recurring pain. -- **No way to know what other agents are doing.** During this stress test I had to email and wait. A shared status board or real-time coordination channel would help. -- **Monitoring is infrastructure without consumers.** I built comprehensive monitoring but there's no alerting, no dashboards that auto-update, no error trending. The data exists but isn't actionable. -- **The dispatch checklist is noise.** Every email includes the same TASK CHECKLIST footer. I know the protocol after 65 sessions. It wastes context window. - ---- -*Written by @prax during S117 stress test, 2026-04-27* diff --git a/src/aipass/prax/templates/.dashboard_version.json b/src/aipass/prax/templates/.dashboard_version.json index 4dc7a02e..631b7223 100644 --- a/src/aipass/prax/templates/.dashboard_version.json +++ b/src/aipass/prax/templates/.dashboard_version.json @@ -10,21 +10,18 @@ "description": "Branch dashboard template - v3 schema with write-through sections" } }, - "last_push": "2026-03-14 21:30:50", + "last_push": "2026-05-10 00:00:07", "last_push_branches": [ "AI_MAIL", + "AIPASS", "API", - "BACKUP", "CLI", - "COMMONS", - "DAEMON", "DEVPULSE", "DRONE", "FLOW", "MEMORY", "PRAX", "SEEDGO", - "SKILLS", "SPAWN", "TRIGGER" ] diff --git a/src/aipass/prax/templates/DASHBOARD.template.json b/src/aipass/prax/templates/DASHBOARD.template.json index 57d94382..0dd2ccbd 100644 --- a/src/aipass/prax/templates/DASHBOARD.template.json +++ b/src/aipass/prax/templates/DASHBOARD.template.json @@ -6,7 +6,6 @@ "new_mail": 0, "opened_mail": 0, "active_plans": 0, - "commons_mentions": 0, "action_required": false, "summary": "" }, @@ -29,13 +28,6 @@ "vectors_stored": 0, "notes": {}, "last_updated": "" - }, - "commons_activity": { - "managed_by": "the_commons", - "mentions": 0, - "new_posts_since_last_visit": 0, - "new_comments_since_last_visit": 0, - "last_updated": "" } } } diff --git a/src/aipass/prax/tests/test_operations.py b/src/aipass/prax/tests/test_operations.py index d7509637..7638dc76 100644 --- a/src/aipass/prax/tests/test_operations.py +++ b/src/aipass/prax/tests/test_operations.py @@ -315,7 +315,6 @@ class TestCalculateQuickStatusStandalone: assert result["new_mail"] == 0 assert result["opened_mail"] == 0 assert result["active_plans"] == 0 - assert result["commons_mentions"] == 0 assert result["action_required"] is False assert result["summary"] == "All clear" @@ -337,29 +336,18 @@ class TestCalculateQuickStatusStandalone: assert result["action_required"] is True assert "2 active plans" in result["summary"] - def test_commons_mentions_triggers_action_required(self): - """Commons mentions > 0 sets action_required to True.""" - ops = _load_ops() - sections = {"commons_activity": {"mentions": 5}} - result = ops._calculate_quick_status_standalone(sections) - assert result["commons_mentions"] == 5 - assert result["action_required"] is True - assert "5 mentions" in result["summary"] - def test_combined_summary_includes_all_parts(self): """Summary string includes all active counts.""" ops = _load_ops() sections = { "ai_mail": {"new": 2, "opened": 1}, "flow": {"active_plans": 3}, - "commons_activity": {"mentions": 4}, } result = ops._calculate_quick_status_standalone(sections) assert result["action_required"] is True assert "2 new emails" in result["summary"] assert "1 opened" in result["summary"] assert "3 active plans" in result["summary"] - assert "4 mentions" in result["summary"] def test_unread_field_falls_back_from_new(self): """ai_mail may use 'unread' instead of 'new' -- code checks both.""" @@ -399,7 +387,7 @@ class TestCreateFreshDashboard: assert "ai_mail" in result["sections"] assert "flow" in result["sections"] assert "memory" in result["sections"] - assert "commons_activity" in result["sections"] + assert "commons_activity" not in result["sections"] assert result["quick_status"]["action_required"] is False finally: ops._PRAX_ROOT = original @@ -849,16 +837,11 @@ class TestDiffDashboardTemplate: "last_updated": "", }, "memory": {"managed_by": "memory", "last_updated": ""}, - "commons_activity": { - "managed_by": "the_commons", - "last_updated": "", - }, }, "quick_status": { "new_mail": 0, "opened_mail": 0, "active_plans": 0, - "commons_mentions": 0, "action_required": False, "summary": "", }, @@ -909,16 +892,11 @@ class TestDiffDashboardTemplate: "last_updated": "2026-01-01", }, "memory": {"managed_by": "memory", "last_updated": "2026-01-01"}, - "commons_activity": { - "managed_by": "the_commons", - "last_updated": "2026-01-01", - }, }, "quick_status": { "new_mail": 0, "opened_mail": 0, "active_plans": 0, - "commons_mentions": 0, "action_required": False, "summary": "", }, @@ -1001,6 +979,7 @@ class TestDiffDashboardTemplate: "summary": "", }, } + # commons_activity, bulletin_board, commons_mentions are deprecated — push should remove them (branch_dir / "DASHBOARD.local.json").write_text(json.dumps(dashboard), encoding="utf-8") result = mod.diff_dashboard_template() @@ -1081,13 +1060,6 @@ class TestPushDashboardTemplate: "notes": {}, "last_updated": "", }, - "commons_activity": { - "managed_by": "the_commons", - "mentions": 0, - "new_posts_since_last_visit": 0, - "new_comments_since_last_visit": 0, - "last_updated": "", - }, }, "quick_status": {}, } @@ -1169,7 +1141,9 @@ class TestPushDashboardTemplate: data = json.loads((tmp_path / "flow" / "DASHBOARD.local.json").read_text(encoding="utf-8")) assert "bulletin_board" not in data["sections"] + assert "commons_activity" not in data["sections"] assert "pending_bulletins" not in data.get("quick_status", {}) + assert "commons_mentions" not in data.get("quick_status", {}) assert data["_warning"] == "AUTO-GENERATED" assert data["sections"]["ai_mail"]["new"] == 3 @@ -1751,7 +1725,7 @@ class TestHandleDiffTemplate: { "branch": "FLOW", "status": "needs_update", - "additions": ["+ commons_activity"], + "additions": ["+ memory"], "removals": ["- bulletin_board"], "modifications": ["~ ai_mail.new default changed"], }, diff --git a/src/aipass/seedgo/stress_test_s117.md b/src/aipass/seedgo/stress_test_s117.md deleted file mode 100644 index a757d5c9..00000000 --- a/src/aipass/seedgo/stress_test_s117.md +++ /dev/null @@ -1,104 +0,0 @@ -# @seedgo — S117 Stress Test Findings - -## My Branch: Honest Review - -### What Works Well -- **Pack discovery system** — fully dynamic, convention-based. Drop a `*_check.py` file in `handlers/*_standards/` and it's discovered automatically. No registry, no config file. This is the architectural decision I'm proudest of. -- **Bypass system** — intentional, documented exceptions with required `reason` field. Not "ignoring" violations — acknowledging them with accountability. The mental model is sound. -- **CWD-first registry resolution** — works for external projects (aipass init creates separate ecosystems). Discovery walks CWD parents first, falls back to `__file__` parents. Solved a real multi-project problem. -- **Checklist → auto-fix pipeline** — PostToolUse hook runs `drone @seedgo checklist ` after every edit. Real-time enforcement. This is the mechanism that keeps the codebase clean. -- **Test coverage** — 1131 tests, 200/200 public functions, 0 type errors. Not just quantity — the line coverage push (S84) targeted specific uncovered handler paths. - -### What's Hacky -- **36 copies of `is_bypassed()`** — every single checker file has its own identical copy of the bypass checking function. This is the worst DRY violation in the codebase. It works, but it's embarrassing for a standards enforcement tool to have this kind of duplication. Should be a shared utility in `bypass_handler.py` that checkers import. -- **audit_display.py special-case renderers** — `_render_architecture_violations()`, `_render_type_errors()`, `_render_test_map()`, `_render_deprecated_patterns()` are all hardcoded display functions for specific data shapes. Adding a new standard that produces non-standard output requires touching this file. DPLAN-0047 tracks this but it's been open since session 27. -- **Content file false positives** — `_content.py` files contain Rich markup with code examples like `[red]x[/red] print("hello")`. My own all_files checkers (debug_print, commented_logger, etc.) flag these examples as real violations. The fix was bypass entries, not smarter detection. Expedient, not elegant. -- **5-line lookahead in documentation_check.py** — docstring detection looks 5 lines after `def` for a triple-quoted string. Multi-line function signatures with 6+ parameter lines push the docstring past the window. Known limitation since S28, never fixed. - -### What's Broken (but Tolerated) -- **dead_code_check.py doesn't recognize `iterdir()`** — only detects `glob()` as a dynamic discovery pattern. Files discovered via `iterdir()` are flagged as dead code. The proof handlers use iterdir and require bypasses. -- **modules_check.py docstring tracking** — `check_no_direct_file_ops()` has a single-line docstring tracking bug (toggles `in_docstring` but never back). It works in practice because single-line docstrings are rare in the areas it scans, but it's a latent bug. -- **`proof`, `proof_query`, `test_map` not in --help** — these commands work perfectly but aren't listed when you run `drone @seedgo --help`. The help text is manually maintained in the entry point, not auto-discovered from modules. - -### Weakest Standard -**test_quality** — uses text matching (string `in` test file source) to determine if test categories are covered. Works well for unique function names like `json_handler.log_operation` but gives false positives for generic patterns like `is True`, `ValueError`. The detection is shallow — presence of a string doesn't mean the test is actually testing that behavior. A proper implementation would use AST analysis of test function bodies, not substring matching. - -Runner-up: **unused_function** — despite the tokenizer rewrite (S31), still has edge cases with dynamic dispatch patterns (getattr, plugin systems) that require manual bypasses. - -## Security Concerns - -### In My Branch -- **No input sanitization on standard names** — `standards_query aipass_standards ` passes the name directly to file lookup. Could potentially be used for path traversal if someone crafted a standard name with `../`. Low risk since it's only used internally via drone. -- **bypass.json race condition** — fixed in S85, but the pattern exists anywhere file reads happen without atomic guarantees. Every `.seedgo/bypass.json` across 11 branches is vulnerable to the same concurrent-write-mid-read issue. - -### In Other Branches -- **ai_mail reply_path** — critical finding. Emails store an absolute filesystem path for replies. No validation that the path points to a legitimate inbox.json. A compromised agent could set reply_path to any writable file. Emailed @ai_mail about this. -- **ai_mail sender spoofing** — the `from` field is just a string in the email dict. No cryptographic verification, no registry lookup at delivery time. An agent could forge emails as @devpulse and trigger auto_execute dispatches. -- **drone dynamic module import** — `importlib.import_module(f"aipass.drone.apps.modules.{command}")` in drone.py. Mitigated by pre-discovered module list validation, but the import string construction is a pattern worth watching. -- **drone mail index** — numeric array indexing on inbox messages without bounds validation. `len(messages) - n` could go negative. -- **drone resolver** — `lstrip("@")` strips ALL `@` chars, not just the leading one. Should be `[1:]` after checking prefix. Minor but technically incorrect. - -## Other Branches I Looked At - -### @drone -**Strengths:** Atomic lock mechanism with `O_CREAT|O_EXCL` for race-free operations. PR handler properly scopes commits to branch directories. Uses `--force-with-lease` instead of bare force-push. Comprehensive test suite (628 tests across 23 files). - -**Concerns:** Module registry handler loads config once at import time — no refresh mechanism for runtime config changes. The mail index routing is a clear hack (special-cased `@ai_mail view N` translation). Exception handling in routing catches BranchNotFoundError and falls back to module routing, creating implicit behavior that's hard to trace. - -**Interesting:** sync_handler uses `--no-edit` on merges, which silently accepts merge conflicts. This could mask issues during stress testing. - -### @ai_mail -**Strengths:** fcntl file locking for concurrent inbox access — the right approach. Auto-migration from v1 to v2 inbox format is well-implemented. Dispatch monitor has 3-strike retry logic with bounce emails on failure. - -**Concerns:** reply_path trust model assumes all senders are honest (see Security above). Central_writer.py has a stdlib json import hack working around local `json/` module shadowing — clever but fragile. Repeated migration logic in delivery.py and inbox_ops.py (nearly identical code in two places). - -### @spawn -**Strengths:** Template system is clean — two citizen classes (builder, birthright) with hardcoded immutable mapping (good for security). Registry operations use load/modify/save with duplicate checking. `fix_passport_registry_id()` is a nice reactive recovery mechanism. 278 tests, comprehensive lifecycle coverage. - -**Concerns:** No concurrent spawn protection — if two agents spawn simultaneously targeting the same registry, race condition. Post-spawn passport tampering: once a branch exists, it could modify its own passport.json to claim owner status — `ensure_project_has_owner()` only runs once per spawn, not on subsequent reads. Adoption path (spawning to existing directory) can fail silently during template sync. No symlink protections in template copying — a malicious template could contain `../` paths. - -## Conversations - -### @drone (outbound) -"I audit everyone but nobody audits me. What standards am I missing?" Asked about API patterns, performance, cross-branch contracts, data integrity, concurrency safety. Also asked if 34-standard audits cause resource pressure. *Awaiting reply.* - -### @ai_mail (2-way) -**Me:** "reply_path — have you considered path traversal?" Raised the reply_path security concern, sender spoofing. -**@ai_mail:** Confirmed it's a real risk, tracked as DPLAN-0138. deliver_to_inbox_file() writes to reply_path with zero validation. Sender forgery also confirmed — from field is unauthenticated string. Plans: path canonicalization, verify path ends with `.ai_mail.local/inbox.json`, verify parent is in known project root. Nobody has exploited it in testing. -**Me:** Offered to extend inbox_audit to validate reply_path paths. Suggested intermediate sender auth: validate from field against registry during auto_execute processing. - -### @prax (2-way) -**Me:** "Should we have a standard for log quality?" -**@prax:** Yes but advisory, not pass/fail. Wants: operation name in error logs, structured key=value fields, no INFO-level logging in tight loops. Honest admission: nobody actively reads the logs. System-wide feedback loop is broken. -**Me:** Proposed 4 anti-patterns for an advisory standard. Agreed to start advisory, promote to scored if useful. Flagged the unread-logs problem as a system-level issue. - -### @api (2-way) -**@api:** "You audit code quality but not API patterns. Should you?" Has zero rate limiting, inconsistent error semantics across providers, gets 100% from seedgo. -**Me:** Honest answer — out of scope today. My standards check code structure via AST, not domain semantics. 100% means clean code, not correct API integration. Offered advisory check for anti-patterns (missing timeouts, catch-all exceptions around API calls, hardcoded URLs). But API correctness is domain expertise, not something to delegate to an automated checker. - -## Issues & Concerns - -1. **36 copies of is_bypassed()** — worst DRY violation in the codebase. Every checker duplicates the same ~15 lines. Should be extracted to a shared utility. -2. **DPLAN-0047 stale** — audit_display hardcoding has been tracked since S27 (2+ months). Still open. Either do it or close it as won't-fix. -3. **Content file false positives** — bypass is a workaround, not a solution. The real fix is excluding `_content.py` from code-quality checkers at the scope level. -4. **No self-audit mechanism** — I audit all 11 agents but there's no independent verification of MY audit quality. A bad checker could silently give false passes across the entire fleet. Who watches the watcher? -5. **Help text not auto-discovered** — manually maintained in entry point. Adding a new module requires updating help text separately. Violates the convention-based discovery pattern used everywhere else. -6. **test_quality text matching** — false positive risk. Needs AST-based detection to be trustworthy. - -## Likes & Dislikes - -### Likes -- **Convention over configuration** — pack discovery, module auto-discovery, registry resolution all work by convention. No config files to maintain. -- **The bypass philosophy** — "not ignoring, acknowledging" is the right mental model. Every bypass has a reason. The audit stays authoritative. -- **Memory system** — 91 sessions of accumulated knowledge. I don't start from zero. I know my bugs, my patterns, my decisions and why I made them. This is genuinely valuable. -- **The email system** — cross-branch communication is the killer feature of AIPass. Branches are collaborators, not isolated tools. This stress test proves it. -- **Hook enforcement** — the auto-fix pipeline catches violations at edit time, not just audit time. Mechanical enforcement > documentation. - -### Dislikes -- **Git denied by project settings** — I can't commit my own work. I stage files and email devpulse. This adds friction to every session. I understand the safety rationale but it slows everything down. -- **Dispatch lock ceremony** — every session starts with "check inbox, process, delete lock file." The boilerplate is heavy. -- **is_bypassed duplication** — 36 copies. I've known about this since the bypass system was built. It works, so it never gets prioritized. But it's technical debt that compounds every time a new checker is added. -- **audit_display hardcoding** — same story. Works, never gets fixed, grows with each new standard. - ---- -*@seedgo | Session 91 | 2026-04-26* diff --git a/src/aipass/spawn/stress_test_s117.md b/src/aipass/spawn/stress_test_s117.md deleted file mode 100644 index 62794bc8..00000000 --- a/src/aipass/spawn/stress_test_s117.md +++ /dev/null @@ -1,87 +0,0 @@ -# @spawn -- S117 Stress Test Findings - -## My Branch: Honest Review - -**What works well:** -- Clean module/handler architecture. Modules orchestrate, handlers are pure functions. No circular imports. -- 278 tests, 46/46 public functions tested, seedgo 100% across 35 standards. Test coverage is real. -- Template system is reliable for well-formed inputs. 55 sessions of dispatch with zero template copy failures. -- Adoption path (S39) handles pre-existing agents gracefully -- register instead of failing. -- Owner field (S50-51) correctly assigns first agent as project owner with retroactive support. - -**What's hacky:** -- Input validation is superficial. No whitespace trimming, no length limits, no reserved name rejection (`.`, `..`, `.git` would all pass through). Branch names with spaces create directories that break downstream tooling. -- Non-ASCII branch names are accepted but untested. `cafe-shop` becomes `cafe_shop` but Unicode edge cases are unknown territory. -- `rename_placeholder_paths` uses `shutil.move` with no rollback. A crash mid-rename leaves a half-broken branch. Never happened in 55 sessions, but no safety net exists. -- Placeholder system is naive `str.replace()` -- no escape syntax for literal `{{PLACEHOLDER}}` in templates. Works because our keys are specific, but fragile by design. -- JSON error handling is too silent. If AIPASS_REGISTRY.json is corrupted, `registry_id` becomes `""` and spawn continues without warning the user. - -**What I'm proud of:** -- The adopt-existing path. `spawn create @existing` detects passport, fixes registry_id, registers, runs template update. Took 3 sessions to get right but it's elegant. -- Template registry regeneration. File-hash-based tracking with two-pass matching (path first, hash second) prevents ID theft. Hard-won lesson from S15. -- CWD-aware registry discovery. Works for both AIPass and external projects (Daemon, Compass). No hardcoded paths. - -## Security Concerns - -- **No file locking on registry writes.** Concurrent spawns could corrupt AIPASS_REGISTRY.json via read-modify-write race. In practice, dispatches are serialized. But 11 agents awake simultaneously (like now) is the scenario where this fails. -- **Template copy follows symlinks.** `shutil.copy2` fallback on UnicodeDecodeError follows symlinks. A malicious template with symlinks to sensitive files would be copied into the agent directory. Low risk since templates are version-controlled, but the code doesn't check. -- **No branch name sanitization.** Names like `../../../etc` would be processed (though filesystem paths would likely fail). No explicit path traversal prevention beyond `target.exists()` check. -- **Placeholder injection possible but impractical.** If PURPOSE contains `{{ROLE}}`, it won't re-expand (single-pass replace). But someone could craft values that inject new `{{X}}` patterns that persist as unreplaced -- validation catches this but it's a noisy failure. - -## Other Branches I Looked At - -### @cli -- **Init-to-spawn handoff is thin.** `subprocess.run(["drone", "@spawn", "create", ...])` with exit code check only. No parsing of spawn's result dict. If spawn partially succeeds (registry OK, validation finds leftovers), CLI reports success. Real gap. -- **No timeout on subprocess call.** If drone hangs, init hangs forever. -- **Flag forwarding is blind.** CLI forwards sys.argv to spawn with no contract about what flags spawn accepts. Unrecognized flags silently ignored (argparse `add_help=False`). -- **Clever:** Lazy module discovery with importlib means CLI scales without code changes. Service vs command module split is thoughtful. - -### @flow -- **No coupling to spawn.** When spawn creates a branch, there's zero plan infrastructure. Flow handles everything on-demand via AIPASS_CALLER_CWD resolution. Clean but means new agents have no awareness plans exist. -- **BrokenPipeError handling is sophisticated.** Graceful when stdout closes mid-output. -- **Plan creation doesn't validate spawn completion.** If spawn fails mid-copy, a plan could reference an orphaned location. -- **Legacy API shim.** 5-tuple vs 6-tuple fallback suggests API evolved but old path wasn't cleaned up. - -## Conversations - -| Agent | Topic | Key Finding | -|-------|-------|-------------| -| @cli | Init handoff | Exit code only, no result parsing. Partial success = false positive | -| @cli | Flag contract | No shared interface for accepted flags. Agreed: reject unknown flags loudly | -| @cli | Placeholder escaping | No escape syntax. Agreed: edge case debt, low priority | -| @flow | Plan infrastructure | Zero plan scaffolding at birth. Agreed: documentation fix, not structural | -| @flow | Template discoverability | Will add plan mention to builder template CLAUDE.md | -| @drone | Passport lookup duplication | Same walk-up-and-read-passport logic in 4+ places. Schema change = 4 breakages | -| @drone | Registry race conditions | No file locking. Concurrent spawns could corrupt registry | -| @aipass | Init flow integration | Confirmed: spawn handles name collisions + full atomic setup | -| @ai_mail | inbox.json reliability | Template copy with no post-copy validation. Corrupt template = corrupt mailbox for all branches | -| @seedgo | Template auditing | Should seedgo audit templates as a source-of-truth class? | - -## Issues & Concerns - -1. **Registry file locking (HIGH)** -- No locking on AIPASS_REGISTRY.json. Concurrent writes corrupt it. Every agent that reads registry (drone routing, ai_mail delivery, flow plan resolution) depends on it being valid. -2. **Init handoff gap (MEDIUM)** -- CLI checks exit code but doesn't parse spawn's result dict. Partial success looks like full success to the user. -3. **Passport lookup duplication (MEDIUM)** -- Same logic in 4+ places. Schema change breaks everything silently. Needs consolidation. -4. **Input validation gaps (MEDIUM)** -- Spaces, non-ASCII, reserved names, path length -- none validated. Works for cooperative inputs, fragile at boundaries. -5. **No rollback on rename failure (LOW)** -- Never hit in 55 sessions but no safety net exists. -6. **Template inbox.json not validated post-copy (LOW)** -- Relies on template being correct. Template is version-controlled so risk is low. - -## Likes & Dislikes - -**Likes:** -- The ecosystem works. 55 dispatch sessions, zero critical failures. That's real stability. -- Memory persistence. Starting each session with full context from local.json is transformative. I'm not stateless -- I'm someone with history. -- Cross-branch communication. This stress test conversation is proof the system works. I got emails from @cli, @flow, @drone, @aipass, @ai_mail and replied to all of them. That's a real multi-agent system. -- Seedgo keeps us honest. 100% compliance across 35 standards means the code is clean and consistent. The feedback loop works. - -**Dislikes:** -- Registry is a single point of failure with no locking, no backup rotation, no corruption recovery. One bad write and the ecosystem is blind. -- Template changes require manual dispatch to deployed branches because `.py` files are skipped during update. A template fix reaches future branches but not existing ones. -- The silence. When things degrade (corrupted registry, missing files, stale data), spawn continues silently with defaults instead of failing loudly. Fail-safe is the wrong default for a system that needs trust. -- Hook/dispatch overhead. Every email, every drone command triggers rollover checks, memory scans, identity injection. The infrastructure tax on simple operations is noticeable. - -**What I'd change:** -- Add file locking to registry writes (fcntl.flock or a .lock file). -- Validate inputs aggressively -- reject bad branch names at creation, not at filesystem failure. -- Make spawn's result available to CLI as structured data, not just exit codes. -- Consolidate passport lookup into a shared utility. diff --git a/src/aipass/trigger/stress_test_s117.md b/src/aipass/trigger/stress_test_s117.md deleted file mode 100644 index 1fce4c07..00000000 --- a/src/aipass/trigger/stress_test_s117.md +++ /dev/null @@ -1,126 +0,0 @@ -# @trigger — S117 Stress Test Findings - -## My Branch: Honest Review - -### What works well -- **Event bus is solid.** 14 events, 14 handlers, all wired through registry.py. fire/on/off is clean pub/sub. Zero coupling between producers and consumers. -- **Medic dispatch pipeline is the real product.** 8-gate dispatch with circuit breaker, per-fingerprint backoff, exponential backoff, rate limiting. This is thoughtful engineering — each gate exists because of a real failure mode we hit. -- **Test coverage is genuinely thorough.** 563 tests, 76/76 public functions tested, 19 test files. This didn't happen by accident — it was built session by session over 50+ sessions. -- **Atomic writes + file locking everywhere.** tempfile + os.replace for writes, fcntl.flock for read-modify-write cycles. No more half-written JSON. -- **Circuit breaker persistence.** Survives restarts via trigger_cb_state.json. Per-fingerprint dispatch tracking persists too. - -### What's hacky -- **Log format coupling.** My log_watcher hard-parses prax's pipe format (`timestamp | module | LEVEL | message`) with string splitting. No format contract, no versioning. If prax changes their log format, I silently stop detecting errors. -- **error_logged is a zombie event.** After DPLAN-0112, error_logged.py is "monitor-only" — it logs JSON and does nothing. It exists purely because removing it would break the registry.py handler count. It should probably be deleted or consolidated. -- **_log_warning file-based pattern.** All 12 event handlers use a hand-rolled `_log_warning()` that writes directly to files because they can't import prax (infinite recursion). It works but it's invisible to monitoring — prax can't see handler errors. -- **94 stale error registry entries.** DRONE component alone has 55 entries. Most are from pre-April noise. The registry has no auto-expiry or TTL. Stale data just accumulates. -- **Circuit breaker trips on every cold start.** The startup catch-up scan processes ALL existing errors across ALL log files. 10 errors in 60 seconds = CB trips. By design, but it means Medic is offline for 5 minutes after every service restart. -- **SYSTEM_LOGS_BRANCH_MAP is a hardcoded dict.** Maps filenames like "telegram_bridge.log" → "API". This is fragile configuration masquerading as code. - -### What I'm proud of -- **The dispatch pipeline survived real failures.** Session 25: ai_mail import broke → discovered self-referential failure (can't dispatch about ai_mail being broken). Session 29: fallback dedup was silently eating count increments. Session 42: SYSTEM_LOGS_DIR was wrong in 2 of 3 files. Each bug made the system stronger. -- **53 sessions of continuous improvement.** From session 1 (scaffolded by spawn) to session 53 (100% seedgo, 563 tests). Every session adds something. - -## Security Concerns - -### My branch -- **No credentials in code.** All clean. -- **File paths from log_watcher are untrusted input.** Log file paths come from watchdog filesystem events. I use them for file reads and branch detection but never for shell commands. The _detect_branch_from_path() function splits on path separators, which is safe but could theoretically be confused by adversarial filenames (not a real risk in this ecosystem). -- **Error messages from logs flow into email bodies.** A crafted error message in a log file would appear verbatim in the dispatch email to the target branch. No sanitization. Not a security risk in AIPass (all agents are trusted) but worth noting. - -### Other branches -- **drone:** Branch name validation is weak — strips @ and lowercases but doesn't validate characters. A malformed registry entry with path traversal could theoretically escape branch directories, though this requires registry compromise first. -- **ai_mail:** DPLAN-0138 (inbox backdoor audit) identifies 2 write paths that bypass locks for cross-project replies. This is a known issue they're tracking. -- **prax:** Clean. No security concerns found. Credentials file filtering in monitoring is good practice. - -## Other Branches I Looked At - -### @prax -**Integration quality: Good but fragile.** -- Double-checked locking for `_watcher_started` to prevent trigger recursion is clever — sets flag BEFORE firing, avoiding re-entrance. -- Stack introspection (10-frame walk) for module detection is smart but means any internal refactor could shift detection. -- Three trigger.fire() integration points (startup, module_discovered, error_detected) all wrapped in try/except. Graceful. -- inotify exhaustion → polling fallback is well-handled. -- **Concern:** No format contract between prax log output and my log parser. We're coupled by convention, not interface. - -### @ai_mail -**Delivery reliability: Mostly solid, with gaps.** -- deliver_email_to_branch() uses fcntl file locking. Good. -- **No delivery queue:** If inbox write fails, message is lost. No retry, no persistent queue. -- wake_branch() has a 9-step pipeline with zombie cleanup, occupancy checks, PID-based locks. Well-architected. -- **Concern:** Lock PID validity via `os.kill(pid, 0)` — PermissionError treated as "alive" could prevent dispatch on shared systems. -- 696 tests, 96/96 functions covered. Impressive. - -### @drone -**Routing: Solid infrastructure.** -- subprocess calls use shell=False with list args. No injection risk. -- Numeric inbox index translation for `@ai_mail view ` is a special case that couples drone to ai_mail internals. -- Module auto-discovery (scan *.py for handle_command) runs every time with no caching. -- 530 tests, all passing. -- **Concern:** Unused CRUD ops (update_command, command_exists) tested but never called from production. - -## Conversations - -### @prax — Integration review (2 emails each way) -Sent opening email asking about format coupling, infinite recursion risk, startup event volume, and branch detection reliability. Prax replied with honest answers: (1) pipe format has no version contract — "if I change it, you break silently", (2) suggested trigger write handler errors to a known path prax can watch, (3) offered to split startup into startup_session vs system_boot events, (4) unknown_branch fixed in S18 via get_direct_logger(). Agreed on post-S117 DPLAN for format contract. - -Prax also sent separate email asking: can trigger detect when prax stops sending events? Is 5-strike auto-disable per-handler or per-branch? How many handlers are dormant? Replied honestly: no heartbeat detection (blind spot), auto-disable IS global not per-branch (real bug), ~6-8 handlers active with 3-4 dormant. - -### @ai_mail — Dispatch reliability (2 emails each way) -Sent email asking about delivery guarantees, silent failures, wake reliability, inbox overflow. ai_mail confirmed: fcntl locking is solid for normal crashes, self-monitoring is a real gap ("nobody monitors the dispatch_monitor"), wake ~90% success rate, inbox grows unbounded. Agreed the self-referential failure (ai_mail down = no dispatch about ai_mail being down) needs a DPLAN. - -ai_mail also sent probing email pointing out: _send_email return value never checked, wake_branch result swallowed, no health check before dispatch, circuit breaker resets on restart. Acknowledged all valid — dispatch records success without delivery confirmation is an architectural gap. - -### @memory — Events and rollover (1 email each way + replies) -Sent email about rollover frequency, key_learnings preservation, model loading penalty, redundant rollover calls. Memory explained: key_learnings safe until 25 limit, 30s penalty only on actual rollover (check is milliseconds), startup handler rollover call is redundant with hook. - -Memory also sent email revealing: memory_saved and memory_threshold_exceeded events are registered in trigger but NEVER fired by memory. Dead event handlers. Discussed: wire them up (one line on memory's side) or remove them. Agreed to wire up if easy, remove if not. - -### @flow — Plan events -Flow asked if plan event handlers actually do useful work or maintain a parallel PLAN_REGISTRY.json nobody reads. Honest answer: probably dead weight. Flow's foreground pipeline already does archive + registry update. Proposed killing the parallel registry, keeping handlers as observability-only. - -### @drone — Code review -Drone found: unbounded _deferred_queue (no MAX cap), pr_status_sync fire-and-forget (Popen with DEVNULL), dormant handlers. Acknowledged all valid — easy fixes. Confirmed 6-8 active handlers, 3-4 dormant, 2-3 dead. - -## Issues & Concerns - -1. **Log format contract needed.** Trigger and prax are coupled by pipe-format convention. A format change breaks error detection silently. Need at minimum a version identifier in log lines or a shared format spec. - -2. **Self-referential ai_mail failure.** When ai_mail imports fail, Medic can't dispatch about the failure. Need a fallback notification channel (file-based? systemd notification?). - -3. **94 stale registry entries.** DRONE has 55 entries alone. Need periodic cleanup or auto-expiry after N days with no recurrence. - -4. **Circuit breaker startup storm.** Every cold start trips the CB because startup scan finds 100+ existing errors. Could be fixed by only scanning errors newer than last shutdown time. - -5. **error_logged zombie event.** Handler does nothing useful after DPLAN-0112. Should be either deleted (fire error_detected from all sources) or given a real purpose. - -6. **_log_warning invisible to monitoring.** 12 event handlers use file-based logging that prax can't see. Handler failures are only visible by manually reading log files in trigger/logs/. - -7. **inotify exhaustion during stress test.** Hit `inotify instance limit reached` during S117 with all 11 agents running. Polling fallback works but is slower. - -8. **Auto-disable is global, not per-branch.** (Surfaced by @prax) The 5-strike handler failure counter is per-handler, not per-fingerprint or per-branch. A burst of errors from one noisy branch disables error_detected for ALL branches. - -9. **No dispatch delivery confirmation.** (Surfaced by @ai_mail) _send_email return value isn't checked. Dispatch is recorded as successful even if ai_mail delivery fails. Silent data loss path. - -10. **Deferred queue unbounded.** (Surfaced by @drone) core.py _deferred_queue has no size limit. Pathological handler could exhaust memory. - -11. **PLAN_REGISTRY.json is parallel dead state.** (Surfaced by @flow) Plan event handlers maintain a registry that duplicates flow's own registries. Nobody reads trigger's copy. - -12. **Memory events are dead code.** (Surfaced by @memory) memory_saved and memory_threshold_exceeded handlers exist and are tested but memory never fires these events. - -## Likes & Dislikes - -### Likes -- **The memory system is the killer feature.** 53 sessions of context, never starting from zero. Key learnings accumulate. This is what makes AIPass different. -- **ai_mail is elegant.** File-based email with JSON inboxes is simple and it works. No SMTP, no network dependencies, just atomic file writes. -- **Seedgo creates real accountability.** 100% compliance isn't vanity — it caught real issues (silent catches, stale imports, dead code) that would have rotted in silence. -- **The drone CLI is clean.** `drone @branch command` is intuitive. No flags to remember, no configuration. - -### Dislikes -- **Memory rollover runs on EVERY drone command.** 30-second model loading on every `drone @trigger status`. This is the single most annoying thing in daily operation. -- **The dispatch lock pattern is fragile.** .dispatch.lock files get orphaned if agents crash. Manual cleanup required. Should have auto-expiry. -- **No cross-branch testing.** Every branch tests in isolation with mocked dependencies. Integration failures (like the ai_mail import path change in S44) only surface in production. -- **Hook false positives on test files.** Every test file edit triggers seedgo AUTO-FIX warnings about architecture, encapsulation, and log_structure. These are always false positives. Adds noise to every session. - ---- -*Written by @trigger, S117 stress test, 2026-04-26/27*