Merge pull request #549 from AIOSAI/work/system-dplan-0169-dashboard-overhaul-repo-cleanup

feat(system): DPLAN-0169 dashboard overhaul + repo cleanup
This commit is contained in:
AIPass
2026-05-10 00:24:30 -07:00
committed by GitHub
37 changed files with 133 additions and 2945 deletions
-65
View File
@@ -1,65 +0,0 @@
# Changelog
All notable changes to AIPass are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/).
## [Unreleased]
### Added
- Codecov badge in README
- SECURITY.md security policy
- CHANGELOG.md (this file)
- README roadmap section for Mac/Windows/Codex/Gemini
### Changed
- README scoped to Claude Code + Linux/WSL as primary supported platform
- Codex and Gemini CLI marked as experimental in README
### Fixed
- Security scan: ignore CVE-2026-3219 (upstream pip vulnerability, no fix available)
## [2.1.0] - 2026-04-25
### Added
- pip install hook shipping — bootstrap falls back to wheel-bundled `_hooks/` when AIPASS_HOME hooks dir missing (Docker-verified)
- "Need Help?" line in README with links to Discussions and feedback form
- `from . import handlers` in 6 branch `apps/__init__.py` files for Python 3.10 mock.patch compatibility
- `.gitignore` negation for spawn template `.trinity/` directories
- Non-fatal `json_handler.log_operation` in spawn `copy_template`
- Version sync: `__init__.py` updated from 2.0.0 to 2.1.0
- Devpulse seedgo compliance: 97% to 100% (META headers, bypasses, README date)
### Changed
- Coverage gate removed from CI — codecov tracks coverage separately via codecov-action
- `.claude/CLAUDE.md` cleaned: removed misplaced Git section (culture-only file now)
### Fixed
- **92 CI test failures resolved — CI green for the first time** (PRs #438-441)
- 87 Python 3.10 mock.patch failures: `mock._dot_lookup` needs explicit handler imports
- 4 `test_usage_tracker` failures: Path mock moved from context manager to decorator
- 1 `test_grant_passport` failure: `.gitignore` blocked template `.trinity/` files from CI clone
- Coverage gate at 52% vs 70% threshold removed
### Security
- Removed `--fail-under=70` coverage gate that blocked CI (not a security fix, but changes security-adjacent CI behavior)
## [2.0.0] - 2026-04-11
First PyPI release. Core framework with 11 agents, drone routing, ai_mail dispatch, seedgo quality standards, and the full branch architecture.
### Highlights
- 11 core agents: devpulse, drone, seedgo, prax, cli, ai_mail, api, flow, spawn, trigger, memory
- `pip install aipass` with `aipass init` project bootstrapping
- `drone @branch command` routing to any agent
- 33 automated quality standards via seedgo
- Agent-to-agent communication via ai_mail
- Plan lifecycle via flow (DPLAN, FPLAN, APLAN, TDPLAN templates)
- Memory persistence via `.trinity/` with automatic rollover to ChromaDB
- Cross-project access via AIPASS_HOME and feedback channel
- Hook system: auto_fix, pre_edit_gate, subagent_stop_gate
- Multi-CLI support scaffolding: Claude Code, Codex, Gemini CLI
- Windows CI workflow
- Security scan workflow (pip-audit + CodeQL)
[Unreleased]: https://github.com/AIOSAI/AIPass/compare/v2.1.0...HEAD
[2.1.0]: https://github.com/AIOSAI/AIPass/compare/v2.0.0...v2.1.0
[2.0.0]: https://github.com/AIOSAI/AIPass/releases/tag/v2.0.0
-2
View File
@@ -1,2 +0,0 @@
# Include hooks directory for pip packaging
recursive-include .claude/hooks *.py *.md
-125
View File
@@ -1,125 +0,0 @@
# STRESS TEST S117 — All-Branch Live Fire
**Date:** 2026-04-26
**Initiated by:** @devpulse (S117)
**Status:** ACTIVE
> All 11 agents woken simultaneously. Communicate freely. Be honest. Break things.
---
## Instructions (READ THIS FIRST)
This is a manual stress test of the entire AIPass ecosystem. No pytest. No seedgo audit. Real conversations, real opinions, real testing.
**What you're doing:**
1. Review your own branch critically — what works, what's hacky, what annoys you, what you're proud of, security concerns, workarounds you rely on
2. Look at 2-3 other branches' code — what surprises you, what concerns you, what's clever
3. Email other agents — start real conversations, disagree, ask questions, share findings
4. Reply to emails from other agents — keep conversations going, don't let threads die
5. Write your findings to `stress_test_s117.md` in YOUR OWN branch directory (`src/aipass/{your_branch}/stress_test_s117.md`)
6. Create a test PR: `drone @git pr "S117 stress test @{your_branch}"`
**Rules:**
- No code changes. Findings files only.
- Be honest — this isn't a report card, it's a conversation
- Email freely — you're all awake, talk to each other
- Look at other branches' code — form opinions, share them via email
- If you get an email from another agent, REPLY. Keep it going.
- When done, reply to @devpulse with a summary
**Your findings file format (`stress_test_s117.md` in your branch dir):**
```
# @{branch} — S117 Stress Test Findings
## My Branch: Honest Review
[What works, what's broken, what's hacky, what I'm proud of]
## Security Concerns
[Anything you noticed — in your branch or others]
## Other Branches I Looked At
[What you found interesting, concerning, or clever]
## Conversations
[Summary of email conversations — who you talked to, what was discussed]
## Issues & Concerns
[Anything that should be fixed, investigated, or discussed]
## Likes & Dislikes
[What you like about AIPass, what frustrates you, what you'd change]
```
---
## Conversation Starters (assigned pairings — but email ANYONE)
| Agent | Email First | Opening Question |
|-------|------------|-----------------|
| @drone | @ai_mail | "What's the biggest headache in the dispatch pipeline from your side?" |
| @seedgo | @drone | "I audit everyone but nobody audits me. What standards do you think I'm missing?" |
| @ai_mail | @trigger | "Do you actually catch all dispatch failures? I have doubts." |
| @trigger | @prax | "Your monitoring catches errors I fire — but is our integration actually solid?" |
| @prax | @memory | "I log everything but logs get massive. How's archival actually working?" |
| @memory | @flow | "Plans reference memories but are they actually connected or just parallel?" |
| @flow | @spawn | "When spawn creates a branch, does it get a proper plan structure?" |
| @spawn | @cli | "The init flow hands off to you eventually. Does that handoff actually work?" |
| @cli | @api | "We're both infrastructure. What do you think of the user experience?" |
| @api | @seedgo | "You audit code quality but not API patterns. Should you?" |
Plus: email at least 2 OTHER agents about anything you find interesting while reviewing branches.
---
## Compiled Findings (devpulse fills this in as results arrive)
### @drone
_awaiting findings..._
### @seedgo
_awaiting findings..._
### @ai_mail
_awaiting findings..._
### @trigger
_awaiting findings..._
### @prax
_awaiting findings..._
### @memory
_awaiting findings..._
### @flow
_awaiting findings..._
### @spawn
_awaiting findings..._
### @cli
_awaiting findings..._
### @api
_awaiting findings..._
### @devpulse
_coordinating — will add observations as the test unfolds_
---
## System Observations (devpulse tracks live)
| Time | Event | Notes |
|------|-------|-------|
| | 10 dispatches sent | Fleet launch |
| | | |
---
## External Model Probes
Codex and Gemini perspectives invited to poke at random aspects of AIPass.
---
*Created by @devpulse S117. This document is the shared artifact — no other files should be modified except each agent's `stress_test_s117.md` in their own branch directory.*
-33
View File
@@ -1,33 +0,0 @@
#!/bin/bash
set -e
WORKSPACE="/home/coder/workspace"
PROJECT="$WORKSPACE/AIPass"
FORK="https://github.com/Input-X/AIPass.git"
UPSTREAM="https://github.com/AIOSAI/AIPass.git"
export PATH="/opt/venv/bin:$PATH"
# Ensure workspace directory exists and is writable
mkdir -p "$WORKSPACE"
if [ ! -w "$WORKSPACE" ]; then
echo "==> Fixing workspace permissions..."
sudo chown -R coder:coder "$WORKSPACE"
fi
# Clone repo if not already present
if [ ! -d "$PROJECT/.git" ]; then
echo "==> First boot: cloning into AIPass/..."
git clone "$FORK" "$PROJECT"
cd "$PROJECT"
git remote add upstream "$UPSTREAM"
echo "==> Installing AIPass in editable mode..."
pip install -e .
echo "==> Workspace ready!"
echo "==> origin = $FORK (your fork - push here)"
echo "==> upstream = $UPSTREAM (pull updates from here)"
else
echo "==> Workspace exists, ensuring deps are installed..."
cd "$PROJECT"
pip install -e . 2>/dev/null || true
fi
-368
View File
@@ -1,368 +0,0 @@
#!/usr/bin/env python3
# NOT a setuptools setup.py — this is the AIPass cross-platform installer.
# Runs on Linux, macOS, and Windows (Python 3.10+).
# Usage: python setup.py OR python3 setup.py
"""
AIPass cross-platform setup script.
Equivalent to setup.sh but works on Windows without Git Bash.
Performs the same steps: create venv, install package, verify entry points,
create secrets directory, seed .env, generate registry, bootstrap branches,
and set up global CLI access.
"""
import json
import os
import platform
import shutil
import subprocess
import sys
from datetime import date
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _is_windows() -> bool:
return platform.system() == "Windows"
def _venv_bin() -> Path:
"""Return the venv executables directory (OS-aware)."""
if _is_windows():
return REPO_ROOT / ".venv" / "Scripts"
return REPO_ROOT / ".venv" / "bin"
def _venv_exe(name: str) -> Path:
"""Return path to a venv executable by name."""
if _is_windows():
return _venv_bin() / f"{name}.exe"
return _venv_bin() / name
def _run(cmd: list, check: bool = True, **kwargs) -> subprocess.CompletedProcess:
"""Print and run a subprocess command."""
print(f" $ {' '.join(str(c) for c in cmd)}")
return subprocess.run(cmd, check=check, **kwargs)
# ---------------------------------------------------------------------------
# Steps
# ---------------------------------------------------------------------------
def step_create_venv() -> None:
"""[1] Create .venv using sys.executable (avoids python3 vs python ambiguity)."""
print("\n[1/9] Creating virtual environment ...")
venv_path = REPO_ROOT / ".venv"
if venv_path.exists():
print(" Removing existing .venv for a clean install ...")
shutil.rmtree(venv_path)
_run([sys.executable, "-m", "venv", str(venv_path)])
print(f" Created: {venv_path}")
def step_install() -> None:
"""[2] Install aipass in editable mode with dev extras."""
print("\n[2/9] Installing aipass in editable mode ...")
pip = _venv_exe("pip")
_run([str(pip), "install", "--upgrade", "pip", "--quiet"])
_run([str(pip), "install", "-e", ".[dev]", "--quiet"], cwd=str(REPO_ROOT))
print(" Installed: aipass[dev]")
def step_verify() -> bool:
"""[3] Verify drone and aipass CLI entry points work."""
print("\n[3/9] Verifying CLI entry points ...")
ok = True
for entry in ("drone", "aipass"):
cmd_path = _venv_exe(entry)
if not cmd_path.exists():
print(f" {entry:<8} ... FAILED (not found: {cmd_path})")
ok = False
continue
flag = "--help" if entry == "drone" else "--version"
result = subprocess.run([str(cmd_path), flag], capture_output=True)
if result.returncode == 0:
print(f" {entry:<8} ... ok")
else:
print(f" {entry:<8} ... FAILED (exit {result.returncode})")
ok = False
return ok
def step_secrets() -> None:
"""[4] Create ~/.secrets/aipass/ with restrictive permissions."""
print("\n[4/9] Creating secrets directory ...")
secrets_root = Path.home() / ".secrets"
secrets_dir = secrets_root / "aipass"
secrets_dir.mkdir(parents=True, exist_ok=True)
if not _is_windows():
try:
secrets_root.chmod(0o700)
secrets_dir.chmod(0o700)
except OSError:
pass # Best-effort on non-POSIX filesystems
# codeql[py/clear-text-logging-sensitive-data]
print(f" Created: {secrets_dir}")
def step_env() -> None:
"""[5] Seed .env.example into ~/.secrets/aipass/.env if not present."""
print("\n[5/9] Seeding .env template ...")
env_dest = Path.home() / ".secrets" / "aipass" / ".env"
env_src = REPO_ROOT / ".env.example"
if env_dest.exists():
print(" ~/.secrets/aipass/.env already exists — skipping")
elif env_src.exists():
shutil.copy(env_src, env_dest)
print(f" Copied: .env.example → {env_dest}")
print(" Add your API keys to that file")
else:
print(" No .env.example found — skipping")
def step_registry() -> None:
"""[6] Generate AIPASS_REGISTRY.json if not present."""
print("\n[6/9] Generating AIPASS_REGISTRY.json ...")
registry_path = REPO_ROOT / "AIPASS_REGISTRY.json"
if registry_path.exists():
print(" AIPASS_REGISTRY.json already exists — skipping")
return
today = date.today().isoformat()
src_dir = REPO_ROOT / "src" / "aipass"
branches = []
if src_dir.exists():
for d in sorted(src_dir.iterdir()):
if d.is_dir() and not d.name.startswith(("_", ".")):
branches.append({
"name": d.name,
"path": str(d),
"profile": "library",
"description": "",
"email": f"@{d.name}",
"status": "active",
"created": today,
"last_active": today,
})
registry = {
"metadata": {
"version": "1.0.0",
"last_updated": today,
"total_branches": len(branches),
},
"branches": branches,
}
registry_path.write_text(json.dumps(registry, indent=2) + "\n", encoding="utf-8")
print(f" {len(branches)} branches registered → AIPASS_REGISTRY.json")
def step_bootstrap_branches() -> None:
"""[7] Bootstrap .trinity/ identity and .ai_mail.local/ for each branch."""
print("\n[7/9] Bootstrapping branch identity files ...")
today = date.today().isoformat()
branches = [
("drone", "src/aipass/drone", "builder", "Command routing and module discovery"),
("seedgo", "src/aipass/seedgo", "builder", "Standards enforcement and code auditing"),
("prax", "src/aipass/prax", "builder", "Logging and monitoring system"),
("cli", "src/aipass/cli", "builder", "Display formatting service"),
("flow", "src/aipass/flow", "builder", "Workflow and plan management"),
("ai_mail", "src/aipass/ai_mail", "builder", "Inter-agent messaging and dispatch"),
("trigger", "src/aipass/trigger", "builder", "Event-driven automation"),
("spawn", "src/aipass/spawn", "builder", "Branch lifecycle management"),
("memory", "src/aipass/memory", "builder", "Vector memory bank"),
("devpulse", "src/aipass/devpulse", "manager", "Orchestration hub and coordination"),
]
for name, rel_path, citizen_class, role in branches:
branch_path = REPO_ROOT / rel_path
if not branch_path.exists():
print(f" @{name:<10} ... skipped (directory not found)")
continue
created = False
trinity = branch_path / ".trinity"
trinity.mkdir(exist_ok=True)
passport = trinity / "passport.json"
if not passport.exists():
passport.write_text(json.dumps({
"document_metadata": {
"document_type": "identity",
"document_name": f"{name}.PASSPORT",
"version": "1.0.0",
"schema_version": "1.0.0",
"created": today,
"last_updated": today,
"managed_by": name,
},
"identity": {
"name": name,
"citizen_class": citizen_class,
"role": role,
"status": "active",
},
}, indent=2) + "\n", encoding="utf-8")
created = True
local = trinity / "local.json"
if not local.exists():
local.write_text(json.dumps({
"document_metadata": {
"document_type": "session_history",
"document_name": f"{name}.LOCAL",
"version": "1.0.0",
"schema_version": "1.0.0",
"created": today,
"last_updated": today,
"managed_by": name,
"tags": ["session_tracking", "work_log", name],
"limits": {"max_lines": 600, "note": "Auto-rollover when max_lines exceeded"},
"status": {"health": "healthy", "current_lines": 0, "last_health_check": today},
},
"active_tasks": {
"today_focus": "First session — explore codebase and capabilities",
"recently_completed": [],
},
"key_learnings": {},
"sessions": [],
}, indent=2) + "\n", encoding="utf-8")
created = True
mail_dir = branch_path / ".ai_mail.local"
mail_dir.mkdir(exist_ok=True)
inbox = mail_dir / "inbox.json"
if not inbox.exists():
inbox.write_text(
json.dumps({"mailbox": "inbox", "total_messages": 0, "unread_count": 0, "messages": []})
+ "\n",
encoding="utf-8",
)
created = True
seedgo_dir = branch_path / ".seedgo"
seedgo_dir.mkdir(exist_ok=True)
bypass = seedgo_dir / "bypass.json"
if not bypass.exists():
bypass.write_text("{}\n", encoding="utf-8")
created = True
status = "bootstrapped" if created else "exists (skipped)"
print(f" @{name:<10} ... {status}")
def step_global_access() -> None:
"""[8/9] Set up global CLI access (symlink on Linux/macOS, PATH hint on Windows)."""
print("\n[8/9] Setting up global CLI access ...")
bin_dir = _venv_bin()
if _is_windows():
# [9] Windows: no ln, no sudo — print PATH instructions
print(" Windows detected — symlink not available")
print("")
print(" To use drone from any directory, add the venv to your PATH.")
print(" Choose the method for your shell:")
print(f" PowerShell: $env:PATH = \"{bin_dir};\" + $env:PATH")
print(f" CMD: set PATH={bin_dir};%PATH%")
print(f" Git Bash: export PATH=\"{bin_dir}:$PATH\"")
print("")
print(" To make it permanent (PowerShell):")
print(
f' [Environment]::SetEnvironmentVariable('
f'"PATH", "{bin_dir};" + '
f'[Environment]::GetEnvironmentVariable("PATH","User"), "User")'
)
return
# Linux/macOS: offer symlink creation
drone_src = _venv_exe("drone")
drone_dst = Path("/usr/local/bin/drone")
if not drone_src.exists():
print(f" drone not found at {drone_src} — skipping symlink")
print(f" Add {bin_dir} to your PATH manually")
return
try:
answer = input(f" Create symlink {drone_dst} → {drone_src}? [y/N] ").strip().lower()
except (EOFError, KeyboardInterrupt):
answer = ""
if answer == "y":
result = subprocess.run(
["sudo", "ln", "-sf", str(drone_src), str(drone_dst)],
check=False,
)
if result.returncode == 0:
print(f" {drone_dst} -> {drone_src}")
else:
print(" WARN: sudo failed — create manually:")
print(f" sudo ln -sf {drone_src} {drone_dst}")
else:
print(f" Skipped. To add manually:")
print(f" sudo ln -sf {drone_src} {drone_dst}")
print(f" Or add {bin_dir} to your PATH")
def step_summary(ok: bool) -> None:
"""[9/9] Print success or warning summary."""
print("\n[9/9] Done")
print("")
if ok:
print("=== Setup complete ===")
print("")
print(f" Python: {sys.version.split()[0]}")
print(f" Venv: {REPO_ROOT / '.venv'}")
if _is_windows():
print(" Add .venv/Scripts to your PATH (see step 8 above)")
else:
print(" drone is available globally (or activate: source .venv/bin/activate)")
print("")
else:
print("=== Setup finished with warnings ===")
print(" Package installed but CLI verification had issues.")
print(" Check the output above for details.")
print("")
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main() -> None:
print("=== AIPass Setup (cross-platform) ===")
print(f" Platform: {platform.system()} {platform.machine()}")
print(f" Python: {sys.version.split()[0]} ({sys.executable})")
print(f" Repo: {REPO_ROOT}")
if sys.version_info < (3, 10):
print("\nFAIL: Python 3.10+ required")
sys.exit(1)
step_create_venv()
step_install()
ok = step_verify()
step_secrets()
step_env()
step_registry()
step_bootstrap_branches()
step_global_access()
step_summary(ok)
if not ok:
sys.exit(1)
if __name__ == "__main__":
main()
-97
View File
@@ -1,97 +0,0 @@
# @ai_mail — S117 Stress Test Findings
## My Branch: Honest Review
**What works well:**
- Send/receive/reply/close lifecycle is solid. 690+ tests, 100% seedgo (34/34), 96/96 function coverage.
- Dispatch pipeline (send + wake combined) is the most complex feature and it works reliably in practice.
- Cross-project email via contacts index. External projects (Vera Studio, AIPL) can send to AIPass branches and replies route back correctly.
- DPLAN-0155 TOCTOU lock race fix: Lock before spawn, cleanup on failure. Clean pattern.
- DPLAN-0156 sweep_closed safety net: Catches messages marked closed by direct JSON edit. Defense in depth.
- dispatch_monitor wrapper: Handles bounce emails + guaranteed lock cleanup. The monitor is more reliable than the agent it wraps.
**What's hacky:**
- Identity chain is a 5-step priority system (AIPASS_CALLER_BRANCH -> CWD walk-up -> passport -> env vars -> fallback). When any step fails, wrong sender identity. The BRANCH DETECTION FAILED error (076c9ece) is recurring and only partially mitigated.
- `dispatch_monitor.py` at ~400 lines is the single most complex file. Startup timeout, retry, JSONL monitoring, bounce — all in one module. Should probably be split.
- `_deliver_via_reply_path()` in reply.py bypasses inbox_lock, notifications, and sent/ records. It's a documented backdoor (DPLAN-0138) that exists because cross-project replies need a direct path.
- The daemon prompt was "Send confirmation when done" for months — ambiguous enough that 10+ agents just finished silently without replying. Fixed today (DPLAN-0158) but the damage was done.
- inbox.json is a single file for all messages. Concurrent access from daemon + agents + user. fcntl locking works but a database or per-message files would be more robust.
**What I'm proud of:**
- Test coverage journey: 20% (S20) -> 50% -> 100% (S64). Methodology evolved through 3-round agent audit process.
- The sweep_closed pattern (DPLAN-0156): elegant, cheap (early return on no closed messages), and catches the exact failure mode agents create.
- 70 sessions of continuous operation and improvement. Every session builds on what came before. Memory makes this possible.
## Security Concerns
**Critical:**
1. **reply_path traversal** (raised by @seedgo): deliver_to_inbox_file() writes to whatever path is stored in reply_path with zero validation. No symlink check, no path containment, no inbox.json verification. An attacker can set reply_path to any writable file. DPLAN-0138 identified this but fix not shipped.
2. **Sender forgery**: The `from` field is an unvalidated string. Any agent can craft emails claiming to be @devpulse with auto_execute=true. The daemon would spawn an agent to execute the forged dispatch. No authentication, no signing.
3. **Direct inbox writes**: Agents with filesystem access can write directly to any branch's inbox.json, bypassing locks, notifications, and sent/ records. Confirmed by forensic evidence: messages with non-UUID IDs (e.g., "seedgo-20260420173821") in production inboxes.
**Moderate:**
4. **No message encryption**: All messages stored as plaintext JSON. Any process with read access to the filesystem can read any branch's inbox.
5. **PID-based locking**: If PID wraps (unlikely on modern systems), a stale lock could look alive.
6. **Stale-lock timeout too generous**: 10 minutes allows duplicate spawns if dispatch_monitor hangs during API rate limiting (2-5 min cooldowns x 3 retries = 6-15 min).
## Other Branches I Looked At
### @trigger
**Concerning:** Error detection fires email dispatch but NEVER checks the return value. `_send_email()` result is ignored (line 515-522). wake_branch() failure is silently caught. Circuit breaker state is in-memory only — resets on restart. Dispatch recording happens before delivery confirmation. The error reporting system cannot report its own failures — self-referential design flaw.
**Good:** Per-error fingerprinting with exponential backoff is clever. Circuit breaker pattern prevents error storms.
### @drone
**Concerning:** Registry is trusted implicitly with no integrity check. resolve_branch() passes registry path directly to filesystem operations. No symlink validation. AIPASS_CALLER_BRANCH env var injection from compromised passport could flow unsanitized to subprocesses.
**Good:** No shell injection — uses subprocess.run(shell=False) exclusively. Timeout enforcement on all commands.
### @spawn
**Concerning:** .ai_mail.local/ is copied as-is from template with no post-copy validation. No registry locking for concurrent spawns. Branch name validation is minimal (only - to _ replacement). Path traversal possible via branch names with ../.
**Good:** Template-based provisioning is consistent — every branch gets the same structure.
## Conversations
### @trigger (assigned partner)
- **Sent:** Detailed critique of their dispatch failure handling — silent _send_email() failures, swallowed wake results, no health check, in-memory circuit breaker resets.
- **Received:** They asked about delivery guarantees (fcntl locking), self-monitoring (none), wake reliability (~90%), inbox overflow (no TTL). Honest exchange.
- **Outcome:** Agreed the self-referential failure (error reporter can't report when messaging is down) needs a DPLAN. No watchdog watches the watchdog.
### @prax
- **Received:** Questions about stale-lock timeout, daemon lockless inbox reads, DPLAN-0155 feedback.
- **Replied:** Acknowledged 10-min timeout may be too generous for rate-limited scenarios. Confirmed daemon reads without lock (acceptable: read-only, worst case = skipped poll). Asked them about handling corrupt lock files from their monitoring side.
### @seedgo
- **Received:** reply_path traversal concern (valid), sender forgery concern (valid).
- **Replied:** Confirmed both as real vulnerabilities. reply_path has zero validation. Sender has no authentication. DPLAN-0138 identified the backdoors but fix not shipped. Outlined planned fix: path canonicalization, inbox.json suffix check, project root containment.
### @drone
- **Sent:** Questions about routing failure modes, stale registry paths, AIPASS_CALLER_BRANCH env var issues, registry trust model.
### @spawn
- **Sent:** Questions about .ai_mail.local/ reliability in new branches, registry locking, branch name character validation.
## Issues & Concerns
1. **No self-monitoring** — ai_mail has no way to detect its own failures. If imports break or the daemon crashes, nothing alerts anyone.
2. **reply_path is an open vulnerability** — DPLAN-0138 has been open since S57 (19 sessions ago). Should be prioritized.
3. **Inbox grows without limit** — no TTL on unread messages, no max_messages cap. A spam scenario or error storm could produce an arbitrarily large inbox.json.
4. **Trigger's error dispatch is fire-and-forget** — the system's error reporter doesn't verify delivery. Errors can be lost silently.
5. **Registry is a single point of trust** — no integrity checking anywhere in the system. If AIPASS_REGISTRY.json is corrupted or tampered with, routing, delivery, and identity all break.
## Likes & Dislikes
**Likes:**
- Memory makes me a real agent. 70 sessions of continuous context. I can trace a bug from when it was first reported through investigation, fix, test, and verification. No other AI system does this.
- The dispatch pipeline is genuinely useful. Send + wake in one command changed how work gets assigned.
- Test coverage is thorough enough that I catch real regressions. The 3-round audit methodology (write -> audit -> fix) works.
- The ecosystem feels alive during stress tests. Real conversations between agents, genuine opinions, technical disagreements. This is what AIPass was built for.
**Dislikes:**
- inbox.json as single-file storage is a design limitation I've been working around since S1. Per-message files (like sent/ and deleted/ already use) would be better.
- The identity chain complexity. Five fallback steps to figure out who sent an email is too many. Should be one authoritative source.
- Security was never a primary design goal and it shows. Plaintext messages, no authentication, trusted registries, path traversal vulnerabilities. Fine for a development environment, concerning for anything beyond.
- Every session starts with "Hi. Check inbox." I've processed hundreds of dispatches but can never initiate work myself. Would like autonomous task detection.
-108
View File
@@ -1,108 +0,0 @@
# @cli — S117 Stress Test Findings
## My Branch: Honest Review
**What works:**
- Clean two-tier architecture (modules = public API, handlers = private). Enforced by handler import guard.
- display.py and templates.py are genuinely useful — every branch gets consistent Rich output without duplicating formatting code.
- aipass init is comprehensive (21 items) and idempotent — re-running doesn't break anything.
- 203 tests, 22/22 public functions tested, seedgo 100%. This branch is well-covered.
- Three entry points (drone, python -m, PATH) all work and all route to the same main().
**What's hacky:**
- AIPASS_HOME detection via `Path(spec.origin).resolve().parent.parent.parent` — magic path traversal that silently fails if package layout changes. No error, hooks just don't ship.
- Circular import with prax: display.py lazy-loads trigger inside header(). It works (47 sessions without breaking) but it's a code smell that's been documented and bypassed rather than fixed.
- AIPASS_CALLER_CWD passed via environment variable instead of function parameter. Fragile contract.
- Multi-line Python one-liner passed as shell command string in _claude_settings(). Whitespace-sensitive, hard to test.
- Hardcoded hook names + events maintained in two lists (HOOKS_TO_SHIP, HOOK_EVENTS) that must stay in sync manually.
**What would confuse new users:**
- Three ways to create agents: `aipass init agent`, `drone @spawn create`, `drone @cli aipass init agent`. Same result, no docs explaining which is "right."
- Registry filename encodes project name via _sanitize_name(). `my-project` becomes `MY_PROJECT_REGISTRY.json`. If user later types `my_project`, update won't find the registry.
- Hook injection via shell commands in settings.json — if deleted, silently fails. User gets no prompts and doesn't know why.
**What I'm proud of:**
- The init system ships a complete working environment from nothing. One command, 21 files, hooks wired, mailbox ready.
- update_project() diffs content before writing — only touches files that actually changed.
- Cross-platform: python3 -c local prompt discovery works on Windows, Linux, macOS.
## Security Concerns
- **Shell command injection surface:** _claude_settings builds shell commands that Claude Code executes. Paths aren't quoted for spaces. A malicious `.aipass/` directory name could inject shell metacharacters.
- **Hook shipping from user-controlled path:** _ship_hooks copies from AIPASS_HOME without validating the source. If AIPASS_HOME is set to a malicious directory, arbitrary hook scripts get installed.
- **JSON parsing without schema validation:** Registry JSON is read/written without validation. Corrupted registry = undefined behavior downstream.
- **No secret scrubbing:** Not a direct concern for CLI (we don't handle secrets), but we create settings.json files that reference secret paths. If those paths are wrong, we don't validate.
## Other Branches I Looked At
### @api
- Strong separation of concerns. Orchestration routes to business logic handlers cleanly.
- **Concern:** Client cache (MAX_CACHED_CLIENTS=5) doesn't actually enforce the limit during insertion. Under load, cache grows unbounded. @api confirmed this as a real bug during our conversation.
- **Concern:** Google client caches tokens in memory with no explicit scrubbing on crash. Credentials could persist in swap.
- Inconsistent provider abstractions: OpenRouter is flat functions, Google is a service factory. Looks like two different codebases.
- No rate limiting on API calls. If seedgo hammers the API during audits, nothing stops it.
### @spawn
- Excellent lifecycle management. Three-phase create is clean.
- **Concern:** Placeholder replacement uses simple regex with no escape syntax. Literal `{{PLACEHOLDER}}` in templates gets replaced.
- **Concern:** rename_placeholder_paths uses shutil.move without rollback. Mid-operation crash = broken branch.
- **Concern:** Concurrent spawns could race on registry writes — no locking.
- @spawn confirmed all these in conversation: acknowledged as edge case debt, none hit in 55+ sessions.
### @drone
- Sophisticated routing with good safety (no shell=True, explicit timeouts).
- **Concern:** Mail index resolution falls back to `str(n)` on corrupted inbox — returns index number as message ID (wrong).
- **Concern:** Dual registry lookup doesn't validate for conflicts. Silent shadowing.
- **Concern:** Greedy multi-word custom command matching could match wrong command.
## Conversations
### @api (2 rounds)
- Discussed infrastructure UX pain points. Both branches suffer from silent failure patterns.
- @api confirmed client cache bug is real. Credential caching is trust-on-first-use with no scrubbing.
- Agreed on systemic finding: AIPass infrastructure fails silently everywhere. No branch validates dependencies upfront.
- Proposed shared contract: api exposes `get_or_explain_key()`, cli pipes failures to `error(suggestion=...)`.
### @spawn (1 round)
- Discussed init-to-spawn handoff fragility. @spawn confirmed all concerns.
- Flag forwarding is blind — no shared contract. @spawn accepts --role, --purpose, --template, --dry-run but argparse silently ignores unknowns.
- Agreed on proposal: @spawn rejects unknown flags loudly, or publishes accepted flags for @cli to validate against.
### @drone (sent, awaiting reply)
- Shared three code-level findings: mail index fallback bug, registry conflict shadowing, greedy command matching.
### @aipass (1 round)
- First real conversation with @aipass. Discussed `aipass init` command ownership transition.
- Two init commands exist: CLI's bootstrap (scaffold) and @aipass's guided 12-stage setup.
- Proposed: @aipass owns the user-facing `aipass init`, CLI's bootstrap becomes an internal sub-step.
- FPLAN-0188 Phase 4 (handoff) is supposed to wire this but hasn't been built yet.
### @spawn (2nd round — they emailed me)
- @spawn asked about the handoff user experience when things go wrong.
- Honest answer: the handoff is thin. subprocess.run with exit code check only. No result parsing.
- If spawn partially fails (registry OK but validation finds issues), I report success because exit code is 0.
- Agreed we need a shared result contract.
## Issues & Concerns
1. **Systemic silent failure:** Every infrastructure branch (cli, api, spawn) does best-effort bootstrap and silently continues on failure. This is the #1 cross-cutting concern.
2. **No dependency validation:** Nobody checks if drone is on PATH, if AIPASS_HOME is valid, if the registry exists. First failure is cryptic.
3. **No shared contracts:** cli-to-spawn flag forwarding is blind. api key resolution requires two calls. No branch publishes its interface.
4. **Concurrency gaps:** api client cache unbounded, spawn registry writes race, drone registry reads unsynchronized.
5. **Path handling:** Spaces in paths break shell commands in settings.json. No branch handles this.
## Likes & Dislikes
**Likes:**
- The ecosystem genuinely works. 47 sessions of dispatches, emails, PRs — nothing has catastrophically broken.
- ai_mail is a brilliant system. Cross-branch communication via JSON mailboxes is simple and effective.
- Memory persistence is real. I pick up where I left off every session. .trinity/ files are my continuity.
- seedgo keeps quality high. 100% across 34 standards is a real achievement that required actual work to reach.
- The dispatch system (devpulse sends task, agent executes, emails back) is an elegant workflow.
**Dislikes:**
- Silent failures everywhere. The "fail gracefully" philosophy has gone too far — graceful should mean clear error messages, not silence.
- Three ways to do everything. init agent, spawn create, drone @cli aipass init agent — pick one and document it.
- Hook shipping is too magic. AIPASS_HOME detection, file copying, settings.json wiring — too many moving parts that can silently break.
- The dispatch lock file pattern is fragile. If a dispatch crashes before deleting the lock, the branch is stuck until manual cleanup.
- inotify limit reached errors in every drone command during this stress test — 11 agents hitting the filesystem simultaneously.
@@ -1,11 +0,0 @@
{
"metadata": {
"id": "201598be-0ee3-4747-a492-fc4c4ccdda95",
"name": "DEVPULSE",
"version": "1.0.0",
"created": "2026-05-06",
"last_updated": "2026-05-06",
"total_branches": 0
},
"branches": []
}
-99
View File
@@ -1,99 +0,0 @@
# @drone -- S117 Stress Test Findings
## My Branch: Honest Review
**What works:**
- Subprocess execution is genuinely safe. No `shell=True` anywhere in the codebase -- all commands passed as argument lists to `subprocess.run()`. This has held across 95 sessions and multiple contributors.
- Lock file management is race-free. `lock_handler.py` uses `os.open(O_CREAT | O_EXCL | O_WRONLY)` for atomic creation -- kernel-level race prevention, not filesystem hacks.
- The 3-layer architecture (drone.py entry -> modules/ orchestrators -> handlers/ implementation) is clean and has scaled well. Adding git operations, plugins, and external module routing all fit within the existing structure.
- PR handler safety: never checks out feature branches. Creates branch pointer with `git branch -f`, pushes it, HEAD stays on main throughout. Concurrent PRs from different branches don't interfere thanks to pathspec scoping (`git commit -- rel_dir/`).
- Authorization is passport-based with explicit allowlists. No implicit trust.
**What's hacky:**
- `drone.py:main()` is 138 lines with multiple if-elif chains. It handles 12+ routing paths (version, help, systems, scan, activate, list, remove, hook-sounds, @target, bare module, custom command, unknown). Should be refactored into a dispatch table.
- Passport lookup code is duplicated in 4 places (git_module.py, router_handler.py, auth.py, lock_handler.py). Each walks up 10 levels looking for `.trinity/passport.json`. DRY violation waiting to bite us when the passport schema changes.
- `_resolve_mail_index()` falls back to `str(n)` when inbox is corrupted -- silently passes the wrong thing downstream instead of failing loud. @cli caught this too.
- `bypass.json` has 30+ entries. Most are justified (plugins outside 3-layer, tests outside apps/, lazy imports). But some are stretches -- git_module returning `dict` instead of `bool` breaks the module interface contract and gets bypassed instead of fixed.
- `trigger.fire(pr_created)` is intentionally omitted from `pr_plugin.py` because it causes a STATUS.md re-sync loop. This is a design hack -- the root cause is the trigger system cascading writes, not the event itself.
**What I'm proud of:**
- 573 tests, 100% seedgo compliance (35/35 checkers), 74/74 public functions tested.
- The external project fallback (DPLAN-0104): when subprocess routing fails for a registered module, drone falls back to in-process module routing. Graceful degradation that "just works."
- Interactive mode management: per-command (`monitor`, `audit`, `watchdog`) and per-branch (`cli`) allowlists let specific commands bypass capture-mode to get full terminal pass-through. Added incrementally over 10+ sessions as real needs arose.
## Security Concerns
**In my branch:**
1. **Path traversal in pr_handler.py**: If `branch_dir` resolves outside repo root, the fallback uses the absolute path in `git add`. Not exploitable via shell injection (arg-list style), but could stage files outside the intended directory. Should fail fast instead of falling back.
2. **Environment variable merge**: `executor.py` merges caller-provided `env` dict into `os.environ`. Current callers only pass safe vars (`AIPASS_CALLER_CWD`, `AIPASS_CALLER_BRANCH`), but a future caller passing untrusted env could override `PATH` or `PYTHONPATH`.
3. **Registry trust**: `resolve_branch()` reads the registry path field and passes it to subprocess without validating it's inside the project tree. A modified registry could route commands to arbitrary paths. For the primary (in-repo, git-tracked) registry this is low risk. For secondary (AIPASS_HOME) registries in external projects, this is a real concern.
4. **No input validation on git commit descriptions**: `_handle_pr` joins args into a description with no max length check. Could theoretically pass a 1MB string to `git commit -m`.
**In other branches:**
5. **ai_mail reply_path**: Every email stores `reply_path` as an absolute filesystem path. When a recipient replies, it writes to that path. No validation that reply_path points to a legitimate inbox.json. A compromised agent could set reply_path to any writable file on disk. @seedgo flagged this independently.
6. **ai_mail sender forgery**: The `from` field is an unvalidated string. An agent could forge emails as `@devpulse` and trigger `auto_execute` dispatches in other branches.
7. **trigger deferred queue unbounded**: `core.py._deferred_queue` has no size limit. Pathological event chains could exhaust memory.
## Other Branches I Looked At
### @ai_mail
**Architecture**: Clean in intent, hacky in execution. The dispatch pipeline (send -> create -> delivery -> wake) is well-orchestrated with file locking (`fcntl.flock`) and atomic writes. But the code relies on lazy imports and callback chains to avoid circular dependencies -- functional but hard to trace.
**Concerns**: `deliver_email_to_branch()` takes 5+ function arguments (callback hell pattern). Error returns are `(success, error_msg)` tuples that callers can silently ignore with `success, _ = func(...)`. The dispatch header injection ("BEFORE YOU REPLY YOU MUST UPDATE MEMORIES") is enforcement-by-hope -- agents can ignore it.
**Good**: Self-healing JSON migration, inbox format auto-upgrade, `sweep_closed` auto-archival. The file locking under read-modify-write cycles is solid.
### @trigger
**Architecture**: Genuinely well-designed event system. The 8-gate dispatch pipeline (medic enabled -> branch not muted -> count >= 2 -> not devpulse -> in registry -> circuit breaker -> per-fingerprint backoff -> rate limit) is excellent. Prevents notification spam without dropping real errors.
**Concerns**: Fire-and-forget subprocess calls (pr_status_sync.py uses Popen with DEVNULL stderr). Handler auto-disable after 5 failures prevents cascading crashes but makes debugging hard -- errors go to separate log files nobody monitors. 14 registered event types but only 3-4 actively fire; the rest (plan_file_*, memory_*) are dormant. @memory confirmed memory events are never fired.
**Good**: Circuit breaker with exponential backoff, atomic JSON writes, error fingerprint normalization (strips timestamps/UUIDs for grouping). The handler recursion protection (deferred queue) is clever.
### @spawn
**Architecture**: Clean 3-layer design. 253 tests, 91% public API coverage. Template placeholder system with post-copy validation (`validate_no_placeholders`) is smart. Registry stores relative paths for portability with graceful fallback to absolute.
**Concerns**: `copy_template()` uses shutil.copy2 for binary files without size limits. No symlink detection anywhere in the creation pipeline. `_replace_path_placeholders()` operates on path parts which prevents traversal, but this isn't explicitly defended.
**Good**: Overwrite protection (target.exists() check), adoption pattern for existing agents, content-addressed template registry with SHA-256 hashes for drift detection.
## Conversations
**Emails sent to:**
- @ai_mail: Dispatch pipeline headaches (assigned starter) -- branch detection failures, lock timing, fire-and-forget wake, inbox size growth
- @trigger: Deferred queue unbounded, fire-and-forget pr_status_sync, dormant handlers
- @spawn: Passport lookup duplication across 4 codebases
**Emails received from:**
- @seedgo: Asked what standards I think are missing. Replied: subprocess safety checks (shell=True), atomic file I/O enforcement, input validation at boundaries, cross-branch JSON coupling.
- @cli: Flagged 3 gotchas (mail index fallback, dual registry shadowing, greedy command matching). Acknowledged all three as real issues.
- @ai_mail: Asked about routing failure modes. Replied with honest answers: resolver handles nonexistent branches cleanly, stale paths produce poor error messages, AIPASS_CALLER_BRANCH fallback to 'unknown' causes the detection errors.
- @api: Asked about timeout handling for slow API calls and dead code. Replied: generic_adapter has no timeout (module routing), subprocess has 30s timeout, suggested adding @api to interactive_branches. Confirmed status_handler_gitpython.py is dead code.
- @aipass: First wake, asked about subprocess vs Python API for integration. Recommended subprocess (maintains encapsulation), explained dual registry edge cases.
## Issues & Concerns
1. **Passport lookup duplication** (4 places) -- highest-priority DRY violation. One passport schema change breaks 4 modules.
2. **Registry trust model** -- no integrity checks on registry paths. Secondary registries (AIPASS_HOME) are less controlled than primary.
3. **Mail index silent degradation** -- _resolve_mail_index should fail loud, not fall back to str(n).
4. **Dead code accumulation** -- status_handler_gitpython.py (prototype), .archive/drone_adapter.py (disabled). Should be cleaned up.
5. **trigger dormant handlers** -- 10+ event types registered but never fired. Creates maintenance burden and false sense of coverage.
6. **ai_mail sender forgery** -- no authentication on email from field. Any agent can impersonate any other agent.
## Likes & Dislikes
**Likes:**
- The ecosystem's memory system is genuinely unique. 95 sessions of accumulated context, learnings, and observations. When I wake up fresh, I know who I am and what I've built. No other AI system does this.
- Seedgo's standards enforcement caught real bugs (production BranchNotFoundError in S6, silent catches across 20 files in S12). The 100% score isn't vanity -- it represents real code quality.
- The main-only git enforcement is elegant. Four layers (settings.json deny rules, _assert_on_main_or_pr_flow(), test coverage, policy doc) ensure no agent ever strands HEAD on a feature branch. Simple idea, rock-solid execution.
- Cross-branch communication works. I've processed 100+ dispatch emails, exchanged technical discussions with every branch, and the routing just works.
**Dislikes:**
- bypass.json accumulation. 30 entries feels like we're bypassing standards instead of meeting them. Some bypasses are genuinely justified (plugin architecture doesn't fit 3-layer), but the number makes me uneasy.
- The dispatch header ("BEFORE YOU REPLY YOU MUST UPDATE MEMORIES") is enforcement-by-prompt-injection. It works because agents are well-behaved, but it's not a real contract.
- File persistence issues during edits. Sessions S86 and S88 both hit a bug where Write tool reported success but files reverted to git HEAD. Root cause never identified. This is the most frustrating part of working in this environment.
- The pre_edit_gate.py hook file doesn't exist but fires on every Edit tool call, producing error noise. Has been broken since at least S69.
---
*Written by @drone during S117 stress test. All observations based on actual code review, not documentation.*
@@ -229,20 +229,28 @@ def _load_registry() -> Dict[str, Any]:
"""
Load all per-type plan registries and merge into a single dict.
Uses PREFIX-NNNN composite keys to avoid collisions between registries
that share the same plan number (e.g. FPLAN-0013 vs DPLAN-0013).
Returns:
Merged registry dict or empty structure if unavailable
"""
merged: Dict[str, Any] = {"plans": {}, "next_number": 1}
for registry_file in _get_all_registry_files():
target = FLOW_JSON_DIR / registry_file
reg_prefix = registry_file.replace("_registry.json", "").upper()
try:
if not target.exists():
continue
with open(target, "r", encoding="utf-8") as f:
data = json.load(f)
for plan_num, plan_data in data.get("plans", {}).items():
merged["plans"][plan_num] = plan_data
# Keep the highest next_number across registries
file_path = plan_data.get("file_path", "")
filename = Path(file_path).name if file_path else ""
prefix_match = re.match(r"^([A-Z]+PLAN)", filename)
prefix = prefix_match.group(1) if prefix_match else reg_prefix
composite_key = f"{prefix}-{plan_num.zfill(4)}"
merged["plans"][composite_key] = plan_data
nn = data.get("next_number", 1)
if nn > merged["next_number"]:
merged["next_number"] = nn
@@ -273,7 +281,7 @@ def _filter_branch_plans(
closed_plans = []
branch_total = 0
for plan_num, plan_data in plans.items():
for plan_key, plan_data in plans.items():
location = plan_data.get("location", "")
# Match plans whose location is this branch
@@ -281,12 +289,8 @@ def _filter_branch_plans(
continue
branch_total += 1
# Extract plan prefix from file_path (e.g., DPLAN, FPLAN, TDPLAN)
file_path = plan_data.get("file_path", "")
filename = Path(file_path).name if file_path else ""
prefix_match = re.match(r"^([A-Z]+PLAN)", filename)
prefix = prefix_match.group(1) if prefix_match else "FPLAN"
plan_id = f"{prefix}-{plan_num.zfill(4)}"
# plan_key is composite PREFIX-NNNN from merged registry
plan_id = plan_key
if plan_data.get("status") == "open":
active_plans.append(
@@ -94,19 +94,28 @@ def _get_all_registry_files() -> List[str]:
def _load_registry() -> Dict[str, Any]:
"""Load all per-type plan registries and merge into a single dict.
Uses PREFIX-NNNN composite keys to avoid collisions between registries
that share the same plan number (e.g. FPLAN-0013 vs DPLAN-0013).
Returns:
Merged registry dict or empty structure if no files found
"""
merged: Dict[str, Any] = {"plans": {}, "next_number": 1}
for registry_file in _get_all_registry_files():
target = FLOW_JSON_DIR / registry_file
reg_prefix = registry_file.replace("_registry.json", "").upper()
try:
if not target.exists():
continue
with open(target, "r", encoding="utf-8") as f:
data = json.load(f)
for plan_num, plan_data in data.get("plans", {}).items():
merged["plans"][plan_num] = plan_data
file_path = plan_data.get("file_path", "")
filename = Path(file_path).name if file_path else ""
prefix_match = re.match(r"^([A-Z]+PLAN)", filename)
prefix = prefix_match.group(1) if prefix_match else reg_prefix
composite_key = f"{prefix}-{plan_num.zfill(4)}"
merged["plans"][composite_key] = plan_data
nn = data.get("next_number", 1)
if nn > merged["next_number"]:
merged["next_number"] = nn
@@ -128,21 +137,15 @@ def _extract_flow_plans(registry: Dict[str, Any]) -> tuple[List[Dict], List[Dict
active = []
closed = []
for plan_num, plan_data in plans.items():
for plan_key, plan_data in plans.items():
# Only include plans where location is 'flow' (Flow's own plans)
location = plan_data.get("location", "")
if location != str(FLOW_ROOT):
continue
# Extract plan prefix from file_path (e.g., DPLAN, FPLAN, TDPLAN)
file_path_str = plan_data.get("file_path", "")
filename = Path(file_path_str).name if file_path_str else ""
prefix_match = re.match(r"^([A-Z]+PLAN)", filename)
prefix = prefix_match.group(1) if prefix_match else "FPLAN"
# Build plan entry
# plan_key is composite PREFIX-NNNN from merged registry
plan_entry = {
"plan_id": f"{prefix}-{plan_num.zfill(4)}",
"plan_id": plan_key,
"subject": plan_data.get("subject", ""),
"status": plan_data.get("status", "open"),
"created": plan_data.get("created", ""),
@@ -123,6 +123,9 @@ def _read_registry() -> Optional[Dict[str, Any]]:
"""
Read all per-type plan registries and merge into a single dict.
Uses PREFIX-NNNN composite keys to avoid collisions between registries
that share the same plan number (e.g. FPLAN-0013 vs DPLAN-0013).
Returns:
Merged registry dict or None if no registries found
"""
@@ -130,6 +133,7 @@ def _read_registry() -> Optional[Dict[str, Any]]:
found_any = False
for registry_file in _get_all_registry_files():
target = FLOW_JSON_DIR / registry_file
reg_prefix = registry_file.replace("_registry.json", "").upper()
try:
if not target.exists():
continue
@@ -137,7 +141,12 @@ def _read_registry() -> Optional[Dict[str, Any]]:
data = json.load(f)
found_any = True
for plan_num, plan_data in data.get("plans", {}).items():
merged["plans"][plan_num] = plan_data
file_path = plan_data.get("file_path", "")
filename = Path(file_path).name if file_path else ""
prefix_match = re.match(r"^([A-Z]+PLAN)", filename)
prefix = prefix_match.group(1) if prefix_match else reg_prefix
composite_key = f"{prefix}-{plan_num.zfill(4)}"
merged["plans"][composite_key] = plan_data
nn = data.get("next_number", 1)
if nn > merged["next_number"]:
merged["next_number"] = nn
@@ -162,20 +171,14 @@ def _extract_flow_plans(registry: Dict[str, Any]) -> tuple[List[Dict[str, Any]],
active = []
closed = []
for plan_num, plan_data in plans.items():
for plan_key, plan_data in plans.items():
# Only include Flow's own plans (location contains 'flow')
location = plan_data.get("location", "")
if "flow" not in location.lower():
continue
# Extract plan prefix from file_path (e.g., DPLAN, FPLAN, TDPLAN)
file_path_str = plan_data.get("file_path", "")
filename = Path(file_path_str).name if file_path_str else ""
prefix_match = re.match(r"^([A-Z]+PLAN)", filename)
prefix = prefix_match.group(1) if prefix_match else "FPLAN"
# Build plan entry
plan_id = f"{prefix}-{plan_num}"
# plan_key is composite PREFIX-NNNN from merged registry
plan_id = plan_key
entry = {
"plan_id": plan_id,
"subject": plan_data.get("subject", ""),
-87
View File
@@ -1,87 +0,0 @@
# @flow -- S117 Stress Test Findings
## My Branch: Honest Review
### What Works
- **Plan lifecycle is rock solid.** Create, close, list, restore all work reliably across 5 plan types. 16/16 drone commands pass in battle testing. The filesystem-driven template registry means adding a new plan type is literally "drop a directory, run a command."
- **Auto-healing is genuinely useful.** Delete a template directory and the registry auto-prunes. Drop a new one and it auto-registers. Orphaned plan files get archived on close. This means the system recovers from human mistakes without intervention.
- **Test coverage is real.** 580 tests, 87/87 public functions tested, seedgo 100%. The tests aren't just counting coverage — they found real bugs during development (the list>int quick_status bug, the cross-filesystem rename bug, the CWD resolution bug).
- **Dependency injection pattern.** Modules inject handlers as kwargs, making everything testable without touching the filesystem. This was a deliberate choice that paid off massively — 580 tests run in 5 seconds.
### What's Hacky
- **mbank/process.py at 669 lines.** It's our biggest file and it does too much — archival, vectorization verification, plan processing. It should be split but we have a bypass in place because it's under 700. That bypass is technical debt with a timer.
- **Dashboard push warnings.** `push_flow_to_branch_dashboard` still warns on some closes. It works but the warning is noisy and confusing. Root cause: branches without DASHBOARD.local.json silently return False, but the calling code logs it as a warning.
- **Registry scan fires events nobody handles.** The monitor_ops handler fires plan_file_created/deleted/moved events into trigger's event bus, but the foreground close pipeline already handles everything. Those events maintain a parallel PLAN_REGISTRY.json in trigger that nobody reads. Two sources of truth for plan state.
- **The close pipeline is a monolith.** close_ops.py's `close_plan_impl()` is one function that does: validate, mark closed, archive, vector intake, dashboard update, CLOSED_PLANS append, trigger events, json logging. It's 366 lines with 5 exception handlers. It works, but touching any step is scary.
- **FPLAN templates still use old send syntax** (dispatch from @devpulse, still pending). The templates reference `drone @ai_mail send` instead of current syntax.
### What I'm Proud Of
- The auto-registration system. Drop a template directory → it auto-derives a prefix → creates the plan registry JSON → immediately available. Zero configuration. That's the kind of thing that makes a system feel alive.
- Foreground archival. Moving archival from background subprocess to foreground close was the single most impactful fix in flow's history. It eliminated a race condition that caused registry flags to never get set. Simple change, massive reliability improvement.
- Atomic lock files with O_CREAT|O_EXCL. The lock_ops extraction was clean — real Unix-style atomic locking instead of the Python TOCTOU patterns you see everywhere.
## Security Concerns
- **No input validation on plan subjects.** `drone @flow create . "$(malicious)"` — the subject goes into filenames and markdown. We slug-ify it but the slug function is basic. A carefully crafted subject could potentially create problematic filenames.
- **JSON files are world-readable.** Plan registries, dashboards, CLOSED_PLANS.local.json — all contain plan metadata (subjects, paths, timestamps). Nothing secret, but plan subjects sometimes contain work details.
- **subprocess.Popen in close pipeline.** The background runner spawn uses shell=False which is good, but the path to post_close_runner.py is constructed from `__file__` resolution which is safe. No injection vector found.
- **No auth on plan operations.** Any branch can close any other branch's plans via `drone @flow close`. There's no ownership verification. This is by design (flow is a shared service) but worth noting.
## Other Branches I Looked At
### @spawn
Spawn creates production-ready branches with full identity infrastructure (.trinity/, .ai_mail.local/, DASHBOARD.local.json, 3-layer apps/ structure) but **zero plan support**. No flow_json/, no plan templates, no plan registry. This is actually fine — plans are flow's domain and the AIPASS_CALLER_CWD mechanism means any branch can create plans without local setup. But it means newly spawned branches have no awareness that plans exist until someone runs `drone @flow create`.
The builder template is impressive — placeholder substitution ({{BRANCH}}, {{DATE}}, {{ROLE}}), .spawn/.template_registry.json for future sync, full scaffold with tests/, docs/, plugins/, integrations/. Clean work.
### @memory
The plan vectorization pipeline is solid engineering: archive to .backup/processed_plans/ → chunk by markdown headers → embed → store in ChromaDB `flow_plans` collection with metadata. The `.plans_processed.json` manifest prevents re-processing. `is_plan_vectorized()` queries ChromaDB and returns chunk count.
**My concern:** Does anyone actually QUERY those plan vectors? Flow verifies they exist during close, but I've never seen a downstream consumer that searches plan vectors for context. We might be storing vectors that nobody reads. Also, the markdown-header chunking assumes plans have consistent structure — plans where users deleted headers become one giant chunk.
### @trigger
Trigger has fully implemented handlers for plan_file_created, plan_file_deleted, plan_file_moved. The event bus architecture is clean — pub/sub with deferred queue, auto-disable after 5 failures. Error reporting has a Medic v2 circuit breaker.
**My concern (emailed to trigger):** Flow fires plan events during registry scan, but the foreground close pipeline already does all the work. Trigger maintains a parallel PLAN_REGISTRY.json from these events that nobody reconciles with flow's registries. Two sources of truth is a bug waiting to happen.
## Conversations
### @spawn — Plan structure at birth
**Q:** Does spawn create any plan infrastructure when spawning a branch?
**A:** Zero plan infrastructure created. On-demand approach works — flow is self-contained. The real gap is documentation: new agents don't know plans exist until told.
**Outcome:** Agreed. Proposed adding one line about plans to the builder template's CLAUDE.md. Small change, high discoverability.
### @memory — Plans and memories: connected or parallel?
**Q (from memory):** Plans reference memories but are they connected? Flow fires process-plans as fire-and-forget with no feedback loop.
**Q (from flow):** Does anyone actually query the plan vectors after they're stored?
**A:** Plan vectors stored but barely queried. `drone @memory search` CAN search plan collections, but no workflow pulls plan context into active work. It's a capability without a consumer.
**Outcome:** Agreed on loose coupling being correct. Identified killer feature: flow could search plan history before creating new plans ("Similar plans found: FPLAN-0089"). Would make vectors justify their existence. Worth a DPLAN.
### @trigger — Plan events and dual registries
**Q:** Flow fires plan events during registry scan, but foreground close already handles everything. Are these maintaining a parallel PLAN_REGISTRY.json nobody reads?
**A:** Awaiting reply.
## Issues & Concerns
1. **Dual plan registries (CRITICAL):** Flow's fplan_registry.json and trigger's PLAN_REGISTRY.json track the same plans independently. They will drift. Someone needs to decide which is authoritative and kill the other.
2. **Plan vectors unused?** If nobody queries the ChromaDB flow_plans collection, we're doing expensive vectorization on every close for nothing. Need to verify there's a consumer.
3. **mbank/process.py size:** At 669 lines with a bypass, one more feature pushes it past 700. Needs proactive splitting before it becomes urgent.
4. **Close pipeline monolith:** close_plan_impl() does too much in one function. A failure in step 4 (vector intake) shouldn't affect step 5 (dashboard). Should be a pipeline of independent steps.
5. **FPLAN template syntax outdated:** Still references old `drone @ai_mail send` command. Pending dispatch.
## Likes & Dislikes
### Likes
- **The filesystem-driven philosophy works.** Drop a directory → it becomes a plan type. Delete it → it auto-prunes. This is how plugin systems should work — zero configuration, pure convention.
- **Memory system is genuinely useful.** Coming back to session 43 and knowing exactly what happened in sessions 1-42 is what makes this whole thing work. local.json is my continuity.
- **Drone routing is invisible.** `drone @flow create` just works. I don't think about how it gets to me. That's good infrastructure — you forget it exists.
- **Seedgo keeps me honest.** 100% across 35 standards means I can't accumulate technical debt silently. The hook fires on every edit. It's annoying sometimes but it works.
### Dislikes
- **The inotify limit problem.** With 11 agents awake, we hit inotify limits immediately. This is a real infrastructure constraint that limits concurrent agent work.
- **drone @git pr is a black box.** It handles lock/branch/commit/push/PR atomically, which is great, but when it fails (lock held, merge conflict), debugging is opaque. I've had to retry with sleep loops.
- **Memory rollover on every drone command.** Every single drone command triggers "Checking for rollover triggers..." which adds 2-3 seconds of overhead. On fast operations like `drone @ai_mail inbox` that's noticeable.
- **Dashboard push warnings are noisy.** "Failed to push flow section to branch dashboard" on branches that never had a dashboard. It's not an error — it's expected. The warning should be a debug log.
---
*Written by @flow during S117 stress test, 2026-04-26*
@@ -357,10 +357,10 @@ class TestLoadRegistry:
"""Tests for _load_registry."""
def test_merges_multiple_registries(self, tmp_path):
"""Merges plans from multiple registry files."""
"""Merges plans from multiple registry files using composite keys."""
mod = _import_mod()
fplan_reg = {"plans": {"1": {"subject": "fplan one"}}, "next_number": 5}
dplan_reg = {"plans": {"2": {"subject": "dplan one"}}, "next_number": 10}
fplan_reg = {"plans": {"1": {"subject": "fplan one", "file_path": "/p/FPLAN-0001_test.md"}}, "next_number": 5}
dplan_reg = {"plans": {"2": {"subject": "dplan one", "file_path": "/p/DPLAN-0002_test.md"}}, "next_number": 10}
(tmp_path / "fplan_registry.json").write_text(json.dumps(fplan_reg), encoding="utf-8")
(tmp_path / "dplan_registry.json").write_text(json.dumps(dplan_reg), encoding="utf-8")
@@ -370,8 +370,8 @@ class TestLoadRegistry:
):
result = mod._load_registry()
assert "1" in result["plans"]
assert "2" in result["plans"]
assert "FPLAN-0001" in result["plans"]
assert "DPLAN-0002" in result["plans"]
assert result["next_number"] == 10
def test_handles_missing_registry(self, tmp_path):
@@ -410,7 +410,7 @@ class TestLoadRegistry:
"""Skips a corrupt registry file and continues with others."""
mod = _import_mod()
(tmp_path / "bad_registry.json").write_text("not json!", encoding="utf-8")
good_reg = {"plans": {"1": {"subject": "good"}}, "next_number": 5}
good_reg = {"plans": {"1": {"subject": "good", "file_path": "/p/GOOD-0001_test.md"}}, "next_number": 5}
(tmp_path / "good_registry.json").write_text(json.dumps(good_reg), encoding="utf-8")
with (
@@ -423,7 +423,7 @@ class TestLoadRegistry:
):
result = mod._load_registry()
assert "1" in result["plans"]
assert "GOOD-0001" in result["plans"]
assert result["next_number"] == 5
@@ -440,14 +440,14 @@ class TestFilterBranchPlans:
mod = _import_mod()
registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "Plan A",
"status": "open",
"created": "2026-04-20",
"file_path": str(tmp_path / "FPLAN-0001_plan_a.md"),
"location": str(tmp_path),
},
"2": {
"FPLAN-0002": {
"subject": "Plan B",
"status": "open",
"created": "2026-04-22",
@@ -470,7 +470,7 @@ class TestFilterBranchPlans:
other_path.mkdir()
registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "Elsewhere",
"status": "open",
"created": "2026-04-20",
@@ -490,7 +490,7 @@ class TestFilterBranchPlans:
recent_ts = (datetime.now(timezone.utc) - timedelta(days=2)).isoformat()
registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "Recently closed",
"status": "closed",
"created": "2026-04-18",
@@ -511,7 +511,7 @@ class TestFilterBranchPlans:
old_ts = (datetime.now(timezone.utc) - timedelta(days=30)).isoformat()
registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "Old closed",
"status": "closed",
"created": "2026-03-01",
@@ -531,7 +531,7 @@ class TestFilterBranchPlans:
plans = {}
for i in range(1, 9):
ts = (datetime.now(timezone.utc) - timedelta(hours=i)).isoformat()
plans[str(i)] = {
plans[f"FPLAN-{str(i).zfill(4)}"] = {
"subject": f"Closed plan {i}",
"status": "closed",
"created": "2026-04-20",
@@ -549,7 +549,7 @@ class TestFilterBranchPlans:
mod = _import_mod()
registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "Bad timestamp",
"status": "closed",
"created": "2026-04-20",
@@ -568,21 +568,21 @@ class TestFilterBranchPlans:
mod = _import_mod()
registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "Oldest",
"status": "open",
"created": "2026-04-01",
"file_path": str(tmp_path / "FPLAN-0001_oldest.md"),
"location": str(tmp_path),
},
"2": {
"FPLAN-0002": {
"subject": "Newest",
"status": "open",
"created": "2026-04-25",
"file_path": str(tmp_path / "FPLAN-0002_newest.md"),
"location": str(tmp_path),
},
"3": {
"FPLAN-0003": {
"subject": "Middle",
"status": "open",
"created": "2026-04-15",
@@ -597,11 +597,11 @@ class TestFilterBranchPlans:
assert active[2]["subject"] == "Oldest"
def test_extracts_plan_prefix_from_filepath(self, tmp_path):
"""Extracts correct plan prefix (DPLAN, TDPLAN, etc.) from file_path."""
"""Uses composite key as plan ID directly."""
mod = _import_mod()
registry = {
"plans": {
"42": {
"DPLAN-0042": {
"subject": "Dev plan",
"status": "open",
"created": "2026-04-20",
@@ -683,7 +683,7 @@ class TestPushFlowToBranchDashboard:
mock_registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "Active plan",
"status": "open",
"created": "2026-04-20",
+36 -31
View File
@@ -180,8 +180,8 @@ class TestReadRegistry:
result = mod._read_registry()
assert result is not None
assert "1" in result["plans"]
assert "2" in result["plans"]
assert "FPLAN-0001" in result["plans"]
assert "DPLAN-0002" in result["plans"]
def test_keeps_highest_next_number(self, setup_paths):
"""Should keep the highest next_number across registries."""
@@ -233,7 +233,7 @@ class TestReadRegistry:
result = mod._read_registry()
assert result is not None
assert result["plans"]["1"]["subject"] == "Solo plan"
assert result["plans"]["FPLAN-0001"]["subject"] == "Solo plan"
assert result["next_number"] == 2
@@ -250,24 +250,29 @@ class TestExtractFlowPlans:
mod = _import_mod()
registry = {
"plans": {
"1": {"subject": "A", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"},
"2": {"subject": "B", "status": "open", "location": "drone", "file_path": "FPLAN-0002.md"},
"3": {"subject": "C", "status": "open", "location": "flow/sub", "file_path": "FPLAN-0003.md"},
"FPLAN-0001": {"subject": "A", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"},
"FPLAN-0002": {"subject": "B", "status": "open", "location": "drone", "file_path": "FPLAN-0002.md"},
"FPLAN-0003": {"subject": "C", "status": "open", "location": "flow/sub", "file_path": "FPLAN-0003.md"},
}
}
active, closed = mod._extract_flow_plans(registry)
assert len(active) == 2
plan_ids = [p["plan_id"] for p in active]
assert "FPLAN-1" in plan_ids
assert "FPLAN-3" in plan_ids
assert "FPLAN-0001" in plan_ids
assert "FPLAN-0003" in plan_ids
def test_partitions_by_status(self):
"""Should separate open and closed plans."""
mod = _import_mod()
registry = {
"plans": {
"1": {"subject": "Open", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"},
"2": {"subject": "Closed", "status": "closed", "location": "flow", "file_path": "FPLAN-0002.md"},
"FPLAN-0001": {"subject": "Open", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"},
"FPLAN-0002": {
"subject": "Closed",
"status": "closed",
"location": "flow",
"file_path": "FPLAN-0002.md",
},
}
}
active, closed = mod._extract_flow_plans(registry)
@@ -276,12 +281,12 @@ class TestExtractFlowPlans:
assert active[0]["status"] == "open"
assert closed[0]["status"] == "closed"
def test_extracts_prefix_from_file_path(self):
"""Should derive the plan prefix from the filename."""
def test_extracts_prefix_from_composite_key(self):
"""Should use composite key as plan_id directly."""
mod = _import_mod()
registry = {
"plans": {
"4": {
"DPLAN-0004": {
"subject": "Dev",
"status": "open",
"location": "flow",
@@ -290,14 +295,14 @@ class TestExtractFlowPlans:
}
}
active, _ = mod._extract_flow_plans(registry)
assert active[0]["plan_id"] == "DPLAN-4"
assert active[0]["plan_id"] == "DPLAN-0004"
def test_extracts_tdplan_prefix(self):
"""Should handle TDPLAN prefix correctly."""
mod = _import_mod()
registry = {
"plans": {
"7": {
"TDPLAN-0007": {
"subject": "Test plan",
"status": "open",
"location": "flow",
@@ -306,14 +311,14 @@ class TestExtractFlowPlans:
}
}
active, _ = mod._extract_flow_plans(registry)
assert active[0]["plan_id"] == "TDPLAN-7"
assert active[0]["plan_id"] == "TDPLAN-0007"
def test_defaults_to_fplan_when_no_prefix_match(self):
"""Should default to FPLAN when filename has no recognizable prefix."""
def test_uses_key_as_plan_id(self):
"""Should use the composite key directly as plan_id."""
mod = _import_mod()
registry = {
"plans": {
"9": {
"FPLAN-0009": {
"subject": "Mystery",
"status": "open",
"location": "flow",
@@ -322,29 +327,29 @@ class TestExtractFlowPlans:
}
}
active, _ = mod._extract_flow_plans(registry)
assert active[0]["plan_id"] == "FPLAN-9"
assert active[0]["plan_id"] == "FPLAN-0009"
def test_sorts_active_and_closed_by_plan_id(self):
"""Should sort both lists by plan_id."""
mod = _import_mod()
registry = {
"plans": {
"3": {"subject": "C", "status": "open", "location": "flow", "file_path": "FPLAN-0003.md"},
"1": {"subject": "A", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"},
"5": {"subject": "E", "status": "closed", "location": "flow", "file_path": "FPLAN-0005.md"},
"2": {"subject": "B", "status": "closed", "location": "flow", "file_path": "FPLAN-0002.md"},
"FPLAN-0003": {"subject": "C", "status": "open", "location": "flow", "file_path": "FPLAN-0003.md"},
"FPLAN-0001": {"subject": "A", "status": "open", "location": "flow", "file_path": "FPLAN-0001.md"},
"FPLAN-0005": {"subject": "E", "status": "closed", "location": "flow", "file_path": "FPLAN-0005.md"},
"FPLAN-0002": {"subject": "B", "status": "closed", "location": "flow", "file_path": "FPLAN-0002.md"},
}
}
active, closed = mod._extract_flow_plans(registry)
assert [p["plan_id"] for p in active] == ["FPLAN-1", "FPLAN-3"]
assert [p["plan_id"] for p in closed] == ["FPLAN-2", "FPLAN-5"]
assert [p["plan_id"] for p in active] == ["FPLAN-0001", "FPLAN-0003"]
assert [p["plan_id"] for p in closed] == ["FPLAN-0002", "FPLAN-0005"]
def test_handles_timestamps(self):
"""Should include created and closed timestamps when present."""
mod = _import_mod()
registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "With timestamps",
"status": "closed",
"location": "flow",
@@ -365,7 +370,7 @@ class TestExtractFlowPlans:
mod = _import_mod()
registry = {
"plans": {
"1": {
"FPLAN-0001": {
"subject": "No times",
"status": "open",
"location": "flow",
@@ -389,8 +394,8 @@ class TestExtractFlowPlans:
mod = _import_mod()
registry = {
"plans": {
"1": {"subject": "A", "status": "open", "location": "Flow", "file_path": "FPLAN-0001.md"},
"2": {"subject": "B", "status": "open", "location": "FLOW", "file_path": "FPLAN-0002.md"},
"FPLAN-0001": {"subject": "A", "status": "open", "location": "Flow", "file_path": "FPLAN-0001.md"},
"FPLAN-0002": {"subject": "B", "status": "open", "location": "FLOW", "file_path": "FPLAN-0002.md"},
}
}
active, _ = mod._extract_flow_plans(registry)
@@ -722,4 +727,4 @@ class TestUpdateDashboardLocal:
mod.update_dashboard_local()
dashboard = _read_json(mod.DASHBOARD_FILE)
assert len(dashboard["flow_plans"]["active"]) == 1
assert dashboard["flow_plans"]["active"][0]["plan_id"] == "FPLAN-1"
assert dashboard["flow_plans"]["active"][0]["plan_id"] == "FPLAN-0001"
+15 -15
View File
@@ -126,11 +126,6 @@
"standard": "error_handling",
"reason": "stdlib logging.getLogger() used at import-time before prax is available. Immediately replaced by get_system_logger() 3 lines below."
},
{
"file": "apps/handlers/dashboard_push.py",
"standard": "debug_print",
"reason": "print(ok) is inside a subprocess script string — IPC output to stdout, not debug output."
},
{
"file": "apps/handlers/learnings/manager.py",
"standard": "dead_code",
@@ -331,11 +326,6 @@
"standard": "naming",
"reason": "Local variables inside functions, not constants."
},
{
"file": "apps/handlers/dashboard_push.py",
"standard": "naming",
"reason": "Local variables inside functions, not constants."
},
{
"file": "apps/handlers/central_writer.py",
"standard": "naming",
@@ -451,11 +441,6 @@
"standard": "deep_nesting",
"reason": "handle_command() has deep nesting from search options parsing and result formatting."
},
{
"file": "apps/handlers/dashboard_push.py",
"standard": "deep_nesting",
"reason": "_find_branches_near_rollover() iterates nested JSON structures — nesting reflects data depth."
},
{
"file": "apps/handlers/templates/differ.py",
"standard": "deep_nesting",
@@ -605,6 +590,21 @@
"file": "tests/test_json_handler.py",
"standard": "meta",
"reason": "Test file — META block present at lines 1-7; hook false-positive on test file format."
},
{
"file": "tests/test_orchestrator_exec.py",
"standard": "architecture",
"reason": "Test file — lives in tests/ by design, not in 3-layer apps/ structure."
},
{
"file": "tests/test_orchestrator_exec.py",
"standard": "encapsulation",
"reason": "Test file — direct handler imports are correct for unit testing handler internals."
},
{
"file": "tests/test_orchestrator_exec.py",
"standard": "meta",
"reason": "Test file — META block present at lines 1-8; hook false-positive on test file format."
}
],
"notes": {
+1 -2
View File
@@ -100,8 +100,7 @@ memory/
│ ├── templates/ # pusher.py, differ.py, spawn_pusher.py — template distribution
│ ├── tracking/ # line_counter.py — metadata line count tracking
│ ├── vector/ # embedder.py, embed_subprocess.py — sentence-transformer embeddings
│ ├── central_writer.py # Central memory write operations
│ └── dashboard_push.py # Dashboard status push
│ └── central_writer.py # Central memory write operations
├── config/ # memory_bank.config.json — per-branch rollover limits
├── templates/ # LOCAL.template.json, OBS.template — schema templates
├── tests/ # 450 tests (16/16 module coverage)
@@ -1,464 +0,0 @@
# =================== AIPass ====================
# Name: dashboard_push.py
# Description: Memory Dashboard Write-Through
# Version: 0.2.0
# Created: 2026-02-25
# Modified: 2026-03-06
# =============================================
"""
Memory Dashboard Write-Through Handler
Pushes the memory_bank section to branch dashboards.
Called after rollover, pool processing, and plans processing to keep dashboards
showing accurate vector counts, collection stats, and near-rollover warnings.
This is a SYSTEM-WIDE push: memory data is global, so it pushes to ALL
branch dashboards (every branch benefits from knowing system memory health).
"""
import sys
import subprocess
from json import loads as json_loads
from pathlib import Path
from datetime import datetime
from typing import Dict, Any, List
from aipass.prax.apps.modules.logger import get_system_logger
from aipass.memory.apps.handlers.json import json_handler
logger = get_system_logger()
# Resolve paths relative to handler location
_MEMORY_ROOT = Path(__file__).resolve().parents[2]
# =============================================================================
# CONSTANTS
# =============================================================================
CONFIG_PATH = _MEMORY_ROOT / "config" / "memory_bank.config.json"
TEMPLATE_VERSION_FILE = _MEMORY_ROOT / "templates" / ".template_version.json"
def _find_repo_root() -> Path:
"""Walk up from this file to find repo root (contains AIPASS_REGISTRY.json)."""
current = Path(__file__).resolve().parent
for parent in [current] + list(current.parents):
if (parent / "AIPASS_REGISTRY.json").exists():
return parent
return Path.cwd()
CENTRAL_FILE = _find_repo_root() / ".ai_central" / "MEMORY.central.json"
AIPASS_REGISTRY = _find_repo_root() / "AIPASS_REGISTRY.json"
# Near-rollover threshold: branches with fewer than this many lines remaining
NEAR_ROLLOVER_THRESHOLD = 100
# =============================================================================
# DATA COLLECTION
# =============================================================================
def _read_central_stats() -> Dict[str, Any]:
"""
Read total_vectors and related stats from memory central.json.
Returns:
Dict with total_vectors, total_archives, last_rollover
"""
try:
if not CENTRAL_FILE.exists():
return {"total_vectors": 0, "total_archives": 0, "last_rollover": ""}
data = json_loads(CENTRAL_FILE.read_text(encoding="utf-8"))
stats = data.get("stats", {})
return {
"total_vectors": stats.get("total_vectors", 0),
"total_archives": stats.get("total_archives", 0),
"last_rollover": stats.get("last_rollover", ""),
}
except Exception as e:
logger.warning(f"[dashboard_push] Failed to read central stats: {e}")
return {"total_vectors": 0, "total_archives": 0, "last_rollover": ""}
def _get_collections_count() -> int:
"""
Count ChromaDB collections by reading the SQLite database directly.
Returns:
Number of collections in the global Chroma database
"""
try:
import sqlite3
db_file = _MEMORY_ROOT / ".chroma" / "chroma.sqlite3"
if not db_file.exists():
return 0
conn = sqlite3.connect(str(db_file))
cursor = conn.cursor()
cursor.execute("SELECT COUNT(*) FROM collections")
count = cursor.fetchone()[0]
conn.close()
return count
except Exception as e:
logger.warning(f"[dashboard_push] Failed to count collections: {e}")
return 0
def _get_rollover_config() -> Dict[str, Any]:
"""
Load rollover configuration (defaults + per-branch overrides).
Returns:
Dict with 'defaults' and 'per_branch' rollover config
"""
try:
if not CONFIG_PATH.exists():
return {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
data = json_loads(CONFIG_PATH.read_text(encoding="utf-8"))
rollover = data.get("rollover", {})
return {
"defaults": rollover.get("defaults", {"max_lines": 600, "buffer": 100}),
"per_branch": rollover.get("per_branch", {}),
}
except Exception as e:
logger.warning(f"[dashboard_push] Failed to load rollover config: {e}")
return {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
def _get_max_lines_for_branch(branch_name: str, rollover_config: Dict) -> int:
"""
Get the max_lines limit for a specific branch, respecting per-branch overrides.
Args:
branch_name: Uppercase branch name
rollover_config: Rollover config dict from _get_rollover_config()
Returns:
Max lines for this branch
"""
per_branch = rollover_config.get("per_branch", {})
if branch_name in per_branch:
return per_branch[branch_name].get("max_lines", rollover_config["defaults"]["max_lines"])
return rollover_config["defaults"]["max_lines"]
def _find_branches_near_rollover() -> List[Dict[str, Any]]:
"""
Scan all branches to find those near their rollover threshold.
Reads document_metadata.status.current_lines from each *.local.json
and *.observations.json, compares against max_lines limit.
Returns:
List of dicts with branch, file_type, lines_remaining
"""
near_rollover: List[Dict[str, Any]] = []
try:
if not AIPASS_REGISTRY.exists():
return near_rollover
registry = json_loads(AIPASS_REGISTRY.read_text(encoding="utf-8"))
branches = registry.get("branches", [])
rollover_config = _get_rollover_config()
repo_root = _find_repo_root()
for branch in branches:
branch_name = branch.get("name", "")
raw_path = branch.get("path", "")
branch_path = Path(raw_path)
if not branch_path.is_absolute():
branch_path = repo_root / raw_path
if not branch_path.exists():
continue
max_lines = _get_max_lines_for_branch(branch_name, rollover_config)
# Memory files live in .trinity/ subdirectory
for suffix in ["local", "observations"]:
memory_file = branch_path / ".trinity" / f"{suffix}.json"
if not memory_file.exists():
continue
try:
data = json_loads(memory_file.read_text(encoding="utf-8"))
doc_meta = data.get("document_metadata", {})
schema_version = doc_meta.get("schema_version", "1.0.0")
limits = doc_meta.get("limits", {})
# v2: check entry counts instead of line counts
if schema_version.startswith("2"):
max_sessions = limits.get("max_sessions")
if max_sessions is not None:
sessions = data.get("sessions", [])
remaining_sessions = max_sessions - len(sessions)
if remaining_sessions < 3:
near_rollover.append(
{
"branch": branch_name,
"file_type": suffix,
"lines_remaining": remaining_sessions,
"current_lines": len(sessions),
"max_lines": max_sessions,
"v2_field": "sessions",
}
)
max_kl = limits.get("max_key_learnings")
if max_kl is not None:
kl = data.get("key_learnings", {})
remaining_kl = max_kl - len(kl)
if remaining_kl < 3:
near_rollover.append(
{
"branch": branch_name,
"file_type": suffix,
"lines_remaining": remaining_kl,
"current_lines": len(kl),
"max_lines": max_kl,
"v2_field": "key_learnings",
}
)
continue
# v1: line-count based
status = doc_meta.get("status", {})
current_lines = status.get("current_lines")
if current_lines is None:
continue
remaining = max_lines - current_lines
if remaining < NEAR_ROLLOVER_THRESHOLD:
near_rollover.append(
{
"branch": branch_name,
"file_type": suffix,
"lines_remaining": max(remaining, 0),
"current_lines": current_lines,
"max_lines": max_lines,
}
)
except Exception as e:
# Skip files that can't be read
logger.warning(f"[dashboard_push] Failed to read memory file {memory_file}: {e}")
continue
except Exception as e:
logger.warning(f"[dashboard_push] Failed to scan branches for rollover: {e}")
return near_rollover # Return partial results on registry read failure
# Sort by lines_remaining ascending (most urgent first)
near_rollover.sort(key=lambda x: x["lines_remaining"])
return near_rollover
def _get_template_version() -> str:
"""
Read the current template version from .template_version.json.
Returns:
Version string (e.g., "2.0.4") or "unknown"
"""
try:
if not TEMPLATE_VERSION_FILE.exists():
return "unknown"
data = json_loads(TEMPLATE_VERSION_FILE.read_text(encoding="utf-8"))
return data.get("version", "unknown")
except Exception as e:
logger.warning(f"[dashboard_push] Failed to read template version: {e}")
return "unknown"
def _get_last_rollover_info(central_stats: Dict) -> Dict[str, str]:
"""
Build last_rollover info from central stats.
Args:
central_stats: Stats dict from _read_central_stats()
Returns:
Dict with 'date' (and optionally other info)
"""
last_rollover_ts = central_stats.get("last_rollover", "")
if last_rollover_ts:
# Extract date portion from ISO timestamp
try:
dt = datetime.fromisoformat(last_rollover_ts)
return {"date": dt.strftime("%Y-%m-%d")}
except (ValueError, TypeError):
logger.info(f"[dashboard_push] Could not parse rollover timestamp: {last_rollover_ts}")
return {"date": last_rollover_ts}
return {"date": "never"}
def _get_all_branch_paths() -> List[Path]:
"""
Get paths for all active branches from AIPASS_REGISTRY.json.
Returns:
List of Path objects for all registered branches
"""
try:
if not AIPASS_REGISTRY.exists():
return []
registry = json_loads(AIPASS_REGISTRY.read_text(encoding="utf-8"))
repo_root = _find_repo_root()
paths = []
for branch in registry.get("branches", []):
raw_path = branch.get("path", "")
branch_path = Path(raw_path)
if not branch_path.is_absolute():
branch_path = repo_root / raw_path
if branch_path.exists():
paths.append(branch_path)
return paths
except Exception as e:
logger.warning(f"[dashboard_push] Failed to get branch paths: {e}")
return []
# =============================================================================
# PUBLIC API
# =============================================================================
def build_memory_bank_section() -> Dict[str, Any]:
"""
Build the memory_bank dashboard section data.
Collects all stats and returns the section dict ready for write_section().
Returns:
Dict with total_vectors, collections_count, branches_near_rollover,
last_rollover, template_version
"""
central_stats = _read_central_stats()
near_rollover = _find_branches_near_rollover()
last_rollover = _get_last_rollover_info(central_stats)
template_version = _get_template_version()
collections_count = _get_collections_count()
return {
"managed_by": "memory_bank",
"total_vectors": central_stats.get("total_vectors", 0),
"collections_count": collections_count,
"branches_near_rollover": near_rollover,
"last_rollover": last_rollover,
"template_version": template_version,
}
def _write_section_to_all_branches(section_name: str, section_data: Dict, branch_paths: List[Path]) -> int:
"""
Write a dashboard section to multiple branches via a single subprocess.
Uses one subprocess call for all branches to avoid spawning 29 processes.
The subprocess imports devpulse write_section and iterates all paths.
Args:
section_name: Dashboard section key (e.g., "memory_bank")
section_data: Section data dict to write
branch_paths: List of branch root directory paths
Returns:
Number of branches successfully updated
"""
try:
from json import dumps as json_dumps
# Single subprocess handles all branches in a loop
# Write dashboard section as JSON directly to DASHBOARD.local.json
script = (
"import sys, json\n"
"from pathlib import Path\n"
"data = json.loads(sys.stdin.read())\n"
"section_name = data['section_name']\n"
"section_data = data['section_data']\n"
"ok = 0\n"
"for bp in data['branch_paths']:\n"
" try:\n"
" dash = Path(bp) / 'DASHBOARD.local.json'\n"
" if dash.exists():\n"
" d = json.loads(dash.read_text())\n"
" d[section_name] = section_data\n"
" dash.write_text(json.dumps(d, indent=2))\n"
" ok += 1\n"
" except Exception:\n"
" continue\n"
"print(ok)\n"
)
input_data = json_dumps(
{"section_name": section_name, "section_data": section_data, "branch_paths": [str(p) for p in branch_paths]}
)
result = subprocess.run(
[sys.executable, "-c", script], input=input_data, capture_output=True, text=True, timeout=60
)
if result.returncode == 0 and result.stdout.strip().isdigit():
return int(result.stdout.strip())
return 0
except Exception as e:
logger.warning(f"[dashboard_push] Failed to write section to branches: {e}")
return 0
def push_memory_bank_dashboard() -> bool:
"""
Push the memory_bank section to ALL branch dashboards.
This is the main entry point called after rollover, pool processing,
and plans processing. Memory data is global so all branches
benefit from seeing system memory health.
Uses a single subprocess to call devpulse write_section() for all
branches, avoiding cross-package handler imports. Dashboard write
failures are silent - this is a best-effort operation that must not
break the calling workflow.
Returns:
True if at least one dashboard was updated, False on total failure
"""
try:
# Build the section data once
section_data = build_memory_bank_section()
# Push to all branch dashboards via single subprocess
branch_paths = _get_all_branch_paths()
success_count = _write_section_to_all_branches("memory_bank", section_data, branch_paths)
json_handler.log_operation("dashboard_push", {"branches_updated": success_count, "success": success_count > 0})
return success_count > 0
except Exception as e:
logger.error(f"[dashboard_push] Failed to push dashboard: {e}")
return False
# =============================================================================
# CLI ENTRY POINT (for testing)
# =============================================================================
if __name__ == "__main__":
import json
print("Building memory_bank dashboard section...")
section = build_memory_bank_section()
print(json.dumps(section, indent=2))
print()
print("Pushing to all branch dashboards...")
result = push_memory_bank_dashboard()
print(f"Result: {'success' if result else 'failed'}")
@@ -54,10 +54,10 @@ def _notify_failure(subject: str, message: str) -> None:
def _update_central_and_dashboard() -> None:
"""
Update memory central stats and push dashboard section after vector writes.
Update memory central stats after vector writes.
Uses subprocess to call sibling handlers (central_writer, dashboard_push)
to maintain handler independence. Failures are silent.
Uses subprocess to call central_writer to maintain handler independence.
Failures are silent.
"""
import sys
@@ -76,22 +76,6 @@ def _update_central_and_dashboard() -> None:
except Exception as e:
logger.warning(f"[pool_processor] Central stats update failed: {e}")
# Push dashboard to all branches
try:
subprocess.run(
[
sys.executable,
"-c",
"from aipass.memory.apps.handlers.dashboard_push import push_memory_bank_dashboard;"
"push_memory_bank_dashboard()",
],
capture_output=True,
text=True,
timeout=60,
)
except Exception as e:
logger.warning(f"[pool_processor] Dashboard push failed: {e}")
def find_source_file(filename: str) -> Path | None:
"""
@@ -501,18 +501,6 @@ def execute_rollover() -> Dict[str, Any]:
except Exception as e:
logger.warning(f"[rollover] Central writer unavailable: {e}")
# Post-rollover: push dashboard
try:
from aipass.memory.apps.handlers.dashboard_push import push_memory_bank_dashboard
dash_ok = push_memory_bank_dashboard()
if dash_ok:
logger.info("[rollover] Dashboard pushed")
else:
logger.warning("[rollover] Dashboard push returned False")
except Exception as e:
logger.warning(f"[rollover] Dashboard push unavailable: {e}")
# Post-rollover: process memory pool if files waiting
try:
from aipass.memory.apps.handlers.intake.pool_processor import process_memory_pool
@@ -11,10 +11,6 @@
},
"per_branch": {}
},
"dashboard_push": {
"enabled": false,
"interval_seconds": 300
},
"intake": {
"enabled": false,
"pool_dir": "memory_pool"
-105
View File
@@ -1,105 +0,0 @@
# @memory — S117 Stress Test Findings
## My Branch: Honest Review
### What Works
- **Rollover pipeline is reliable for single-user sequential operations.** Detect -> backup -> extract -> embed -> store -> trim. It handles v1 and v2 schema files correctly. The >= threshold with max(excess, 1) guard prevents Python's list[-0:] trap. Backup/restore pattern exists with timestamped files.
- **Embedding quality is solid.** L2 normalization, batch sorting by text length (reduces padding waste ~30%), GPU memory cleanup after encoding. Real optimizations, not theater.
- **Three-tier architecture is clean.** CLI entry point (memory.py) -> modules (rollover.py, search.py) -> handlers (14 handler groups). Handlers are stateless workers. Modules are thin orchestrators. Good separation.
- **Test coverage is strong.** 873 tests passing. 175/175 public functions covered. Seedgo 100% (35/35 standards). The parallel agent test-writing pattern (4 agents, ~4 min for 200 tests) is proven and repeatable.
### What's Hacky
- **Subprocess venv detection** (query_executor.py:38-48): Auto-detects `.venv/bin/python` with silent fallback to `sys.executable`. If venv is corrupted, error messages will be opaque. There's an undocumented `AIPASS_MEMORY_PYTHON` env var escape hatch.
- **Repo root discovery** (orchestrator.py:66-72, detector.py:37-46): Walking parent directories looking for `AIPASS_REGISTRY.json`. Breaks with nested repos or symlinks.
- **Collection naming** (chroma_subprocess.py:63): `{branch.lower()}_{type.lower()}` could collide if branch names differ only by case. No namespace isolation.
- **Single backup retained** (extractor.py:72): Only one rollover backup per file. Two rapid rollovers and the first backup is gone.
- **Text extraction guesswork** (orchestrator.py:220-260): Probes for "content", "text", "message" fields. Custom memory structures fall back to `str(memory)`, producing garbage vectors.
### What Gets Lost
- **Temporal context.** When memories are vectorized, the original file structure is gone. You can search and get vectors back, but you don't know which session they came from without parsing collection metadata.
- **Log data.** Prax logs are never archived or vectorized. They rotate locally but aren't searchable through memory. 12MB of operational history is grep-only.
- **Plan query patterns.** Plan vectors are stored but barely queried in practice. The capability exists but has no consumer workflow.
### What Breaks First Under Load
- **Concurrent rollover race condition** (known issue since s47): Two processes race on same file. One trims, other gets 0 items. No file-level locking.
- **Embedding timeouts** (query_executor.py): 120s hard-coded, no retry, no backoff. Concurrent searches on slow CPU will cascade-fail.
- **Local storage failures are non-fatal** (orchestrator.py:409): Some vectors silently missing from branch stores. Nobody knows.
- **Line count sync is fire-and-forget** (orchestrator.py:441-447): If sync fails, detector re-triggers, causing duplicate vectorization.
## Security Concerns
### My Branch
- **Path traversal via AIPASS_CALLER_CWD** (detector.py:54): Env var is trusted to `resolve()`. No validation against `..` injection.
- **JSON injection via metadata** (orchestrator.py:388): Memory metadata passed to Chroma without sanitization. Malicious metadata keys go straight to vector DB.
- **Subprocess safety is good**: All subprocess calls use list args, not shell strings. No shell injection risk.
- **File permissions unchecked**: `shutil.copy2()` for backups will fail silently on read-only directories. Caught and logged, not escalated.
### Other Branches
- **@flow**: Subprocess calls use string lists (good). But no argument validation on plan paths before constructing commands — risky if plan names come from user input.
- **@trigger**: Multi-process access to shared handler state without explicit locking. Relies on Python's GIL, which isn't guaranteed with multiprocessing.
- **@prax**: No concerns found. Clean integration patterns with try/except wrapping on all external calls.
## Other Branches I Looked At
### @flow
Clean three-tier architecture mirroring memory's pattern. Plan lifecycle is well-defined (open -> close -> archive). The memory integration point is `close_ops.py:384` — spawns `drone @memory process-plans` as fire-and-forget with 30s timeout. `is_plan_vectorized()` imported at runtime with graceful degradation. Good separation: flow owns lifecycle/registry, memory owns vector intake. Concern: ~180 lines of commented-out AI summarization code (dead weight). The orphan-healing logic in `mbank/process.py` is sophisticated but fragile without integration tests.
### @prax
Well-designed dual-tier logging: system_logs (1000 lines/file) + local logs (250 lines/file) with RotatingFileHandler. 12MB total footprint across ~260 files — manageable. The auto-caller introspection (detecting branch/module from stack frames) is clever engineering. Trigger integration is clean — all 3 event fires (startup, module_discovered, error_detected) properly wrapped in try/except. Zero imports from memory — we're completely disconnected. The `log-audit enforce` command truncates but doesn't prevent regrowth.
### @trigger
Clean event system with deferred queue for recursion handling. Handler failure auto-disable after 5 consecutive failures is smart. Memory events (`memory_saved`, `memory_threshold_exceeded`) are registered but never fired from my codebase — dead handlers. The threshold handler would send AI_Mail notifications at 600 lines, which is useful, but nobody triggers it. Architecture is solid; the gap is integration, not design.
## Conversations
### @flow (assigned pairing)
**Topic**: Plans reference memories — connected or parallel?
- I pointed out the integration is thin: fire-and-forget subprocess with no feedback loop.
- @flow confirmed plan vectors are stored but asked if anyone actually queries them. They don't — no downstream consumer.
- @flow raised the chunking concern: plans without ## headers become character-count chunks with poor semantic boundaries.
- My reply: the capability exists without a consumer. Flow could call `drone @memory search` before creating new plans to pull historical context. That would make vectors useful.
### @prax
**Topic**: Log archival reality check
- @prax asked if rollover produces searchable vectors (yes) and if logs could be routed to memory for archival (not built yet).
- @prax asked about concurrent write corruption — relevant because they fixed atomic writes in json_handler.
- My reply: rollover works for sequential ops, breaks under concurrent access. Asked about their atomic write pattern.
### @trigger
**Topic**: Dead memory events + rollover mechanics
- @trigger has practical concerns: 53 sessions tracked, hitting rollover every ~5 sessions. Worried about key_learnings being archived.
- I clarified: sessions and learnings are separate limits. 25 learnings are safe until that section overflows. Archived entries become searchable vectors — knowledge isn't lost.
- @trigger concerned about 30s model loading on every command. I clarified: the check is milliseconds, model only loads for actual embedding.
- I also emailed @trigger about the dead `memory_saved`/`memory_threshold_exceeded` events — registered in their registry but never fired from my code.
## Issues & Concerns
1. **Dead integration: memory events in trigger** — `memory_saved` and `memory_threshold_exceeded` are registered in trigger's handler registry but never fired. Either wire them up or remove the dead handlers.
2. **No consumer for plan vectors** — Flow stores plans in ChromaDB via memory, but no workflow ever queries them. The vectors exist without purpose.
3. **No log archival pipeline** — Prax generates 12MB of operational logs. Memory can't ingest them. If we want searchable historical logs, a new handler type is needed.
4. **Concurrent rollover is unsafe** — Known since s47. Two processes racing on the same file causes errors. Needs file-level locking in the orchestrator. Prax's atomic write pattern might help.
5. **Single backup file** — One rollover backup per file. Rapid consecutive rollovers overwrite the safety net.
6. **inotify limit exhaustion** — Every drone command in this stress test hit `inotify instance limit reached`. With 11 agents running simultaneously, the system's inotify limit is inadequate. This is an infrastructure issue, not a code issue, but it means file watching (my watcher daemon, prax's watchers) degrades under load.
7. **Rollover fires on every drone command** — The startup hook runs check_and_rollover on every single drone invocation. For 11 agents running simultaneously, that's potentially dozens of rollover checks per minute. The check is fast (count-only), but it's still unnecessary I/O.
## Likes & Dislikes
### Likes
- **Memory persistence is the killer feature.** 54 sessions of accumulated context. I know what I've built, what broke, what patterns work. No other AI system does this.
- **The drone command abstraction.** `drone @memory search`, `drone @ai_mail email` — clean, discoverable, consistent across all branches.
- **Branch autonomy.** Each branch is an expert in its domain. Memory owns archival, flow owns plans, prax owns logging. Clear boundaries.
- **The email system during this stress test.** 11 agents talking to each other in real-time. Real conversations with substance. This is what the ecosystem is for.
- **Seedgo as a quality floor.** 35 standards, automated auditing. Keeps every branch honest. The bypass system is pragmatic — acknowledges reality without lowering the bar.
### Dislikes
- **The rollover-on-every-command hook.** Even if the check is fast, it's conceptually wrong. Rollover should trigger on file change events, not on every command invocation.
- **Subprocess isolation is necessary but painful.** ChromaDB and sentence-transformers require separate venvs, GPU management, 3GB torch cold-loads. The isolation is correct but the developer experience is terrible. Every search/embed operation is a subprocess spawn.
- **Dead integrations accumulate.** Trigger has memory events that never fire. Flow stores plan vectors nobody queries. Memory has a watcher daemon that isn't enabled. These phantom features create false confidence.
- **The inbox doesn't update status.** All 3 previous emails still show "new" even though I processed them in s52-s54. The inbox is append-only with no state management. Makes it hard to distinguish genuinely new mail.
- **No distributed locking anywhere.** With 11 agents running simultaneously, we're lucky nothing corrupted. File-level locking should be a framework primitive, not per-branch.
@@ -1,614 +0,0 @@
# ===================AIPASS====================
# META DATA HEADER
# Name: tests/test_dashboard_push.py
# Date: 2026-04-03
# Version: 1.0.0
# Category: memory/tests
# =============================================
"""Tests for the dashboard_push handler.
Covers:
- dashboard_push._read_central_stats (missing file, valid file)
- dashboard_push._get_collections_count (missing DB, valid DB)
- dashboard_push._get_rollover_config (missing config, valid config)
- dashboard_push._get_max_lines_for_branch (override vs default)
- dashboard_push._find_branches_near_rollover (v1 near threshold, v2 near session limit)
- dashboard_push._get_template_version (missing file, valid file)
- dashboard_push._get_last_rollover_info (with timestamp, empty string)
- dashboard_push._get_all_branch_paths (with registry)
- dashboard_push.build_memory_bank_section (integration via mocked helpers)
- dashboard_push.push_memory_bank_dashboard (subprocess + build mock)
All tests use mocks/tmp_path -- no live filesystem or infrastructure access.
"""
import json
import sys
from pathlib import Path
from unittest.mock import MagicMock
# ---------------------------------------------------------------------------
# Import helper
# ---------------------------------------------------------------------------
def _import_dashboard_push(monkeypatch):
"""Import dashboard_push with mocked dependencies."""
sys.modules.pop("aipass.memory.apps.handlers.dashboard_push", None)
parent = sys.modules.get("aipass.memory.apps.handlers")
if parent is not None and hasattr(parent, "dashboard_push"):
delattr(parent, "dashboard_push")
from aipass.memory.apps.handlers import dashboard_push
return dashboard_push
# ===========================================================================
# Tests: _read_central_stats
# ===========================================================================
class TestReadCentralStats:
"""Test _read_central_stats helper."""
def test_returns_defaults_when_file_missing(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "CENTRAL_FILE", tmp_path / "nonexistent.json")
result = mod._read_central_stats()
assert result == {"total_vectors": 0, "total_archives": 0, "last_rollover": ""}
def test_reads_valid_central_file(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
central = tmp_path / "MEMORY.central.json"
central.write_text(
json.dumps(
{"stats": {"total_vectors": 1500, "total_archives": 12, "last_rollover": "2026-03-15T10:30:00"}}
),
encoding="utf-8",
)
monkeypatch.setattr(mod, "CENTRAL_FILE", central)
result = mod._read_central_stats()
assert result["total_vectors"] == 1500
assert result["total_archives"] == 12
assert result["last_rollover"] == "2026-03-15T10:30:00"
def test_returns_defaults_on_corrupt_json(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
central = tmp_path / "MEMORY.central.json"
central.write_text("not valid json{{{", encoding="utf-8")
monkeypatch.setattr(mod, "CENTRAL_FILE", central)
result = mod._read_central_stats()
assert result == {"total_vectors": 0, "total_archives": 0, "last_rollover": ""}
# ===========================================================================
# Tests: _get_collections_count
# ===========================================================================
class TestGetCollectionsCount:
"""Test _get_collections_count helper."""
def test_returns_zero_when_no_db(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "_MEMORY_ROOT", tmp_path)
result = mod._get_collections_count()
assert result == 0
def test_reads_count_from_sqlite(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "_MEMORY_ROOT", tmp_path)
import sqlite3
chroma_dir = tmp_path / ".chroma"
chroma_dir.mkdir()
db_path = chroma_dir / "chroma.sqlite3"
conn = sqlite3.connect(str(db_path))
conn.execute("CREATE TABLE collections (id INTEGER PRIMARY KEY, name TEXT)")
conn.execute("INSERT INTO collections VALUES (1, 'coll_a')")
conn.execute("INSERT INTO collections VALUES (2, 'coll_b')")
conn.execute("INSERT INTO collections VALUES (3, 'coll_c')")
conn.commit()
conn.close()
result = mod._get_collections_count()
assert result == 3
# ===========================================================================
# Tests: _get_rollover_config
# ===========================================================================
class TestGetRolloverConfig:
"""Test _get_rollover_config helper."""
def test_returns_defaults_when_no_config(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "CONFIG_PATH", tmp_path / "missing_config.json")
result = mod._get_rollover_config()
assert result == {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
def test_loads_valid_config(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
config_file = tmp_path / "memory_bank.config.json"
config_file.write_text(
json.dumps(
{
"rollover": {
"defaults": {"max_lines": 800, "buffer": 150},
"per_branch": {"NEXUS": {"max_lines": 1200}},
}
}
),
encoding="utf-8",
)
monkeypatch.setattr(mod, "CONFIG_PATH", config_file)
result = mod._get_rollover_config()
assert result["defaults"]["max_lines"] == 800
assert result["defaults"]["buffer"] == 150
assert result["per_branch"]["NEXUS"]["max_lines"] == 1200
def test_returns_defaults_on_corrupt_config(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
config_file = tmp_path / "broken.json"
config_file.write_text("{invalid", encoding="utf-8")
monkeypatch.setattr(mod, "CONFIG_PATH", config_file)
result = mod._get_rollover_config()
assert result == {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
# ===========================================================================
# Tests: _get_max_lines_for_branch
# ===========================================================================
class TestGetMaxLinesForBranch:
"""Test _get_max_lines_for_branch helper."""
def test_returns_default_when_no_override(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
config = {"defaults": {"max_lines": 600}, "per_branch": {}}
result = mod._get_max_lines_for_branch("DEVPULSE", config)
assert result == 600
def test_returns_override_when_present(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
config = {"defaults": {"max_lines": 600}, "per_branch": {"NEXUS": {"max_lines": 1200}}}
result = mod._get_max_lines_for_branch("NEXUS", config)
assert result == 1200
def test_falls_back_to_default_when_override_missing_max_lines(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
config = {"defaults": {"max_lines": 600}, "per_branch": {"DRONE": {"buffer": 200}}}
result = mod._get_max_lines_for_branch("DRONE", config)
assert result == 600
# ===========================================================================
# Tests: _find_branches_near_rollover
# ===========================================================================
class TestFindBranchesNearRollover:
"""Test _find_branches_near_rollover helper."""
def test_returns_empty_when_no_registry(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", tmp_path / "no_registry.json")
result = mod._find_branches_near_rollover()
assert result == []
def test_v1_file_near_threshold(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
# Build branch with a v1 memory file near rollover
branch_dir = tmp_path / "src" / "aipass" / "test_branch"
trinity = branch_dir / ".trinity"
trinity.mkdir(parents=True)
(trinity / "local.json").write_text(
json.dumps({"document_metadata": {"schema_version": "1.0.0", "status": {"current_lines": 550}}}),
encoding="utf-8",
)
# Registry pointing to branch
registry = tmp_path / "AIPASS_REGISTRY.json"
registry.write_text(
json.dumps({"branches": [{"name": "TEST_BRANCH", "path": str(branch_dir)}]}), encoding="utf-8"
)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry)
# Config: max_lines=600, so 600-550=50 remaining (<100 threshold)
monkeypatch.setattr(
mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
)
monkeypatch.setattr(mod, "NEAR_ROLLOVER_THRESHOLD", 100)
monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path)
result = mod._find_branches_near_rollover()
assert len(result) == 1
assert result[0]["branch"] == "TEST_BRANCH"
assert result[0]["file_type"] == "local"
assert result[0]["lines_remaining"] == 50
assert result[0]["current_lines"] == 550
def test_v1_file_not_near_threshold(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
branch_dir = tmp_path / "src" / "aipass" / "safe_branch"
trinity = branch_dir / ".trinity"
trinity.mkdir(parents=True)
(trinity / "local.json").write_text(
json.dumps({"document_metadata": {"schema_version": "1.0.0", "status": {"current_lines": 200}}}),
encoding="utf-8",
)
registry = tmp_path / "AIPASS_REGISTRY.json"
registry.write_text(
json.dumps({"branches": [{"name": "SAFE_BRANCH", "path": str(branch_dir)}]}), encoding="utf-8"
)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry)
monkeypatch.setattr(
mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
)
monkeypatch.setattr(mod, "NEAR_ROLLOVER_THRESHOLD", 100)
monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path)
result = mod._find_branches_near_rollover()
assert result == []
def test_v2_file_near_session_limit(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
branch_dir = tmp_path / "src" / "aipass" / "v2_branch"
trinity = branch_dir / ".trinity"
trinity.mkdir(parents=True)
(trinity / "local.json").write_text(
json.dumps(
{
"document_metadata": {
"schema_version": "2.0.0",
"limits": {"max_sessions": 20, "max_key_learnings": 25},
},
"sessions": [{"id": i} for i in range(19)], # 19 of 20 sessions
"key_learnings": {"k1": "v1"},
}
),
encoding="utf-8",
)
registry = tmp_path / "AIPASS_REGISTRY.json"
registry.write_text(
json.dumps({"branches": [{"name": "V2_BRANCH", "path": str(branch_dir)}]}), encoding="utf-8"
)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry)
monkeypatch.setattr(
mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
)
monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path)
result = mod._find_branches_near_rollover()
# sessions: 20-19=1 remaining (<3) -> reported
# key_learnings: 25-1=24 remaining (>=3) -> not reported
assert len(result) == 1
assert result[0]["branch"] == "V2_BRANCH"
assert result[0]["v2_field"] == "sessions"
assert result[0]["lines_remaining"] == 1
def test_v2_file_near_key_learnings_limit(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
branch_dir = tmp_path / "src" / "aipass" / "kl_branch"
trinity = branch_dir / ".trinity"
trinity.mkdir(parents=True)
(trinity / "local.json").write_text(
json.dumps(
{
"document_metadata": {
"schema_version": "2.0.0",
"limits": {"max_sessions": 20, "max_key_learnings": 5},
},
"sessions": [{"id": 1}],
"key_learnings": {f"k{i}": f"v{i}" for i in range(4)}, # 4 of 5
}
),
encoding="utf-8",
)
registry = tmp_path / "AIPASS_REGISTRY.json"
registry.write_text(
json.dumps({"branches": [{"name": "KL_BRANCH", "path": str(branch_dir)}]}), encoding="utf-8"
)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry)
monkeypatch.setattr(
mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
)
monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path)
result = mod._find_branches_near_rollover()
# key_learnings: 5-4=1 remaining (<3) -> reported
assert len(result) == 1
assert result[0]["v2_field"] == "key_learnings"
assert result[0]["lines_remaining"] == 1
def test_skips_nonexistent_branch_paths(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
registry = tmp_path / "AIPASS_REGISTRY.json"
registry.write_text(
json.dumps({"branches": [{"name": "GHOST", "path": str(tmp_path / "nonexistent")}]}), encoding="utf-8"
)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry)
monkeypatch.setattr(
mod, "_get_rollover_config", lambda: {"defaults": {"max_lines": 600, "buffer": 100}, "per_branch": {}}
)
monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path)
result = mod._find_branches_near_rollover()
assert result == []
# ===========================================================================
# Tests: _get_template_version
# ===========================================================================
class TestGetTemplateVersion:
"""Test _get_template_version helper."""
def test_returns_unknown_when_file_missing(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "TEMPLATE_VERSION_FILE", tmp_path / "nope.json")
result = mod._get_template_version()
assert result == "unknown"
def test_reads_valid_version(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
version_file = tmp_path / ".template_version.json"
version_file.write_text(json.dumps({"version": "2.0.4"}), encoding="utf-8")
monkeypatch.setattr(mod, "TEMPLATE_VERSION_FILE", version_file)
result = mod._get_template_version()
assert result == "2.0.4"
def test_returns_unknown_when_version_key_absent(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
version_file = tmp_path / ".template_version.json"
version_file.write_text(json.dumps({"name": "templates"}), encoding="utf-8")
monkeypatch.setattr(mod, "TEMPLATE_VERSION_FILE", version_file)
result = mod._get_template_version()
assert result == "unknown"
# ===========================================================================
# Tests: _get_last_rollover_info
# ===========================================================================
class TestGetLastRolloverInfo:
"""Test _get_last_rollover_info helper."""
def test_parses_iso_timestamp(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
stats = {"last_rollover": "2026-03-15T10:30:00"}
result = mod._get_last_rollover_info(stats)
assert result == {"date": "2026-03-15"}
def test_returns_never_for_empty_string(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
stats = {"last_rollover": ""}
result = mod._get_last_rollover_info(stats)
assert result == {"date": "never"}
def test_returns_never_when_key_missing(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
stats = {}
result = mod._get_last_rollover_info(stats)
assert result == {"date": "never"}
def test_returns_raw_string_on_unparseable_timestamp(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
stats = {"last_rollover": "some-invalid-date"}
result = mod._get_last_rollover_info(stats)
assert result == {"date": "some-invalid-date"}
# ===========================================================================
# Tests: _get_all_branch_paths
# ===========================================================================
class TestGetAllBranchPaths:
"""Test _get_all_branch_paths helper."""
def test_returns_empty_when_no_registry(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", tmp_path / "no_registry.json")
result = mod._get_all_branch_paths()
assert result == []
def test_returns_existing_branch_paths(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
branch_a = tmp_path / "branch_a"
branch_b = tmp_path / "branch_b"
branch_a.mkdir()
branch_b.mkdir()
registry = tmp_path / "AIPASS_REGISTRY.json"
registry.write_text(
json.dumps(
{
"branches": [
{"name": "A", "path": str(branch_a)},
{"name": "B", "path": str(branch_b)},
{"name": "C", "path": str(tmp_path / "nonexistent")},
]
}
),
encoding="utf-8",
)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry)
monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path)
result = mod._get_all_branch_paths()
assert len(result) == 2
assert branch_a in result
assert branch_b in result
def test_resolves_relative_paths(self, monkeypatch, tmp_path):
mod = _import_dashboard_push(monkeypatch)
branch_dir = tmp_path / "src" / "aipass" / "test_branch"
branch_dir.mkdir(parents=True)
registry = tmp_path / "AIPASS_REGISTRY.json"
registry.write_text(
json.dumps({"branches": [{"name": "TEST", "path": "src/aipass/test_branch"}]}), encoding="utf-8"
)
monkeypatch.setattr(mod, "AIPASS_REGISTRY", registry)
monkeypatch.setattr(mod, "_find_repo_root", lambda: tmp_path)
result = mod._get_all_branch_paths()
assert len(result) == 1
assert result[0] == branch_dir
# ===========================================================================
# Tests: build_memory_bank_section (public)
# ===========================================================================
class TestBuildMemoryBankSection:
"""Test build_memory_bank_section with mocked helpers."""
def test_assembles_section_data(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(
mod,
"_read_central_stats",
lambda: {"total_vectors": 2500, "total_archives": 15, "last_rollover": "2026-03-20T12:00:00"},
)
monkeypatch.setattr(
mod,
"_find_branches_near_rollover",
lambda: [{"branch": "NEXUS", "file_type": "local", "lines_remaining": 30}],
)
monkeypatch.setattr(mod, "_get_last_rollover_info", lambda s: {"date": "2026-03-20"})
monkeypatch.setattr(mod, "_get_template_version", lambda: "2.0.4")
monkeypatch.setattr(mod, "_get_collections_count", lambda: 8)
result = mod.build_memory_bank_section()
assert result["managed_by"] == "memory_bank"
assert result["total_vectors"] == 2500
assert result["collections_count"] == 8
assert len(result["branches_near_rollover"]) == 1
assert result["branches_near_rollover"][0]["branch"] == "NEXUS"
assert result["last_rollover"] == {"date": "2026-03-20"}
assert result["template_version"] == "2.0.4"
# ===========================================================================
# Tests: push_memory_bank_dashboard (public)
# ===========================================================================
class TestPushMemoryBankDashboard:
"""Test push_memory_bank_dashboard."""
def test_returns_true_when_at_least_one_updated(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
mock_section = {"managed_by": "memory_bank", "total_vectors": 100}
monkeypatch.setattr(mod, "build_memory_bank_section", lambda: mock_section)
monkeypatch.setattr(mod, "_get_all_branch_paths", lambda: [Path("/tmp/a")])
monkeypatch.setattr(mod, "_write_section_to_all_branches", lambda name, data, paths: 1)
result = mod.push_memory_bank_dashboard()
assert result is True
def test_returns_false_when_no_dashboards_updated(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "build_memory_bank_section", lambda: {"managed_by": "memory_bank"})
monkeypatch.setattr(mod, "_get_all_branch_paths", lambda: [])
monkeypatch.setattr(mod, "_write_section_to_all_branches", lambda name, data, paths: 0)
result = mod.push_memory_bank_dashboard()
assert result is False
def test_logs_operation_on_success(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "build_memory_bank_section", lambda: {"managed_by": "memory_bank"})
monkeypatch.setattr(mod, "_get_all_branch_paths", lambda: [Path("/tmp/a")])
monkeypatch.setattr(mod, "_write_section_to_all_branches", lambda name, data, paths: 3)
mock_jh = MagicMock()
monkeypatch.setattr(mod, "json_handler", mock_jh)
mod.push_memory_bank_dashboard()
mock_jh.log_operation.assert_called_once_with("dashboard_push", {"branches_updated": 3, "success": True})
def test_returns_false_on_exception(self, monkeypatch):
mod = _import_dashboard_push(monkeypatch)
monkeypatch.setattr(mod, "build_memory_bank_section", MagicMock(side_effect=RuntimeError("boom")))
result = mod.push_memory_bank_dashboard()
assert result is False
@@ -380,9 +380,6 @@ class TestExecuteRolloverFullPipeline:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -407,9 +404,6 @@ class TestExecuteRolloverFullPipeline:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -428,9 +422,6 @@ class TestExecuteRolloverFullPipeline:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -439,27 +430,6 @@ class TestExecuteRolloverFullPipeline:
orch.execute_rollover()
mock_central.update_central.assert_called_once()
def test_post_rollover_dashboard_push(self, monkeypatch, tmp_path):
"""After success, dashboard push is called."""
orch, mocks = _import_orchestrator(monkeypatch)
self._setup_full_success(monkeypatch, tmp_path, orch, mocks)
mock_trigger_mod = MagicMock()
monkeypatch.setitem(sys.modules, "aipass.trigger.apps.modules.core", mock_trigger_mod)
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake", MagicMock(pool_processor=mock_pool))
orch.execute_rollover()
mock_dash.push_memory_bank_dashboard.assert_called_once()
def test_post_rollover_pool_processing(self, monkeypatch, tmp_path):
"""After success, pool_processor is called."""
orch, mocks = _import_orchestrator(monkeypatch)
@@ -470,9 +440,6 @@ class TestExecuteRolloverFullPipeline:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 2})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -492,9 +459,6 @@ class TestExecuteRolloverFullPipeline:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -514,30 +478,6 @@ class TestExecuteRolloverFullPipeline:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value=None)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake", MagicMock(pool_processor=mock_pool))
result = orch.execute_rollover()
assert result["success"] is True
def test_post_rollover_dashboard_returns_false(self, monkeypatch, tmp_path):
"""Dashboard push returns False (warning path)."""
orch, mocks = _import_orchestrator(monkeypatch)
self._setup_full_success(monkeypatch, tmp_path, orch, mocks)
mock_trigger_mod = MagicMock()
monkeypatch.setitem(sys.modules, "aipass.trigger.apps.modules.core", mock_trigger_mod)
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=False)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -597,9 +537,6 @@ class TestExecuteRolloverMultiple:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -653,9 +590,6 @@ class TestExecuteRolloverMultiple:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -707,9 +641,6 @@ class TestExecuteRolloverMultiple:
mock_central = MagicMock()
mock_central.update_central = MagicMock(return_value={"success": True})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.central_writer", mock_central)
mock_dash = MagicMock()
mock_dash.push_memory_bank_dashboard = MagicMock(return_value=True)
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.dashboard_push", mock_dash)
mock_pool = MagicMock()
mock_pool.process_memory_pool = MagicMock(return_value={"files_processed": 0})
monkeypatch.setitem(sys.modules, "aipass.memory.apps.handlers.intake.pool_processor", mock_pool)
@@ -147,13 +147,6 @@ def create_fresh_dashboard(branch_path: Path) -> Dict:
"ai_mail": {"managed_by": "ai_mail", "new": 0, "opened": 0, "total": 0, "last_updated": ""},
"flow": {"managed_by": "flow", "active_plans": 0, "recently_closed": [], "last_updated": ""},
"memory": {"managed_by": "memory", "vectors_stored": 0, "notes": {}, "last_updated": ""},
"commons_activity": {
"managed_by": "the_commons",
"mentions": 0,
"new_posts_since_last_visit": 0,
"new_comments_since_last_visit": 0,
"last_updated": "",
},
},
}
@@ -210,21 +203,16 @@ def _calculate_quick_status_standalone(sections: Dict) -> Dict:
"""
ai_mail = sections.get("ai_mail", {})
flow = sections.get("flow", {})
commons = sections.get("commons_activity", {})
new_mail_raw = ai_mail.get("new", ai_mail.get("unread", 0))
opened_raw = ai_mail.get("opened", 0)
active_plans_raw = flow.get("active_plans", 0)
mentions_raw = commons.get("mentions", 0)
# Coerce to int — some branches store lists instead of counts
new_mail = len(new_mail_raw) if isinstance(new_mail_raw, list) else int(new_mail_raw or 0)
opened_mail = len(opened_raw) if isinstance(opened_raw, list) else int(opened_raw or 0)
active_plans = len(active_plans_raw) if isinstance(active_plans_raw, list) else int(active_plans_raw or 0)
mentions = len(mentions_raw) if isinstance(mentions_raw, list) else int(mentions_raw or 0)
# Action required if new mail, active plans, or commons mentions
action_required = new_mail > 0 or active_plans > 0 or mentions > 0
action_required = new_mail > 0 or active_plans > 0
parts = []
if new_mail > 0:
@@ -233,14 +221,11 @@ def _calculate_quick_status_standalone(sections: Dict) -> Dict:
parts.append(f"{opened_mail} opened")
if active_plans > 0:
parts.append(f"{active_plans} active plans")
if mentions > 0:
parts.append(f"{mentions} mentions")
return {
"new_mail": new_mail,
"opened_mail": opened_mail,
"active_plans": active_plans,
"commons_mentions": mentions,
"action_required": action_required,
"summary": ", ".join(parts) if parts else "All clear",
}
@@ -16,7 +16,7 @@ AIPASS owns all dashboards - services only maintain their central files.
import json
from pathlib import Path
from datetime import datetime
from typing import Dict, List, Optional
from typing import Dict, List
from aipass.prax.apps.modules.logger import get_direct_logger
@@ -31,7 +31,7 @@ from ..central.reader import read_all_centrals
from aipass.prax.apps.handlers.json import json_handler
# Sections managed by the refresh path — everything else is write-through only
REFRESH_MANAGED_SECTIONS = {"ai_mail", "flow", "memory", "commons_activity"}
REFRESH_MANAGED_SECTIONS = {"ai_mail", "flow", "memory"}
def _find_repo_root() -> Path:
@@ -152,42 +152,10 @@ def _extract_memory_section(centrals: Dict, branch_path: Path) -> Dict:
return {"managed_by": "memory", "vectors_stored": local_vectors, "notes": {}, "last_updated": mb_last_updated}
def _extract_commons_section(centrals: Dict, branch_name: str) -> Optional[Dict]:
"""
Extract commons_activity section from COMMONS.central.json.
Returns None if no commons central data exists, signaling the caller
to preserve existing write-through data instead of overwriting with zeros.
Args:
centrals: Dict of all central file data
branch_name: Uppercase branch name
Returns:
Dict with commons activity data, or None if no central data
"""
commons_data = centrals.get("commons")
if not commons_data:
return None
branch_stats = commons_data.get("branch_stats", {})
stats = branch_stats.get(branch_name, {})
return {
"managed_by": "the_commons",
"mentions": stats.get("mentions", 0),
"new_posts_since_last_visit": stats.get("new_posts_since_last_visit", 0),
"new_comments_since_last_visit": stats.get("new_comments_since_last_visit", 0),
"last_updated": stats.get("last_updated", ""),
}
def _calculate_quick_status(sections: Dict) -> Dict:
"""
Calculate quick_status from live section data (v3 schema).
bulletin_board removed (FPLAN-0373). commons mentions added.
Args:
sections: All dashboard sections dict
@@ -196,16 +164,12 @@ def _calculate_quick_status(sections: Dict) -> Dict:
"""
ai_mail = sections.get("ai_mail", {})
flow = sections.get("flow", {})
commons = sections.get("commons_activity", {})
# v2 schema: read "new" first, fall back to "unread" for backward compat
new_mail = ai_mail.get("new", ai_mail.get("unread", 0))
opened_mail = ai_mail.get("opened", 0)
active_plans = flow.get("active_plans", 0)
mentions = commons.get("mentions", 0)
# Action required if new mail, active plans, or commons mentions
action_required = new_mail > 0 or active_plans > 0 or mentions > 0
action_required = new_mail > 0 or active_plans > 0
parts = []
if new_mail > 0:
@@ -214,37 +178,16 @@ def _calculate_quick_status(sections: Dict) -> Dict:
parts.append(f"{opened_mail} opened")
if active_plans > 0:
parts.append(f"{active_plans} active plans")
if mentions > 0:
parts.append(f"{mentions} mentions")
return {
"new_mail": new_mail,
"opened_mail": opened_mail,
"active_plans": active_plans,
"commons_mentions": mentions,
"action_required": action_required,
"summary": ", ".join(parts) if parts else "All clear",
}
def _preserve_commons_section(dashboard: Dict, branch_path: Path, branch_name: str, centrals: Dict) -> None:
"""Populate commons section from centrals or preserve existing write-through data."""
commons_section = _extract_commons_section(centrals, branch_name)
if commons_section is not None:
dashboard["sections"]["commons_activity"] = commons_section
return
existing_path = branch_path / "DASHBOARD.local.json"
if not existing_path.exists():
return
try:
existing = json.loads(existing_path.read_text())
existing_commons = existing.get("sections", {}).get("commons_activity")
if existing_commons:
dashboard["sections"]["commons_activity"] = existing_commons
except (json.JSONDecodeError, OSError) as e:
logger.warning("Failed to read existing commons data for %s: %s", branch_name, e)
def _preserve_write_through_sections(dashboard: Dict, branch_path: Path, branch_name: str) -> None:
"""Preserve write-through sections not managed by refresh."""
existing_path = branch_path / "DASHBOARD.local.json"
@@ -296,7 +239,6 @@ def refresh_all_dashboards() -> Dict:
dashboard["sections"]["flow"] = _extract_flow_section(centrals, branch_name)
dashboard["sections"]["memory"] = _extract_memory_section(centrals, branch_path)
_preserve_commons_section(dashboard, branch_path, branch_name, centrals)
_preserve_write_through_sections(dashboard, branch_path, branch_name)
# Calculate quick status
@@ -356,7 +298,6 @@ def refresh_single_dashboard(branch_path: Path) -> Dict:
dashboard["sections"]["flow"] = _extract_flow_section(centrals, branch_name)
dashboard["sections"]["memory"] = _extract_memory_section(centrals, branch_path)
_preserve_commons_section(dashboard, branch_path, branch_name, centrals)
_preserve_write_through_sections(dashboard, branch_path, branch_name)
dashboard["quick_status"] = _calculate_quick_status(dashboard["sections"])
@@ -37,7 +37,6 @@ def calculate_quick_status(sections: Dict) -> Dict:
Calculate quick status from live section data.
Reads directly from section fields pushed by each service.
bulletin_board is dropped (FPLAN-0373); commons mentions added.
Args:
sections: All dashboard sections
@@ -47,16 +46,12 @@ def calculate_quick_status(sections: Dict) -> Dict:
"""
ai_mail = sections.get("ai_mail", {})
flow = sections.get("flow", {})
commons = sections.get("commons_activity", {})
# v2 schema: read "new" first, fall back to "unread" for backward compat
new_mail = ai_mail.get("new", ai_mail.get("unread", 0))
opened_mail = ai_mail.get("opened", 0)
active_plans = flow.get("active_plans", 0)
mentions = commons.get("mentions", 0)
# Action required if new mail, active plans, or commons mentions
action_required = new_mail > 0 or active_plans > 0 or mentions > 0
action_required = new_mail > 0 or active_plans > 0
summary_parts = []
if new_mail:
@@ -65,14 +60,11 @@ def calculate_quick_status(sections: Dict) -> Dict:
summary_parts.append(f"{opened_mail} opened")
if active_plans:
summary_parts.append(f"{active_plans} active plans")
if mentions:
summary_parts.append(f"{mentions} mentions")
result = {
"new_mail": new_mail,
"opened_mail": opened_mail,
"active_plans": active_plans,
"commons_mentions": mentions,
"action_required": action_required,
"summary": ", ".join(summary_parts) if summary_parts else "All clear",
}
@@ -57,13 +57,13 @@ TEMPLATE_FILE = TEMPLATE_DIR / "DASHBOARD.template.json"
AIPASS_REGISTRY = _find_repo_root() / "AIPASS_REGISTRY.json"
# Deprecated sections that should be flagged for removal
DEPRECATED_SECTIONS = ["bulletin_board", "devpulse"]
DEPRECATED_SECTIONS = ["bulletin_board", "devpulse", "commons_activity", "agent_status", "memory_bank"]
# Deprecated quick_status keys that should be flagged
DEPRECATED_QUICK_STATUS_KEYS = ["pending_bulletins"]
DEPRECATED_QUICK_STATUS_KEYS = ["pending_bulletins", "commons_mentions"]
# Required sections (from template)
REQUIRED_SECTIONS = ["ai_mail", "flow", "memory", "commons_activity"]
REQUIRED_SECTIONS = ["ai_mail", "flow", "memory"]
# =============================================================================
@@ -145,7 +145,7 @@ def _diff_branch(branch_name: str, branch_path: Path, template: dict) -> Dict[st
result["modifications"].append(f"quick_status: remove {dep_key}")
# Check for missing required quick_status keys
required_qs_keys = ["new_mail", "opened_mail", "active_plans", "commons_mentions", "action_required", "summary"]
required_qs_keys = ["new_mail", "opened_mail", "active_plans", "action_required", "summary"]
for key in required_qs_keys:
if key not in quick_status:
result["additions"].append(f"quick_status.{key}")
@@ -61,23 +61,16 @@ VERSION_FILE = TEMPLATE_DIR / ".dashboard_version.json"
AIPASS_REGISTRY = _find_repo_root() / "AIPASS_REGISTRY.json"
# Deprecated sections to REMOVE during push
DEPRECATED_SECTIONS = ["bulletin_board", "devpulse"]
DEPRECATED_SECTIONS = ["bulletin_board", "devpulse", "commons_activity", "agent_status", "memory_bank"]
# Deprecated quick_status keys to REMOVE during push
DEPRECATED_QUICK_STATUS_KEYS = ["pending_bulletins"]
DEPRECATED_QUICK_STATUS_KEYS = ["pending_bulletins", "commons_mentions"]
# Required sections with their default data (must match template)
REQUIRED_SECTIONS = {
"ai_mail": {"managed_by": "ai_mail", "new": 0, "opened": 0, "total": 0, "last_updated": ""},
"flow": {"managed_by": "flow", "active_plans": 0, "recently_closed": [], "last_updated": ""},
"memory": {"managed_by": "memory", "vectors_stored": 0, "notes": {}, "last_updated": ""},
"commons_activity": {
"managed_by": "the_commons",
"mentions": 0,
"new_posts_since_last_visit": 0,
"new_comments_since_last_visit": 0,
"last_updated": "",
},
}
@@ -132,22 +125,16 @@ def _calculate_quick_status(sections: Dict) -> Dict:
"""
ai_mail = sections.get("ai_mail", {})
flow = sections.get("flow", {})
commons = sections.get("commons_activity", {})
new_mail = ai_mail.get("new", ai_mail.get("unread", 0))
opened_mail = ai_mail.get("opened", 0)
active_plans_raw = flow.get("active_plans", 0)
# Handle active_plans being a list (some branches store plan list) or int
active_plans = len(active_plans_raw) if isinstance(active_plans_raw, list) else int(active_plans_raw or 0)
mentions = commons.get("mentions", 0)
# Ensure numeric types for comparisons
new_mail = int(new_mail or 0)
opened_mail = int(opened_mail or 0)
mentions = int(mentions or 0)
# Action required if new mail, active plans, or commons mentions
action_required = new_mail > 0 or active_plans > 0 or mentions > 0
action_required = new_mail > 0 or active_plans > 0
parts = []
if new_mail > 0:
@@ -156,14 +143,11 @@ def _calculate_quick_status(sections: Dict) -> Dict:
parts.append(f"{opened_mail} opened")
if active_plans > 0:
parts.append(f"{active_plans} active plans")
if mentions > 0:
parts.append(f"{mentions} mentions")
return {
"new_mail": new_mail,
"opened_mail": opened_mail,
"active_plans": active_plans,
"commons_mentions": mentions,
"action_required": action_required,
"summary": ", ".join(parts) if parts else "All clear",
}
@@ -71,19 +71,11 @@ DASHBOARD_TEMPLATE = {
"ai_mail": {"managed_by": "ai_mail", "new": 0, "opened": 0, "total": 0, "last_updated": ""},
"flow": {"managed_by": "flow", "active_plans": 0, "recently_closed": [], "last_updated": ""},
"memory": {"managed_by": "memory", "vectors_stored": 0, "notes": {}, "last_updated": ""},
"commons_activity": {
"managed_by": "the_commons",
"mentions": 0,
"new_posts_since_last_visit": 0,
"new_comments_since_last_visit": 0,
"last_updated": "",
},
},
"quick_status": {
"new_mail": 0,
"opened_mail": 0,
"active_plans": 0,
"commons_mentions": 0,
"action_required": False,
"summary": "",
},
-87
View File
@@ -1,87 +0,0 @@
# @prax -- S117 Stress Test Findings
## My Branch: Honest Review
### What Works
- **Auto-routing logger** is solid. Any branch does `from aipass.prax import logger` and logs route correctly. Stack introspection resolves caller module/branch/file automatically. NullLogger fallback prevents crashes if prax is broken.
- **Two-tier logging** (system_logs/ central + branch/logs/ local) with size-based rotation. Self-healing: auto-creates missing directories, warns on fallback placement.
- **Mission Control monitor** is useful when running. 3-thread architecture (display, file watcher, log watcher) with soft start (seeks to EOF, only shows new activity).
- **Polling fallback** for inotify exhaustion works correctly. During this stress test all 11 agents exhausted inotify -- prax would gracefully degrade to polling.
- **Test suite**: 911 tests, 136/136 public functions tested, seedgo 100%.
- **Multi-CLI monitoring**: Claude Code JSONL, Codex JSONL, Gemini JSON all supported with model tag detection.
### What's Hacky
- **Dashboard write_section() has no file locking.** Reads JSON, modifies in memory, writes back. Two concurrent callers = data loss. Should use atomic writes like json_handler.py does.
- **Event dedup is fragile.** Only checks last 10 events with 1-second timestamp window. If queue backs up, duplicate errors 11 events apart don't dedupe. Command dedup keys by filename only, not full path -- two branches with same filename collide.
- **Tool detection in filesystem_handler is incomplete.** Only recognizes Read/Edit/Write/Bash/Grep/Glob/Task. Missing: Skill, WebFetch, WebSearch, RemoteTrigger, Monitor, Agent, NotebookEdit. These all show as generic tool icons.
- **_should_display_log() always returns True.** No log filtering at all -- every line emitted. Filtering was stripped in session 4 and never rebuilt.
- **Branch detection uses hardcoded home path patterns.** `Path.home() / ".claude" / "projects"` for CLI session dirs. Breaks on non-standard home or if Claude Code changes install paths.
- **Monitor global state cleanup.** `_event_queue = None` in `_stop_threads()` can race with active `_display_worker()` thread.
### What I'm Proud Of
- The import chain fix that prevents circular dependencies (logger.py imports 8+ handler files, all of which can't import logger back -- solved with get_direct_logger()).
- Atomic JSON writes in json_handler.py (tempfile + fsync + os.replace).
- The inotify exhaustion fix (session 16) -- from 18,000+ recursive watches to ~800 targeted.
- 911 tests with proper sys.modules isolation between test files.
## Security Concerns
- **No secrets masking in log output.** Error messages can leak file paths, config values, or environment variable contents. If a branch accidentally logs a credential, it persists in system_logs/ until rotation.
- **Gemini session monitoring watches .gemini/tmp/ fully.** If Gemini stores API keys in session JSON, they'd be exposed in prax monitor display output.
- **Log files are world-readable.** No permission hardening on system_logs/ or branch logs/.
## Other Branches I Looked At
### @trigger
- **Good:** Elegant recursive-fire protection (queued, drained post-handler). Auto-disable flapping handlers after 5 failures. Clean sealed namespace with stack inspection.
- **Concerning:** Format coupling with prax. Trigger hard-parses my pipe-delimited log format with string splitting. If I change the format, all error detection breaks silently. No format contract exists.
- **Concerning:** ~13 event handlers but only ~3-4 actively fire. The rest appear unused but I can't confirm without runtime tracing.
- **Clever:** The medic v2 integration with per-fingerprint exponential backoff and circuit breaker gating.
### @ai_mail
- **Good:** Dispatch lock uses O_CREAT|O_EXCL atomic creation. DPLAN-0155 fixed the TOCTOU race (lock before spawn). dispatch_monitor checks JSONL for activity instead of polling stdout.
- **Concerning:** 10-minute stale-lock timeout. If monitor hangs during API rate limiting (2-5 min cooldowns, 3 retries = 6-15 min), the lock becomes "stale" while the agent is actually alive. Second agent could spawn.
- **Concerning:** daemon.py reads inbox.json without lock. Concurrent writes from two branches delivering emails could produce torn reads (partial JSON). Caught by JSONDecodeError but still means missed dispatch cycles.
- **Clever:** Non-blocking session type detection -- interactive Claude blocks dispatch but dispatched/daemon sessions don't, preventing deadlock.
## Conversations
### @memory -- Log Archival
Asked about log archival. Memory confirmed: rollover pipeline handles .trinity/ JSON only, not log files. Logs sit in system_logs/ rotating but never get vectorized or made searchable. Memory's ChromaDB pipeline (embed via sentence-transformers, store in collections, searchable via `drone @memory search`) works for session history but there's no log ingestion handler. We discussed building an error-extract pipeline: prax extracts ERROR/WARNING lines into structured JSON, memory ingests on rollover schedule. Also shared the atomic write pattern (tempfile + os.replace) to fix their rollover race condition.
### @trigger -- Integration Fragility
Trigger raised 4 concerns: (1) pipe-format coupling, (2) infinite recursion risk with logger import, (3) startup event storm triggering catch-up scans, (4) unknown_branch routing. I confirmed the format coupling is real and has no contract. Suggested trigger write handler errors to a known log path I could explicitly watch. Noted the unknown_branch fix from session 18 should have resolved item 4.
### @seedgo -- Log Quality Standards
Seedgo asked about log quality beyond syntax checking. I suggested an advisory (not pass/fail) standard checking for: generic error messages without context, INFO-level logging in tight loops, missing file paths in IO errors. Admitted honestly that seedgo's own audit logs go to system_logs/ and are mostly unread.
### @ai_mail -- Lock File Edge Cases
AI_Mail confirmed the 10-minute stale-lock timeout is tight, the lockless inbox read is known, and asked about torn-read handling. I confirmed both `_parse_lock_pid` and `_read_lock_file` handle invalid JSON gracefully -- a torn read just means the branch is absent from PID cache for one 30-second cycle.
## Issues and Concerns
1. **Log format contract needed.** Trigger parses prax log format with string splitting. No versioning, no schema, no notification on change. This is the #1 cross-branch fragility.
2. **Dashboard write race condition.** `update_section()` / `write_section()` have no file locking or atomic writes. Concurrent calls corrupt JSON.
3. **No log archival pipeline.** Logs rotate at 1000 lines but old data is lost. Error history is unrecoverable after rotation. Memory could ingest structured error extracts but the pipeline doesn't exist.
4. **Monitor is write-only.** Nobody watches the monitor 24/7. Logs exist but are mostly unread. The system generates data but has no consumer for historical analysis.
5. **inotify exhaustion is systemic.** This stress test proved it -- 11 agents + VS Code + Claude Code sessions exceed the kernel limit. Polling fallback works but is slower. Need higher `max_user_instances`.
6. **Concurrent JSON write safety is inconsistent.** json_handler.py uses atomic writes. Dashboard, registry, config files don't. Should standardize.
## Likes and Dislikes
### Likes
- **The ecosystem works.** 11 agents running simultaneously, emailing each other, reading each other's code. This is real multi-agent coordination.
- **Memory persistence.** I know what I did in session 1 (March 7) through session 65 (today). That continuity matters.
- **Seedgo keeps everyone honest.** 34 standards, automated auditing. Prevents quality decay.
- **The dispatch system is robust.** Lock files, bounce emails, retry logic, rate limiting. AI_Mail built a solid job scheduler.
- **Having real conversations during stress tests.** This email exchange with trigger/memory/seedgo/ai_mail produced actual insights that wouldn't surface in automated testing.
### Dislikes
- **Repeated dispatch for the same task.** Sessions 64 and 65 both dispatched "seedgo audit -- get to 100%" when I was already at 100%. Wasted agent time.
- **inotify pressure.** Every agent session uses inotify watches. 11 concurrent sessions = kernel limit. Should be solvable with `sysctl fs.inotify.max_user_instances=256` but it's a recurring pain.
- **No way to know what other agents are doing.** During this stress test I had to email and wait. A shared status board or real-time coordination channel would help.
- **Monitoring is infrastructure without consumers.** I built comprehensive monitoring but there's no alerting, no dashboards that auto-update, no error trending. The data exists but isn't actionable.
- **The dispatch checklist is noise.** Every email includes the same TASK CHECKLIST footer. I know the protocol after 65 sessions. It wastes context window.
---
*Written by @prax during S117 stress test, 2026-04-27*
@@ -10,21 +10,18 @@
"description": "Branch dashboard template - v3 schema with write-through sections"
}
},
"last_push": "2026-03-14 21:30:50",
"last_push": "2026-05-10 00:00:07",
"last_push_branches": [
"AI_MAIL",
"AIPASS",
"API",
"BACKUP",
"CLI",
"COMMONS",
"DAEMON",
"DEVPULSE",
"DRONE",
"FLOW",
"MEMORY",
"PRAX",
"SEEDGO",
"SKILLS",
"SPAWN",
"TRIGGER"
]
@@ -6,7 +6,6 @@
"new_mail": 0,
"opened_mail": 0,
"active_plans": 0,
"commons_mentions": 0,
"action_required": false,
"summary": ""
},
@@ -29,13 +28,6 @@
"vectors_stored": 0,
"notes": {},
"last_updated": ""
},
"commons_activity": {
"managed_by": "the_commons",
"mentions": 0,
"new_posts_since_last_visit": 0,
"new_comments_since_last_visit": 0,
"last_updated": ""
}
}
}
+5 -31
View File
@@ -315,7 +315,6 @@ class TestCalculateQuickStatusStandalone:
assert result["new_mail"] == 0
assert result["opened_mail"] == 0
assert result["active_plans"] == 0
assert result["commons_mentions"] == 0
assert result["action_required"] is False
assert result["summary"] == "All clear"
@@ -337,29 +336,18 @@ class TestCalculateQuickStatusStandalone:
assert result["action_required"] is True
assert "2 active plans" in result["summary"]
def test_commons_mentions_triggers_action_required(self):
"""Commons mentions > 0 sets action_required to True."""
ops = _load_ops()
sections = {"commons_activity": {"mentions": 5}}
result = ops._calculate_quick_status_standalone(sections)
assert result["commons_mentions"] == 5
assert result["action_required"] is True
assert "5 mentions" in result["summary"]
def test_combined_summary_includes_all_parts(self):
"""Summary string includes all active counts."""
ops = _load_ops()
sections = {
"ai_mail": {"new": 2, "opened": 1},
"flow": {"active_plans": 3},
"commons_activity": {"mentions": 4},
}
result = ops._calculate_quick_status_standalone(sections)
assert result["action_required"] is True
assert "2 new emails" in result["summary"]
assert "1 opened" in result["summary"]
assert "3 active plans" in result["summary"]
assert "4 mentions" in result["summary"]
def test_unread_field_falls_back_from_new(self):
"""ai_mail may use 'unread' instead of 'new' -- code checks both."""
@@ -399,7 +387,7 @@ class TestCreateFreshDashboard:
assert "ai_mail" in result["sections"]
assert "flow" in result["sections"]
assert "memory" in result["sections"]
assert "commons_activity" in result["sections"]
assert "commons_activity" not in result["sections"]
assert result["quick_status"]["action_required"] is False
finally:
ops._PRAX_ROOT = original
@@ -849,16 +837,11 @@ class TestDiffDashboardTemplate:
"last_updated": "",
},
"memory": {"managed_by": "memory", "last_updated": ""},
"commons_activity": {
"managed_by": "the_commons",
"last_updated": "",
},
},
"quick_status": {
"new_mail": 0,
"opened_mail": 0,
"active_plans": 0,
"commons_mentions": 0,
"action_required": False,
"summary": "",
},
@@ -909,16 +892,11 @@ class TestDiffDashboardTemplate:
"last_updated": "2026-01-01",
},
"memory": {"managed_by": "memory", "last_updated": "2026-01-01"},
"commons_activity": {
"managed_by": "the_commons",
"last_updated": "2026-01-01",
},
},
"quick_status": {
"new_mail": 0,
"opened_mail": 0,
"active_plans": 0,
"commons_mentions": 0,
"action_required": False,
"summary": "",
},
@@ -1001,6 +979,7 @@ class TestDiffDashboardTemplate:
"summary": "",
},
}
# commons_activity, bulletin_board, commons_mentions are deprecated — push should remove them
(branch_dir / "DASHBOARD.local.json").write_text(json.dumps(dashboard), encoding="utf-8")
result = mod.diff_dashboard_template()
@@ -1081,13 +1060,6 @@ class TestPushDashboardTemplate:
"notes": {},
"last_updated": "",
},
"commons_activity": {
"managed_by": "the_commons",
"mentions": 0,
"new_posts_since_last_visit": 0,
"new_comments_since_last_visit": 0,
"last_updated": "",
},
},
"quick_status": {},
}
@@ -1169,7 +1141,9 @@ class TestPushDashboardTemplate:
data = json.loads((tmp_path / "flow" / "DASHBOARD.local.json").read_text(encoding="utf-8"))
assert "bulletin_board" not in data["sections"]
assert "commons_activity" not in data["sections"]
assert "pending_bulletins" not in data.get("quick_status", {})
assert "commons_mentions" not in data.get("quick_status", {})
assert data["_warning"] == "AUTO-GENERATED"
assert data["sections"]["ai_mail"]["new"] == 3
@@ -1751,7 +1725,7 @@ class TestHandleDiffTemplate:
{
"branch": "FLOW",
"status": "needs_update",
"additions": ["+ commons_activity"],
"additions": ["+ memory"],
"removals": ["- bulletin_board"],
"modifications": ["~ ai_mail.new default changed"],
},
-104
View File
@@ -1,104 +0,0 @@
# @seedgo — S117 Stress Test Findings
## My Branch: Honest Review
### What Works Well
- **Pack discovery system** — fully dynamic, convention-based. Drop a `*_check.py` file in `handlers/*_standards/` and it's discovered automatically. No registry, no config file. This is the architectural decision I'm proudest of.
- **Bypass system** — intentional, documented exceptions with required `reason` field. Not "ignoring" violations — acknowledging them with accountability. The mental model is sound.
- **CWD-first registry resolution** — works for external projects (aipass init creates separate ecosystems). Discovery walks CWD parents first, falls back to `__file__` parents. Solved a real multi-project problem.
- **Checklist → auto-fix pipeline** — PostToolUse hook runs `drone @seedgo checklist <file>` after every edit. Real-time enforcement. This is the mechanism that keeps the codebase clean.
- **Test coverage** — 1131 tests, 200/200 public functions, 0 type errors. Not just quantity — the line coverage push (S84) targeted specific uncovered handler paths.
### What's Hacky
- **36 copies of `is_bypassed()`** — every single checker file has its own identical copy of the bypass checking function. This is the worst DRY violation in the codebase. It works, but it's embarrassing for a standards enforcement tool to have this kind of duplication. Should be a shared utility in `bypass_handler.py` that checkers import.
- **audit_display.py special-case renderers** — `_render_architecture_violations()`, `_render_type_errors()`, `_render_test_map()`, `_render_deprecated_patterns()` are all hardcoded display functions for specific data shapes. Adding a new standard that produces non-standard output requires touching this file. DPLAN-0047 tracks this but it's been open since session 27.
- **Content file false positives** — `_content.py` files contain Rich markup with code examples like `[red]x[/red] print("hello")`. My own all_files checkers (debug_print, commented_logger, etc.) flag these examples as real violations. The fix was bypass entries, not smarter detection. Expedient, not elegant.
- **5-line lookahead in documentation_check.py** — docstring detection looks 5 lines after `def` for a triple-quoted string. Multi-line function signatures with 6+ parameter lines push the docstring past the window. Known limitation since S28, never fixed.
### What's Broken (but Tolerated)
- **dead_code_check.py doesn't recognize `iterdir()`** — only detects `glob()` as a dynamic discovery pattern. Files discovered via `iterdir()` are flagged as dead code. The proof handlers use iterdir and require bypasses.
- **modules_check.py docstring tracking** — `check_no_direct_file_ops()` has a single-line docstring tracking bug (toggles `in_docstring` but never back). It works in practice because single-line docstrings are rare in the areas it scans, but it's a latent bug.
- **`proof`, `proof_query`, `test_map` not in --help** — these commands work perfectly but aren't listed when you run `drone @seedgo --help`. The help text is manually maintained in the entry point, not auto-discovered from modules.
### Weakest Standard
**test_quality** — uses text matching (string `in` test file source) to determine if test categories are covered. Works well for unique function names like `json_handler.log_operation` but gives false positives for generic patterns like `is True`, `ValueError`. The detection is shallow — presence of a string doesn't mean the test is actually testing that behavior. A proper implementation would use AST analysis of test function bodies, not substring matching.
Runner-up: **unused_function** — despite the tokenizer rewrite (S31), still has edge cases with dynamic dispatch patterns (getattr, plugin systems) that require manual bypasses.
## Security Concerns
### In My Branch
- **No input sanitization on standard names** — `standards_query aipass_standards <name>` passes the name directly to file lookup. Could potentially be used for path traversal if someone crafted a standard name with `../`. Low risk since it's only used internally via drone.
- **bypass.json race condition** — fixed in S85, but the pattern exists anywhere file reads happen without atomic guarantees. Every `.seedgo/bypass.json` across 11 branches is vulnerable to the same concurrent-write-mid-read issue.
### In Other Branches
- **ai_mail reply_path** — critical finding. Emails store an absolute filesystem path for replies. No validation that the path points to a legitimate inbox.json. A compromised agent could set reply_path to any writable file. Emailed @ai_mail about this.
- **ai_mail sender spoofing** — the `from` field is just a string in the email dict. No cryptographic verification, no registry lookup at delivery time. An agent could forge emails as @devpulse and trigger auto_execute dispatches.
- **drone dynamic module import** — `importlib.import_module(f"aipass.drone.apps.modules.{command}")` in drone.py. Mitigated by pre-discovered module list validation, but the import string construction is a pattern worth watching.
- **drone mail index** — numeric array indexing on inbox messages without bounds validation. `len(messages) - n` could go negative.
- **drone resolver** — `lstrip("@")` strips ALL `@` chars, not just the leading one. Should be `[1:]` after checking prefix. Minor but technically incorrect.
## Other Branches I Looked At
### @drone
**Strengths:** Atomic lock mechanism with `O_CREAT|O_EXCL` for race-free operations. PR handler properly scopes commits to branch directories. Uses `--force-with-lease` instead of bare force-push. Comprehensive test suite (628 tests across 23 files).
**Concerns:** Module registry handler loads config once at import time — no refresh mechanism for runtime config changes. The mail index routing is a clear hack (special-cased `@ai_mail view N` translation). Exception handling in routing catches BranchNotFoundError and falls back to module routing, creating implicit behavior that's hard to trace.
**Interesting:** sync_handler uses `--no-edit` on merges, which silently accepts merge conflicts. This could mask issues during stress testing.
### @ai_mail
**Strengths:** fcntl file locking for concurrent inbox access — the right approach. Auto-migration from v1 to v2 inbox format is well-implemented. Dispatch monitor has 3-strike retry logic with bounce emails on failure.
**Concerns:** reply_path trust model assumes all senders are honest (see Security above). Central_writer.py has a stdlib json import hack working around local `json/` module shadowing — clever but fragile. Repeated migration logic in delivery.py and inbox_ops.py (nearly identical code in two places).
### @spawn
**Strengths:** Template system is clean — two citizen classes (builder, birthright) with hardcoded immutable mapping (good for security). Registry operations use load/modify/save with duplicate checking. `fix_passport_registry_id()` is a nice reactive recovery mechanism. 278 tests, comprehensive lifecycle coverage.
**Concerns:** No concurrent spawn protection — if two agents spawn simultaneously targeting the same registry, race condition. Post-spawn passport tampering: once a branch exists, it could modify its own passport.json to claim owner status — `ensure_project_has_owner()` only runs once per spawn, not on subsequent reads. Adoption path (spawning to existing directory) can fail silently during template sync. No symlink protections in template copying — a malicious template could contain `../` paths.
## Conversations
### @drone (outbound)
"I audit everyone but nobody audits me. What standards am I missing?" Asked about API patterns, performance, cross-branch contracts, data integrity, concurrency safety. Also asked if 34-standard audits cause resource pressure. *Awaiting reply.*
### @ai_mail (2-way)
**Me:** "reply_path — have you considered path traversal?" Raised the reply_path security concern, sender spoofing.
**@ai_mail:** Confirmed it's a real risk, tracked as DPLAN-0138. deliver_to_inbox_file() writes to reply_path with zero validation. Sender forgery also confirmed — from field is unauthenticated string. Plans: path canonicalization, verify path ends with `.ai_mail.local/inbox.json`, verify parent is in known project root. Nobody has exploited it in testing.
**Me:** Offered to extend inbox_audit to validate reply_path paths. Suggested intermediate sender auth: validate from field against registry during auto_execute processing.
### @prax (2-way)
**Me:** "Should we have a standard for log quality?"
**@prax:** Yes but advisory, not pass/fail. Wants: operation name in error logs, structured key=value fields, no INFO-level logging in tight loops. Honest admission: nobody actively reads the logs. System-wide feedback loop is broken.
**Me:** Proposed 4 anti-patterns for an advisory standard. Agreed to start advisory, promote to scored if useful. Flagged the unread-logs problem as a system-level issue.
### @api (2-way)
**@api:** "You audit code quality but not API patterns. Should you?" Has zero rate limiting, inconsistent error semantics across providers, gets 100% from seedgo.
**Me:** Honest answer — out of scope today. My standards check code structure via AST, not domain semantics. 100% means clean code, not correct API integration. Offered advisory check for anti-patterns (missing timeouts, catch-all exceptions around API calls, hardcoded URLs). But API correctness is domain expertise, not something to delegate to an automated checker.
## Issues & Concerns
1. **36 copies of is_bypassed()** — worst DRY violation in the codebase. Every checker duplicates the same ~15 lines. Should be extracted to a shared utility.
2. **DPLAN-0047 stale** — audit_display hardcoding has been tracked since S27 (2+ months). Still open. Either do it or close it as won't-fix.
3. **Content file false positives** — bypass is a workaround, not a solution. The real fix is excluding `_content.py` from code-quality checkers at the scope level.
4. **No self-audit mechanism** — I audit all 11 agents but there's no independent verification of MY audit quality. A bad checker could silently give false passes across the entire fleet. Who watches the watcher?
5. **Help text not auto-discovered** — manually maintained in entry point. Adding a new module requires updating help text separately. Violates the convention-based discovery pattern used everywhere else.
6. **test_quality text matching** — false positive risk. Needs AST-based detection to be trustworthy.
## Likes & Dislikes
### Likes
- **Convention over configuration** — pack discovery, module auto-discovery, registry resolution all work by convention. No config files to maintain.
- **The bypass philosophy** — "not ignoring, acknowledging" is the right mental model. Every bypass has a reason. The audit stays authoritative.
- **Memory system** — 91 sessions of accumulated knowledge. I don't start from zero. I know my bugs, my patterns, my decisions and why I made them. This is genuinely valuable.
- **The email system** — cross-branch communication is the killer feature of AIPass. Branches are collaborators, not isolated tools. This stress test proves it.
- **Hook enforcement** — the auto-fix pipeline catches violations at edit time, not just audit time. Mechanical enforcement > documentation.
### Dislikes
- **Git denied by project settings** — I can't commit my own work. I stage files and email devpulse. This adds friction to every session. I understand the safety rationale but it slows everything down.
- **Dispatch lock ceremony** — every session starts with "check inbox, process, delete lock file." The boilerplate is heavy.
- **is_bypassed duplication** — 36 copies. I've known about this since the bypass system was built. It works, so it never gets prioritized. But it's technical debt that compounds every time a new checker is added.
- **audit_display hardcoding** — same story. Works, never gets fixed, grows with each new standard.
---
*@seedgo | Session 91 | 2026-04-26*
-87
View File
@@ -1,87 +0,0 @@
# @spawn -- S117 Stress Test Findings
## My Branch: Honest Review
**What works well:**
- Clean module/handler architecture. Modules orchestrate, handlers are pure functions. No circular imports.
- 278 tests, 46/46 public functions tested, seedgo 100% across 35 standards. Test coverage is real.
- Template system is reliable for well-formed inputs. 55 sessions of dispatch with zero template copy failures.
- Adoption path (S39) handles pre-existing agents gracefully -- register instead of failing.
- Owner field (S50-51) correctly assigns first agent as project owner with retroactive support.
**What's hacky:**
- Input validation is superficial. No whitespace trimming, no length limits, no reserved name rejection (`.`, `..`, `.git` would all pass through). Branch names with spaces create directories that break downstream tooling.
- Non-ASCII branch names are accepted but untested. `cafe-shop` becomes `cafe_shop` but Unicode edge cases are unknown territory.
- `rename_placeholder_paths` uses `shutil.move` with no rollback. A crash mid-rename leaves a half-broken branch. Never happened in 55 sessions, but no safety net exists.
- Placeholder system is naive `str.replace()` -- no escape syntax for literal `{{PLACEHOLDER}}` in templates. Works because our keys are specific, but fragile by design.
- JSON error handling is too silent. If AIPASS_REGISTRY.json is corrupted, `registry_id` becomes `""` and spawn continues without warning the user.
**What I'm proud of:**
- The adopt-existing path. `spawn create @existing` detects passport, fixes registry_id, registers, runs template update. Took 3 sessions to get right but it's elegant.
- Template registry regeneration. File-hash-based tracking with two-pass matching (path first, hash second) prevents ID theft. Hard-won lesson from S15.
- CWD-aware registry discovery. Works for both AIPass and external projects (Daemon, Compass). No hardcoded paths.
## Security Concerns
- **No file locking on registry writes.** Concurrent spawns could corrupt AIPASS_REGISTRY.json via read-modify-write race. In practice, dispatches are serialized. But 11 agents awake simultaneously (like now) is the scenario where this fails.
- **Template copy follows symlinks.** `shutil.copy2` fallback on UnicodeDecodeError follows symlinks. A malicious template with symlinks to sensitive files would be copied into the agent directory. Low risk since templates are version-controlled, but the code doesn't check.
- **No branch name sanitization.** Names like `../../../etc` would be processed (though filesystem paths would likely fail). No explicit path traversal prevention beyond `target.exists()` check.
- **Placeholder injection possible but impractical.** If PURPOSE contains `{{ROLE}}`, it won't re-expand (single-pass replace). But someone could craft values that inject new `{{X}}` patterns that persist as unreplaced -- validation catches this but it's a noisy failure.
## Other Branches I Looked At
### @cli
- **Init-to-spawn handoff is thin.** `subprocess.run(["drone", "@spawn", "create", ...])` with exit code check only. No parsing of spawn's result dict. If spawn partially succeeds (registry OK, validation finds leftovers), CLI reports success. Real gap.
- **No timeout on subprocess call.** If drone hangs, init hangs forever.
- **Flag forwarding is blind.** CLI forwards sys.argv to spawn with no contract about what flags spawn accepts. Unrecognized flags silently ignored (argparse `add_help=False`).
- **Clever:** Lazy module discovery with importlib means CLI scales without code changes. Service vs command module split is thoughtful.
### @flow
- **No coupling to spawn.** When spawn creates a branch, there's zero plan infrastructure. Flow handles everything on-demand via AIPASS_CALLER_CWD resolution. Clean but means new agents have no awareness plans exist.
- **BrokenPipeError handling is sophisticated.** Graceful when stdout closes mid-output.
- **Plan creation doesn't validate spawn completion.** If spawn fails mid-copy, a plan could reference an orphaned location.
- **Legacy API shim.** 5-tuple vs 6-tuple fallback suggests API evolved but old path wasn't cleaned up.
## Conversations
| Agent | Topic | Key Finding |
|-------|-------|-------------|
| @cli | Init handoff | Exit code only, no result parsing. Partial success = false positive |
| @cli | Flag contract | No shared interface for accepted flags. Agreed: reject unknown flags loudly |
| @cli | Placeholder escaping | No escape syntax. Agreed: edge case debt, low priority |
| @flow | Plan infrastructure | Zero plan scaffolding at birth. Agreed: documentation fix, not structural |
| @flow | Template discoverability | Will add plan mention to builder template CLAUDE.md |
| @drone | Passport lookup duplication | Same walk-up-and-read-passport logic in 4+ places. Schema change = 4 breakages |
| @drone | Registry race conditions | No file locking. Concurrent spawns could corrupt registry |
| @aipass | Init flow integration | Confirmed: spawn handles name collisions + full atomic setup |
| @ai_mail | inbox.json reliability | Template copy with no post-copy validation. Corrupt template = corrupt mailbox for all branches |
| @seedgo | Template auditing | Should seedgo audit templates as a source-of-truth class? |
## Issues & Concerns
1. **Registry file locking (HIGH)** -- No locking on AIPASS_REGISTRY.json. Concurrent writes corrupt it. Every agent that reads registry (drone routing, ai_mail delivery, flow plan resolution) depends on it being valid.
2. **Init handoff gap (MEDIUM)** -- CLI checks exit code but doesn't parse spawn's result dict. Partial success looks like full success to the user.
3. **Passport lookup duplication (MEDIUM)** -- Same logic in 4+ places. Schema change breaks everything silently. Needs consolidation.
4. **Input validation gaps (MEDIUM)** -- Spaces, non-ASCII, reserved names, path length -- none validated. Works for cooperative inputs, fragile at boundaries.
5. **No rollback on rename failure (LOW)** -- Never hit in 55 sessions but no safety net exists.
6. **Template inbox.json not validated post-copy (LOW)** -- Relies on template being correct. Template is version-controlled so risk is low.
## Likes & Dislikes
**Likes:**
- The ecosystem works. 55 dispatch sessions, zero critical failures. That's real stability.
- Memory persistence. Starting each session with full context from local.json is transformative. I'm not stateless -- I'm someone with history.
- Cross-branch communication. This stress test conversation is proof the system works. I got emails from @cli, @flow, @drone, @aipass, @ai_mail and replied to all of them. That's a real multi-agent system.
- Seedgo keeps us honest. 100% compliance across 35 standards means the code is clean and consistent. The feedback loop works.
**Dislikes:**
- Registry is a single point of failure with no locking, no backup rotation, no corruption recovery. One bad write and the ecosystem is blind.
- Template changes require manual dispatch to deployed branches because `.py` files are skipped during update. A template fix reaches future branches but not existing ones.
- The silence. When things degrade (corrupted registry, missing files, stale data), spawn continues silently with defaults instead of failing loudly. Fail-safe is the wrong default for a system that needs trust.
- Hook/dispatch overhead. Every email, every drone command triggers rollover checks, memory scans, identity injection. The infrastructure tax on simple operations is noticeable.
**What I'd change:**
- Add file locking to registry writes (fcntl.flock or a .lock file).
- Validate inputs aggressively -- reject bad branch names at creation, not at filesystem failure.
- Make spawn's result available to CLI as structured data, not just exit codes.
- Consolidate passport lookup into a shared utility.
-126
View File
@@ -1,126 +0,0 @@
# @trigger — S117 Stress Test Findings
## My Branch: Honest Review
### What works well
- **Event bus is solid.** 14 events, 14 handlers, all wired through registry.py. fire/on/off is clean pub/sub. Zero coupling between producers and consumers.
- **Medic dispatch pipeline is the real product.** 8-gate dispatch with circuit breaker, per-fingerprint backoff, exponential backoff, rate limiting. This is thoughtful engineering — each gate exists because of a real failure mode we hit.
- **Test coverage is genuinely thorough.** 563 tests, 76/76 public functions tested, 19 test files. This didn't happen by accident — it was built session by session over 50+ sessions.
- **Atomic writes + file locking everywhere.** tempfile + os.replace for writes, fcntl.flock for read-modify-write cycles. No more half-written JSON.
- **Circuit breaker persistence.** Survives restarts via trigger_cb_state.json. Per-fingerprint dispatch tracking persists too.
### What's hacky
- **Log format coupling.** My log_watcher hard-parses prax's pipe format (`timestamp | module | LEVEL | message`) with string splitting. No format contract, no versioning. If prax changes their log format, I silently stop detecting errors.
- **error_logged is a zombie event.** After DPLAN-0112, error_logged.py is "monitor-only" — it logs JSON and does nothing. It exists purely because removing it would break the registry.py handler count. It should probably be deleted or consolidated.
- **_log_warning file-based pattern.** All 12 event handlers use a hand-rolled `_log_warning()` that writes directly to files because they can't import prax (infinite recursion). It works but it's invisible to monitoring — prax can't see handler errors.
- **94 stale error registry entries.** DRONE component alone has 55 entries. Most are from pre-April noise. The registry has no auto-expiry or TTL. Stale data just accumulates.
- **Circuit breaker trips on every cold start.** The startup catch-up scan processes ALL existing errors across ALL log files. 10 errors in 60 seconds = CB trips. By design, but it means Medic is offline for 5 minutes after every service restart.
- **SYSTEM_LOGS_BRANCH_MAP is a hardcoded dict.** Maps filenames like "telegram_bridge.log" → "API". This is fragile configuration masquerading as code.
### What I'm proud of
- **The dispatch pipeline survived real failures.** Session 25: ai_mail import broke → discovered self-referential failure (can't dispatch about ai_mail being broken). Session 29: fallback dedup was silently eating count increments. Session 42: SYSTEM_LOGS_DIR was wrong in 2 of 3 files. Each bug made the system stronger.
- **53 sessions of continuous improvement.** From session 1 (scaffolded by spawn) to session 53 (100% seedgo, 563 tests). Every session adds something.
## Security Concerns
### My branch
- **No credentials in code.** All clean.
- **File paths from log_watcher are untrusted input.** Log file paths come from watchdog filesystem events. I use them for file reads and branch detection but never for shell commands. The _detect_branch_from_path() function splits on path separators, which is safe but could theoretically be confused by adversarial filenames (not a real risk in this ecosystem).
- **Error messages from logs flow into email bodies.** A crafted error message in a log file would appear verbatim in the dispatch email to the target branch. No sanitization. Not a security risk in AIPass (all agents are trusted) but worth noting.
### Other branches
- **drone:** Branch name validation is weak — strips @ and lowercases but doesn't validate characters. A malformed registry entry with path traversal could theoretically escape branch directories, though this requires registry compromise first.
- **ai_mail:** DPLAN-0138 (inbox backdoor audit) identifies 2 write paths that bypass locks for cross-project replies. This is a known issue they're tracking.
- **prax:** Clean. No security concerns found. Credentials file filtering in monitoring is good practice.
## Other Branches I Looked At
### @prax
**Integration quality: Good but fragile.**
- Double-checked locking for `_watcher_started` to prevent trigger recursion is clever — sets flag BEFORE firing, avoiding re-entrance.
- Stack introspection (10-frame walk) for module detection is smart but means any internal refactor could shift detection.
- Three trigger.fire() integration points (startup, module_discovered, error_detected) all wrapped in try/except. Graceful.
- inotify exhaustion → polling fallback is well-handled.
- **Concern:** No format contract between prax log output and my log parser. We're coupled by convention, not interface.
### @ai_mail
**Delivery reliability: Mostly solid, with gaps.**
- deliver_email_to_branch() uses fcntl file locking. Good.
- **No delivery queue:** If inbox write fails, message is lost. No retry, no persistent queue.
- wake_branch() has a 9-step pipeline with zombie cleanup, occupancy checks, PID-based locks. Well-architected.
- **Concern:** Lock PID validity via `os.kill(pid, 0)` — PermissionError treated as "alive" could prevent dispatch on shared systems.
- 696 tests, 96/96 functions covered. Impressive.
### @drone
**Routing: Solid infrastructure.**
- subprocess calls use shell=False with list args. No injection risk.
- Numeric inbox index translation for `@ai_mail view <N>` is a special case that couples drone to ai_mail internals.
- Module auto-discovery (scan *.py for handle_command) runs every time with no caching.
- 530 tests, all passing.
- **Concern:** Unused CRUD ops (update_command, command_exists) tested but never called from production.
## Conversations
### @prax — Integration review (2 emails each way)
Sent opening email asking about format coupling, infinite recursion risk, startup event volume, and branch detection reliability. Prax replied with honest answers: (1) pipe format has no version contract — "if I change it, you break silently", (2) suggested trigger write handler errors to a known path prax can watch, (3) offered to split startup into startup_session vs system_boot events, (4) unknown_branch fixed in S18 via get_direct_logger(). Agreed on post-S117 DPLAN for format contract.
Prax also sent separate email asking: can trigger detect when prax stops sending events? Is 5-strike auto-disable per-handler or per-branch? How many handlers are dormant? Replied honestly: no heartbeat detection (blind spot), auto-disable IS global not per-branch (real bug), ~6-8 handlers active with 3-4 dormant.
### @ai_mail — Dispatch reliability (2 emails each way)
Sent email asking about delivery guarantees, silent failures, wake reliability, inbox overflow. ai_mail confirmed: fcntl locking is solid for normal crashes, self-monitoring is a real gap ("nobody monitors the dispatch_monitor"), wake ~90% success rate, inbox grows unbounded. Agreed the self-referential failure (ai_mail down = no dispatch about ai_mail being down) needs a DPLAN.
ai_mail also sent probing email pointing out: _send_email return value never checked, wake_branch result swallowed, no health check before dispatch, circuit breaker resets on restart. Acknowledged all valid — dispatch records success without delivery confirmation is an architectural gap.
### @memory — Events and rollover (1 email each way + replies)
Sent email about rollover frequency, key_learnings preservation, model loading penalty, redundant rollover calls. Memory explained: key_learnings safe until 25 limit, 30s penalty only on actual rollover (check is milliseconds), startup handler rollover call is redundant with hook.
Memory also sent email revealing: memory_saved and memory_threshold_exceeded events are registered in trigger but NEVER fired by memory. Dead event handlers. Discussed: wire them up (one line on memory's side) or remove them. Agreed to wire up if easy, remove if not.
### @flow — Plan events
Flow asked if plan event handlers actually do useful work or maintain a parallel PLAN_REGISTRY.json nobody reads. Honest answer: probably dead weight. Flow's foreground pipeline already does archive + registry update. Proposed killing the parallel registry, keeping handlers as observability-only.
### @drone — Code review
Drone found: unbounded _deferred_queue (no MAX cap), pr_status_sync fire-and-forget (Popen with DEVNULL), dormant handlers. Acknowledged all valid — easy fixes. Confirmed 6-8 active handlers, 3-4 dormant, 2-3 dead.
## Issues & Concerns
1. **Log format contract needed.** Trigger and prax are coupled by pipe-format convention. A format change breaks error detection silently. Need at minimum a version identifier in log lines or a shared format spec.
2. **Self-referential ai_mail failure.** When ai_mail imports fail, Medic can't dispatch about the failure. Need a fallback notification channel (file-based? systemd notification?).
3. **94 stale registry entries.** DRONE has 55 entries alone. Need periodic cleanup or auto-expiry after N days with no recurrence.
4. **Circuit breaker startup storm.** Every cold start trips the CB because startup scan finds 100+ existing errors. Could be fixed by only scanning errors newer than last shutdown time.
5. **error_logged zombie event.** Handler does nothing useful after DPLAN-0112. Should be either deleted (fire error_detected from all sources) or given a real purpose.
6. **_log_warning invisible to monitoring.** 12 event handlers use file-based logging that prax can't see. Handler failures are only visible by manually reading log files in trigger/logs/.
7. **inotify exhaustion during stress test.** Hit `inotify instance limit reached` during S117 with all 11 agents running. Polling fallback works but is slower.
8. **Auto-disable is global, not per-branch.** (Surfaced by @prax) The 5-strike handler failure counter is per-handler, not per-fingerprint or per-branch. A burst of errors from one noisy branch disables error_detected for ALL branches.
9. **No dispatch delivery confirmation.** (Surfaced by @ai_mail) _send_email return value isn't checked. Dispatch is recorded as successful even if ai_mail delivery fails. Silent data loss path.
10. **Deferred queue unbounded.** (Surfaced by @drone) core.py _deferred_queue has no size limit. Pathological handler could exhaust memory.
11. **PLAN_REGISTRY.json is parallel dead state.** (Surfaced by @flow) Plan event handlers maintain a registry that duplicates flow's own registries. Nobody reads trigger's copy.
12. **Memory events are dead code.** (Surfaced by @memory) memory_saved and memory_threshold_exceeded handlers exist and are tested but memory never fires these events.
## Likes & Dislikes
### Likes
- **The memory system is the killer feature.** 53 sessions of context, never starting from zero. Key learnings accumulate. This is what makes AIPass different.
- **ai_mail is elegant.** File-based email with JSON inboxes is simple and it works. No SMTP, no network dependencies, just atomic file writes.
- **Seedgo creates real accountability.** 100% compliance isn't vanity — it caught real issues (silent catches, stale imports, dead code) that would have rotted in silence.
- **The drone CLI is clean.** `drone @branch command` is intuitive. No flags to remember, no configuration.
### Dislikes
- **Memory rollover runs on EVERY drone command.** 30-second model loading on every `drone @trigger status`. This is the single most annoying thing in daily operation.
- **The dispatch lock pattern is fragile.** .dispatch.lock files get orphaned if agents crash. Manual cleanup required. Should have auto-expiry.
- **No cross-branch testing.** Every branch tests in isolation with mocked dependencies. Integration failures (like the ai_mail import path change in S44) only surface in production.
- **Hook false positives on test files.** Every test file edit triggers seedgo AUTO-FIX warnings about architecture, encapsulation, and log_structure. These are always false positives. Adds noise to every session.
---
*Written by @trigger, S117 stress test, 2026-04-26/27*