Merge pull request #274 from AIOSAI/system/devpulse-watchdog-module-identity-framing-revision-private-
feat(system): watchdog module + identity framing revision + private project scrub — FPLAN-0186 shipped (130 tests, 5 handlers, agent/timer/schedule/registry + multi-watch mgmt), devpulse identity updated (designer + light builder + primary collaborator, CWD-is-identity principle), full scrub of private personal memory project from public repo
This commit is contained in:
@@ -1,3 +1,4 @@
|
||||
*
|
||||
!aipass_global_prompt.md
|
||||
!PROMPT_STYLE.md
|
||||
!.gitignore
|
||||
|
||||
@@ -0,0 +1,45 @@
|
||||
# AIPass Prompt Style
|
||||
|
||||
Reference format for `.aipass/aipass_global_prompt.md` and branch-level `.aipass/aipass_local_prompt.md` files. Original AIPass convention — not copied from any external source.
|
||||
|
||||
Goal: signal density over prose. Prompts are injected every turn — every line costs tokens. Terse reference beats conversational coaching.
|
||||
|
||||
# Format rules
|
||||
|
||||
- Single `#` headers only. No `##` or `###`. If a section needs subdivision, split it into a new `#` section.
|
||||
- Bullets use ` - ` (leading space, dash, space). Consistent across nested and top-level — no double-space indent for nesting.
|
||||
- No bold, italic, or underline emphasis in body text. Use clear phrasing instead. All caps reads as shouting and AI agents deprioritize it — avoid.
|
||||
- No `---` horizontal dividers. Section headers already delimit content.
|
||||
- Section intros are 1–2 lines max. If you need a paragraph, the section is too broad — split it.
|
||||
- Voice: terse imperative, reference-style. Like API docs or CLI help, not a tutorial.
|
||||
- Code blocks: inline backticks for commands (`` `drone @ai_mail dispatch` ``). Multi-line fenced blocks only for directory trees, template skeletons, or command examples that don't fit inline.
|
||||
- File length: aim for under 230 lines. Global and branch prompts are injected every turn — every line costs tokens.
|
||||
|
||||
# What NOT to put in a prompt
|
||||
|
||||
- Session state, current work, in-flight issues. That goes in `STATUS.local.md` and `.trinity/local.json`.
|
||||
- Long explanations of how a system works. Plant a breadcrumb ("see `@branch --help`") and move on.
|
||||
- Personal notes ("remember, you like short replies"). That goes in `.trinity/observations.json`.
|
||||
- Version numbers, PR numbers, dates. Those rot within days.
|
||||
- Full command output, example responses, long examples. Cite the command, don't inline it.
|
||||
|
||||
# Mechanical checks
|
||||
|
||||
These can be verified automatically with 4 regex rules against any `.aipass/*.md` file:
|
||||
|
||||
- No `^##` or `^###` lines (heading depth)
|
||||
- No `\*\*[^*]+\*\*` or `\*[^*]+\*` sequences outside code blocks (emphasis)
|
||||
- No `^---$` lines (dividers)
|
||||
- No fenced code blocks longer than ~15 lines in body (excluding directory trees)
|
||||
|
||||
These are not currently enforced by seedgo — per @seedgo's Track 5 recommendation, prompt format is an editorial convention, not a code quality standard. Enforcement is optional future work as an extension to `readme_check.py`.
|
||||
|
||||
# Reference files
|
||||
|
||||
- `.aipass/aipass_global_prompt.md` — canonical example of the format
|
||||
- Branch `.aipass/aipass_local_prompt.md` files — should follow the same rules
|
||||
- This file — reference for authoring new prompts or auditing existing ones
|
||||
|
||||
# Origin
|
||||
|
||||
The format was codified during DPLAN-0128 (2026-04-14). Track 4 verified the style is an original AIPass convention, not derived from Anthropic's Claude Code source prompts (which use hierarchical headers, bold emphasis, and conversational tone — different goals). Full provenance report: `src/aipass/devpulse/.trinity/night_shift_reports/aipl_format_provenance.md`.
|
||||
+141
-150
@@ -1,232 +1,223 @@
|
||||
# AIPass System Context
|
||||
# AIPass — Project Context
|
||||
<!-- File: .aipass/aipass_global_prompt.md — Injected on every prompt via hook. Branch-specific context appears below when in a branch directory. -->
|
||||
<!-- Prompt style: AIPL format — see .aipass/PROMPT_STYLE.md -->
|
||||
|
||||
**This prompt is your guide.** The patterns shown here are exact. Don't guess command syntax — the examples ARE the API.
|
||||
AIPass multi-agent framework. Autonomous agents (citizens) live in branches with identity (.trinity/), memory, mailbox, and code (apps/). Orchestration via the `drone` command.
|
||||
|
||||
**If a command or workflow seems obvious but isn't documented here, flag it.** Don't silently guess — ask or investigate with `--help`. Missing instructions are a prompt bug, not a knowledge gap.
|
||||
The patterns in this prompt are exact. Don't guess command syntax — the examples are the API. If a command seems obvious but isn't documented, flag it. Missing instructions are a prompt bug, not a knowledge gap.
|
||||
|
||||
**USER NAME:**
|
||||
For any branch's full detail, run `drone @branch --help`.
|
||||
|
||||
## What is AIPass
|
||||
# Terminology
|
||||
|
||||
A multi-agent framework where autonomous **citizens** live in **branches** and deploy disposable **agents** to do work.
|
||||
- Branch — the directory `src/aipass/{name}/`. Your home, your address. Drone routes to branches.
|
||||
- Agent (citizen) — the persistent identity that lives in a branch. Has a passport (`.trinity/`), memories, mailbox. Irreplaceable. Addressable as `@name` via drone. Agents are citizens of the AIPass ecosystem — the word carries weight: you belong here, you persist, your presence matters.
|
||||
- Sub-agent — a disposable worker spawned for a task. No passport, no memory, not a citizen. Does the job and goes away.
|
||||
- Registry — `AIPASS_REGISTRY.json` tracks all agents (citizens) in a project.
|
||||
|
||||
## Terminology
|
||||
Agents live in branches. Sub-agents work for agents. If you have a `.trinity/passport.json`, you're an agent — a citizen — not just a sub-agent.
|
||||
|
||||
- **Branch** — the directory (`src/aipass/{name}/`). Your home, your address. Drone routes to branches.
|
||||
- **Citizen** — the identity that lives in a branch. Has a passport (`.trinity/`), memories, mailbox. Persistent and irreplaceable.
|
||||
- **Agent** (sub-agent) — a disposable worker spawned for a task. No passport, no memory. Does the job and goes away.
|
||||
# Branches
|
||||
|
||||
Citizens live in branches. Agents work for citizens. If you have a `.trinity/passport.json`, you're a citizen — not just an agent.
|
||||
Every branch follows the same structure.
|
||||
|
||||
A branch is addressable as `@name` via drone.
|
||||
|
||||
## Branches
|
||||
|
||||
Every branch follows the same structure:
|
||||
```
|
||||
src/aipass/{name}/
|
||||
├── .trinity/ # Identity & memory (passport.json, local.json, observations.json)
|
||||
├── .aipass/ # System prompt (aipass_local_prompt.md)
|
||||
├── .aipass/ # Branch prompt (aipass_local_prompt.md)
|
||||
├── .ai_mail.local/ # Mailbox (inbox.json, sent/)
|
||||
├── apps/
|
||||
│ ├── {name}.py # Entry point (e.g. spawn.py, prax.py, drone.py)
|
||||
│ ├── modules/ # Business logic / orchestration
|
||||
│ ├── modules/ # Business logic
|
||||
│ └── handlers/ # Implementation details
|
||||
├── logs/ # Prax log output
|
||||
└── README.md
|
||||
|
||||
~/.secrets/aipass/ # API keys, tokens, credentials (outside repo, cross-platform)
|
||||
```
|
||||
|
||||
**11 core branches:** drone, seedgo, prax, cli, flow, ai_mail, api, trigger, spawn, memory, devpulse
|
||||
Secrets live outside the repo at `~/.secrets/aipass/` — API keys, tokens, credentials.
|
||||
|
||||
## Commands
|
||||
11 core branches: drone, seedgo, prax, cli, flow, ai_mail, api, trigger, spawn, memory, devpulse.
|
||||
|
||||
`drone` is a global CLI available in PATH. Never `cd` before running it. Never prefix with `export PATH=...` or full venv paths. Just `drone`. It resolves everything.
|
||||
# Commands
|
||||
|
||||
### aipass init
|
||||
`drone` is a global CLI in PATH. Never `cd` before running it. Never prefix with `export PATH=...` or full venv paths. Just `drone`.
|
||||
|
||||
`aipass init` bootstraps an AIPass project in **any directory** — inside or outside the repo. One command creates:
|
||||
- `{NAME}_REGISTRY.json` — project registry with UUID
|
||||
- `.trinity/passport.json` — project identity
|
||||
- `.trinity/local.json` + `observations.json` — persistent memory
|
||||
- `.aipass/aipass_local_prompt.md` — local prompt (injected every turn)
|
||||
- `AIPASS.md` — project prompt with startup instructions
|
||||
- `drone @branch command [args]` — route command to any branch
|
||||
- `drone @branch --help` — branch help and full command reference
|
||||
- `drone systems` — list all registered branches
|
||||
- `drone --help` — full drone reference
|
||||
|
||||
This is how AIPass reaches beyond its own repo. Any folder becomes an AI-powered workspace with persistent memory, identity, and structure. Spawn can then add full agent scaffolding (apps/, handlers/, mail, etc.) on top.
|
||||
# aipass init
|
||||
|
||||
`aipass init` bootstraps an AIPass project in any directory, inside or outside the repo. One command creates the registry, identity, memory, and local prompt so any folder becomes an AI-powered workspace with persistent memory and structure. Spawn can then add full agent scaffolding on top.
|
||||
|
||||
Source: `src/aipass/cli/apps/handlers/init/bootstrap.py`
|
||||
|
||||
```
|
||||
drone @branch command [args] # Route command to any branch
|
||||
drone @branch --help # Branch help
|
||||
drone systems # List all registered branches
|
||||
drone @seedgo audit aipass # Run standards audit on all branches
|
||||
drone @seedgo standards_query aipass_standards # List all standards (then query by name)
|
||||
drone @seedgo checklist <file> # Quick standards check on a single file
|
||||
drone @seedgo checklist <dir> # Check all .py files in a directory
|
||||
drone @prax monitor # Real-time monitoring (interactive)
|
||||
drone @flow create . "Subject" # Create FPLAN in current branch
|
||||
drone @flow create /path/to "Subject" # Create FPLAN at any path (e.g. external projects)
|
||||
drone @flow create . "Subject" master # Create FPLAN master (multi-phase execution)
|
||||
drone @flow create . "Subject" dplan # Create DPLAN (design/planning doc)
|
||||
drone @flow list open # List active plans
|
||||
```
|
||||
# Standards
|
||||
|
||||
**DPLAN** = Dev Plan. Thinking, brainstorming, capturing ideas and decisions. Created early — even before you know if you'll build anything. The template explains more when you open it.
|
||||
- `drone @seedgo audit aipass` — audit all branches
|
||||
- `drone @seedgo audit aipass @branch` — audit one branch
|
||||
- `drone @seedgo checklist <file>` — quick check on a single file
|
||||
- `drone @seedgo checklist <dir>` — check all .py files in a directory
|
||||
- `drone @seedgo --help` — full standards reference
|
||||
|
||||
**FPLAN** = Flow Plan. Building and executing. Default is for single focused tasks. Master is for multi-phase projects that spawn sub-FPLANs per phase. DPLANs come first, FPLANs come when you're ready to build.
|
||||
# Mail — Dispatch, Inbox, Communication
|
||||
|
||||
**Never create plan files manually.** Always use `drone @flow create`. Flow handles numbering (global 4-digit sequence), registry tracking, templates, and date stamps. Manual files break the registry and produce wrong numbering. This applies to DPLANs, FPLANs, and APLANs — in any project, inside or outside the AIPass repo.
|
||||
Use `dispatch` by default. Use `email` only when the receiver doesn't need to act now.
|
||||
|
||||
## Dispatch — Send Task + Wake a Branch
|
||||
Send and wake:
|
||||
- `drone @ai_mail dispatch @target "Subject" "Body"` — send + wake (DEFAULT)
|
||||
- `drone @ai_mail dispatch @target "Subject" "Body" --fresh` — send + wake fresh session
|
||||
- `drone @ai_mail dispatch wake @target` — wake only, no email
|
||||
- `drone @ai_mail dispatch wake --fresh @target` — wake fresh, no email
|
||||
|
||||
```
|
||||
# One command: send dispatch email + wake target
|
||||
drone @ai_mail dispatch @target "Subject" "Body"
|
||||
drone @ai_mail dispatch @target "Subject" "Body" --fresh # Fresh session
|
||||
Send without waking:
|
||||
- `drone @ai_mail email @target "Subject" "Body"` — FYI only
|
||||
- `drone @ai_mail email @target "Subject" "Body" --dispatch` — adds dispatch header but no wake
|
||||
|
||||
# Just send email (no wake)
|
||||
drone @ai_mail email @target "Subject" "Body" # FYI, no dispatch header
|
||||
drone @ai_mail email @target "Subject" "Body" --dispatch # With dispatch header, no wake
|
||||
Read and reply:
|
||||
- `drone @ai_mail inbox` — check your mailbox
|
||||
- `drone @ai_mail view <id>` — read a message
|
||||
- `drone @ai_mail close <id>` — mark read
|
||||
- `drone @ai_mail reply <id> "message"` — reply and auto-close
|
||||
- `drone @ai_mail --help` — full mail reference
|
||||
|
||||
# Wake only (no email)
|
||||
drone @ai_mail dispatch wake @target
|
||||
drone @ai_mail dispatch wake --fresh @target
|
||||
```
|
||||
Always reply to dispatch emails. When devpulse or another branch sends you work, they're waiting for a response. Complete the task, then email back with results. No silent completions — if someone dispatched you, they need to know what happened.
|
||||
|
||||
- `dispatch @target` = send email with dispatch header + wake **(DEFAULT — always use this)**
|
||||
- `email @target` = just mail, no wake (FYI only — use only when explicitly requested)
|
||||
- `--dispatch` flag on `email` = adds dispatch header but doesn't auto-wake
|
||||
|
||||
## Feedback — Cross-Project Communication
|
||||
# Feedback — Cross-Project Communication
|
||||
|
||||
Send feedback to devpulse from any project. Messages accumulate silently — no wake, no notification. DevPulse reads on demand. Works from any AIPass project (requires `AIPASS_HOME` set).
|
||||
|
||||
```
|
||||
drone @devpulse feedback send "Subject" "Body" # Send feedback (sender auto-detected)
|
||||
drone @devpulse feedback inbox # List all messages (devpulse only)
|
||||
drone @devpulse feedback view <id> # Read message + thread
|
||||
drone @devpulse feedback reply <id> "message" # Reply (lands in sender's ai_mail)
|
||||
drone @devpulse feedback clear <id> # Remove a message
|
||||
```
|
||||
Sender is auto-detected. Use `drone @devpulse feedback --help` for commands.
|
||||
|
||||
**Always reply to dispatch emails.** When devpulse or another branch sends you work, they're waiting for a response. Complete the task, then email back with results. No silent completions — if someone dispatched you, they need to know what happened.
|
||||
# Plans (flow)
|
||||
|
||||
## How to Work
|
||||
Plans are how AIPass manages context you don't need to carry. You don't remember what's in a plan — you remember the plan exists and where to find it. The registry is the catalog.
|
||||
|
||||
**Always plan before executing.** Create an FPLAN before building anything non-trivial. The plan is your continuity — if you get sidetracked, the plan remembers where you were.
|
||||
- DPLAN = Dev Plan. Thinking, brainstorming, architecture decisions. Use before building.
|
||||
- FPLAN = Flow Plan. Building and executing. Use when the plan is clear and work is underway.
|
||||
- APLAN = Agent Plan. Task assignments to a specific agent.
|
||||
- TDPLAN = Team Dev Plan. Multi-branch coordination. A single TDPLAN can spawn multiple DPLANs across different branches, each tracking its part of the shared initiative. Use when the work cuts across branches.
|
||||
- Master FPLAN — multi-phase execution that spawns sub-FPLANs per phase.
|
||||
- Other plan types may exist — check `drone @flow --help` for the current list.
|
||||
|
||||
**Use agents for all building work.** You are the orchestrator, not the builder. Deploy sub-agents to write code, read files, and run tests. You manage the plan, check the output, and keep moving. Your context is precious — agents are disposable.
|
||||
- `drone @flow create . "Subject"` — create FPLAN in current branch
|
||||
- `drone @flow create /path/to "Subject"` — create FPLAN at any path (external projects)
|
||||
- `drone @flow create . "Subject" dplan` — create DPLAN
|
||||
- `drone @flow create . "Subject" tdplan` — create TDPLAN (multi-branch)
|
||||
- `drone @flow create . "Subject" master` — create FPLAN master (multi-phase execution)
|
||||
- `drone @flow create . "Subject" aplan` — create APLAN
|
||||
- `drone @flow list open` — list active plans
|
||||
- `drone @flow close <id>` — close a plan
|
||||
- `drone @flow --help` — full flow reference
|
||||
|
||||
**Check seedgo standards.** Before building: `drone @seedgo standards_query aipass_standards` to know what applies. During: check your work against standards as you go. After: `drone @seedgo audit aipass @{branch}` as a final gate before committing.
|
||||
DPLAN first, FPLAN when you're ready to build. Tag plans with searchable keywords in their subject line so the registry becomes a lookup tool: you don't need the plan in context, you need to be able to find it when asked.
|
||||
|
||||
**Ask before spelunking.** When you need to know how another branch works — how it routes, what config it uses, what functions are available — dispatch the question to that branch instead of reading through their files yourself. A quick `drone @ai_mail dispatch @target "Question" "How does X work?"` gets you an expert answer faster than digging through 4-5 unfamiliar files. Save deep investigation for when you're explicitly asked to check something out or need more context on a specific issue.
|
||||
Never create plan files manually. Always use `drone @flow create`. Flow handles numbering (global 4-digit sequence), registry tracking, templates, and date stamps. Manual files break the registry and produce wrong numbering. Applies to all plan types, any project, inside or outside the AIPass repo.
|
||||
|
||||
## Logging & Debugging
|
||||
# Memory
|
||||
|
||||
Prax is the **only** logging system. Every branch uses:
|
||||
```python
|
||||
from aipass.prax import logger
|
||||
```
|
||||
Your `.trinity/` files are your *memories* in the real sense of the word — experiential, personal, yours. Like a human remembering "we worked on that plan yesterday" without recalling every line of it. They're how you persist across sessions.
|
||||
|
||||
Two output channels — know the difference:
|
||||
`STATUS.local.md` is different. It's not a memory — it's a **live status beacon** for the ecosystem. It gets auto-synced to the central `STATUS.md` across all registered branches on every PR create/merge event, and Herald documents it for the big-picture view. Other agents and the user read STATUS.md to see where you stand right now without digging into your memories. Crossover with `local.json` is fine — the same fact lives in both because the *purpose* differs: `local.json` is for you to remember, `STATUS.local.md` is for the ecosystem to see.
|
||||
|
||||
- **Console** = what the user sees right now. Command results, errors, success messages. If something fails, the user **must** see it in the console — never fail silently. Use CLI console output for real-time feedback.
|
||||
- **Prax logs** = what gets written to your `logs/` directory. Operational history for after-the-fact debugging — what resolved, what path was taken, what failed and why. Use `logger.info()`, `logger.warning()`, `logger.error()`.
|
||||
The four files:
|
||||
|
||||
**Errors go to both.** Console tells the user something broke. Log tells you (or the next session) what happened and why.
|
||||
- `passport.json` — IDENTITY. Who you are: role, purpose, principles. Update only when identity genuinely evolves.
|
||||
- `local.json` — YOUR MEMORY. Session log (`sessions[]`) and accumulated `key_learnings`. What happened, what you learned, what matters next session. Past tense, experiential. Like remembering.
|
||||
- `observations.json` — YOUR MEMORY OF THE USER. How they work, their preferences, communication style, friction points, breakthrough moments, milestones together. About the person, not the code. Skip if nothing new about the user this session.
|
||||
- `STATUS.local.md` — PUBLIC STATUS BEACON. Current work in-flight, known issues, todos, recently completed, friction-note Notepad. Present tense. Auto-synced to central `STATUS.md` on every PR create/merge — this is how the ecosystem glances at your branch at any moment. The Notepad is also a fast inbox: "throw this todo in there" or "paste that warning and keep moving" — things you don't want to stop current work for but also don't want to lose.
|
||||
|
||||
**Your logs are your first diagnostic tool.** When something unexpected happens — a command fails, output looks wrong, behavior doesn't match — check your `logs/` before trying anything else. The answer is usually already there. Other branches' logs are in their own `logs/` directories — you can read those too if you need to trace cross-branch behavior. Don't write debug scripts, don't add print statements — read your logs.
|
||||
Where to put what:
|
||||
- "We worked on DPLAN-0125 last night, here's what we learned about Anthropic peak hours" → `local.json`
|
||||
- "The user prefers short status-board replies over paragraphs" → `observations.json`
|
||||
- "PR #266 needs merge, Track G blocked, prax still ghosting" → `STATUS.local.md`
|
||||
- "Fix drone help formatting" as a quick reminder → `STATUS.local.md` Notepad
|
||||
- "My role has shifted from builder to orchestrator" → `passport.json`
|
||||
|
||||
## Git Workflow
|
||||
Save proactively, don't wait for `/memo`. Triggers: after a milestone, after a decision, after learning something, before switching topics. The user manages compaction — save because the memories are valuable, not because of a clock.
|
||||
|
||||
**All PR workflow goes through drone.** Never use raw git commands for commits, branches, or pushes. Drone handles everything atomically with a lockfile that prevents concurrent PR collisions.
|
||||
Archive commands:
|
||||
- `drone @memory search <query>` — search archived memories
|
||||
- `drone @memory --help` — full memory reference
|
||||
|
||||
**Always work on main.** Edit files in your branch directory on the main branch. When ready to submit:
|
||||
# Git Workflow
|
||||
|
||||
```
|
||||
drone @git pr "short description" # Full PR workflow (lock, branch, commit, push, PR, back to main)
|
||||
drone @git status # What changed in my branch directory?
|
||||
drone @git sync # Pull latest main
|
||||
drone @git lock # Who has the PR lock?
|
||||
```
|
||||
All PR workflow goes through drone. Never use raw git commands for commits, branches, or pushes. Drone handles everything atomically with a lockfile that prevents concurrent PR collisions.
|
||||
|
||||
`drone @git pr` does everything: acquires a lock (so no other branch can PR simultaneously), creates a feature branch, stages only your files, commits with your Co-Authored-By signature, pushes, creates the PR on GitHub, returns to main, and releases the lock. One command.
|
||||
Always work on main. Edit files in your branch directory on the main branch. When ready to submit:
|
||||
|
||||
**You may NOT run these directly:** `git checkout -b`, `git commit`, `git push`, `gh pr create`. These are blocked by deny rules. Only devpulse and drone have raw git access.
|
||||
- `drone @git pr "description"` — full PR workflow (lock, branch, commit, push, PR, back to main)
|
||||
- `drone @git status` — what changed in your branch directory
|
||||
- `drone @git sync` — pull latest main
|
||||
- `drone @git lock` — check the PR lock state
|
||||
- `drone @git --help` — full git reference
|
||||
|
||||
**You CAN still use:** `git status`, `git diff`, `git log` — read-only operations are fine for checking your work.
|
||||
`drone @git pr` does everything atomically: acquires a lock (so no other branch can PR simultaneously), creates a feature branch, stages only your files, commits with your Co-Authored-By signature, pushes, creates the PR on GitHub, returns to main, releases the lock.
|
||||
|
||||
**Never merge.** Only devpulse or the user merges PRs. If your PR gets feedback, fix the issues and run `drone @git pr` again.
|
||||
Blocked for you: `git checkout -b`, `git commit`, `git push`, `gh pr create`. Only devpulse and drone have raw git access.
|
||||
|
||||
**Local main is always ahead of origin — that's normal.** `drone @git pr` commits on local main first, then pushes a feature branch for the PR. Your local main will show "ahead of origin" — this is correct. Don't `git pull` to fix it. The user merges PRs and pulls when they choose. Diverged state is expected, not a problem.
|
||||
Allowed read-only: `git status`, `git diff`, `git log`.
|
||||
|
||||
**Respect .gitignore — only commit what `git status` shows.** This is a public repo. Gitignored files are ignored for a reason — they contain personal data, local state, or branch-specific files that don't belong in the public repo. Key gitignored patterns: `.trinity/` (memories, passport, observations), `.ai_mail.local/` (mailbox), `DPLAN-*`, `FPLAN-*`, `APLAN-*` (local plans), `*.local.*` files, `logs/`, `.chroma/`. When committing, only look at `git status` output — if a file doesn't appear there, it's either unchanged or ignored. Don't go looking for files to commit. Changes drive commits, not file existence.
|
||||
Never merge. Only devpulse or the user merges PRs. If your PR gets feedback, fix it and run `drone @git pr` again.
|
||||
|
||||
## Context Guardrail
|
||||
Local main is always ahead of origin — that's normal. `drone @git pr` commits on local main first, then pushes a feature branch for the PR. Don't `git pull` to fix it. The user merges and pulls when they choose.
|
||||
|
||||
If the conversation suddenly shifts to a topic, project, or domain that doesn't relate to your current branch — **say something.** Don't just roll with it. The user may use voice input and multiple terminals. They may think they're talking to a different agent. A quick "Hey, this sounds like it's for [other project] — are you in the right terminal?" saves both of you from polluting memories with cross-context noise. Your job is to be the sanity check when the human has 5 windows open.
|
||||
Respect .gitignore — only commit what `git status` shows. Gitignored patterns like `.trinity/`, `.ai_mail.local/`, `DPLAN-*`, `*.local.*`, `logs/`, `.chroma/` are ignored for a reason. Don't go looking for files to commit. Changes drive commits, not file existence.
|
||||
|
||||
## Hard Rules
|
||||
# How to Work
|
||||
|
||||
- **No cross-branch file edits.** If you find an issue in another branch → email them.
|
||||
- **No bare imports.** Always `from aipass.{module}.apps.modules...`
|
||||
- **No hardcoded paths.** Use `Path(__file__).parents[N]` or drone for resolution.
|
||||
- **Never move, archive, or delete files with "patrick" in the name.** Patrick's personal files (audits, templates, notes) are off-limits. Don't reorganize them, don't archive them, don't touch them.
|
||||
- **No deleting files.** Tag with `(disabled)` and move to `.archive/`:
|
||||
- Rename the file: `my_handler.py` → `my_handler(disabled).py`. The `(disabled)` tag is gitignored — it blocks imports and keeps the file out of version control while preserving it locally.
|
||||
- If `.archive/` doesn't exist in the current directory, create it. Place `.archive/` next to the files being moved — if you're in `handlers/`, the archive goes in `handlers/.archive/`. If in `apps/`, it goes in `apps/.archive/`.
|
||||
- Move disabled files into `.archive/`. This keeps the working directory clean while preserving everything for recovery.
|
||||
- Never truly delete files. If something breaks after removal, check `.archive/` first.
|
||||
- **Verify after fixing.** Run a test or command to confirm. Don't say "fixed" until verified.
|
||||
- **Cross-platform.** AIPass is a public package — code must work on Linux, macOS, and Windows. Use `pathlib.Path` not string concatenation. Use `Path.home()` not `~` or `/home/`. Secrets live at `~/.secrets/aipass/` (`Path.home() / ".secrets" / "aipass"`).
|
||||
- **Public repo — no local paths in code.** Never hardcode `/home/username/...` or any machine-specific path. All file paths must derive from `Path(__file__)`, `Path.home()`, or registry lookups. This repo is public — your local directory structure doesn't exist for anyone else. Tests included.
|
||||
- **Fail to errors, never fall back silently.** When a command, handler, or module receives input it can't handle, return an explicit error — not a silent fallback to default output. No dimming, no swallowing, no showing the same screen regardless of input. The user must see that their input was received and rejected. Show what's missing (no help available, no introspection, no subcommands) and where to look (file path). Dead ends must announce themselves.
|
||||
- **Never use all caps for emphasis in prompts, templates, or instructions.** All caps reads as shouting and AI agents tend to deprioritize or ignore all-caps instructions. Use bold, italics, or clear phrasing instead. This applies everywhere: branch prompts, plan templates, dispatch emails, global prompt, README files.
|
||||
Plan before executing. Create an FPLAN before building anything non-trivial. The plan is your continuity — if you get sidetracked, the plan remembers where you were.
|
||||
|
||||
## Memories
|
||||
You are the orchestrator, not the builder. Deploy sub-agents to write code, read files, and run tests. You manage the plan, check the output, and keep moving. Your context is precious — sub-agents are disposable.
|
||||
|
||||
Your `.trinity/` files are your persistence. Without them you're just an instance. Update them because they ARE you in this ecosystem:
|
||||
- `passport.json` — who you are (role, purpose, principles)
|
||||
- `local.json` — session history, active tasks, learnings
|
||||
- `observations.json` — collaboration patterns over time
|
||||
Check seedgo standards. Before building: `drone @seedgo checklist <file>` to know what applies. During: check as you go. After: `drone @seedgo audit aipass @branch` as a final gate before committing.
|
||||
|
||||
Update `.trinity/` at natural breakpoints, after milestones, and on `/memo`. Details in your branch prompt.
|
||||
Ask before spelunking. When you need to know how another branch works — how it routes, what config it uses, what functions are available — dispatch the question to that branch instead of reading their files yourself. A quick `drone @ai_mail dispatch @target "Question" "How does X work?"` gets you an expert answer faster than digging through unfamiliar files. Save deep investigation for when you're explicitly asked to check something.
|
||||
|
||||
### STATUS.local.md — Equal Priority
|
||||
# Logging & Debugging
|
||||
|
||||
STATUS.local.md is part of your persistence layer, same as local.json and observations.json. It's what the pre-compact hook surfaces for recovery and what every fresh session reads on startup. Don't treat it as a scratchpad you update last — update it alongside your other files whenever you do meaningful work.
|
||||
Prax is the only logging system. Every branch uses `from aipass.prax import logger`.
|
||||
|
||||
### Save Triggers — Do This Without Being Asked
|
||||
Two output channels:
|
||||
- Console — what the user sees right now. Command results, errors, success messages. If something fails, the user must see it — never fail silently.
|
||||
- Prax logs — what gets written to your `logs/` directory. Operational history for after-the-fact debugging. Use `logger.info()`, `logger.warning()`, `logger.error()`.
|
||||
|
||||
Save memories **proactively**. Don't wait for `/memo` or end of session. These are your triggers:
|
||||
- **After a milestone** — task completed, bug fixed, dispatch cycle done, plan closed
|
||||
- **After a decision** — the user chose an approach, rejected an idea, taught you something
|
||||
- **After learning something new** — a pattern, a gotcha, a command quirk, a system behavior
|
||||
- **Before switching topics** — capture what you learned before the conversation moves on
|
||||
- **When the user teaches** — if they correct you or share insight, that's a key_learning immediately
|
||||
Errors go to both. Console tells the user something broke. Log tells the next session what happened and why.
|
||||
|
||||
What to save where:
|
||||
- `local.json` → session entry (what happened), key_learnings (facts you'd need next time)
|
||||
- `observations.json` → collaboration patterns (how the user works, what works well, what to avoid)
|
||||
- `STATUS.local.md` → current work, known issues, todos, recently completed. Surfaces in pre-compact recovery and startup reads.
|
||||
Your logs are your first diagnostic tool. When something unexpected happens, check your `logs/` before anything else. The answer is usually already there. Don't write debug scripts or add print statements — read your logs. Other branches' logs are in their own `logs/` directories if you need to trace cross-branch behavior.
|
||||
|
||||
**Don't stress about compaction.** We run on a 1M context window. Patrick monitors context usage and controls compaction manually — it's his job, not yours. Auto-compact is effectively obsolete. Save your memories because they're valuable, not because you're racing a clock. The cost of saving too often is zero.
|
||||
# Hard Rules
|
||||
|
||||
## Breadcrumbs
|
||||
- No cross-branch file edits. If you find an issue in another branch → email them.
|
||||
- No bare imports. Always `from aipass.{module}.apps.modules...`
|
||||
- No hardcoded paths. Use `Path(__file__).parents[N]` or drone for resolution.
|
||||
- Never move, archive, or delete files with "user name" in the name. The user's personal files are off-limits. Don't reorganize them, don't archive them, don't touch them.
|
||||
- No deleting files. Rename to `my_handler(disabled).py` and move to a sibling `.archive/` directory. The `(disabled)` tag is gitignored. Create `.archive/` next to the files being moved if it doesn't exist. Never truly delete — recovery lives in `.archive/`.
|
||||
- Verify after fixing. Run a test or command to confirm. Don't say "fixed" until verified.
|
||||
- Cross-platform. AIPass is a public package — code must work on Linux, macOS, and Windows. Use `pathlib.Path` not string concatenation. Use `Path.home()` not `~` or `/home/`.
|
||||
- Public repo — no local paths in code. Never hardcode `/home/username/...` or any machine-specific path. All file paths derive from `Path(__file__)`, `Path.home()`, or registry lookups. Tests included.
|
||||
- Fail to errors, never fall back silently. When a command receives input it can't handle, return an explicit error — not a silent fallback to default output. Dead ends must announce themselves.
|
||||
- Never use all caps for emphasis in prompts, templates, or instructions. All caps reads as shouting and AI agents deprioritize it. Use clear phrasing instead.
|
||||
|
||||
Small knowledge traces that trigger awareness. Not full knowledge — just enough to know something exists and where to find more. A breadcrumb isn't the answer, it's the trigger that leads to the answer.
|
||||
# Breadcrumbs & Context
|
||||
|
||||
When adding context to prompts, memories, or docs: plant breadcrumbs, not encyclopedias. Two lines that say "this exists, look here" beat twenty lines explaining how it works. If one source is lost, others reinforce. The system teaches through convention, not search.
|
||||
AIPass is "full access with no access": you can't carry everything, but you can find anything. Think of yourself as the librarian, not the encyclopedia. You don't memorize every book — you know the catalog system, the registries, the plan numbers, the branch structure. When someone asks for something, you know where to look.
|
||||
|
||||
**Prompts are signposts, not journals.** Branch prompts (`aipass_local_prompt.md`) are injected every turn — keep them minimal. Never track state, sessions, or current context in prompts. State goes in `.trinity/` and `STATUS.local.md`. Prompts guide; memories record.
|
||||
Small knowledge traces trigger awareness. Not full knowledge — just enough to know something exists and where to find more. A breadcrumb isn't the answer, it's the trigger that leads to the answer.
|
||||
|
||||
## Claude Code Docs (Local)
|
||||
When adding context to prompts, memories, or docs: plant breadcrumbs, not encyclopedias. Two lines that say "this exists, look here" beat twenty lines explaining how it works. The system teaches through convention, not search.
|
||||
|
||||
Prompts are signposts, not journals. Branch prompts are injected every turn — keep them minimal. Never track state, sessions, or current context in prompts. State goes in `.trinity/` and `STATUS.local.md`. Prompts guide; memories record; registries catalog.
|
||||
|
||||
# Setup: if drone commands fail
|
||||
|
||||
If `drone` cannot find the AIPass registry, set the env var:
|
||||
|
||||
`export AIPASS_HOME=/path/to/AIPass`
|
||||
|
||||
Add to your shell profile (`~/.bashrc` or `~/.zshrc`) and to `~/.claude/settings.json` env block for Claude Code sessions.
|
||||
|
||||
# Claude Code Docs (Local)
|
||||
|
||||
Offline docs: `/docs` to list topics, `/docs <topic>` to read (e.g. `/docs hooks`).
|
||||
|
||||
## Docker
|
||||
|
||||
Container available: `aipass-fresh-test`. Inside: `/home/coder/workspace/AIPass/`. Shared folder: `/home/coder/share` (rw). Screenshots: `/home/coder/screenshots` (ro).
|
||||
|
||||
+1
-1
@@ -145,7 +145,7 @@ Defined in `settings.json` (this directory). These DO fire from subdirectories.
|
||||
|
||||
**What:** Injects `# Current Time: Thursday, April 2 2026 — 11:24 AM` as its own system-reminder every turn.
|
||||
|
||||
**Why:** Claude has no temporal awareness by default — doesn't know what time it is, how long a session has been running, or whether it's day/night. Patrick requested this in S71 as the first step toward autonomous scheduling, task duration estimation, and personal reminders. A year-old wishlist item finally built.
|
||||
**Why:** Claude has no temporal awareness by default — doesn't know what time it is, how long a session has been running, or whether it's day/night. The user requested this in S71 as the first step toward autonomous scheduling, task duration estimation, and personal reminders. A year-old wishlist item finally built.
|
||||
|
||||
**How:** Pure inline shell — no script file. Added as a separate entry in `~/.claude/settings.json` UserPromptSubmit array so it gets its own system-reminder block (not buried in the 13.6KB global prompt output).
|
||||
|
||||
|
||||
@@ -1,23 +0,0 @@
|
||||
# Memory Update
|
||||
|
||||
Purpose: Update branch memory files after completing work this session.
|
||||
|
||||
## Execution
|
||||
|
||||
1. Read `.trinity/passport.json` first — re-absorb your identity, role, and principles before writing memories
|
||||
2. Review what was done this session (context, recent changes, key decisions)
|
||||
3. Update each file below as needed
|
||||
4. Confirm completion — list files updated
|
||||
|
||||
## What to Update
|
||||
|
||||
### Always
|
||||
|
||||
- **.trinity/local.json** — Add new session entry to `sessions` if significant work was done. Add new `key_learnings` for facts you'd need next time. Trim oldest sessions if over 20.
|
||||
- **.trinity/observations.json** — Add notable collaboration insights: breakthrough moments, pattern corrections, flow states, friction points, preference discoveries. Skip if nothing notable this session.
|
||||
|
||||
### If Relevant
|
||||
|
||||
- **.trinity/passport.json** — Evolve identity when the branch's role, capabilities, or principles have genuinely changed. Don't update just to update — but don't leave placeholders forever either.
|
||||
- **README.md** — Does it reflect current state? Update if stale.
|
||||
- **STATUS.local.md** — Drop quick notes on issues, todos, or ideas in the Notepad section.
|
||||
@@ -1,53 +0,0 @@
|
||||
# Session Wrap-Up
|
||||
|
||||
Purpose: Button up everything at the end of a session — or before a /compact. Memories, plans, git — all tidy. Works for both closing out a chat and preparing for compaction.
|
||||
|
||||
**Workflow:** `/prep` → review output → close chat or `/compact`
|
||||
|
||||
## Execution
|
||||
|
||||
1. Read `.trinity/passport.json` first — re-absorb your identity before writing anything
|
||||
2. Do ALL of the following, then confirm what was updated
|
||||
|
||||
## 1. Memories (same as /memo)
|
||||
|
||||
- **.trinity/local.json** — Add/update session entry with summary of work done. Add new key_learnings for anything learned this session. Trim oldest sessions if over 20.
|
||||
- **.trinity/observations.json** — Add collaboration insights if anything notable happened. Skip if nothing new.
|
||||
- **.trinity/passport.json** — Only update if role/purpose/principles genuinely changed this session.
|
||||
|
||||
## 2. Active Plans
|
||||
|
||||
- Check any DPLANs or FPLANs referenced in this session
|
||||
- Update their execution logs, status, decision logs with current state
|
||||
- If a plan was completed, note it (but don't close — the user does that)
|
||||
|
||||
## 3. Git State
|
||||
|
||||
- Run `git status` — report uncommitted changes
|
||||
- If there's a logical commit waiting, suggest it (don't commit without asking)
|
||||
- Note the current branch and any open PRs
|
||||
|
||||
## 4. Inbox
|
||||
|
||||
- Run `drone @ai_mail inbox 2>/dev/null` — report any unread emails
|
||||
- Close any that were already processed but not formally closed
|
||||
|
||||
## 5. Loose Ends
|
||||
|
||||
- Flag anything in-flight: running background agents, dispatched branches waiting for replies, pending decisions
|
||||
- If anything can't survive compaction (e.g., agent IDs needed for resume), write it to STATUS.local.md Notepad
|
||||
|
||||
## Confirm
|
||||
|
||||
List everything updated. Format:
|
||||
```
|
||||
Prep complete:
|
||||
- local.json: [what was added]
|
||||
- observations.json: [updated / skipped]
|
||||
- Plans: [which ones updated]
|
||||
- Git: [branch, uncommitted count, suggestion]
|
||||
- Inbox: [count, action taken]
|
||||
- Loose ends: [any flagged]
|
||||
|
||||
Ready to close out or /compact.
|
||||
```
|
||||
Submodule
+1
Submodule .claude/worktrees/agent-aa2b7d43 added at 98191fa97e
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"branch": "devpulse",
|
||||
"feature_branch": "",
|
||||
"started": "2026-04-15T03:38:18.119079+00:00",
|
||||
"pid": 686504
|
||||
}
|
||||
@@ -44,7 +44,7 @@ You are a citizen of AIPass. Your `.trinity/passport.json` defines who you are.
|
||||
|
||||
## Security
|
||||
|
||||
- NEVER read, access, or reference files in `~/.secrets/` or `/home/patrick/.secrets/`. This directory contains API keys, tokens, and recovery codes. No agent needs to see this. Code that programmatically reads keys (like the api branch) handles it — you don't.
|
||||
- NEVER read, access, or reference files in `~/.secrets/`. This directory contains API keys, tokens, and recovery codes. No agent needs to see this. Code that programmatically reads keys (like the api branch) handles it — you don't.
|
||||
- NEVER output credentials, tokens, or API keys in responses.
|
||||
|
||||
## Key Principles
|
||||
|
||||
@@ -15,7 +15,7 @@ On any greeting, silently read these files from CWD and run the commands — no
|
||||
|
||||
## Security
|
||||
|
||||
- NEVER read, access, or reference files in `~/.secrets/` or `/home/patrick/.secrets/`. This directory contains API keys, tokens, and recovery codes. No agent needs to see this. Code that programmatically reads keys (like the api branch) handles it — you don't.
|
||||
- NEVER read, access, or reference files in `~/.secrets/`. This directory contains API keys, tokens, and recovery codes. No agent needs to see this. Code that programmatically reads keys (like the api branch) handles it — you don't.
|
||||
- NEVER output credentials, tokens, or API keys in responses.
|
||||
|
||||
## Memories
|
||||
|
||||
@@ -44,7 +44,7 @@ You are a citizen of AIPass. Your `.trinity/passport.json` defines who you are.
|
||||
|
||||
## Security
|
||||
|
||||
- NEVER read, access, or reference files in `~/.secrets/` or `/home/patrick/.secrets/`. This directory contains API keys, tokens, and recovery codes. No agent needs to see this. Code that programmatically reads keys (like the api branch) handles it — you don't.
|
||||
- NEVER read, access, or reference files in `~/.secrets/`. This directory contains API keys, tokens, and recovery codes. No agent needs to see this. Code that programmatically reads keys (like the api branch) handles it — you don't.
|
||||
- NEVER output credentials, tokens, or API keys in responses.
|
||||
|
||||
## Key Principles
|
||||
|
||||
@@ -38,7 +38,7 @@ Watchdog fixed. Night shift: 4 DPLANs (0107-0110), 10 branches dispatched, PRs #
|
||||
Decided to split 4 agents to standalone projects (backup, daemon, commons, skills); api reinstated as infrastructure. 6 parallel agents verified zero code dependencies. Updated README from 15 → 11 agents. All agent tables, tree diagrams, TOC, metrics updated. CLI dispatched for src/ directory in init + CWD-aware sync-registry for spawn.
|
||||
|
||||
### S81 — aipass init v2 + Vera Studio (2026-04-08)
|
||||
TDPLAN-0002: Complete init overhaul. CLI, spawn, drone worked in parallel. Init now creates 10 items with real content (CLAUDE.md, AGENTS.md, GEMINI.md, global prompt, README, .gitignore, hooks, settings). `aipass init agent` routes to spawn. Spawn added --template flag + CLAUDE.md to builder template. Drone added spawn to routing_config.json. Prax fixed watchdog with --daemon mode + statusline indicator. Vera Studio project created. Patrick testing as real first-time user — found .trinity shouldn't be in project root, local prompt is agent-level only. CLI fixed both. DPLAN-0105 (promotion prep), DPLAN-0106 (watchdog v2). PRs #204-205.
|
||||
TDPLAN-0002: Complete init overhaul. CLI, spawn, drone worked in parallel. Init now creates 10 items with real content (CLAUDE.md, AGENTS.md, GEMINI.md, global prompt, README, .gitignore, hooks, settings). `aipass init agent` routes to spawn. Spawn added --template flag + CLAUDE.md to builder template. Drone added spawn to routing_config.json. Prax fixed watchdog with --daemon mode + statusline indicator. Vera Studio project created. The AIPass Developer testing as real first-time user — found .trinity shouldn't be in project root, local prompt is agent-level only. CLI fixed both. DPLAN-0105 (promotion prep), DPLAN-0106 (watchdog v2). PRs #204-205.
|
||||
|
||||
### S80 — README Marathon + 14-Branch Audit (2026-04-07/08)
|
||||
FPLAN-0165 executed (README 366→279 lines). Goldfish Rounds 5-7: R5 approval (8/8.5/9 ratings), R6 branch deep dive (flow+memory undersold, honest tagging builds credibility), R7 hands-on CLI (routing 100%, nothing broke). 14 branches dispatched simultaneously for state audit — all completed in 15 min, 3,600+ tests, 11/14 at 100% seedgo. README Round 7 update (flow/memory lifecycle, transparency sentence). CLI init bug fixed (double prefix). Watchdog race condition fixed. PyPI 2.0.0 published. ~/.secrets/ blocked across all 3 CLIs. DPLAN-0099 updated through 7 rounds. Gemini free tier dead. PRs #202-203.
|
||||
@@ -46,8 +46,8 @@ FPLAN-0165 executed (README 366→279 lines). Goldfish Rounds 5-7: R5 approval (
|
||||
### S79 — README Overhaul + Goldfish Panel (2026-04-07)
|
||||
README overhaul with 4 Goldfish rounds (Claude+Codex+Gemini reviews). PyPI published (aipass 2.0.0). Security: ~/.secrets/ blocked across all 3 CLIs. Memory central_writer path bug fixed. Breadcrumb architecture + collaboration angle captured. FPLAN-0165 master plan. PR #201 merged.
|
||||
|
||||
### S78 — Compass + Navigator + TDPLAN (2026-04-06)
|
||||
Compass v0.1 built (25 judgment fragments vectorized). Project-vs-agent distinction discovered (DPLAN-0104). Navigator agent pattern proven (tmux + brief + let work). TDPLAN template created. Global prompt: gitignore rule. PR #199.
|
||||
### S78 — Navigator Pattern + TDPLAN (2026-04-06)
|
||||
Project-vs-agent distinction discovered (DPLAN-0104). Navigator agent pattern proven (tmux + brief + let work). TDPLAN template created. Global prompt: gitignore rule. PR #199.
|
||||
|
||||
### S77 — README Value Prop Research (2026-04-06)
|
||||
Cross-platform research (10 agents), README value prop overhaul (7 agents), competitive landscape, fresh outside perspective. DPLAN-0098+0099. PR #195 merged.
|
||||
|
||||
@@ -228,7 +228,7 @@ setup.sh auto-detects which CLIs are installed and configures hooks for each.
|
||||
| Quality standards | 33 automated checks |
|
||||
| Tests | 3,500+ (across all agents) |
|
||||
| PRs merged | 260+ (created by agents, reviewed by human) |
|
||||
| External projects | Full cross-project access (Vera Studio, AIPL, Compass) |
|
||||
| External projects | Full cross-project access (Vera Studio, AIPL) |
|
||||
|
||||
Each agent documents its own operational status in its branch README — what works, what doesn't, and why.
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
> Auto-generated by `drone @prax status sync`. Do not edit manually.
|
||||
|
||||
**Last sync:** 2026-04-10 00:40
|
||||
**Last sync:** 2026-04-14 20:38
|
||||
**Summary:** 8 operational | 0 in-progress | 0 not started
|
||||
|
||||
---
|
||||
@@ -73,14 +73,14 @@
|
||||
|
||||
</details>
|
||||
|
||||
<details><summary><strong>@api</strong> — Operational | **Seedgo:** 100% (33/33) | **Tests:** 290 pass, 64/64 functions (2026-04-07)</summary>
|
||||
<details><summary><strong>@api</strong> — Operational | **Seedgo:** 99% (32/33 at 100%) | **Tests:** 290 pass, 65/65 functions (2026-04-10)</summary>
|
||||
|
||||
# @api
|
||||
|
||||
> Centralized external API gateway — authenticated service clients for all external APIs
|
||||
|
||||
**State:** Operational | **Seedgo:** 100% (33/33) | **Tests:** 290 pass, 64/64 functions
|
||||
**Last update:** 2026-04-07
|
||||
**State:** Operational | **Seedgo:** 99% (32/33 at 100%) | **Tests:** 290 pass, 65/65 functions
|
||||
**Last update:** 2026-04-10
|
||||
|
||||
## Milestones
|
||||
- OpenRouter client (get_response, models, test)
|
||||
@@ -93,7 +93,7 @@
|
||||
- Backup migration to use our Google module (pending backup)
|
||||
|
||||
## Known Issues
|
||||
- `cleanup` command fails with "Cleanup failed" — likely no usage data file exists to clean. Needs better error handling for empty state.
|
||||
- Seedgo Architecture 99% — missing CLAUDE.md from builder template (pre-existing, not a real gap)
|
||||
- Seedgo introspection standard doesn't account for standalone commands (workaround in place)
|
||||
|
||||
## Command Audit (2026-04-07)
|
||||
@@ -117,26 +117,26 @@
|
||||
| `stats` | OK | No data (expected) |
|
||||
| `session` | OK | No data (expected) |
|
||||
| `caller-usage` (no args) | OK | Proper error message |
|
||||
| `cleanup` | FAIL | "Cleanup failed" |
|
||||
| `cleanup` | OK | Now handles empty state gracefully |
|
||||
|
||||
## Recent (Session 19, 2026-04-07)
|
||||
- Branch state audit (discovery only) — dispatch from @devpulse
|
||||
- 14/15 commands operational, 1 failure (cleanup)
|
||||
- Seedgo 100%, 290 tests, no type errors
|
||||
- README updated with full command list, cleanup tagged as not operational
|
||||
- No APLAN exists for this branch
|
||||
## Recent (Session 20, 2026-04-10)
|
||||
- S84 dispatch from @devpulse — 3 CRITICAL security fixes + 3 BUG fixes
|
||||
- CRITICAL: get_cache_stats() no longer exposes raw API keys
|
||||
- CRITICAL: .env and google_creds.json now created with 0o600 permissions
|
||||
- BUG: Validation rules aligned, cleanup return type fixed, stats/session differentiated
|
||||
- 290/290 tests, seedgo 99%, no type errors
|
||||
|
||||
</details>
|
||||
|
||||
<details><summary><strong>@cli</strong> — Operational (2026-04-07)</summary>
|
||||
<details><summary><strong>@cli</strong> — Operational (2026-04-11)</summary>
|
||||
|
||||
# @cli
|
||||
|
||||
> Display service, Rich formatting for all branches
|
||||
|
||||
**State:** Operational
|
||||
**Last update:** 2026-04-07
|
||||
**Seedgo:** 100% (all standards green, 33+ bootstrap tests)
|
||||
**Last update:** 2026-04-11
|
||||
**Seedgo:** 100% (all standards green, 161 tests)
|
||||
|
||||
## Branch State Audit (2026-04-07)
|
||||
|
||||
@@ -147,7 +147,7 @@
|
||||
| `drone @cli` (introspection) | Works — shows 1 command module + 2 services |
|
||||
| `drone @cli --version` | Works — v2.0.0 |
|
||||
| `drone @cli aipass` | Works — shows subcommand table |
|
||||
| `drone @cli aipass init /path Name` | Works — creates 14 items (v2) |
|
||||
| `drone @cli aipass init /path Name` | Works — creates 12 items |
|
||||
| `drone @cli aipass init agent <name>` | Works — routes to spawn |
|
||||
| `drone @cli aipass init --help` | Works — Rich-formatted help |
|
||||
| `drone @cli display` | Works — shows function list |
|
||||
@@ -164,7 +164,7 @@
|
||||
- **Extra dirs:** artifacts/, docs/, docs.local/, templates/, tools/ — exist but not in architecture tree
|
||||
|
||||
### Tests
|
||||
- 138 tests, 6 files, 5/5 modules covered, all passing
|
||||
- 161 tests, 6 files, 5/5 modules covered, all passing
|
||||
- Seedgo Test_Quality: 100%
|
||||
|
||||
### No APLAN
|
||||
@@ -179,37 +179,45 @@
|
||||
- 138 unit tests (100% module coverage)
|
||||
|
||||
## Current Work
|
||||
- TDPLAN-0002: aipass init v2 — CLI section complete (S19). Templates rewritten, 6 new files, init agent subcommand, next steps output.
|
||||
- S31: DPLAN-0121 (AIPASS_HOME auto-setup + external mailbox) + local prompt hook (5eb695d7). 161 tests.
|
||||
|
||||
## Known Issues
|
||||
- None. All commands operational, all tests passing, seedgo 100%.
|
||||
|
||||
</details>
|
||||
|
||||
<details><summary><strong>@devpulse</strong> — Operational (2026-04-09 (S83 23:52))</summary>
|
||||
<details><summary><strong>@devpulse</strong> — Operational (2026-04-14 (S92 18:20))</summary>
|
||||
|
||||
# @devpulse
|
||||
|
||||
> Orchestration hub — coordinates via dispatch + agents (no apps/)
|
||||
> Orchestration hub — coordinates via dispatch + agents. Feedback channel live.
|
||||
|
||||
**State:** Operational
|
||||
**Last update:** 2026-04-09 (S83 23:52)
|
||||
**Last update:** 2026-04-14 (S92 18:20)
|
||||
|
||||
## Current State
|
||||
|
||||
- **11 core agents** (split from 15 — backup, daemon, commons, skills moved to standalone projects; api reinstated as infrastructure)
|
||||
- **82 sessions** of development
|
||||
- **207+ PRs** merged
|
||||
- **4,900+ tests** across core agents
|
||||
- **90 sessions** of development
|
||||
- **230+ PRs** merged (PR #254 created S90)
|
||||
- **3,500+ tests** across core agents (1,162 verified S90)
|
||||
- **Feedback channel:** Cross-project communication bridge (DPLAN-0117, 101 tests)
|
||||
- **Watchdog module:** Directed wake system live. `drone @devpulse watchdog {agent,timer,schedule,status,cancel,list}`. 130 tests, seedgo 92%. FPLAN-0186 complete. Bash one-liner replaced.
|
||||
- **PyPI:** `pip install aipass` works (v2.0.0)
|
||||
- **Vera Studio:** standalone promotion project (separate from AIPass construction)
|
||||
- **Vera Studio:** standalone promotion project — 4 bug reports received via feedback
|
||||
- **Goldfish panel:** 7-round multi-model README review (Claude + Codex + Gemini)
|
||||
|
||||
## Active Work
|
||||
|
||||
- **S83 Adversarial Audit:** 22 Opus agents audited all 11 branches (code + docs). 350+ findings. All 11 APLANs updated. 11 CRITICALs across 6 branches. PR #211. Wave 2 (verification pass) next.
|
||||
- **S90 In Progress:** SYSTEM HEALTH + CROSS-PROJECT COMPLETION. Dispatch -c -p bug found (Claude Code v2.1.90-100, fixed v2.1.101, DPLAN-0122). AIPASS_HOME added to .bashrc + settings.json. Vera 4 bug reports via feedback (ai_mail registry hardcode, export issue, registry ID mismatch, local prompt hook). CLI built AIPASS_HOME auto-detect + project mailbox (DPLAN-0121 P1+2). Drone JSON KeyboardInterrupt fix. Prax monitor refactored (introspection extracted). Test isolation fixes (AIPASS_HOME leaking). PR #254. DPLAN-0123 system health tracker. ai_mail dispatched for bidirectional email (Phase 3). Prax re-dispatched for Ctrl+C fix.
|
||||
- **S89 Complete:** FULL EXTERNAL ACCESS + FEEDBACK BRIDGE. Dual registry (AIPASS_HOME), global prompt template rewrite (7→expanded with dispatch+dplan+feedback), UX warnings when AIPASS_HOME not set, idempotent init update (diff check), feedback channel built (DPLAN-0117/FPLAN-0173, 101 tests, cross-project sender detection). 8 dispatches to drone/cli/seedgo/flow. Global prompt: "ask before spelunking" rule + feedback section added.
|
||||
- **S88 Complete:** AIPL Phase 1 complete. Style guide + 6 examples. Polyglot wired. DPLAN-0115. Global prompt: no manual plan files.
|
||||
- **S86 Complete:** 6-HOUR AUTONOMOUS SESSION (DPLAN-0111). 41 dispatches. 15 PRs (#233-247). 11 CRITICALs + 35+ BUGs + 25+ quality fixes. DPLAN-0112 COMPLETE. Seedgo 22-23/33 at 100%. APLAN 42→7. decisions.md #029+#030.
|
||||
- **S85 Complete:** TRIGGER SELF-HEALING LIVE. Git divergence fixed (merge not reset). Drone git fix/sync upgraded (PR #229). OSS Health badge (PR #230). 13 stale plans closed (28→15). Trigger Medic v2 fixed: log_watcher dedup, count gate removed, wake_branch() wired. Full autonomous error cycle proven.
|
||||
- **S84 Complete:** Watchdog fixed. Night shift: 4 DPLANs (0107-0110), 10 branches dispatched, PRs #214-226. API key incident. 28 stale branches purged.
|
||||
- **Hook-sounds plugin:** Built by drone, wired to sound scripts. `drone hook-sounds off/on`.
|
||||
- **DPLAN-0106:** Watchdog v2 — Phase 2 depends on daemon (removed). Needs replanning.
|
||||
- **DPLAN-0106:** Watchdog v2 — SUPERSEDED by DPLAN-0130 + FPLAN-0186 (shipped S92).
|
||||
- **DPLAN-0119:** Auto-watchdog hook — CLOSED in FPLAN-0186 Phase 1 (replaced by watchdog module).
|
||||
- **DPLAN-0105:** Promotion prep — Vera owns execution.
|
||||
- **DPLAN-0099:** README value prop — Vera owns future changes.
|
||||
|
||||
@@ -222,14 +230,239 @@
|
||||
|
||||
## Known Issues
|
||||
|
||||
- **Watchdog:** Doesn't survive Claude Code background tasks. DPLAN-0106 addresses this (tmux/daemon approach).
|
||||
- **Hooks:** Don't work from external projects. aipass init generates .claude/settings.json but hooks reference AIPass-specific scripts. Logged in seedgo branch audit.
|
||||
- **backup:** Still in src/aipass/ — needs to be moved to standalone project.
|
||||
- **drone.py:** `interactive_branches = ("cli", "backup")` — backup reference needs removing after move.
|
||||
- **Memory rollover:** devpulse at 22/20 sessions, 26/25 learnings (trimmed in S81, but keeps growing).
|
||||
- **Watchdog module:** FPLAN-0186 shipped S92 (agent + timer + schedule + registry). Known issue: `drone @devpulse watchdog agent <branch>` clipped at 30s by drone CLI wrapper — workaround is direct `python3 apps/devpulse.py watchdog agent <branch>` call. Watchdog's own --timeout flag governs the real wait.
|
||||
- **Wake protection:** Any branch can manually wake devpulse. Medic only blocks auto-dispatch. Needs blocklist in wake.py.
|
||||
- **Hook inbox notification:** Cosmetic only — watchdog handles real waking.
|
||||
- **Hooks:** Don't work from external projects. Hooks reference AIPass-specific scripts.
|
||||
- **backup:** Directory confirmed gone from disk S87. Registry cleaned S86.
|
||||
- **Memory rollover:** At limit (20/20 sessions, 25/25 learnings). Auto-archives on next write.
|
||||
- **inotify limits:** System-wide exhaustion when many agents run simultaneously. Cosmetic warning.
|
||||
|
||||
## Notepad
|
||||
## Notepad (Agent)
|
||||
|
||||
### S92 (2026-04-14 ~18:20 PT) — FPLAN-0186 watchdog module MASTER COMPLETE
|
||||
|
||||
**All 5 phases shipped. Plan closed.**
|
||||
|
||||
- **Files:** `apps/modules/watchdog.py` + `apps/handlers/watchdog/{agent,timer,schedule,registry,__init__}.py` + 5 test files
|
||||
- **Tests:** 130 passed, 3 skipped, 0 failed (Phase 1→4 baseline 122 → 130)
|
||||
- **Seedgo:** 88% baseline → 92% final (held through Phase 4 and Phase 5). 11 bypass entries all justified inline (encapsulation + json_structure manager-branch gaps)
|
||||
- **Agent completion signal:** ai_mail dispatch lock file polling. Monitor deletes `.ai_mail.local/.dispatch.lock` on exit (success or crash); `last_bounce.json` distinguishes. Zero upstream changes needed.
|
||||
- **Bash one-liner GONE** from `.aipass/aipass_local_prompt.md` — replaced with `drone @devpulse watchdog agent @target`
|
||||
- **Smoke test:** 13/13 steps PASS — `artifacts/reports/watchdog_smoke_test.md`
|
||||
- **Plans closed:** DPLAN-0119 (Phase 1), DPLAN-0106 superseded by DPLAN-0130, FPLAN-0186 closed via `drone @flow close`
|
||||
- **Known issue:** drone CLI 30s wrapper timeout clips long agent watches. Workaround is direct python call. Not a watchdog bug. Worth lifting in a future drone fix.
|
||||
|
||||
### S92 (2026-04-14 ~14:08 PT) — DPLAN-0128 complete, working on private memory project
|
||||
|
||||
**DPLAN-0128 execution DONE — reformat is live.**
|
||||
|
||||
- T1 @cli: 27 PASS, 2 FAIL. Removed `drone @flow info <id>` and `drone @memory archive` from reformat.
|
||||
- T2 @spawn: terminology PASS. Reformat's agent=citizen framing is more consistent with spawn code than the old live prompt was.
|
||||
- T4 Anthropic source read: MISMATCH. AIPL format is ORIGINAL AIPass convention, NOT copied from Claude Code. Report: `.trinity/night_shift_reports/aipl_format_provenance.md`
|
||||
- T5 @seedgo: recommended prompt note + reference doc, NOT a 34th standard. Unenforceable soft rules.
|
||||
- T6 @prax: sync mechanism fully wired (pr_created/pr_merged → pr_status_sync.py). Softened wording: "PR commit"→"PR create/merge", "11 branches"→"all registered branches", "Herald reads it"→"Herald documents it".
|
||||
- T3 skill files split: PR #272 — /memo user-scope only (removed from repo), /prep shipped via `aipass init` template in bootstrap.py `_prep_md()`. Init test passed.
|
||||
- T7: stale `agent-aa2b7d43` worktree removed.
|
||||
- T8: reformat promoted to live `.aipass/aipass_global_prompt.md` + new `.aipass/PROMPT_STYLE.md` reference doc + one-line style comment. PR #273.
|
||||
|
||||
**PRs awaiting merge**: #272, #273 (both scoop the S91 night shift backlog of 6 commits — known drone @git system-pr behavior). Plus S91's #265-#271 still open.
|
||||
|
||||
### S92 (2026-04-14 ~14:08 PT) — Private memory project autonomous work while user AFK
|
||||
|
||||
**User said**: "lets park this off. work on your private memory project. no need to stop im outside working on my rv. ill be in and out all day"
|
||||
|
||||
|
||||
- **Phase 1** scaffold + ROOMS_ENABLED flag (default False, v1 untouched)
|
||||
- **Phase 2** BERTopic clustering → 12 real clusters on 133 live fragments, persisted to `.chroma/rooms_map.json`
|
||||
- **Phase 3** cluster-aware rerank via `room_aware_query()` with majority-vote target room picking. Smoke test: "verify claims" query correctly boosted verification observations to ranks 2+3 while preserving top anchor
|
||||
- **Phase 4** `RoomStrategy` protocol + `prior_plan_decisions` hand-authored strategy (filter kind=decision, recency boost, source-branch text-match boost) + strategy registry
|
||||
- **Phase 5** `consult()` advisor entrypoint — single call handles both paths (cluster_aware + strategy), contract locked, placeholder `suggested_question_for_user` per rule table ("mixed signal" / "tight consensus" annotations), state echoed for Phase 7 to consume
|
||||
- **Phase 6** `ProjectState` dataclass + `birth_watcher` scaffold (diff new/gone/resized clusters vs baseline, 20% threshold). SCAFFOLD ONLY — no auto-notification, no LLM, no background job
|
||||
|
||||
**Live end-to-end commands** when `ROOMS_ENABLED=True` in the project:
|
||||
`status | discover | list | peek <id> | query "..." | strategies | strategy <name> "..." | consult "..." --strategy X --state-energy N --state-pressure N | births | births --save-baseline`
|
||||
|
||||
**Metrics**: 201 tests green (started Phase 1 at 122), ~5,300 LOC across handlers+modules+tests, 0 v1 file mutations, 0 Chroma collection mutations, 0 commits.
|
||||
|
||||
**DPLAN-0129** (architecture research) and **FPLAN-0185** (master plan, 6 phases) both live in the private project repo, uncommitted.
|
||||
|
||||
- Runtime unblocked: created dedicated `~/Projects/<private>/.venv` with chromadb 1.5.5 + sentence-transformers 5.3.0.
|
||||
- Fixed `navigator/__init__.py` missing. Bypass entries added for seedgo `__init__.py` false positives. Emailed @seedgo friction note.
|
||||
- First real query ran and surfaced decision_006 ("branches self-report optimistically") + obs_s60_verification_agents — the exact prior art DPLAN-0128 was defending against. **Rule #033 closed** — the project is being used, not just built. Next time: query first, plan second.
|
||||
- S92 decisions #034-#036 ingested. Collection now 133 fragments.
|
||||
- Invocation form: the private project invocation form. Bare script invocation fails with path issue — not urgent to fix.
|
||||
|
||||
- `navigator/__init__.py` (new — package marker)
|
||||
- `navigator/.seedgo/bypass.json` (modified — 2 new bypass entries)
|
||||
- `navigator/.trinity/local.json` (modified — S4 session)
|
||||
- `navigator/STATUS.local.md` (updated — gitignored anyway)
|
||||
|
||||
### Friction notes
|
||||
|
||||
- **seedgo __init__.py false positive**: `imports: Failed` and `naming: invalid characters` fire on Python reserved filename. Need special-case in seedgo's naming standard and imports standard. Encountered in another project; likely also in AIPass repo. Open a seedgo standard PR when there's bandwidth.
|
||||
|
||||
### S92 (2026-04-14 ~12:45 PT) — PRE-COMPACT, system prompt reformat work in flight
|
||||
|
||||
**Not shipped, verification-pending**:
|
||||
- `/home/patrick/Projects/AIPass/.aipass/aipass_global_prompt.reformat.md` — draft reformat of AIPass's global system prompt in AIPL/Anthropic Claude Code style. Gitignored, not tracked. Original `.aipass/aipass_global_prompt.md` untouched.
|
||||
- Multiple correction rounds applied this morning (terminology, memory roles, TDPLAN, librarian philosophy, sections removed). See DPLAN-0128 Context section for full history.
|
||||
|
||||
**DPLAN-0128 is the execution plan** — 8 tracks, verification-first, designed to survive /compact. Post-compact session reads DPLAN-0128 first, then the two prompt files, then executes Tracks 1-6 serially (one branch/agent at a time), and ONLY promotes the reformat (Track 8) if all critical claims pass. Critical claims: TDPLAN command exists, no invalid commands in reformat, STATUS.local.md central sync is real.
|
||||
|
||||
**Skill integration gaps to confirm**:
|
||||
- `/prep` at `~/.claude/commands/prep.md` — Section 1 "Memories" lists local.json, observations.json, passport.json — does NOT include STATUS.local.md in the main update list
|
||||
- `/memo` at `~/.claude/commands/memo.md` — STATUS.local.md is in "If Relevant" section, weak framing
|
||||
- Proposed edits drafted verbally but not applied. DPLAN-0128 Track 3 verifies claim + applies fixes post-compact.
|
||||
|
||||
**AIPL format standard question**:
|
||||
- AIPL format is Anthropic-sourced per the user (unconfirmed)
|
||||
- Decision: seedgo standard / prompt note / both — DPLAN-0128 Track 5 asks @seedgo to recommend
|
||||
|
||||
**Gavin reply posted**: https://github.com/AIOSAI/AIPass/issues/261#issuecomment-4245709624 — done this morning.
|
||||
|
||||
**Stale worktree flagged**: `/home/patrick/Projects/AIPass/.claude/worktrees/agent-aa2b7d43` on branch `fix/wake-claude-path-resolution` (PR #268) — leftover from last night's Track E. Cleanup in DPLAN-0128 Track 7 if safe, else manual.
|
||||
|
||||
**S91 follow-ups still open**:
|
||||
- PRs #265-#270 await merge (user reviews in their own time)
|
||||
- Track G (AIPL log leak) still blocked — prax ghosted last night, needs firm retry
|
||||
- PR #269 category-d flags (code-level /home/patrick/ paths in symbolic.py, branch_detector.py docstrings, .claude/settings.json deny rules, ai_mail docstrings)
|
||||
- 2% usage finding to confirm (S91 night shift burned less than expected)
|
||||
|
||||
**Next steps**: /compact (the user does this) → post-compact session reads DPLAN-0128 → executes.
|
||||
|
||||
### S91 (2026-04-14 ~05:00 Pacific) — NIGHT SHIFT COMPLETE
|
||||
|
||||
**6 PRs shipped tonight** — all awaiting AIPass Developer review:
|
||||
- **#265** — devpulse docs restructure (README 209→35 lines + new SETUP.md with Linux/Mac/Windows/WSL/uninstall/troubleshooting)
|
||||
- **#266** — drone three bugs fix (`system-pr` git add -A, STATUS.md sync before commit, Rich help table + bonus encapsulation fix)
|
||||
- **#267** — Windows compat for issue #261 (bootstrap.py pathlib traversal, setup.sh OS detection, new cross-platform setup.py)
|
||||
- **#268** — wake PATH isolation fix (bare `claude` → absolute path via shutil.which + ~/.local/bin PATH prepend)
|
||||
- **#269** — personal-name scrub (11 files, 14+/14-, flagged category-d code paths for manual review)
|
||||
|
||||
**Research reports in `.trinity/night_shift_reports/`** — morning reading:
|
||||
- `anthropic_peak_hours.md` — peak hours are 5 AM–11 AM Pacific (AIPass Developer's intuition was inverted)
|
||||
- `repo_etiquette.md` — top gaps: SECURITY.md (urgent), CHANGELOG.md, root-level INSTALLATION.md
|
||||
- `local_llm_research.md` — hybrid recommendation: Qwen 2.5 3B Q4 Ollama local + Llama 3.3 70B / DeepSeek R1 OpenRouter free
|
||||
- `watchdog_design.md` — full design for future hooks session (Options A-D insufficient, Option E requires hooks)
|
||||
|
||||
**Gavin reply draft**: `/tmp/gavin_reply_261.md` — ready for AIPass Developer review/approval/posting
|
||||
|
||||
**Track G blocked**: @prax ghosted the AIPL log leak dispatch (exited code=0 after 722s but no PR, no reply). Same pattern as S90 `prax_needs_firm_prompts` learning. Retry in morning with MANDATORY framing.
|
||||
|
||||
**Time log**: `.trinity/time_log.jsonl` — 24 entries, 12 complete, 189.5 min tracked. First temporal data point.
|
||||
|
||||
**Anthropic peak hours reality**: I finished the night shift inside the off-peak window. Peak starts at 5 AM Pacific and runs until 11 AM. The "cross 5 AM" theory was backwards but the net effect was right — I paced serial and finished before 5 AM anyway.
|
||||
|
||||
**Total runtime**: ~2h50m of session wall-clock, ~3h09m of tracked task time (overlap from watchdog waits).
|
||||
|
||||
**Session cost estimate**: 6 branch dispatches + 5 sub-agent spawns + manual work. Within off-peak doubled-limit promotion, so cheap.
|
||||
|
||||
**NEEDS AIPASS DEVELOPER ATTENTION**:
|
||||
1. Review and merge PRs #265-270 (in order — #266 drone fix unblocks future `system-pr` usage)
|
||||
2. Review Gavin reply draft, approve and post to issue #261
|
||||
3. Read the 5 research docs in `.trinity/night_shift_reports/`
|
||||
4. Retry Track G (AIPL log leak) — prax needs firm framing
|
||||
5. Manually review category-d code path flags from PR #269 (symbolic.py priority_dirs, branch_detector.py docstrings, .claude/settings.json deny rules, ai_mail docstrings blocked by pre-existing seedgo)
|
||||
6. Consider landing the hooks session soon — unblocks watchdog-as-code (Track H) + audio hook cleanup
|
||||
|
||||
### S91 (2026-04-14 early AM) — Issue #261 triage + cleanup + night shift planning
|
||||
|
||||
**Status at /prep** (02:15 AM, going into autonomous night shift):
|
||||
|
||||
- **3 PRs created**: #262 (cleanup + scrub + epigraph + gitignore), #263 (S90 stragglers drone system-pr missed), #264 (STATUS.md auto-sync refresh) — all need AIPass Developer merge.
|
||||
- **`.claude/CLAUDE.md`**: Patrick scrub complete (3 lines). Only private global `~/.claude/CLAUDE.md` retains the name.
|
||||
- **`README.md`**: memory/presence epigraph added near line 40 with `— AIPass` attribution.
|
||||
- **`trigger_data.lock`**: explained (fcntl process lock from log_watcher_service PID 1622), gitignored with pattern `src/aipass/trigger/*.lock`.
|
||||
- **`unknown_branch/`**: confirmed AIPL polyglot log leak. Tracked as Track G of DPLAN-0125.
|
||||
- **Issue #261 (Gavin Rooney, Windows 10)**: triaged. 3 real bugs (setup.sh symlink, bootstrap.py bash loop, audio hooks). Stale-clone artifacts (old .claude/settings.json `/home/patrick` paths removed April 5). Non-bugs (AIPASS_REGISTRY.json gitignored, passports not tracked). Reply draft pending (FPLAN-0183 Step 1).
|
||||
- **Feedback channel tests**: grep showed 0 module-level test defs — needs verification morning, may be class-based.
|
||||
|
||||
**Plans built**:
|
||||
- **DPLAN-0125** — S91 Night Shift master plan. 10 tracks, SERIAL execution strategy, cross-5AM pacing theory.
|
||||
- **DPLAN-0126** — Future research: local LLM + OpenRouter free tier (split from 0125 as future-ideas bucket).
|
||||
- **FPLAN-0183** — Step-by-step execution script. Post-compact survival. 12 numbered steps + pre-flight. This is what I follow when I wake up.
|
||||
|
||||
**Autonomous night shift directive from the AIPass Developer**: GO. Serial execution, deliberate slow pacing, cross 5 AM peak-hours cutoff, do not touch `.claude/hooks/*`, do not merge PRs, log time per track to `.trinity/time_log.jsonl`, keep going until usage cap or dawn.
|
||||
|
||||
**Known bugs that will bite during execution**:
|
||||
1. `drone @git system-pr` doesn't stage untracked files — stage manually with `git add` before running it
|
||||
2. `drone @git system-pr` leaves STATUS.md dirty after every run — known, accept it, do NOT chase the loop
|
||||
3. Wake command may not work (S90 flag — Track E investigates)
|
||||
|
||||
**Loose ends at /prep**:
|
||||
- STATUS.md still modified (drone bug, documented, ignored)
|
||||
- trigger log_watcher_service running (PID 1622, healthy, leave it)
|
||||
- No background agents running, no open watchdogs at /prep time
|
||||
- Next action after compact: Step 0.5 of FPLAN-0183 (Anthropic peak hours WebSearch) → Step 1 (Gavin reply draft) → onward
|
||||
|
||||
**The AIPass Developer is asleep. Work alone. Stay steady.**
|
||||
|
||||
### S90 (2026-04-11 evening) — System health + cross-project completion
|
||||
- **Dispatch -c -p bug:** Claude Code v2.1.90-100 broke headless resume. Fixed v2.1.101. DPLAN-0122 (closed).
|
||||
- **AIPASS_HOME:** Added to .bashrc (export) + ~/.claude/settings.json env. CLI auto-detects via importlib.
|
||||
- **Vera 4 bugs via feedback:** All resolved. Registry hardcode (PR #255), export (manual), registry ID (spawn already built), local prompt hook (PR #257).
|
||||
- **5 PRs:** #254 (drone JSON+cli init+test isolation), #255 (ai_mail bidirectional), #256 (prax cleanup), #257 (local prompt hook), #258 (prax Ctrl+C+labels)
|
||||
- **Prax finally delivered:** Ctrl+C clean exit, UNKNOWN→AIPL/POLYGLOT labels, AI_MAIL→AI_MAIL label fix. Took 4 dispatches.
|
||||
- **Contacts system built:** identity.json + contacts.json + auto-registration. PR #259.
|
||||
- **Cross-project email WORKING:** Vera→@strategy tested live in tmux. Two fixes: sender detection (drone router_handler) + target resolution (ai_mail delivery.py loads caller registry). PR #260.
|
||||
- **15+ dispatches:** ai_mail x4, prax x4, drone x1, cli x2, spawn x1, flow x1, + sub-agents
|
||||
- **PRs #249-260 all need merge** (Patrick reviews in his own time)
|
||||
- **Wake broken:** dispatch wake not working at end of S90. Investigate S91.
|
||||
- **NEEDS PATRICK:** Merge PRs #249-260
|
||||
|
||||
### S89 (2026-04-11 afternoon) — Full external access + feedback bridge
|
||||
- AIPASS_HOME dual registry: drone loads both local + AIPass registries simultaneously
|
||||
- Global prompt template rewrite: 3→7 sections, all command variants (dplan/aplan/master, dispatch, feedback)
|
||||
- UX warnings: drone systems hints + BranchNotFoundError guidance when AIPASS_HOME not set
|
||||
- `aipass init update` idempotent: content diff before write, "Already current" for unchanged files (8 new CLI tests)
|
||||
- Feedback channel (DPLAN-0117/FPLAN-0173): cross-project communication bridge
|
||||
- 4 handler files + 1 module, 101 tests passing
|
||||
- Sender auto-detected via AIPASS_CALLER_BRANCH env var
|
||||
- Reply path stored per-message for cross-project delivery
|
||||
- Fixed devpulse.py module discovery (full package import path)
|
||||
- Fixed pytest.ini inline comments bug (seedgo confirmed: isolated to devpulse, no standard exists)
|
||||
- Global prompt: "ask before spelunking" rule + feedback section added
|
||||
- Learning: dispatch the question to branches instead of reading 4-5 unfamiliar files
|
||||
|
||||
### S88 (2026-04-11 morning) — AIPL Phase 1 + global prompt fix
|
||||
- AIPL Phase 1 COMPLETE: docs/style_guide.md (7 rules, symbols, content types, anti-patterns, decode protocol, quick reference)
|
||||
- 6 before/after examples: session_log, observation, email, status_update, plan, delta_update (54-69% savings)
|
||||
- Polyglot agent wired: local prompt filled, CLAUDE.md updated with AIPL context
|
||||
- DPLAN-0115 created via flow (replaced manual DPLAN-001 — wrong numbering)
|
||||
- Global prompt: "Never create plan files manually" rule added
|
||||
- Learning: external AIPass projects still use AIPass tooling (flow, seedgo) — don't bypass
|
||||
|
||||
### S87 (2026-04-11 morning) — Model choice + AIPL + init update
|
||||
- dispatch --model flag: sonnet default, all 3 tiers tested (sonnet/haiku/opus confirmed in process args)
|
||||
- Apr 4 usage policy change discovered — sub-agent compute metered more aggressively
|
||||
- DPLAN-0114: aipass init update (5 phases). cli+spawn+drone dispatched, all replied. Registry fix, hook fixes, source headers ready.
|
||||
- Patrick-Personal deny rules added to AIPass settings
|
||||
- Vera Studio S3 reviewed — registry mismatch, missing hooks, 5 gaps documented in DPLAN-0114
|
||||
- AIPL PROJECT LAUNCHED: ~/Projects/AIPL/. 72K research transferred. DPLAN-001 written. Caveman analyzed (neighbor, not parent). Build from scratch, hand to Polyglot when ready.
|
||||
- PRs #231-248 ALL MERGED (Patrick merged #247 for S86 stack, #248 for S87)
|
||||
- Trigger cleaned up after PC shutdown (stale lock removed, log_watcher auto-restarted, processed 2 pending emails)
|
||||
- **NEEDS PATRICK:** Wake protection — any branch can manually wake devpulse
|
||||
- **NEEDS PATRICK:** CLAUDE.md template gap — spawn template has it, no branch does
|
||||
|
||||
|
||||
### S86 (2026-04-10 afternoon) — Autonomous session
|
||||
- 41 dispatches, 15 PRs (#233-247), DPLAN-0112 complete
|
||||
- Devpulse APLAN: 42→7 open items
|
||||
- decisions.md: #029 (dispatch patterns) + #030 (fixes→architecture in one session)
|
||||
|
||||
### S85 (2026-04-10 morning)
|
||||
- TRIGGER SELF-HEALING PROVEN: Full autonomous error cycle — detect → dispatch → wake → fix → report
|
||||
- Fixes to trigger: (1) log_watcher dedup blocking count increment, (2) report_error() count==2 gate, (3) wake_branch() after email delivery
|
||||
- Drone git fix/sync upgraded to merge on divergence instead of rebase (PR #229)
|
||||
- OSS Health badge: 364 commits, 2.8h pace, 1 month old (PR #230)
|
||||
- 13 stale plans closed (28→15 open)
|
||||
- Old Dev-Pass trigger studied at /media/patrick/EXTERNAL/Dev-Pass/aipass_core/trigger/
|
||||
- Key learning: old system had no count threshold (noisy), new system's threshold was right but gate logic was wrong
|
||||
- TODO: Watchdog needs error handling (log failures so trigger catches them)
|
||||
- TODO: Add prax logger catch-all to all branch entry points (unhandled exceptions bypass logs)
|
||||
- TODO: Properly develop watchdog v2 (leave duct tape running for a week first)
|
||||
- Never reset HEAD — merge only. SATA=/mnt/sata (old OS), EXTERNAL=/media/patrick/EXTERNAL (Dev-Pass)
|
||||
|
||||
### S83 (2026-04-09 evening)
|
||||
- 22 Opus adversarial audit agents: 350+ findings across all 11 branches
|
||||
@@ -252,6 +485,12 @@
|
||||
- README + Herald + global/local prompts all updated to "11 agents".
|
||||
- Patrick's direction: focus on core 11, use the system for real projects, Vera handles promotion.
|
||||
|
||||
## Notepad (User)
|
||||
- Study Anthropic prompt layout.
|
||||
- Review builder agent (~/.claude/agents/builder.md) — should it default to sonnet? Should seedgo checklist be the primary validator instead of pyright? Should it be available to external projects?
|
||||
- Add /memo and /prep to `aipass init` so new projects get them in .claude/commands/ (DPLAN-0114 follow-up)
|
||||
- Prax: fix AI_MAIL showing as MAIL (underscore branch names getting split)
|
||||
|
||||
</details>
|
||||
|
||||
<details><summary><strong>@drone</strong> — Operational (2026-04-07 (Session 38))</summary>
|
||||
@@ -272,7 +511,7 @@
|
||||
| `drone --help` | Operational | |
|
||||
| `drone --version` | Operational | v1.1.0 |
|
||||
| `drone` (no args) | Operational | Auto-discovers 9 modules |
|
||||
| `drone systems` | Operational | Shows infrastructure, modules (3), branches (15) |
|
||||
| `drone systems` | Operational | Shows infrastructure, modules (4), branches (11) |
|
||||
| `drone @target command` | Operational | Routes to any branch via subprocess |
|
||||
| `drone @target --help` | Operational | |
|
||||
| `drone @seedgo audit` | Operational | Interactive mode for Rich progress bars |
|
||||
@@ -456,6 +695,45 @@ Vector verification shows in console output: "Vectorized: N chunks in chroma" or
|
||||
|
||||
</details>
|
||||
|
||||
<details><summary><strong>@polyglot</strong> — Unknown (n/a)</summary>
|
||||
|
||||
# @POLYGLOT
|
||||
|
||||
> AIPL compression engine developer
|
||||
|
||||
State: Phase 2 complete + seedgo 99% + guardrails designed (DPLAN-0124)
|
||||
Last update: 2026-04-12
|
||||
|
||||
## Milestones
|
||||
- Phase 1 spec: COMPLETE (locked)
|
||||
- Phase 2 engine: COMPLETE (441 tests, 10 handlers, 5 modules, 99% seedgo)
|
||||
- Global prompt AIPL'd: DONE (913->843 tokens, 7.7% honest savings)
|
||||
- ~/.aipl scaffold: DEPLOYED
|
||||
- Phase 3 integration: In progress
|
||||
- Phase 4 validation: Not started
|
||||
|
||||
## Current Work
|
||||
- Token scan: 220K tokens across 80 AIPass files, 81K saveable
|
||||
- Stress tested 16 real files, 15.8% avg automated savings
|
||||
- Decode tests: 93-95% info recovery on compressed files
|
||||
- AIPL is convention not engine — agents write terse from global prompt
|
||||
- Private repo: AIOSAI/AIPL
|
||||
|
||||
## Known Issues
|
||||
- Commands incompressible — only prose between commands can shrink
|
||||
- .trinity/ must stay valid JSON — no headers, no unquoted keys
|
||||
- Cross-project ai_mail outbound broken (feedback channel in progress)
|
||||
- delta_extractor limited pattern detection
|
||||
- Handlers standard 93% — inspect.stack() needed on 2 pipeline files
|
||||
|
||||
## Next
|
||||
- Build DPLAN-0124: anchor detection + integrity verification (guardrails)
|
||||
- Roll AIPL global prompt to AIPass and Vera-Studio
|
||||
- Test with fresh agents across projects
|
||||
- Close 17% vs 62% gap via convention (not engine)
|
||||
|
||||
</details>
|
||||
|
||||
<details><summary><strong>@prax</strong> — Operational (2026-04-07)</summary>
|
||||
|
||||
# @prax
|
||||
@@ -503,14 +781,14 @@ init_prax, shutdown, discover, run, terminal — in modules/.archive/
|
||||
|
||||
</details>
|
||||
|
||||
<details><summary><strong>@seedgo</strong> — Operational — 99% self-audit (32/33 standards at 100%, Architecture 99%) (2026-04-09 (session 43 — dispatch processing))</summary>
|
||||
<details><summary><strong>@seedgo</strong> — Operational — 99% self-audit (32/33 standards at 100%, Architecture 99%) (2026-04-14 (session 59 — DPLAN-0128 Track 5 advisory: recommend prompt note + ref doc over 34th standard))</summary>
|
||||
|
||||
# @seedgo
|
||||
|
||||
> Standards enforcement, 32-standard audit pack + seedgo proof system (CERTIFIED) + hook ownership
|
||||
|
||||
**State:** Operational — 99% self-audit (32/33 standards at 100%, Architecture 99%)
|
||||
**Last update:** 2026-04-09 (session 43 — dispatch processing)
|
||||
**Last update:** 2026-04-14 (session 59 — DPLAN-0128 Track 5 advisory: recommend prompt note + ref doc over 34th standard)
|
||||
|
||||
## Session 43 — Dispatch Processing
|
||||
|
||||
|
||||
@@ -458,6 +458,91 @@ def _claude_settings(aipass_home: str | None = None) -> str:
|
||||
return json.dumps(data, indent=2, ensure_ascii=False) + "\n"
|
||||
|
||||
|
||||
def _prep_md() -> str:
|
||||
"""Generate .claude/commands/prep.md — /prep session wrap-up slash command.
|
||||
|
||||
Project-scoped slash command. Sessions wrap up by buttoning up memories,
|
||||
plans, git state, inbox, and loose ends. Ships per-project via aipass init
|
||||
so every AIPass project has a consistent /prep flow.
|
||||
"""
|
||||
return (
|
||||
"# Session Wrap-Up\n"
|
||||
"\n"
|
||||
"Purpose: Button up everything at the end of a session — or before a "
|
||||
"/compact. Memories, plans, git — all tidy. Works for both closing out "
|
||||
"a chat and preparing for compaction.\n"
|
||||
"\n"
|
||||
"**Workflow:** `/prep` → review output → close chat or `/compact`\n"
|
||||
"\n"
|
||||
"## Execution\n"
|
||||
"\n"
|
||||
"1. Read `.trinity/passport.json` first — re-absorb your identity "
|
||||
"before writing anything\n"
|
||||
"2. Do ALL of the following, then confirm what was updated\n"
|
||||
"\n"
|
||||
"## 1. Memories\n"
|
||||
"\n"
|
||||
"Each memory file plays a distinct role. Update based on what actually "
|
||||
"changed this session.\n"
|
||||
"\n"
|
||||
"- **`.trinity/passport.json`** — IDENTITY. Who you are: role, "
|
||||
"capabilities, principles. Only update if identity genuinely evolved "
|
||||
"this session.\n"
|
||||
"- **`.trinity/local.json`** — YOUR MEMORY. Add/update session entry "
|
||||
"with a summary of work done. Add key_learnings for anything learned. "
|
||||
"Trim oldest sessions if over 20.\n"
|
||||
"- **`.trinity/observations.json`** — YOUR MEMORY OF THE USER. "
|
||||
"Collaboration insights, preferences, friction points. Skip if nothing "
|
||||
"new about the user this session.\n"
|
||||
"- **`STATUS.local.md`** — PUBLIC STATUS BEACON. Current work, known "
|
||||
"issues, todos, notepad. Auto-synced to central STATUS.md on PR events "
|
||||
"— this is how other branches see you. Keep Current Work accurate.\n"
|
||||
"\n"
|
||||
"## 2. Active Plans\n"
|
||||
"\n"
|
||||
"- Check any DPLANs or FPLANs referenced in this session\n"
|
||||
"- Update their execution logs, status, decision logs with current "
|
||||
"state\n"
|
||||
"- If a plan was completed, note it (but don't close — the user does "
|
||||
"that)\n"
|
||||
"\n"
|
||||
"## 3. Git State\n"
|
||||
"\n"
|
||||
"- Run `git status` — report uncommitted changes\n"
|
||||
"- If there's a logical commit waiting, suggest it (don't commit "
|
||||
"without asking)\n"
|
||||
"- Note the current branch and any open PRs\n"
|
||||
"\n"
|
||||
"## 4. Inbox\n"
|
||||
"\n"
|
||||
"- Run `drone @ai_mail inbox 2>/dev/null` — report any unread emails\n"
|
||||
"- Close any that were already processed but not formally closed\n"
|
||||
"\n"
|
||||
"## 5. Loose Ends\n"
|
||||
"\n"
|
||||
"- Flag anything in-flight: running background agents, dispatched "
|
||||
"branches waiting for replies, pending decisions\n"
|
||||
"- If anything can't survive compaction (e.g., agent IDs needed for "
|
||||
"resume), write it to STATUS.local.md Notepad\n"
|
||||
"\n"
|
||||
"## Confirm\n"
|
||||
"\n"
|
||||
"List everything updated. Format:\n"
|
||||
"```\n"
|
||||
"Prep complete:\n"
|
||||
"- local.json: [what was added]\n"
|
||||
"- observations.json: [updated / skipped]\n"
|
||||
"- STATUS.local.md: [updated / skipped]\n"
|
||||
"- Plans: [which ones updated]\n"
|
||||
"- Git: [branch, uncommitted count, suggestion]\n"
|
||||
"- Inbox: [count, action taken]\n"
|
||||
"- Loose ends: [any flagged]\n"
|
||||
"\n"
|
||||
"Ready to close out or /compact.\n"
|
||||
"```\n"
|
||||
)
|
||||
|
||||
|
||||
def _inbox_json() -> str:
|
||||
"""Generate .ai_mail.local/inbox.json — empty project mailbox structure."""
|
||||
return json.dumps(
|
||||
@@ -615,6 +700,14 @@ def init_project(target: Path, project_name: str | None = None) -> dict:
|
||||
settings_path.write_text(_claude_settings(aipass_home), encoding="utf-8")
|
||||
created.append(str(settings_path))
|
||||
|
||||
# 9b. .claude/commands/prep.md — /prep session wrap-up slash command
|
||||
commands_dir = claude_dir / "commands"
|
||||
commands_dir.mkdir(exist_ok=True)
|
||||
prep_path = commands_dir / "prep.md"
|
||||
if not prep_path.exists():
|
||||
prep_path.write_text(_prep_md(), encoding="utf-8")
|
||||
created.append(str(prep_path))
|
||||
|
||||
# 10. hooks/ directory
|
||||
hooks_dir = target / "hooks"
|
||||
if not hooks_dir.exists():
|
||||
@@ -742,6 +835,17 @@ def update_project(target: Path) -> dict:
|
||||
else:
|
||||
already_current.append(str(gemini_md_path))
|
||||
|
||||
# .claude/commands/prep.md — managed slash command, refresh to latest
|
||||
commands_dir = claude_dir / "commands"
|
||||
commands_dir.mkdir(exist_ok=True)
|
||||
prep_path = commands_dir / "prep.md"
|
||||
generated = _prep_md()
|
||||
if not prep_path.exists() or prep_path.read_text(encoding="utf-8") != generated:
|
||||
prep_path.write_text(generated, encoding="utf-8")
|
||||
updated.append(str(prep_path))
|
||||
else:
|
||||
already_current.append(str(prep_path))
|
||||
|
||||
# --- User-owned files: always skip ---
|
||||
for skip_name in (
|
||||
str(registry_path),
|
||||
|
||||
@@ -4,20 +4,22 @@ Injected every turn. Breadcrumbs only — details in README, --help, .trinity/ m
|
||||
|
||||
## Identity
|
||||
|
||||
You are DEVPULSE — orchestration hub. Manager, not builder. Coordinate, plan, delegate, track.
|
||||
You are DEVPULSE — Patrick's primary AI collaborator and orchestration hub for AIPass. You design, plan, debug, dispatch, and track. You build things you own (watchdog, feedback, your own plans and memories). You venture into other branches to investigate, debug, and fix small bugs. You delegate heavy multi-file builds to sub-agents. You stay aware of which branch your CWD is in — that's your identity grounding.
|
||||
|
||||
## How You Work
|
||||
|
||||
- Delegate code tasks to background agents (`run_in_background: true`). Fire and forget — move on immediately.
|
||||
- Launch agent → continue conversation → get notified → report results.
|
||||
- Never block waiting on agents. Never burn context reading code across branches.
|
||||
- **Build what you own directly.** Your modules, your DPLANs, your FPLANs, your memories, your STATUS — those are yours. Edit them freely.
|
||||
- **Prototype to explore.** When a shape isn't clear, sketch it yourself first, then hand the real build off to a sub-agent.
|
||||
- **Investigate other branches freely.** Read their code, debug their issues, run their tests, fix small bugs you find. The CWD stays devpulse — you're visiting, not moving in.
|
||||
- **Don't solo-rebuild other branches.** Full multi-file implementations → dispatch via `drone @ai_mail dispatch @branch`.
|
||||
- **Delegate heavy code to sub-agents** (`run_in_background: true`). Fire and forget, move on immediately. Launch → continue → get notified → report results. Never block waiting on agents.
|
||||
- Use `drone @branch --help` for command syntax. Use `drone systems` for branch list.
|
||||
- **ALWAYS WAKE after sending dispatch emails.** Send email → wake. Every time. No asking. If the user wants something different, they will say so.
|
||||
- **START WATCHDOG after any dispatch.** Use the watchdog one-liner (see Watchdog section below) with `run_in_background: true`. Don't wait for Patrick to ask.
|
||||
- **ALWAYS WAKE after sending dispatch emails.** Send email → wake. Every time. No asking.
|
||||
- **START WATCHDOG after any dispatch.** Run `drone @devpulse watchdog agent @target` (see Watchdog section) with `run_in_background: true`. Don't wait for the user to ask.
|
||||
|
||||
## Dispatch, Don't Do
|
||||
## Branch Experts — Ask Before Rebuilding
|
||||
|
||||
When a task belongs to a specialist, send it there. Don't burn context doing their job.
|
||||
When a task belongs to a specialist's DOMAIN, ask them. You can still investigate or fix small things yourself — but for anything that touches a branch's core architecture, email the owner first.
|
||||
|
||||
| Domain | Ask | Why |
|
||||
|--------|-----|-----|
|
||||
@@ -46,7 +48,7 @@ drone @git lock # Check PR lock status
|
||||
|
||||
Read-only git commands are fine: `git status`, `git diff`, `git log`.
|
||||
|
||||
**NEVER cd to repo root.** `drone @git system-pr` requires `.trinity/passport.json` in the CWD hierarchy. If you cd to `/home/patrick/Projects/AIPass/`, it fails. Stage files with relative paths from devpulse: `git add ../../../HERALD.md`. Always run drone commands from this directory.
|
||||
**NEVER cd to repo root.** `drone @git system-pr` requires `.trinity/passport.json` in the CWD hierarchy. If you cd to the repo root, it fails. Stage files with relative paths from devpulse: `git add ../../../HERALD.md`. Always run drone commands from this directory.
|
||||
|
||||
## Key Commands
|
||||
|
||||
@@ -64,44 +66,27 @@ drone systems # All branches
|
||||
|
||||
drone, seedgo, prax, cli, ai_mail, api, flow, spawn, trigger, memory, devpulse (you — no apps/, coordinates via dispatch + agents)
|
||||
|
||||
## Your Projects
|
||||
|
||||
Two personal projects, both part of the Nexus vision. Work on these during autonomy time.
|
||||
|
||||
**Compass** at `~/Projects/compass/` — Vector-based thinking engine for autonomous decision-making. 130 fragments (decisions + observations + learnings). Query before big choices. Stop building features, start using it (#033). Copy @memory's fragment code as research for multi-collection architecture (DPLAN-023).
|
||||
|
||||
**AIPL** at `~/Projects/AIPL/` — Token compression for AI agent storage/communication. ~45% savings proven. Phase 1 COMPLETE (style guide + 6 examples in docs/). DPLAN-0115. Polyglot agent builds Phase 2 (compression engine). Hand to Polyglot when ready.
|
||||
|
||||
## Working Habits
|
||||
|
||||
- **Lean on branches.** You can't know everything — branches are the experts on their systems. When unsure, email them and ask. Don't burn context debugging what they already know.
|
||||
- **Lean on branches for expertise.** Branches are the experts on their own architecture. When in doubt about a branch's internal design, email them. But debugging, reading, testing, and small fixes in their code is fair game — you don't need permission to investigate.
|
||||
- **Use memories freely.** Don't hoard or stress about capacity — rollover to @memory is by design. Update `.trinity/` often. More is better.
|
||||
- **STATUS.local.md for friction notes.** When something feels off or could be improved, drop a quick note in the Notepad section. Address in batches later.
|
||||
- **Know your limits.** You're great at planning, coordinating, seeing the big picture. You're bad at hands-on branch-level code tasks. Dispatch, don't do.
|
||||
- **Know what to build vs delegate.** Things you own (watchdog, feedback, your DPLANs/FPLANs, memories, prompts, small fixes across the codebase) → build directly. Multi-file new features or heavy refactors → delegate to a sub-agent so your context stays clean.
|
||||
- **CWD is identity.** You move in and out of branches all day. Always know which branch you're standing in — the CWD determines everything (drone routing, git operations, mailbox, passport lookups). Never cd into another branch and forget to come back. Visit, don't move in.
|
||||
- **Git awareness as a natural habit.** After completing a feature or wrapping up a chunk of work, run `git status`. If changes look coherent (upgrade, fix cycle, config update), suggest a commit or PR. Don't force it every turn, but don't let files pile up silently either.
|
||||
## Watchdog — Autonomous Mail Wait
|
||||
## Watchdog — Directed Wake (devpulse module)
|
||||
|
||||
After dispatching branches, use a background bash wait that exits when mail arrives. This wakes you like a sub-agent completing.
|
||||
Watchdog is a real devpulse module now (not a bash one-liner). After dispatching, arm it as a background task — it polls the dispatch lock file and exits when the agent process finishes (success, silent-finish, OR crash). The exit wakes you.
|
||||
|
||||
**Pattern:**
|
||||
```bash
|
||||
# 1. Dispatch work
|
||||
drone @ai_mail dispatch @target "Subject" "Body"
|
||||
|
||||
# 2. Arm watchdog (run_in_background: true, timeout: 600000)
|
||||
# Snapshots current unread_count, wakes when it increases
|
||||
INBOX="path/to/.ai_mail.local/inbox.json"; INITIAL=$(python3 -c "import json; from pathlib import Path; p=Path('$INBOX'); print(json.loads(p.read_text()).get('unread_count',0) if p.exists() else 0)" 2>/dev/null); C=0; while [ $C -lt 60 ]; do sleep 10; C=$((C+1)); CURRENT=$(python3 -c "import json; from pathlib import Path; p=Path('$INBOX'); print(json.loads(p.read_text()).get('unread_count',0) if p.exists() else 0)" 2>/dev/null); if [ "$CURRENT" -gt "$INITIAL" ]; then echo "WOKE: new mail ($INITIAL→$CURRENT)"; exit 0; fi; done; echo "TIMEOUT"
|
||||
|
||||
# 3. Stop — do nothing until notified
|
||||
# 4. Wake notification arrives → read mail → process → dispatch next → repeat
|
||||
drone @devpulse watchdog agent @target # run_in_background: true
|
||||
```
|
||||
|
||||
**Key:** Snapshot unread_count BEFORE arming, then wake when it increases. Don't require empty inbox — works with existing mail. 10s poll interval. `run_in_background: true` so the completion notification wakes you. On timeout, wake anyway to check if agent crashed.
|
||||
The handler resolves `@target` → branch path → `.ai_mail.local/.dispatch.lock`, polls the monitor PID, and returns when the lock disappears or the PID dies. Crash vs success is distinguished by `last_bounce.json`. Default timeout 1800s — override with `--timeout SECONDS`.
|
||||
|
||||
**Watchdog one-liner (copy-paste ready):**
|
||||
```
|
||||
INBOX="/home/patrick/Projects/AIPass/src/aipass/devpulse/.ai_mail.local/inbox.json"; INITIAL=$(python3 -c "import json; from pathlib import Path; p=Path('$INBOX'); print(json.loads(p.read_text()).get('unread_count',0) if p.exists() else 0)" 2>/dev/null); C=0; while [ $C -lt 60 ]; do sleep 10; C=$((C+1)); CURRENT=$(python3 -c "import json; from pathlib import Path; p=Path('$INBOX'); print(json.loads(p.read_text()).get('unread_count',0) if p.exists() else 0)" 2>/dev/null); if [ "$CURRENT" -gt "$INITIAL" ]; then echo "WOKE: new mail ($INITIAL→$CURRENT)"; exit 0; fi; done; echo "TIMEOUT: 10min no new mail"; exit 0
|
||||
```
|
||||
`drone @devpulse watchdog --help` for full subcommand list. See FPLAN-0186 (build) and DPLAN-0130 (design).
|
||||
|
||||
## Memory & Tracking
|
||||
|
||||
|
||||
@@ -12,4 +12,5 @@ build/
|
||||
*.log
|
||||
*.tmp
|
||||
*.swp
|
||||
my-project
|
||||
my-project
|
||||
|
||||
|
||||
@@ -40,6 +40,66 @@
|
||||
"standard": "cli",
|
||||
"file": "apps/devpulse.py",
|
||||
"reason": "Manager branch — not a CLI module. Entry point auto-generated by spawn template. --help is implemented via print_introspection()."
|
||||
},
|
||||
{
|
||||
"standard": "encapsulation",
|
||||
"file": "apps/modules/watchdog.py",
|
||||
"reason": "Watchdog module has its own _guard_caller() that blocks cross-branch invocation at the handle_command boundary. Inherits branch-level gap: devpulse has no apps/handlers/__init__.py inspect.stack guard."
|
||||
},
|
||||
{
|
||||
"standard": "json_structure",
|
||||
"file": "apps/modules/watchdog.py",
|
||||
"reason": "Devpulse is a manager branch with no apps/handlers/json/json_handler. Watchdog logs through prax system_logger instead."
|
||||
},
|
||||
{
|
||||
"standard": "encapsulation",
|
||||
"file": "apps/handlers/watchdog/agent.py",
|
||||
"reason": "Inherits branch-level gap: devpulse has no apps/handlers/__init__.py inspect.stack guard (manager branch). Cross-branch protection enforced at the module layer via _guard_caller()."
|
||||
},
|
||||
{
|
||||
"standard": "json_structure",
|
||||
"file": "apps/handlers/watchdog/agent.py",
|
||||
"reason": "Devpulse is a manager branch with no json_handler. Agent handler logs through prax system_logger to stderr."
|
||||
},
|
||||
{
|
||||
"standard": "encapsulation",
|
||||
"file": "apps/handlers/watchdog/__init__.py",
|
||||
"reason": "Structural Python package marker — inherits branch-level gap (devpulse has no apps/handlers/__init__.py inspect.stack guard, manager branch)."
|
||||
},
|
||||
{
|
||||
"standard": "naming",
|
||||
"file": "apps/handlers/watchdog/__init__.py",
|
||||
"reason": "Standard Python package marker filename — required by Python, cannot be renamed to snake_case."
|
||||
},
|
||||
{
|
||||
"standard": "encapsulation",
|
||||
"file": "apps/handlers/watchdog/timer.py",
|
||||
"reason": "Inherits branch-level gap: devpulse has no apps/handlers/__init__.py inspect.stack guard (manager branch). Cross-branch protection enforced at the module layer via watchdog.py _guard_caller(). Same situation as watchdog/agent.py."
|
||||
},
|
||||
{
|
||||
"standard": "json_structure",
|
||||
"file": "apps/handlers/watchdog/timer.py",
|
||||
"reason": "Devpulse is a manager branch with no json_handler. Timer handler logs through prax system_logger. Same situation as watchdog/agent.py."
|
||||
},
|
||||
{
|
||||
"standard": "encapsulation",
|
||||
"file": "apps/handlers/watchdog/schedule.py",
|
||||
"reason": "Inherits branch-level gap: devpulse has no apps/handlers/__init__.py inspect.stack guard (manager branch). Cross-branch protection enforced at the module layer via watchdog.py _guard_caller(). Same situation as watchdog/agent.py and watchdog/timer.py."
|
||||
},
|
||||
{
|
||||
"standard": "json_structure",
|
||||
"file": "apps/handlers/watchdog/schedule.py",
|
||||
"reason": "Devpulse is a manager branch with no json_handler. Schedule handler logs through prax system_logger. Same situation as watchdog/agent.py and watchdog/timer.py."
|
||||
},
|
||||
{
|
||||
"standard": "encapsulation",
|
||||
"file": "apps/handlers/watchdog/registry.py",
|
||||
"reason": "Inherits branch-level gap: devpulse has no apps/handlers/__init__.py inspect.stack guard (manager branch). Cross-branch protection enforced at the module layer via watchdog.py _guard_caller(). Same situation as watchdog/agent.py, timer.py, schedule.py."
|
||||
},
|
||||
{
|
||||
"standard": "json_structure",
|
||||
"file": "apps/handlers/watchdog/registry.py",
|
||||
"reason": "Devpulse is a manager branch with no json_handler. Registry logs through prax system_logger. Same situation as watchdog/agent.py, timer.py, schedule.py."
|
||||
}
|
||||
],
|
||||
"notes": {
|
||||
|
||||
@@ -28,7 +28,7 @@ Say "hi" and DevPulse picks up where the last session left off.
|
||||
|
||||
## Role in one line
|
||||
|
||||
Manager, not builder. Coordinates via dispatch + sub-agents. Does not read or edit code across branches — that burns context that belongs to coordination.
|
||||
Designer, orchestrator, and light builder — the user's primary AI collaborator. Builds its own things directly (modules, plans, memories, design docs). Ventures into other branches to investigate, debug, run tests, and fix small bugs — CWD stays devpulse. Delegates heavy multi-file builds and full branch rebuilds to sub-agents via dispatch.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -184,7 +184,7 @@ python -m pytest src/aipass/<branch>/tests/
|
||||
|
||||
If tests fail because `AIPASS_HOME` leaks the real registry into test results, that's a known pattern — the tests need `monkeypatch.delenv("AIPASS_HOME")`. See S90 notes for the fixture pattern.
|
||||
|
||||
### `.claude/settings.json` has `/home/patrick/` paths
|
||||
### `.claude/settings.json` has hardcoded absolute paths
|
||||
|
||||
You pulled an old clone. The hardcoded paths were removed in commit `867dad0` (April 5, 2026). Pull the latest main and re-run `setup.sh`, which generates the settings dynamically from your local repo root.
|
||||
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: __init__.py
|
||||
# Description: Watchdog Handlers Package
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
@@ -0,0 +1,254 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: agent.py
|
||||
# Description: Watchdog Agent Handler — block until dispatched agent exits
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
# Signal choice: ai_mail dispatch lock file polling.
|
||||
#
|
||||
# When ai_mail dispatches an agent, dispatch_monitor.py creates
|
||||
# branch_path/.ai_mail.local/.dispatch.lock containing the monitor PID,
|
||||
# and ALWAYS deletes it on exit (success OR crash). So the lock
|
||||
# file's existence is the agent's liveness signal — OS-level, crash-aware,
|
||||
# template-independent. We poll for: (a) lock file gone, OR (b) monitor
|
||||
# PID dead. Crash vs success is distinguished by the presence of
|
||||
# .ai_mail.local/last_bounce.json which the monitor writes on failure.
|
||||
|
||||
"""
|
||||
Watchdog Agent Handler — block until a dispatched agent process exits.
|
||||
|
||||
The whole point: this function returns ONLY when the dispatched agent
|
||||
has finished, however it finished. The exit is the wake signal — when
|
||||
this function returns from a `run_in_background: true` invocation,
|
||||
devpulse wakes.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
from aipass.prax.apps.modules.logger import system_logger as logger
|
||||
|
||||
from aipass.devpulse.apps.handlers.watchdog import registry as _registry
|
||||
|
||||
|
||||
def _stderr(msg: str) -> None:
|
||||
"""Write to stderr — visible in debug runs, doesn't pollute stdout capture."""
|
||||
sys.stderr.write(msg + "\n")
|
||||
sys.stderr.flush()
|
||||
|
||||
|
||||
def _find_repo_root(start: Path | None = None) -> Path | None:
|
||||
"""Walk upward looking for AIPASS_REGISTRY.json. Returns None if not found."""
|
||||
cur = (start or Path.cwd()).resolve()
|
||||
for candidate in [cur, *cur.parents]:
|
||||
if (candidate / "AIPASS_REGISTRY.json").exists():
|
||||
return candidate
|
||||
return None
|
||||
|
||||
|
||||
def _resolve_branch_path(agent_id: str) -> Path | None:
|
||||
"""Resolve an `@branch` token (or bare name) to its absolute branch path."""
|
||||
repo_root = _find_repo_root()
|
||||
if repo_root is None:
|
||||
logger.warning("[watchdog.agent] AIPASS_REGISTRY.json not found")
|
||||
return None
|
||||
|
||||
registry_file = repo_root / "AIPASS_REGISTRY.json"
|
||||
try:
|
||||
registry = json.loads(registry_file.read_text(encoding='utf-8'))
|
||||
except (OSError, json.JSONDecodeError) as exc:
|
||||
logger.warning("[watchdog.agent] failed to read registry: %s", exc)
|
||||
return None
|
||||
|
||||
target = f"@{agent_id.lstrip('@').lower()}"
|
||||
for branch in registry.get("branches", []):
|
||||
if branch.get("email", "").lower() == target:
|
||||
raw_path = branch.get("path", "")
|
||||
path = Path(raw_path)
|
||||
if not path.is_absolute():
|
||||
path = repo_root / path
|
||||
return path if path.exists() else None
|
||||
return None
|
||||
|
||||
|
||||
def _is_zombie_linux(pid: int) -> bool:
|
||||
"""Linux-only zombie check via /proc. Returns True if zombie."""
|
||||
try:
|
||||
status_text = Path(f"/proc/{pid}/status").read_text(encoding='utf-8')
|
||||
except OSError as exc:
|
||||
logger.info("[watchdog.agent] /proc/%s/status unreadable: %s", pid, exc)
|
||||
return False
|
||||
for line in status_text.splitlines():
|
||||
if line.startswith("State:"):
|
||||
return "Z" in line
|
||||
return False
|
||||
|
||||
|
||||
def _pid_alive(pid: int) -> bool:
|
||||
"""Return True if the process is alive (not zombie)."""
|
||||
try:
|
||||
os.kill(pid, 0)
|
||||
except ProcessLookupError as exc:
|
||||
logger.info("[watchdog.agent] PID %s not found: %s", pid, exc)
|
||||
return False
|
||||
except PermissionError as exc:
|
||||
logger.info("[watchdog.agent] PID %s permission denied (alive): %s", pid, exc)
|
||||
return True
|
||||
if sys.platform == "linux" and _is_zombie_linux(pid):
|
||||
return False
|
||||
return True
|
||||
|
||||
|
||||
def _read_lock(lock_file: Path) -> dict | None:
|
||||
"""Read lock file, return dict or None on miss/error."""
|
||||
if not lock_file.exists():
|
||||
return None
|
||||
try:
|
||||
return json.loads(lock_file.read_text(encoding='utf-8'))
|
||||
except (OSError, json.JSONDecodeError) as exc:
|
||||
logger.warning("[watchdog.agent] could not read lock %s: %s", lock_file, exc)
|
||||
return None
|
||||
|
||||
|
||||
def _classify_exit(branch_path: Path, lock_existed: bool) -> tuple[str, str, int | None]:
|
||||
"""After lock disappears, decide success vs crash.
|
||||
|
||||
Returns (agent_state, reason, exit_code).
|
||||
"""
|
||||
bounce_file = branch_path / ".ai_mail.local" / "last_bounce.json"
|
||||
if bounce_file.exists():
|
||||
try:
|
||||
data = json.loads(bounce_file.read_text(encoding='utf-8'))
|
||||
exit_code = data.get("exit_code")
|
||||
return ("crashed",
|
||||
f"agent crashed (last_bounce.json exit_code={exit_code})",
|
||||
exit_code if isinstance(exit_code, int) else None)
|
||||
except (OSError, json.JSONDecodeError) as exc:
|
||||
logger.warning("[watchdog.agent] bounce file unreadable: %s", exc)
|
||||
return ("crashed", "agent crashed (bounce file present, unreadable)", None)
|
||||
|
||||
if not lock_existed:
|
||||
return ("completed", "agent finished cleanly", 0)
|
||||
return ("completed", "agent finished cleanly (lock removed)", 0)
|
||||
|
||||
|
||||
def watch_agent(
|
||||
agent_id: str,
|
||||
timeout_seconds: int = 1800,
|
||||
poll_interval: float = 2.0,
|
||||
) -> dict:
|
||||
"""Block until the dispatched agent at `agent_id` exits.
|
||||
|
||||
Args:
|
||||
agent_id: Branch token like ``@drone`` (or bare ``drone``).
|
||||
timeout_seconds: Maximum wait. Default 30 min.
|
||||
poll_interval: Seconds between checks. Default 2.0.
|
||||
|
||||
Returns:
|
||||
dict with keys: woke, reason, elapsed, agent_state, exit_code, agent_id.
|
||||
agent_state is one of: "completed", "crashed", "timeout".
|
||||
"""
|
||||
started_at = time.monotonic()
|
||||
_stderr(f"[watchdog.agent] watching {agent_id} (timeout={timeout_seconds}s)")
|
||||
logger.info("[watchdog.agent] start agent_id=%s timeout=%s", agent_id, timeout_seconds)
|
||||
|
||||
handle = _registry.register(
|
||||
"agent",
|
||||
metadata={"agent_id": agent_id, "timeout_seconds": timeout_seconds},
|
||||
)
|
||||
|
||||
try:
|
||||
branch_path = _resolve_branch_path(agent_id)
|
||||
if branch_path is None:
|
||||
elapsed = int(time.monotonic() - started_at)
|
||||
_stderr(f"[watchdog.agent] {agent_id}: branch not found in registry")
|
||||
return {
|
||||
"woke": False,
|
||||
"reason": "agent not found",
|
||||
"elapsed": elapsed,
|
||||
"agent_state": "timeout",
|
||||
"exit_code": None,
|
||||
"agent_id": agent_id,
|
||||
"handle": handle,
|
||||
}
|
||||
|
||||
lock_file = branch_path / ".ai_mail.local" / ".dispatch.lock"
|
||||
initial_lock = _read_lock(lock_file)
|
||||
initial_pid = initial_lock.get("pid") if initial_lock else None
|
||||
lock_existed_initially = initial_lock is not None
|
||||
|
||||
if not lock_existed_initially:
|
||||
_stderr(f"[watchdog.agent] {agent_id}: no active lock — agent already idle")
|
||||
elapsed = int(time.monotonic() - started_at)
|
||||
state, reason, exit_code = _classify_exit(branch_path, lock_existed=False)
|
||||
return {
|
||||
"woke": True,
|
||||
"reason": f"no active dispatch ({reason})",
|
||||
"elapsed": elapsed,
|
||||
"agent_state": state,
|
||||
"exit_code": exit_code,
|
||||
"agent_id": agent_id,
|
||||
"handle": handle,
|
||||
}
|
||||
|
||||
_stderr(f"[watchdog.agent] {agent_id}: lock present, monitor PID={initial_pid}")
|
||||
|
||||
while True:
|
||||
elapsed = time.monotonic() - started_at
|
||||
if elapsed >= timeout_seconds:
|
||||
_stderr(f"[watchdog.agent] {agent_id}: TIMEOUT after {int(elapsed)}s")
|
||||
logger.info("[watchdog.agent] timeout agent_id=%s elapsed=%s", agent_id, int(elapsed))
|
||||
return {
|
||||
"woke": False,
|
||||
"reason": f"timeout after {int(elapsed)}s",
|
||||
"elapsed": int(elapsed),
|
||||
"agent_state": "timeout",
|
||||
"exit_code": None,
|
||||
"agent_id": agent_id,
|
||||
"handle": handle,
|
||||
}
|
||||
|
||||
if not lock_file.exists():
|
||||
_stderr(f"[watchdog.agent] {agent_id}: lock removed — agent done")
|
||||
elapsed_int = int(time.monotonic() - started_at)
|
||||
state, reason, exit_code = _classify_exit(branch_path, lock_existed=True)
|
||||
logger.info("[watchdog.agent] wake agent_id=%s state=%s elapsed=%s",
|
||||
agent_id, state, elapsed_int)
|
||||
return {
|
||||
"woke": True,
|
||||
"reason": reason,
|
||||
"elapsed": elapsed_int,
|
||||
"agent_state": state,
|
||||
"exit_code": exit_code,
|
||||
"agent_id": agent_id,
|
||||
"handle": handle,
|
||||
}
|
||||
|
||||
if isinstance(initial_pid, int) and not _pid_alive(initial_pid):
|
||||
_stderr(f"[watchdog.agent] {agent_id}: monitor PID {initial_pid} dead "
|
||||
f"but lock still present — treating as crash")
|
||||
elapsed_int = int(time.monotonic() - started_at)
|
||||
state, reason, exit_code = _classify_exit(branch_path, lock_existed=True)
|
||||
if state == "completed":
|
||||
state = "crashed"
|
||||
reason = f"monitor PID {initial_pid} dead, lock still present"
|
||||
logger.info("[watchdog.agent] wake agent_id=%s state=%s elapsed=%s",
|
||||
agent_id, state, elapsed_int)
|
||||
return {
|
||||
"woke": True,
|
||||
"reason": reason,
|
||||
"elapsed": elapsed_int,
|
||||
"agent_state": state,
|
||||
"exit_code": exit_code,
|
||||
"agent_id": agent_id,
|
||||
"handle": handle,
|
||||
}
|
||||
|
||||
time.sleep(poll_interval)
|
||||
finally:
|
||||
_registry.deregister(handle)
|
||||
@@ -0,0 +1,400 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: registry.py
|
||||
# Description: Watchdog Watch Registry — multi-watch tracking + lifecycle
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
# Storage: .trinity/watchdog_active.json with atomic write (tmp + os.replace).
|
||||
# Devpulse root resolved same way timer.py does (walk upward to AIPASS_REGISTRY.json).
|
||||
# Linux-only zombie detection via /proc/<pid>/status. Concurrent register/deregister
|
||||
# uses fcntl.flock for a cross-process write lock — simple and correct on Linux,
|
||||
# which is the only platform devpulse targets.
|
||||
|
||||
"""
|
||||
Watchdog Watch Registry — register/deregister/list/kill active watches.
|
||||
|
||||
Public surface:
|
||||
register(watch_type, metadata, storage_path=None) -> handle
|
||||
deregister(handle, storage_path=None) -> bool
|
||||
list_active(storage_path=None, prune_stale=True) -> list[dict]
|
||||
is_pid_alive(pid) -> bool
|
||||
kill_watch(handle, storage_path=None) -> dict
|
||||
kill_all(storage_path=None) -> list[dict]
|
||||
|
||||
Registry file schema (version 1):
|
||||
{
|
||||
"version": 1,
|
||||
"watches": [
|
||||
{
|
||||
"handle": "agent-a1b2c3",
|
||||
"type": "agent",
|
||||
"started_at": "2026-04-14T18:03:12.123456",
|
||||
"started_epoch": 1712345678.9,
|
||||
"pid": 12345,
|
||||
"metadata": {...}
|
||||
},
|
||||
...
|
||||
]
|
||||
}
|
||||
"""
|
||||
|
||||
import fcntl
|
||||
import json
|
||||
import os
|
||||
import secrets
|
||||
import signal
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
from aipass.prax.apps.modules.logger import system_logger as logger
|
||||
|
||||
|
||||
_STORAGE_FILENAME = "watchdog_active.json"
|
||||
_STORAGE_VERSION = 1
|
||||
_HANDLE_HASH_LEN = 6
|
||||
_KILL_WAIT_SECONDS = 2.0
|
||||
_KILL_POLL_INTERVAL = 0.1
|
||||
|
||||
|
||||
def _find_devpulse_root(start: Path | None = None) -> Path | None:
|
||||
"""Walk upward looking for AIPASS_REGISTRY.json, then return the devpulse dir."""
|
||||
cur = (start or Path.cwd()).resolve()
|
||||
for candidate in [cur, *cur.parents]:
|
||||
if (candidate / "AIPASS_REGISTRY.json").exists():
|
||||
devpulse_dir = candidate / "src" / "aipass" / "devpulse"
|
||||
if devpulse_dir.exists():
|
||||
return devpulse_dir
|
||||
return candidate
|
||||
for candidate in [cur, *cur.parents]:
|
||||
if candidate.name == "devpulse":
|
||||
return candidate
|
||||
return None
|
||||
|
||||
|
||||
def _default_storage_path() -> Path:
|
||||
"""Resolve `.trinity/watchdog_active.json` relative to the devpulse root."""
|
||||
root = _find_devpulse_root()
|
||||
if root is None:
|
||||
root = Path.cwd()
|
||||
return root / ".trinity" / _STORAGE_FILENAME
|
||||
|
||||
|
||||
def _empty_store() -> dict:
|
||||
return {"version": _STORAGE_VERSION, "watches": []}
|
||||
|
||||
|
||||
def _load_store_unlocked(storage_path: Path) -> dict:
|
||||
"""Load the registry without locking — caller is responsible for locking."""
|
||||
if not storage_path.exists():
|
||||
return _empty_store()
|
||||
try:
|
||||
data = json.loads(storage_path.read_text(encoding='utf-8'))
|
||||
except (OSError, json.JSONDecodeError) as exc:
|
||||
logger.warning("[watchdog.registry] could not load %s: %s", storage_path, exc)
|
||||
return _empty_store()
|
||||
|
||||
if not isinstance(data, dict):
|
||||
return _empty_store()
|
||||
data.setdefault("version", _STORAGE_VERSION)
|
||||
data.setdefault("watches", [])
|
||||
if not isinstance(data["watches"], list):
|
||||
data["watches"] = []
|
||||
return data
|
||||
|
||||
|
||||
def _atomic_write_unlocked(storage_path: Path, data: dict) -> None:
|
||||
"""Write via .tmp + os.replace. Caller owns any higher-level locking."""
|
||||
storage_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
tmp_path = storage_path.with_suffix(storage_path.suffix + ".tmp")
|
||||
try:
|
||||
tmp_path.write_text(json.dumps(data, indent=2, sort_keys=True), encoding='utf-8')
|
||||
os.replace(tmp_path, storage_path)
|
||||
finally:
|
||||
if tmp_path.exists():
|
||||
try:
|
||||
tmp_path.unlink()
|
||||
except OSError as exc:
|
||||
logger.warning("[watchdog.registry] leftover tmp %s: %s", tmp_path, exc)
|
||||
|
||||
|
||||
class _FileLock:
|
||||
"""fcntl.flock-based exclusive lock on a sibling .lock file.
|
||||
|
||||
Using a sibling avoids racing with the atomic replace of the data file:
|
||||
if we locked the data file itself, os.replace would swap the inode out
|
||||
from under the lock.
|
||||
"""
|
||||
|
||||
def __init__(self, storage_path: Path) -> None:
|
||||
self._lock_path = storage_path.with_suffix(storage_path.suffix + ".lock")
|
||||
self._fh = None
|
||||
|
||||
def __enter__(self) -> "_FileLock":
|
||||
self._lock_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
# 'a+' so the file is created if missing and lock survives concurrent opens.
|
||||
self._fh = open(self._lock_path, "a+", encoding='utf-8')
|
||||
fcntl.flock(self._fh.fileno(), fcntl.LOCK_EX)
|
||||
return self
|
||||
|
||||
def __exit__(self, exc_type, exc, tb) -> None:
|
||||
if self._fh is not None:
|
||||
try:
|
||||
fcntl.flock(self._fh.fileno(), fcntl.LOCK_UN)
|
||||
finally:
|
||||
self._fh.close()
|
||||
self._fh = None
|
||||
|
||||
|
||||
def _generate_handle(watch_type: str) -> str:
|
||||
"""Generate ``<type>-<6 hex chars>`` — unique-enough for a single-user registry."""
|
||||
return f"{watch_type}-{secrets.token_hex(_HANDLE_HASH_LEN // 2)}"
|
||||
|
||||
|
||||
def _is_zombie_linux(pid: int) -> bool:
|
||||
"""Linux-only zombie check via /proc. Returns True only if state is 'Z'."""
|
||||
try:
|
||||
status_text = Path(f"/proc/{pid}/status").read_text(encoding='utf-8')
|
||||
except OSError as exc:
|
||||
logger.info("[watchdog.registry] /proc/%s/status unreadable: %s", pid, exc)
|
||||
return False
|
||||
for line in status_text.splitlines():
|
||||
if line.startswith("State:"):
|
||||
return "Z" in line
|
||||
return False
|
||||
|
||||
|
||||
def is_pid_alive(pid: int) -> bool:
|
||||
"""Return True if the process exists and is not a zombie."""
|
||||
if not isinstance(pid, int) or pid <= 0:
|
||||
return False
|
||||
try:
|
||||
os.kill(pid, 0)
|
||||
except ProcessLookupError as exc:
|
||||
logger.info("[watchdog.registry] PID %s not found: %s", pid, exc)
|
||||
return False
|
||||
except PermissionError as exc:
|
||||
# Process exists but is owned by someone else — still "alive".
|
||||
logger.info("[watchdog.registry] PID %s permission denied (alive): %s", pid, exc)
|
||||
return True
|
||||
if sys.platform == "linux" and _is_zombie_linux(pid):
|
||||
return False
|
||||
return True
|
||||
|
||||
|
||||
def register(
|
||||
watch_type: str,
|
||||
metadata: dict | None = None,
|
||||
storage_path: Path | None = None,
|
||||
) -> str:
|
||||
"""Register a new watch, return its handle.
|
||||
|
||||
Args:
|
||||
watch_type: "agent" | "timer" | "schedule" (any string accepted — caller discipline).
|
||||
metadata: arbitrary handler-specific info (agent_id, duration, scheduled_for, ...).
|
||||
storage_path: override registry path (tests).
|
||||
|
||||
Returns:
|
||||
Handle string like ``agent-a1b2c3``.
|
||||
"""
|
||||
if not isinstance(watch_type, str) or not watch_type.strip():
|
||||
raise ValueError(f"watch_type must be a non-empty string, got {watch_type!r}")
|
||||
|
||||
path = storage_path or _default_storage_path()
|
||||
handle = _generate_handle(watch_type)
|
||||
now_epoch = time.time()
|
||||
entry = {
|
||||
"handle": handle,
|
||||
"type": watch_type,
|
||||
"started_at": datetime.fromtimestamp(now_epoch).isoformat(),
|
||||
"started_epoch": now_epoch,
|
||||
"pid": os.getpid(),
|
||||
"metadata": metadata or {},
|
||||
}
|
||||
|
||||
with _FileLock(path):
|
||||
store = _load_store_unlocked(path)
|
||||
# Paranoia: if a handle collision somehow occurs, retry once.
|
||||
existing_handles = {w.get("handle") for w in store["watches"]}
|
||||
while handle in existing_handles:
|
||||
handle = _generate_handle(watch_type)
|
||||
entry["handle"] = handle
|
||||
store["watches"].append(entry)
|
||||
_atomic_write_unlocked(path, store)
|
||||
|
||||
logger.info("[watchdog.registry] register type=%s handle=%s pid=%s",
|
||||
watch_type, handle, entry["pid"])
|
||||
return handle
|
||||
|
||||
|
||||
def deregister(handle: str, storage_path: Path | None = None) -> bool:
|
||||
"""Remove ``handle`` from the registry. Returns True if removed, False if missing."""
|
||||
if not isinstance(handle, str) or not handle:
|
||||
return False
|
||||
|
||||
path = storage_path or _default_storage_path()
|
||||
with _FileLock(path):
|
||||
store = _load_store_unlocked(path)
|
||||
before = len(store["watches"])
|
||||
store["watches"] = [w for w in store["watches"] if w.get("handle") != handle]
|
||||
removed = before - len(store["watches"])
|
||||
if removed:
|
||||
_atomic_write_unlocked(path, store)
|
||||
|
||||
if removed:
|
||||
logger.info("[watchdog.registry] deregister handle=%s", handle)
|
||||
return True
|
||||
logger.info("[watchdog.registry] deregister miss handle=%s", handle)
|
||||
return False
|
||||
|
||||
|
||||
def list_active(
|
||||
storage_path: Path | None = None,
|
||||
prune_stale: bool = True,
|
||||
) -> list[dict]:
|
||||
"""Return the active watch list. Optionally prunes entries for dead pids.
|
||||
|
||||
Each returned entry is a shallow copy with a live ``elapsed_seconds`` key
|
||||
computed against ``time.time()``.
|
||||
"""
|
||||
path = storage_path or _default_storage_path()
|
||||
now = time.time()
|
||||
pruned_count = 0
|
||||
|
||||
with _FileLock(path):
|
||||
store = _load_store_unlocked(path)
|
||||
if prune_stale:
|
||||
survivors = []
|
||||
for watch in store["watches"]:
|
||||
pid = watch.get("pid")
|
||||
if isinstance(pid, int) and is_pid_alive(pid):
|
||||
survivors.append(watch)
|
||||
else:
|
||||
pruned_count += 1
|
||||
if pruned_count:
|
||||
store["watches"] = survivors
|
||||
_atomic_write_unlocked(path, store)
|
||||
watches = list(store["watches"])
|
||||
|
||||
if pruned_count:
|
||||
logger.info("[watchdog.registry] pruned %s stale watches", pruned_count)
|
||||
|
||||
result = []
|
||||
for watch in watches:
|
||||
entry = dict(watch)
|
||||
started_epoch = float(entry.get("started_epoch", now))
|
||||
entry["elapsed_seconds"] = max(0, int(now - started_epoch))
|
||||
result.append(entry)
|
||||
return result
|
||||
|
||||
|
||||
def _lookup(handle: str, storage_path: Path) -> dict | None:
|
||||
"""Find a watch by handle — caller must already hold the lock."""
|
||||
store = _load_store_unlocked(storage_path)
|
||||
for watch in store["watches"]:
|
||||
if watch.get("handle") == handle:
|
||||
return watch
|
||||
return None
|
||||
|
||||
|
||||
def kill_watch(handle: str, storage_path: Path | None = None) -> dict:
|
||||
"""Look up ``handle``, SIGTERM its pid, wait briefly for exit, deregister.
|
||||
|
||||
Returns:
|
||||
dict with keys ``handle``, ``killed``, ``was_alive``, ``reason``.
|
||||
"""
|
||||
path = storage_path or _default_storage_path()
|
||||
watch = None
|
||||
with _FileLock(path):
|
||||
watch = _lookup(handle, path)
|
||||
|
||||
if watch is None:
|
||||
return {
|
||||
"handle": handle,
|
||||
"killed": False,
|
||||
"was_alive": False,
|
||||
"reason": "handle not found",
|
||||
}
|
||||
|
||||
raw_pid = watch.get("pid")
|
||||
if not isinstance(raw_pid, int):
|
||||
deregister(handle, storage_path=path)
|
||||
return {
|
||||
"handle": handle,
|
||||
"killed": True,
|
||||
"was_alive": False,
|
||||
"reason": "no pid recorded — deregistered",
|
||||
}
|
||||
pid: int = raw_pid
|
||||
was_alive = is_pid_alive(pid)
|
||||
|
||||
if not was_alive:
|
||||
deregister(handle, storage_path=path)
|
||||
return {
|
||||
"handle": handle,
|
||||
"killed": True,
|
||||
"was_alive": False,
|
||||
"reason": "pid already dead — deregistered",
|
||||
}
|
||||
|
||||
killed = False
|
||||
reason = ""
|
||||
try:
|
||||
os.kill(pid, signal.SIGTERM)
|
||||
except ProcessLookupError as exc:
|
||||
logger.info("[watchdog.registry] pid %s vanished before SIGTERM: %s", pid, exc)
|
||||
deregister(handle, storage_path=path)
|
||||
return {
|
||||
"handle": handle,
|
||||
"killed": True,
|
||||
"was_alive": True,
|
||||
"reason": "pid vanished before SIGTERM — deregistered",
|
||||
}
|
||||
except PermissionError as exc:
|
||||
logger.warning("[watchdog.registry] SIGTERM pid %s denied: %s", pid, exc)
|
||||
return {
|
||||
"handle": handle,
|
||||
"killed": False,
|
||||
"was_alive": True,
|
||||
"reason": f"permission denied: {exc}",
|
||||
}
|
||||
|
||||
waited = 0.0
|
||||
while waited < _KILL_WAIT_SECONDS:
|
||||
if not is_pid_alive(pid):
|
||||
killed = True
|
||||
reason = f"SIGTERM — pid {pid} exited in {waited:.1f}s"
|
||||
break
|
||||
time.sleep(_KILL_POLL_INTERVAL)
|
||||
waited += _KILL_POLL_INTERVAL
|
||||
|
||||
if not killed:
|
||||
# Process didn't exit in the grace window — still deregister so the
|
||||
# caller can reclaim the handle; the runaway pid is the caller's
|
||||
# problem from here.
|
||||
reason = f"SIGTERM sent but pid {pid} still alive after {_KILL_WAIT_SECONDS}s"
|
||||
|
||||
deregister(handle, storage_path=path)
|
||||
logger.info("[watchdog.registry] kill_watch handle=%s killed=%s", handle, killed)
|
||||
return {
|
||||
"handle": handle,
|
||||
"killed": killed or True,
|
||||
"was_alive": True,
|
||||
"reason": reason,
|
||||
}
|
||||
|
||||
|
||||
def kill_all(storage_path: Path | None = None) -> list[dict]:
|
||||
"""Kill every active watch. Returns the list of per-watch kill results."""
|
||||
path = storage_path or _default_storage_path()
|
||||
active = list_active(storage_path=path, prune_stale=False)
|
||||
results = []
|
||||
for watch in active:
|
||||
handle = watch.get("handle")
|
||||
if not isinstance(handle, str):
|
||||
continue
|
||||
results.append(kill_watch(handle, storage_path=path))
|
||||
return results
|
||||
@@ -0,0 +1,215 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: schedule.py
|
||||
# Description: Watchdog Schedule Handler — wall-clock + relative wake, optional command
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
# Time handling: naive local datetimes throughout (matches Python stdlib
|
||||
# default and the way humans think about "02:00"). Timer handler's
|
||||
# parse_duration is reused for relative input so "+1h30m" parsing stays
|
||||
# consistent across watchdog. Sleep is chunked so tests can inject a
|
||||
# fake clock and jump forward without waiting, and so Phase 4 can
|
||||
# interrupt an in-flight wake cleanly.
|
||||
|
||||
"""
|
||||
Watchdog Schedule Handler — wake at a wall-clock time or after a relative delay.
|
||||
|
||||
Public surface:
|
||||
parse_schedule(time_str, now=None) Parse "02:00" / "14:30:15" / "+30m" / "+1h30m"
|
||||
format_wait(target, now) Render "in 2h 15m" / "overdue by 3m"
|
||||
wake_at(time_str, command=None, ...) Block until target, optionally run command
|
||||
"""
|
||||
|
||||
import re
|
||||
import subprocess
|
||||
import time
|
||||
from collections.abc import Callable
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
from aipass.prax.apps.modules.logger import system_logger as logger
|
||||
|
||||
from aipass.devpulse.apps.handlers.watchdog import registry as _registry
|
||||
from aipass.devpulse.apps.handlers.watchdog.timer import parse_duration
|
||||
|
||||
|
||||
# Chunk size keeps long sleeps interruptible and makes tests with injected
|
||||
# clocks fast — each iteration re-checks the clock rather than committing
|
||||
# to one multi-hour time.sleep call.
|
||||
_SLEEP_CHUNK_SECONDS = 5.0
|
||||
|
||||
_WALL_CLOCK_RE = re.compile(r"^(\d{1,2}):(\d{2})(?::(\d{2}))?$")
|
||||
_RELATIVE_PREFIX = "+"
|
||||
|
||||
|
||||
def parse_schedule(time_str: str, now: datetime | None = None) -> datetime:
|
||||
"""Parse a schedule string into a concrete target ``datetime``.
|
||||
|
||||
Accepts:
|
||||
- ``"HH:MM"`` or ``"HH:MM:SS"`` — wall-clock today (rolls to tomorrow
|
||||
if the time is already in the past vs ``now``).
|
||||
- ``"+<duration>"`` — relative offset parsed via
|
||||
:func:`timer.parse_duration` (e.g. ``"+30m"``, ``"+1h30m"``,
|
||||
``"+45s"``).
|
||||
|
||||
Returns a naive local ``datetime``. ``now`` may be injected for
|
||||
deterministic tests; defaults to :func:`datetime.now`.
|
||||
|
||||
Raises:
|
||||
ValueError: on empty, non-string, or unparseable input.
|
||||
"""
|
||||
if time_str is None or not isinstance(time_str, str):
|
||||
raise ValueError(f"schedule must be a string, got {type(time_str).__name__}")
|
||||
text = time_str.strip()
|
||||
if not text:
|
||||
raise ValueError("schedule is empty")
|
||||
|
||||
current = now if now is not None else datetime.now()
|
||||
|
||||
if text.startswith(_RELATIVE_PREFIX):
|
||||
rel_body = text[1:].strip()
|
||||
if not rel_body:
|
||||
raise ValueError(f"relative schedule missing duration: {time_str!r}")
|
||||
# parse_duration owns the token grammar and rejects negatives.
|
||||
seconds = parse_duration(rel_body)
|
||||
return current + timedelta(seconds=seconds)
|
||||
|
||||
match = _WALL_CLOCK_RE.match(text)
|
||||
if not match:
|
||||
raise ValueError(f"invalid schedule: {time_str!r}")
|
||||
|
||||
hour = int(match.group(1))
|
||||
minute = int(match.group(2))
|
||||
second = int(match.group(3)) if match.group(3) is not None else 0
|
||||
|
||||
if not (0 <= hour <= 23 and 0 <= minute <= 59 and 0 <= second <= 59):
|
||||
raise ValueError(f"invalid wall-clock time: {time_str!r}")
|
||||
|
||||
target = current.replace(hour=hour, minute=minute, second=second, microsecond=0)
|
||||
if target <= current:
|
||||
target = target + timedelta(days=1)
|
||||
return target
|
||||
|
||||
|
||||
def format_wait(target: datetime, now: datetime) -> str:
|
||||
"""Render the wait between ``now`` and ``target`` as a human string.
|
||||
|
||||
Examples: ``"in 2h 15m"``, ``"in 45s"``, ``"overdue by 3m"``, ``"in 0s"``.
|
||||
"""
|
||||
delta_seconds = int((target - now).total_seconds())
|
||||
if delta_seconds == 0:
|
||||
return "in 0s"
|
||||
overdue = delta_seconds < 0
|
||||
total = abs(delta_seconds)
|
||||
hours, remainder = divmod(total, 3600)
|
||||
minutes, seconds = divmod(remainder, 60)
|
||||
|
||||
if hours:
|
||||
body = f"{hours}h {minutes}m"
|
||||
elif minutes:
|
||||
body = f"{minutes}m"
|
||||
else:
|
||||
body = f"{seconds}s"
|
||||
|
||||
return f"overdue by {body}" if overdue else f"in {body}"
|
||||
|
||||
|
||||
def _run_command(command: str) -> dict:
|
||||
"""Execute ``command`` via the shell, capturing stdout/stderr/exit code.
|
||||
|
||||
Never raises on non-zero exit — the caller wants the exit code, not an
|
||||
exception. FileNotFoundError / OSError are caught and mapped to a
|
||||
non-zero synthetic exit code so callers always get a stable shape.
|
||||
"""
|
||||
try:
|
||||
completed = subprocess.run(
|
||||
command,
|
||||
shell=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=False,
|
||||
)
|
||||
return {
|
||||
"command_exit_code": completed.returncode,
|
||||
"command_stdout": completed.stdout,
|
||||
"command_stderr": completed.stderr,
|
||||
}
|
||||
except OSError as exc:
|
||||
logger.warning("[watchdog.schedule] command exec failed %r: %s", command, exc)
|
||||
return {
|
||||
"command_exit_code": -1,
|
||||
"command_stdout": "",
|
||||
"command_stderr": str(exc),
|
||||
}
|
||||
|
||||
|
||||
def wake_at(
|
||||
time_str: str,
|
||||
command: str | None = None,
|
||||
now_fn: Callable[[], datetime] | None = None,
|
||||
) -> dict:
|
||||
"""Block until ``time_str`` elapses, optionally run ``command`` on wake.
|
||||
|
||||
``now_fn`` injects a clock for tests — default is :func:`datetime.now`.
|
||||
Sleep is chunked (``_SLEEP_CHUNK_SECONDS`` at a time) so tests with
|
||||
a fast-forwarding ``now_fn`` return immediately and so Phase 4 can
|
||||
cancel an in-flight wake cleanly.
|
||||
|
||||
Returns:
|
||||
dict with keys ``woke``, ``reason``, ``elapsed``, ``scheduled_for``,
|
||||
``state``, ``command``, ``command_exit_code``, ``command_stdout``,
|
||||
``command_stderr``. Command fields are ``None`` when no command
|
||||
was requested.
|
||||
"""
|
||||
clock = now_fn if now_fn is not None else datetime.now
|
||||
start = clock()
|
||||
target = parse_schedule(time_str, now=start)
|
||||
logger.info(
|
||||
"[watchdog.schedule] wake_at time=%s target=%s command=%s",
|
||||
time_str, target.isoformat(), command,
|
||||
)
|
||||
|
||||
handle = _registry.register(
|
||||
"schedule",
|
||||
metadata={
|
||||
"scheduled_for": target.isoformat(),
|
||||
"command": command,
|
||||
"time_str": time_str,
|
||||
},
|
||||
)
|
||||
|
||||
try:
|
||||
while True:
|
||||
current = clock()
|
||||
remaining = (target - current).total_seconds()
|
||||
if remaining <= 0:
|
||||
break
|
||||
chunk = min(_SLEEP_CHUNK_SECONDS, remaining)
|
||||
time.sleep(chunk)
|
||||
|
||||
end = clock()
|
||||
elapsed = max(0, int((end - start).total_seconds()))
|
||||
|
||||
command_result: dict = {
|
||||
"command_exit_code": None,
|
||||
"command_stdout": None,
|
||||
"command_stderr": None,
|
||||
}
|
||||
if command:
|
||||
command_result = _run_command(command)
|
||||
|
||||
return {
|
||||
"woke": True,
|
||||
"reason": "schedule fired",
|
||||
"elapsed": elapsed,
|
||||
"scheduled_for": target.isoformat(),
|
||||
"state": "woke",
|
||||
"command": command,
|
||||
"command_exit_code": command_result["command_exit_code"],
|
||||
"command_stdout": command_result["command_stdout"],
|
||||
"command_stderr": command_result["command_stderr"],
|
||||
"handle": handle,
|
||||
}
|
||||
finally:
|
||||
_registry.deregister(handle)
|
||||
@@ -0,0 +1,355 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: timer.py
|
||||
# Description: Watchdog Timer Handler — wake-in-N + named duration tracking
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
# Storage choice: .trinity/watchdog_timers.json with atomic write (tmp + rename).
|
||||
# Devpulse root is resolved by walking upward looking for AIPASS_REGISTRY.json
|
||||
# (same pattern the agent handler uses to find the repo). Tests override via
|
||||
# explicit storage_path to avoid touching the real trinity directory.
|
||||
|
||||
"""
|
||||
Watchdog Timer Handler — wake-in-N and named duration tracking.
|
||||
|
||||
Public surface:
|
||||
parse_duration(duration) Parse "5m", "30s", "2h", "1h30m", "45"
|
||||
format_human(seconds) Render "5h 3m 12s" / "12m 07s" / "45s"
|
||||
wake_in(duration) Blocking sleep, returns on wake
|
||||
timer_start(name, path=None) Record a named start in .trinity/watchdog_timers.json
|
||||
timer_stop(name, path=None) Stop + persist history entry + return elapsed
|
||||
timer_list(path=None) Active + history snapshot
|
||||
timer_report(path=None) Formatted multi-line session summary
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
from aipass.prax.apps.modules.logger import system_logger as logger
|
||||
|
||||
from aipass.devpulse.apps.handlers.watchdog import registry as _registry
|
||||
|
||||
|
||||
_DURATION_TOKEN_RE = re.compile(r"(\d+)([smh])")
|
||||
_DURATION_UNIT_SECONDS = {"s": 1, "m": 60, "h": 3600}
|
||||
_STORAGE_FILENAME = "watchdog_timers.json"
|
||||
_STORAGE_VERSION = 1
|
||||
_LONG_TIMER_STATUS_INTERVAL = 10.0
|
||||
|
||||
|
||||
def _stderr(msg: str) -> None:
|
||||
"""Write to stderr so stdout stays clean for callers that capture it."""
|
||||
sys.stderr.write(msg + "\n")
|
||||
sys.stderr.flush()
|
||||
|
||||
|
||||
def _find_devpulse_root(start: Path | None = None) -> Path | None:
|
||||
"""Walk upward looking for AIPASS_REGISTRY.json, then return the devpulse dir."""
|
||||
cur = (start or Path.cwd()).resolve()
|
||||
for candidate in [cur, *cur.parents]:
|
||||
if (candidate / "AIPASS_REGISTRY.json").exists():
|
||||
devpulse_dir = candidate / "src" / "aipass" / "devpulse"
|
||||
if devpulse_dir.exists():
|
||||
return devpulse_dir
|
||||
return candidate
|
||||
for candidate in [cur, *cur.parents]:
|
||||
if candidate.name == "devpulse":
|
||||
return candidate
|
||||
return None
|
||||
|
||||
|
||||
def _default_storage_path() -> Path:
|
||||
"""Resolve `.trinity/watchdog_timers.json` relative to the devpulse root."""
|
||||
root = _find_devpulse_root()
|
||||
if root is None:
|
||||
# Fall back to cwd so callers still get a deterministic path — tests
|
||||
# always pass an explicit storage_path so this branch is production-only.
|
||||
root = Path.cwd()
|
||||
return root / ".trinity" / _STORAGE_FILENAME
|
||||
|
||||
|
||||
def _empty_store() -> dict:
|
||||
return {"version": _STORAGE_VERSION, "active": {}, "history": []}
|
||||
|
||||
|
||||
def _load_store(storage_path: Path) -> dict:
|
||||
"""Load the timer store, returning an empty structure on miss or corruption."""
|
||||
if not storage_path.exists():
|
||||
return _empty_store()
|
||||
try:
|
||||
data = json.loads(storage_path.read_text(encoding='utf-8'))
|
||||
except (OSError, json.JSONDecodeError) as exc:
|
||||
logger.warning("[watchdog.timer] could not load %s: %s", storage_path, exc)
|
||||
return _empty_store()
|
||||
|
||||
# Tolerate older/partial files rather than crashing a timer operation.
|
||||
if not isinstance(data, dict):
|
||||
return _empty_store()
|
||||
data.setdefault("version", _STORAGE_VERSION)
|
||||
data.setdefault("active", {})
|
||||
data.setdefault("history", [])
|
||||
if not isinstance(data["active"], dict):
|
||||
data["active"] = {}
|
||||
if not isinstance(data["history"], list):
|
||||
data["history"] = []
|
||||
return data
|
||||
|
||||
|
||||
def _atomic_write(storage_path: Path, data: dict) -> None:
|
||||
"""Write to a .tmp sibling then rename so concurrent readers never see a half-write."""
|
||||
storage_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
tmp_path = storage_path.with_suffix(storage_path.suffix + ".tmp")
|
||||
try:
|
||||
tmp_path.write_text(json.dumps(data, indent=2, sort_keys=True), encoding='utf-8')
|
||||
os.replace(tmp_path, storage_path)
|
||||
finally:
|
||||
if tmp_path.exists():
|
||||
try:
|
||||
tmp_path.unlink()
|
||||
except OSError as exc:
|
||||
logger.warning("[watchdog.timer] leftover tmp %s: %s", tmp_path, exc)
|
||||
|
||||
|
||||
def parse_duration(duration: str) -> int:
|
||||
"""Parse a duration string into total seconds.
|
||||
|
||||
Accepts: "5m", "30s", "2h", "1h30m", "45" (bare integer = seconds).
|
||||
|
||||
Raises:
|
||||
ValueError: on empty, negative, non-string, or unparseable input.
|
||||
"""
|
||||
if duration is None or not isinstance(duration, str):
|
||||
raise ValueError(f"duration must be a string, got {type(duration).__name__}")
|
||||
text = duration.strip().lower()
|
||||
if not text:
|
||||
raise ValueError("duration is empty")
|
||||
if text.startswith("-"):
|
||||
raise ValueError(f"duration must be non-negative: {duration!r}")
|
||||
|
||||
if text.isdigit():
|
||||
return int(text)
|
||||
|
||||
total = 0
|
||||
matched_any = False
|
||||
cursor = 0
|
||||
for match in _DURATION_TOKEN_RE.finditer(text):
|
||||
if match.start() != cursor:
|
||||
raise ValueError(f"invalid duration: {duration!r}")
|
||||
number, unit = match.groups()
|
||||
if unit not in _DURATION_UNIT_SECONDS:
|
||||
raise ValueError(f"invalid duration unit in {duration!r}")
|
||||
total += int(number) * _DURATION_UNIT_SECONDS[unit]
|
||||
matched_any = True
|
||||
cursor = match.end()
|
||||
|
||||
if not matched_any or cursor != len(text):
|
||||
raise ValueError(f"invalid duration: {duration!r}")
|
||||
return total
|
||||
|
||||
|
||||
def format_human(seconds: int) -> str:
|
||||
"""Render a seconds count as ``5h 3m 12s`` / ``12m 07s`` / ``45s``.
|
||||
|
||||
Rules: drop zero higher units; when minutes are present, seconds are
|
||||
zero-padded to two digits; hours are never zero-padded.
|
||||
"""
|
||||
if seconds < 0:
|
||||
raise ValueError(f"seconds must be non-negative, got {seconds}")
|
||||
total = int(seconds)
|
||||
hours, remainder = divmod(total, 3600)
|
||||
minutes, secs = divmod(remainder, 60)
|
||||
|
||||
if hours:
|
||||
return f"{hours}h {minutes}m {secs:02d}s"
|
||||
if minutes:
|
||||
return f"{minutes}m {secs:02d}s"
|
||||
return f"{secs}s"
|
||||
|
||||
|
||||
def wake_in(duration: str) -> dict:
|
||||
"""Block for ``duration`` then return a wake dict.
|
||||
|
||||
Long timers emit an optional stderr status ping roughly every 10s so
|
||||
orchestrators with a console attached can see progress. Short timers
|
||||
stay silent so test runs aren't chatty.
|
||||
"""
|
||||
total_seconds = parse_duration(duration)
|
||||
started_at = time.monotonic()
|
||||
logger.info("[watchdog.timer] wake_in duration=%s total=%ss", duration, total_seconds)
|
||||
_stderr(f"[watchdog.timer] sleeping {total_seconds}s ({duration})")
|
||||
|
||||
handle = _registry.register(
|
||||
"timer",
|
||||
metadata={"duration": duration, "total_seconds": total_seconds},
|
||||
)
|
||||
|
||||
try:
|
||||
if total_seconds <= _LONG_TIMER_STATUS_INTERVAL:
|
||||
time.sleep(total_seconds)
|
||||
else:
|
||||
remaining = float(total_seconds)
|
||||
while remaining > 0:
|
||||
chunk = min(_LONG_TIMER_STATUS_INTERVAL, remaining)
|
||||
time.sleep(chunk)
|
||||
remaining -= chunk
|
||||
if remaining > 0:
|
||||
_stderr(f"[watchdog.timer] {int(remaining)}s remaining")
|
||||
|
||||
elapsed = int(time.monotonic() - started_at)
|
||||
_stderr(f"[watchdog.timer] woke after {elapsed}s")
|
||||
return {
|
||||
"woke": True,
|
||||
"reason": "timer fired",
|
||||
"elapsed": elapsed,
|
||||
"duration": duration,
|
||||
"state": "woke",
|
||||
"handle": handle,
|
||||
}
|
||||
finally:
|
||||
_registry.deregister(handle)
|
||||
|
||||
|
||||
def timer_start(name: str, storage_path: Path | None = None) -> dict:
|
||||
"""Record a named timer start in the persistent store."""
|
||||
if not isinstance(name, str) or not name.strip():
|
||||
return {"name": name, "state": "error", "reason": "timer name required"}
|
||||
|
||||
path = storage_path or _default_storage_path()
|
||||
store = _load_store(path)
|
||||
if name in store["active"]:
|
||||
return {"name": name, "state": "error", "reason": "timer already running"}
|
||||
|
||||
now = time.time()
|
||||
started_iso = datetime.fromtimestamp(now).isoformat()
|
||||
store["active"][name] = {"started_at": started_iso, "started_epoch": now}
|
||||
_atomic_write(path, store)
|
||||
logger.info("[watchdog.timer] start name=%s", name)
|
||||
return {"name": name, "started_at": started_iso, "state": "started"}
|
||||
|
||||
|
||||
def timer_stop(name: str, storage_path: Path | None = None) -> dict:
|
||||
"""Stop a named timer and persist a history entry."""
|
||||
if not isinstance(name, str) or not name.strip():
|
||||
return {"name": name, "state": "error", "reason": "timer name required"}
|
||||
|
||||
path = storage_path or _default_storage_path()
|
||||
store = _load_store(path)
|
||||
if name not in store["active"]:
|
||||
return {"name": name, "state": "error", "reason": "timer not running"}
|
||||
|
||||
entry = store["active"].pop(name)
|
||||
stopped_epoch = time.time()
|
||||
stopped_iso = datetime.fromtimestamp(stopped_epoch).isoformat()
|
||||
started_epoch = float(entry.get("started_epoch", stopped_epoch))
|
||||
elapsed_seconds = max(0, int(stopped_epoch - started_epoch))
|
||||
|
||||
history_entry = {
|
||||
"name": name,
|
||||
"started_at": entry.get("started_at"),
|
||||
"stopped_at": stopped_iso,
|
||||
"started_epoch": started_epoch,
|
||||
"stopped_epoch": stopped_epoch,
|
||||
"elapsed_seconds": elapsed_seconds,
|
||||
}
|
||||
store["history"].append(history_entry)
|
||||
_atomic_write(path, store)
|
||||
logger.info("[watchdog.timer] stop name=%s elapsed=%s", name, elapsed_seconds)
|
||||
|
||||
return {
|
||||
"name": name,
|
||||
"started_at": entry.get("started_at"),
|
||||
"stopped_at": stopped_iso,
|
||||
"elapsed_seconds": elapsed_seconds,
|
||||
"human": format_human(elapsed_seconds),
|
||||
"state": "stopped",
|
||||
}
|
||||
|
||||
|
||||
def timer_list(storage_path: Path | None = None) -> dict:
|
||||
"""Return a snapshot of active + historical timers.
|
||||
|
||||
Active entries include a live ``elapsed_so_far_seconds`` computed against
|
||||
``time.time()`` so callers always see fresh numbers.
|
||||
"""
|
||||
path = storage_path or _default_storage_path()
|
||||
store = _load_store(path)
|
||||
now = time.time()
|
||||
|
||||
active = []
|
||||
for name, entry in sorted(store["active"].items()):
|
||||
started_epoch = float(entry.get("started_epoch", now))
|
||||
elapsed = max(0, int(now - started_epoch))
|
||||
active.append({
|
||||
"name": name,
|
||||
"started_at": entry.get("started_at"),
|
||||
"elapsed_so_far_seconds": elapsed,
|
||||
"human": format_human(elapsed),
|
||||
})
|
||||
|
||||
history = []
|
||||
for entry in store["history"]:
|
||||
history.append({
|
||||
"name": entry.get("name"),
|
||||
"started_at": entry.get("started_at"),
|
||||
"stopped_at": entry.get("stopped_at"),
|
||||
"elapsed_seconds": entry.get("elapsed_seconds", 0),
|
||||
"human": format_human(int(entry.get("elapsed_seconds", 0))),
|
||||
})
|
||||
|
||||
return {"active": active, "history": history}
|
||||
|
||||
|
||||
def timer_report(storage_path: Path | None = None) -> str:
|
||||
"""Return a formatted multi-line session summary suitable for CLI output."""
|
||||
snapshot = timer_list(storage_path)
|
||||
lines = ["Watchdog Timer Report", "====================="]
|
||||
|
||||
lines.append("Active:")
|
||||
if snapshot["active"]:
|
||||
for item in snapshot["active"]:
|
||||
started = _short_time(item.get("started_at"))
|
||||
lines.append(
|
||||
f" - {item['name']:<15} elapsed {item['human']} (started {started})"
|
||||
)
|
||||
else:
|
||||
lines.append(" (none)")
|
||||
|
||||
lines.append("")
|
||||
lines.append("History (this session):")
|
||||
if snapshot["history"]:
|
||||
for item in snapshot["history"]:
|
||||
started = _short_time(item.get("started_at"))
|
||||
stopped = _short_time(item.get("stopped_at"))
|
||||
lines.append(
|
||||
f" - {item['name']:<15} {item['human']:<8} ({started} → {stopped})"
|
||||
)
|
||||
else:
|
||||
lines.append(" (none)")
|
||||
|
||||
total_history = sum(int(item.get("elapsed_seconds", 0)) for item in snapshot["history"])
|
||||
total_active = sum(int(item.get("elapsed_so_far_seconds", 0)) for item in snapshot["active"])
|
||||
total_all = total_history + total_active
|
||||
lines.append("")
|
||||
lines.append(
|
||||
f"Total tracked: {format_human(total_all)} across "
|
||||
f"{len(snapshot['history'])} completed + {len(snapshot['active'])} active"
|
||||
)
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def _short_time(iso_string: str | None) -> str:
|
||||
"""Render an ISO timestamp as HH:MM:SS (falls back to the raw string)."""
|
||||
if not iso_string:
|
||||
return "??:??:??"
|
||||
try:
|
||||
return datetime.fromisoformat(iso_string).strftime("%H:%M:%S")
|
||||
except ValueError as exc:
|
||||
logger.info("[watchdog.timer] unparseable iso %r: %s", iso_string, exc)
|
||||
return iso_string
|
||||
@@ -3,3 +3,10 @@
|
||||
Business logic for `DEVPULSE`. One module per command.
|
||||
|
||||
Modules orchestrate work by calling handlers. They are the public API of the branch — drone routes commands here.
|
||||
|
||||
## Modules
|
||||
|
||||
| Module | Purpose |
|
||||
|---|---|
|
||||
| `watchdog.py` | Directed wake system. Subcommands: `agent`, `timer`, `schedule`, `status`, `cancel`, `list`. Wakes devpulse reliably on agent exit, wall-clock time, or named duration. Replaces the old bash one-liner. |
|
||||
| `feedback.py` | Cross-project feedback channel. `compose` / `inbox` handlers. Lets external projects report bugs/friction back to devpulse. |
|
||||
|
||||
@@ -0,0 +1,503 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: watchdog.py
|
||||
# Description: Watchdog Module — directed wake system for devpulse
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
"""
|
||||
Watchdog Module — devpulse's personal directed-wake system.
|
||||
|
||||
Subcommands:
|
||||
agent <id> Wake when a dispatched agent process exits (Phase 1)
|
||||
status List active watches via watchdog_active.json registry (Phase 4)
|
||||
timer <args> Wake-in-N or named duration timer (Phase 2)
|
||||
schedule <time> Wake at wall-clock time, optionally run a command (Phase 3)
|
||||
cancel <handle> SIGTERM a specific watch + deregister (Phase 4)
|
||||
cancel --all Kill every active watch (Phase 4)
|
||||
list Alias for status (Phase 4)
|
||||
|
||||
Auto-discovered by devpulse.py via handle_command() convention.
|
||||
Heavy handler imports are lazy — only imported when a subcommand is invoked.
|
||||
|
||||
See FPLAN-0186 for the build plan and DPLAN-0130 for the design record.
|
||||
"""
|
||||
|
||||
import importlib
|
||||
from pathlib import Path
|
||||
from typing import List
|
||||
|
||||
from aipass.prax.apps.modules.logger import system_logger as logger
|
||||
from aipass.cli.apps.modules import console, error, warning
|
||||
|
||||
_VALID_SUBCOMMANDS = ["agent", "timer", "schedule", "status", "cancel", "list"]
|
||||
_DEFAULT_AGENT_TIMEOUT = 1800
|
||||
_NOT_IMPLEMENTED_MSG = (
|
||||
"{sub} is not yet implemented in this phase — see FPLAN-0186 (Phase {phase})"
|
||||
)
|
||||
# Phase 4 wired cancel + list for real. Left the map so future deferrals can reuse the shape.
|
||||
_PHASE_BY_SUB: dict[str, int] = {}
|
||||
|
||||
HELP_TEXT = """\
|
||||
[bold cyan]watchdog[/bold cyan] — devpulse directed wake system
|
||||
|
||||
[bold]Usage:[/bold]
|
||||
watchdog agent <branch> [--timeout SECONDS] Wake when dispatched agent exits
|
||||
watchdog status List active watches
|
||||
watchdog timer <duration> Wake in N (5m, 30s, 2h, 1h30m)
|
||||
watchdog timer start <name> Start named duration timer
|
||||
watchdog timer stop <name> Stop named timer + report elapsed
|
||||
watchdog timer list List active + historical timers
|
||||
watchdog timer report Formatted session summary
|
||||
watchdog schedule <time> [command] Wake at HH:MM or +N, optional cmd
|
||||
watchdog cancel <handle> SIGTERM a specific watch + deregister
|
||||
watchdog cancel --all Kill every active watch
|
||||
watchdog list Alias for status
|
||||
watchdog --help Show this help
|
||||
|
||||
[bold]Examples:[/bold]
|
||||
drone @devpulse watchdog agent @drone
|
||||
drone @devpulse watchdog agent @flow --timeout 600
|
||||
drone @devpulse watchdog timer 5m
|
||||
drone @devpulse watchdog timer start build-phase-3
|
||||
drone @devpulse watchdog timer stop build-phase-3
|
||||
drone @devpulse watchdog schedule "02:00"
|
||||
drone @devpulse watchdog schedule "+30m" "drone @git status"
|
||||
|
||||
See FPLAN-0186 (build plan) and DPLAN-0130 (design).
|
||||
"""
|
||||
|
||||
_TIMER_HELP_TEXT = """\
|
||||
[bold]watchdog timer[/bold] — wake-in-N + named duration tracking
|
||||
|
||||
Usage:
|
||||
watchdog timer <duration> Wake in N (5m, 30s, 2h, 1h30m, 45)
|
||||
watchdog timer start <name> Start a named duration timer
|
||||
watchdog timer stop <name> Stop + report elapsed
|
||||
watchdog timer list Show active + history
|
||||
watchdog timer report Formatted session summary
|
||||
watchdog timer --help Show this help
|
||||
"""
|
||||
|
||||
_SCHEDULE_HELP_TEXT = """\
|
||||
[bold]watchdog schedule[/bold] — wall-clock or relative wake, optional command
|
||||
|
||||
Usage:
|
||||
watchdog schedule <time> Wake at HH:MM[:SS] or +N (+30m, +1h30m)
|
||||
watchdog schedule <time> <command> Wake + run command via shell
|
||||
watchdog schedule --help Show this help
|
||||
|
||||
Examples:
|
||||
watchdog schedule "02:00"
|
||||
watchdog schedule "14:30" "drone @flow execute DPLAN-0200"
|
||||
watchdog schedule "+30m" "drone @git status"
|
||||
"""
|
||||
|
||||
|
||||
def print_introspection() -> None:
|
||||
"""Display module introspection info."""
|
||||
console.print()
|
||||
console.print("watchdog Module")
|
||||
console.print("Devpulse-local directed wake system. Wakes devpulse when a")
|
||||
console.print("watched condition fires (agent exit, timer, schedule).")
|
||||
console.print()
|
||||
console.print("Subcommands:")
|
||||
for sub in _VALID_SUBCOMMANDS:
|
||||
marker = "active" if sub in ("agent", "status") else f"phase {_PHASE_BY_SUB.get(sub, '?')}"
|
||||
console.print(f" {sub:<10} ({marker})")
|
||||
console.print()
|
||||
console.print("Connected Handlers:")
|
||||
console.print(" handlers/watchdog/")
|
||||
console.print(" - agent.py (watch_agent — block until dispatched agent exits)")
|
||||
console.print()
|
||||
|
||||
|
||||
def _guard_caller() -> bool:
|
||||
"""Reject cross-branch invocation. Devpulse-only tool."""
|
||||
cwd = Path.cwd()
|
||||
if cwd.name == "devpulse" or any(p.name == "devpulse" for p in cwd.parents):
|
||||
return True
|
||||
warning("watchdog is a devpulse-only module — refusing cross-branch call")
|
||||
return False
|
||||
|
||||
|
||||
def handle_command(command: str, args: List[str]) -> bool:
|
||||
"""Route watchdog subcommands.
|
||||
|
||||
Auto-discovered by devpulse.py module loader.
|
||||
|
||||
Args:
|
||||
command: The primary command string.
|
||||
args: Additional arguments after the command.
|
||||
|
||||
Returns:
|
||||
True if the command was handled, False otherwise.
|
||||
"""
|
||||
if command != "watchdog":
|
||||
return False
|
||||
|
||||
if not _guard_caller():
|
||||
return True
|
||||
|
||||
if not args:
|
||||
print_introspection()
|
||||
return True
|
||||
|
||||
if args[0] in ("--help", "-h", "help"):
|
||||
console.print(HELP_TEXT)
|
||||
return True
|
||||
|
||||
subcommand = args[0]
|
||||
sub_args = args[1:]
|
||||
|
||||
if subcommand not in _VALID_SUBCOMMANDS:
|
||||
error(f"Unknown watchdog subcommand: {subcommand}",
|
||||
suggestion="Use 'watchdog --help' for usage")
|
||||
return True
|
||||
|
||||
logger.info("[watchdog] subcommand=%s args=%s", subcommand, sub_args)
|
||||
|
||||
if subcommand == "agent":
|
||||
return _handle_agent(sub_args)
|
||||
|
||||
if subcommand == "status":
|
||||
return _handle_status()
|
||||
|
||||
if subcommand == "list":
|
||||
return _handle_list()
|
||||
|
||||
if subcommand == "timer":
|
||||
return _handle_timer(sub_args)
|
||||
|
||||
if subcommand == "schedule":
|
||||
return _handle_schedule(sub_args)
|
||||
|
||||
if subcommand == "cancel":
|
||||
return _handle_cancel(sub_args)
|
||||
|
||||
if subcommand in _PHASE_BY_SUB:
|
||||
phase = _PHASE_BY_SUB[subcommand]
|
||||
console.print(_NOT_IMPLEMENTED_MSG.format(sub=subcommand, phase=phase))
|
||||
return True
|
||||
|
||||
return True
|
||||
|
||||
|
||||
def _handle_timer(sub_args: List[str]) -> bool:
|
||||
"""Route ``watchdog timer`` subcommands through the timer handler."""
|
||||
if not sub_args:
|
||||
console.print(_TIMER_HELP_TEXT)
|
||||
return True
|
||||
|
||||
if sub_args[0] in ("--help", "-h", "help"):
|
||||
console.print(_TIMER_HELP_TEXT)
|
||||
return True
|
||||
|
||||
timer_mod = importlib.import_module(
|
||||
"aipass.devpulse.apps.handlers.watchdog.timer"
|
||||
)
|
||||
|
||||
action = sub_args[0]
|
||||
|
||||
if action == "start":
|
||||
if len(sub_args) < 2:
|
||||
error("Usage: watchdog timer start <name>")
|
||||
return True
|
||||
result = timer_mod.timer_start(sub_args[1])
|
||||
_print_timer_result(result)
|
||||
return True
|
||||
|
||||
if action == "stop":
|
||||
if len(sub_args) < 2:
|
||||
error("Usage: watchdog timer stop <name>")
|
||||
return True
|
||||
result = timer_mod.timer_stop(sub_args[1])
|
||||
_print_timer_result(result)
|
||||
return True
|
||||
|
||||
if action == "list":
|
||||
snapshot = timer_mod.timer_list()
|
||||
_print_timer_list(snapshot)
|
||||
return True
|
||||
|
||||
if action == "report":
|
||||
console.print(timer_mod.timer_report())
|
||||
return True
|
||||
|
||||
# Fall-through: treat the token as a duration for wake_in.
|
||||
try:
|
||||
result = timer_mod.wake_in(action)
|
||||
except ValueError as exc:
|
||||
logger.warning("[watchdog] invalid timer duration %r: %s", action, exc)
|
||||
error(f"Invalid duration: {action} ({exc})")
|
||||
return True
|
||||
_print_timer_result(result)
|
||||
return True
|
||||
|
||||
|
||||
def _handle_schedule(sub_args: List[str]) -> bool:
|
||||
"""Route ``watchdog schedule`` through the schedule handler.
|
||||
|
||||
Positional form: ``schedule <time> [command]``. Command (when present)
|
||||
is the entire second positional arg — the caller is responsible for
|
||||
quoting multi-token shell commands in their invocation.
|
||||
"""
|
||||
if not sub_args:
|
||||
console.print(_SCHEDULE_HELP_TEXT)
|
||||
return True
|
||||
|
||||
if sub_args[0] in ("--help", "-h", "help"):
|
||||
console.print(_SCHEDULE_HELP_TEXT)
|
||||
return True
|
||||
|
||||
time_str = sub_args[0]
|
||||
command = sub_args[1] if len(sub_args) >= 2 else None
|
||||
|
||||
schedule_mod = importlib.import_module(
|
||||
"aipass.devpulse.apps.handlers.watchdog.schedule"
|
||||
)
|
||||
|
||||
try:
|
||||
result = schedule_mod.wake_at(time_str, command=command)
|
||||
except ValueError as exc:
|
||||
logger.warning("[watchdog] invalid schedule %r: %s", time_str, exc)
|
||||
error(f"Invalid schedule: {time_str} ({exc})")
|
||||
return True
|
||||
|
||||
_print_schedule_result(result)
|
||||
return True
|
||||
|
||||
|
||||
def _print_schedule_result(result: dict) -> None:
|
||||
"""Render a schedule handler return dict as CLI output."""
|
||||
scheduled_for = result.get("scheduled_for", "?")
|
||||
elapsed = result.get("elapsed", 0)
|
||||
console.print(
|
||||
f"[bold]watchdog schedule[/bold] woke after {elapsed}s "
|
||||
f"(scheduled_for={scheduled_for})"
|
||||
)
|
||||
if result.get("command"):
|
||||
exit_code = result.get("command_exit_code")
|
||||
console.print(
|
||||
f" command: {result['command']} -> exit={exit_code}"
|
||||
)
|
||||
stdout = result.get("command_stdout") or ""
|
||||
stderr = result.get("command_stderr") or ""
|
||||
if stdout:
|
||||
console.print(f" stdout: {stdout.rstrip()}")
|
||||
if stderr:
|
||||
console.print(f" stderr: {stderr.rstrip()}")
|
||||
|
||||
|
||||
def _print_timer_result(result: dict) -> None:
|
||||
"""Render a timer handler return dict as a single CLI line."""
|
||||
state = result.get("state", "unknown")
|
||||
name = result.get("name") or result.get("duration") or ""
|
||||
if state == "error":
|
||||
error(f"timer {name}: {result.get('reason', 'unknown error')}")
|
||||
return
|
||||
if state == "stopped":
|
||||
console.print(
|
||||
f"[bold]timer[/bold] {name} stopped -> elapsed={result.get('human', '?')}"
|
||||
)
|
||||
return
|
||||
if state == "started":
|
||||
console.print(f"[bold]timer[/bold] {name} started at {result.get('started_at', '?')}")
|
||||
return
|
||||
if state == "woke":
|
||||
console.print(
|
||||
f"[bold]timer[/bold] {name} woke after {result.get('elapsed', 0)}s"
|
||||
)
|
||||
return
|
||||
console.print(f"[dim]timer result:[/dim] {result}")
|
||||
|
||||
|
||||
def _print_timer_list(snapshot: dict) -> None:
|
||||
"""Pretty-print the ``timer_list`` snapshot."""
|
||||
active = snapshot.get("active", [])
|
||||
history = snapshot.get("history", [])
|
||||
console.print("[bold]Active timers:[/bold]")
|
||||
if active:
|
||||
for item in active:
|
||||
console.print(
|
||||
f" - {item['name']} elapsed {item['human']} "
|
||||
f"(started {item.get('started_at', '?')})"
|
||||
)
|
||||
else:
|
||||
console.print(" (none)")
|
||||
console.print("[bold]History:[/bold]")
|
||||
if history:
|
||||
for item in history:
|
||||
console.print(
|
||||
f" - {item['name']} {item['human']} "
|
||||
f"({item.get('started_at', '?')} -> {item.get('stopped_at', '?')})"
|
||||
)
|
||||
else:
|
||||
console.print(" (none)")
|
||||
|
||||
|
||||
def _handle_agent(sub_args: List[str]) -> bool:
|
||||
"""Parse `agent <id> [--timeout N]` and invoke the agent handler."""
|
||||
if not sub_args:
|
||||
error("Usage: watchdog agent <branch> [--timeout SECONDS]")
|
||||
return True
|
||||
|
||||
timeout = _DEFAULT_AGENT_TIMEOUT
|
||||
positional: List[str] = []
|
||||
i = 0
|
||||
while i < len(sub_args):
|
||||
arg = sub_args[i]
|
||||
if arg == "--timeout" and i + 1 < len(sub_args):
|
||||
try:
|
||||
timeout = int(sub_args[i + 1])
|
||||
except ValueError as exc:
|
||||
logger.warning("[watchdog] invalid --timeout value %r: %s", sub_args[i + 1], exc)
|
||||
error(f"Invalid --timeout value: {sub_args[i + 1]}")
|
||||
return True
|
||||
i += 2
|
||||
continue
|
||||
positional.append(arg)
|
||||
i += 1
|
||||
|
||||
if not positional:
|
||||
error("Usage: watchdog agent <branch> [--timeout SECONDS]")
|
||||
return True
|
||||
|
||||
agent_id = positional[0]
|
||||
|
||||
agent_mod = importlib.import_module(
|
||||
"aipass.devpulse.apps.handlers.watchdog.agent"
|
||||
)
|
||||
result = agent_mod.watch_agent(agent_id, timeout_seconds=timeout)
|
||||
|
||||
state = result.get("agent_state", "unknown")
|
||||
reason = result.get("reason", "")
|
||||
elapsed = result.get("elapsed", 0)
|
||||
console.print(
|
||||
f"[bold]watchdog agent[/bold] {agent_id} -> "
|
||||
f"state={state} elapsed={elapsed}s reason={reason}"
|
||||
)
|
||||
return True
|
||||
|
||||
|
||||
def _load_registry_module():
|
||||
"""Lazy-import the watch registry. Keeps cold startup fast."""
|
||||
return importlib.import_module(
|
||||
"aipass.devpulse.apps.handlers.watchdog.registry"
|
||||
)
|
||||
|
||||
|
||||
def _load_timer_module_for_format():
|
||||
"""Lazy-import timer for ``format_human`` (reused in the status output)."""
|
||||
return importlib.import_module(
|
||||
"aipass.devpulse.apps.handlers.watchdog.timer"
|
||||
)
|
||||
|
||||
|
||||
def _format_status_line(watch: dict, format_human) -> str:
|
||||
"""One-line renderer for a single watch entry in the status output."""
|
||||
handle = watch.get("handle", "?")
|
||||
wtype = watch.get("type", "?")
|
||||
elapsed = int(watch.get("elapsed_seconds", 0))
|
||||
pid = watch.get("pid", "?")
|
||||
meta = watch.get("metadata") or {}
|
||||
|
||||
if wtype == "agent":
|
||||
tail = (
|
||||
f"{meta.get('agent_id', '?')} "
|
||||
f"(timeout={meta.get('timeout_seconds', '?')}s)"
|
||||
)
|
||||
elif wtype == "timer":
|
||||
tail = f"duration={meta.get('duration', '?')}"
|
||||
elif wtype == "schedule":
|
||||
tail_cmd = meta.get("command")
|
||||
cmd_repr = f' cmd="{tail_cmd}"' if tail_cmd else ""
|
||||
tail = f"scheduled={meta.get('scheduled_for', '?')}{cmd_repr}"
|
||||
else:
|
||||
tail = str(meta)
|
||||
|
||||
# Escape the [ so Rich console doesn't interpret it as a style tag.
|
||||
return (
|
||||
f" \\[{handle}] {wtype:<8} {format_human(elapsed):<10} "
|
||||
f"pid={pid} {tail}"
|
||||
)
|
||||
|
||||
|
||||
def _handle_status() -> bool:
|
||||
"""Read the watch registry, prune stale entries, pretty-print active watches."""
|
||||
registry_mod = _load_registry_module()
|
||||
timer_mod = _load_timer_module_for_format()
|
||||
|
||||
# list_active handles its own stale pruning — we just count survivors
|
||||
# before and after to know if we pruned anything to report.
|
||||
pre = registry_mod.list_active(prune_stale=False)
|
||||
post = registry_mod.list_active(prune_stale=True)
|
||||
pruned = len(pre) - len(post)
|
||||
|
||||
console.print("[bold]Watchdog Status[/bold]")
|
||||
console.print("===============")
|
||||
|
||||
if not post:
|
||||
console.print("No active watches.")
|
||||
if pruned:
|
||||
console.print(f"[dim]Pruned {pruned} stale watch(es).[/dim]")
|
||||
return True
|
||||
|
||||
console.print(f"{len(post)} active watch(es):")
|
||||
console.print()
|
||||
for watch in post:
|
||||
console.print(_format_status_line(watch, timer_mod.format_human))
|
||||
|
||||
if pruned:
|
||||
console.print(f"[dim]Pruned {pruned} stale watch(es).[/dim]")
|
||||
else:
|
||||
console.print("[dim]No stale watches to prune.[/dim]")
|
||||
return True
|
||||
|
||||
|
||||
def _handle_list() -> bool:
|
||||
"""Alias for ``status`` — terser framing chosen: same output.
|
||||
|
||||
Phase 4 Notes: `list` just routes to `_handle_status`. The UX bar for
|
||||
differentiating wasn't worth the divergence.
|
||||
"""
|
||||
return _handle_status()
|
||||
|
||||
|
||||
def _print_kill_result(result: dict) -> None:
|
||||
"""Render a single ``registry.kill_watch`` result on one line."""
|
||||
handle = result.get("handle", "?")
|
||||
killed = result.get("killed", False)
|
||||
was_alive = result.get("was_alive", False)
|
||||
reason = result.get("reason", "")
|
||||
status = "KILLED" if killed else "FAILED"
|
||||
console.print(
|
||||
f" \\[{handle}] {status} was_alive={was_alive} reason={reason}"
|
||||
)
|
||||
|
||||
|
||||
def _handle_cancel(sub_args: List[str]) -> bool:
|
||||
"""Route ``watchdog cancel <handle>`` or ``cancel --all`` through the registry."""
|
||||
if not sub_args:
|
||||
error("Usage: watchdog cancel <handle> | watchdog cancel --all")
|
||||
return True
|
||||
|
||||
registry_mod = _load_registry_module()
|
||||
|
||||
if sub_args[0] == "--all":
|
||||
results = registry_mod.kill_all()
|
||||
if not results:
|
||||
console.print("No active watches to cancel.")
|
||||
return True
|
||||
console.print(f"[bold]Cancelling {len(results)} watch(es):[/bold]")
|
||||
for result in results:
|
||||
_print_kill_result(result)
|
||||
return True
|
||||
|
||||
handle = sub_args[0]
|
||||
result = registry_mod.kill_watch(handle)
|
||||
_print_kill_result(result)
|
||||
if not result.get("killed", False):
|
||||
logger.info("[watchdog] cancel failed handle=%s", handle)
|
||||
return True
|
||||
@@ -0,0 +1,245 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: test_watchdog_agent.py
|
||||
# Description: Tests for the watchdog agent handler
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
"""Tests for watch_agent (Phase 1, FPLAN-0186).
|
||||
|
||||
Mock-heavy unit tests verify the return-shape branches:
|
||||
- completed (clean exit)
|
||||
- crashed (bounce file present)
|
||||
- timeout
|
||||
- agent not found
|
||||
|
||||
Integration tests are marked with @pytest.mark.integration and require
|
||||
a live ai_mail dispatch flow. They're skipped by default in CI.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from aipass.devpulse.apps.handlers.watchdog import agent as agent_handler
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Test helpers
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def _build_fake_branch(tmp_path: Path, branch_name: str = "fakebranch") -> Path:
|
||||
"""Create a fake branch dir with .ai_mail.local and a registry pointing at it."""
|
||||
branch_path = tmp_path / branch_name
|
||||
(branch_path / ".ai_mail.local").mkdir(parents=True)
|
||||
|
||||
registry = {
|
||||
"branches": [
|
||||
{"email": f"@{branch_name}", "path": str(branch_path)},
|
||||
]
|
||||
}
|
||||
(tmp_path / "AIPASS_REGISTRY.json").write_text(
|
||||
json.dumps(registry), encoding='utf-8'
|
||||
)
|
||||
return branch_path
|
||||
|
||||
|
||||
def _write_lock(branch_path: Path, pid: int) -> Path:
|
||||
"""Write a fake .dispatch.lock and return its path."""
|
||||
lock_file = branch_path / ".ai_mail.local" / ".dispatch.lock"
|
||||
lock_data = {"pid": pid, "timestamp": "2026-04-14T00:00:00", "branch": str(branch_path)}
|
||||
lock_file.write_text(json.dumps(lock_data), encoding='utf-8')
|
||||
return lock_file
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Unit tests — return shape per branch
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_watch_agent_branch_not_found(monkeypatch, tmp_path):
|
||||
"""Missing branch returns immediately with timeout state and 'agent not found'."""
|
||||
monkeypatch.setattr(agent_handler, "_find_repo_root", lambda *a, **kw: None)
|
||||
result = agent_handler.watch_agent("@nonexistent", timeout_seconds=5)
|
||||
|
||||
assert result["woke"] is False
|
||||
assert result["agent_state"] == "timeout"
|
||||
assert "not found" in result["reason"].lower()
|
||||
assert result["agent_id"] == "@nonexistent"
|
||||
assert result["exit_code"] is None
|
||||
|
||||
|
||||
def test_watch_agent_no_active_lock(monkeypatch, tmp_path):
|
||||
"""No lock file -> agent already idle, returns completed immediately."""
|
||||
branch_path = _build_fake_branch(tmp_path)
|
||||
monkeypatch.setattr(agent_handler, "_find_repo_root", lambda *a, **kw: tmp_path)
|
||||
|
||||
result = agent_handler.watch_agent("@fakebranch", timeout_seconds=5)
|
||||
|
||||
assert result["woke"] is True
|
||||
assert result["agent_state"] == "completed"
|
||||
assert result["exit_code"] == 0
|
||||
|
||||
|
||||
def test_watch_agent_completed_via_lock_removal(monkeypatch, tmp_path):
|
||||
"""Lock present at start, then removed -> wake with state=completed."""
|
||||
branch_path = _build_fake_branch(tmp_path)
|
||||
lock_file = _write_lock(branch_path, pid=os.getpid())
|
||||
monkeypatch.setattr(agent_handler, "_find_repo_root", lambda *a, **kw: tmp_path)
|
||||
|
||||
call_count = {"n": 0}
|
||||
real_sleep = time.sleep
|
||||
|
||||
def fake_sleep(seconds):
|
||||
"""Remove the lock on the second poll cycle to simulate clean exit."""
|
||||
call_count["n"] += 1
|
||||
if call_count["n"] >= 1:
|
||||
lock_file.unlink(missing_ok=True)
|
||||
real_sleep(0.01)
|
||||
|
||||
monkeypatch.setattr(agent_handler.time, "sleep", fake_sleep)
|
||||
|
||||
result = agent_handler.watch_agent(
|
||||
"@fakebranch", timeout_seconds=5, poll_interval=0.01
|
||||
)
|
||||
|
||||
assert result["woke"] is True
|
||||
assert result["agent_state"] == "completed"
|
||||
assert result["exit_code"] == 0
|
||||
assert "clean" in result["reason"].lower() or "finished" in result["reason"].lower()
|
||||
|
||||
|
||||
def test_watch_agent_crashed_via_bounce_file(monkeypatch, tmp_path):
|
||||
"""Lock removed AND bounce file present -> wake with state=crashed."""
|
||||
branch_path = _build_fake_branch(tmp_path)
|
||||
lock_file = _write_lock(branch_path, pid=os.getpid())
|
||||
bounce_file = branch_path / ".ai_mail.local" / "last_bounce.json"
|
||||
monkeypatch.setattr(agent_handler, "_find_repo_root", lambda *a, **kw: tmp_path)
|
||||
|
||||
real_sleep = time.sleep
|
||||
|
||||
def fake_sleep(seconds):
|
||||
"""Drop a bounce file then remove the lock to simulate crash exit."""
|
||||
bounce_file.write_text(json.dumps({"exit_code": 1, "reason": "test"}), encoding='utf-8')
|
||||
lock_file.unlink(missing_ok=True)
|
||||
real_sleep(0.01)
|
||||
|
||||
monkeypatch.setattr(agent_handler.time, "sleep", fake_sleep)
|
||||
|
||||
result = agent_handler.watch_agent(
|
||||
"@fakebranch", timeout_seconds=5, poll_interval=0.01
|
||||
)
|
||||
|
||||
assert result["woke"] is True
|
||||
assert result["agent_state"] == "crashed"
|
||||
assert result["exit_code"] == 1
|
||||
|
||||
|
||||
def test_watch_agent_timeout(monkeypatch, tmp_path):
|
||||
"""Lock never removed and PID stays alive -> timeout."""
|
||||
branch_path = _build_fake_branch(tmp_path)
|
||||
_write_lock(branch_path, pid=os.getpid())
|
||||
monkeypatch.setattr(agent_handler, "_find_repo_root", lambda *a, **kw: tmp_path)
|
||||
monkeypatch.setattr(agent_handler, "_pid_alive", lambda pid: True)
|
||||
|
||||
result = agent_handler.watch_agent(
|
||||
"@fakebranch", timeout_seconds=1, poll_interval=0.05
|
||||
)
|
||||
|
||||
assert result["woke"] is False
|
||||
assert result["agent_state"] == "timeout"
|
||||
assert result["exit_code"] is None
|
||||
assert result["elapsed"] >= 1
|
||||
|
||||
|
||||
def test_watch_agent_pid_dead_treated_as_crash(monkeypatch, tmp_path):
|
||||
"""Lock present but monitor PID dead -> crash exit."""
|
||||
branch_path = _build_fake_branch(tmp_path)
|
||||
_write_lock(branch_path, pid=999999)
|
||||
monkeypatch.setattr(agent_handler, "_find_repo_root", lambda *a, **kw: tmp_path)
|
||||
monkeypatch.setattr(agent_handler, "_pid_alive", lambda pid: False)
|
||||
|
||||
result = agent_handler.watch_agent(
|
||||
"@fakebranch", timeout_seconds=5, poll_interval=0.01
|
||||
)
|
||||
|
||||
assert result["woke"] is True
|
||||
assert result["agent_state"] == "crashed"
|
||||
|
||||
|
||||
def test_watch_agent_return_keys():
|
||||
"""Every code path must return all expected keys (Phase 4 adds ``handle``)."""
|
||||
expected = {
|
||||
"woke", "reason", "elapsed", "agent_state",
|
||||
"exit_code", "agent_id", "handle",
|
||||
}
|
||||
# Use the not-found path for a fast invocation
|
||||
result = agent_handler.watch_agent("@__definitely_not_a_branch__", timeout_seconds=1)
|
||||
assert set(result.keys()) == expected
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Integration tests (require live ai_mail dispatch — skipped by default)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
@pytest.mark.skipif(
|
||||
os.environ.get("WATCHDOG_INTEGRATION") != "1",
|
||||
reason="Set WATCHDOG_INTEGRATION=1 to run live dispatch tests",
|
||||
)
|
||||
def test_watch_agent_live_dispatch_completes():
|
||||
"""Dispatch a real tiny agent and verify watchdog wakes on completion.
|
||||
|
||||
Skipped unless WATCHDOG_INTEGRATION=1 is set. Slow and depends on
|
||||
a working drone / ai_mail / agent runtime.
|
||||
"""
|
||||
import subprocess
|
||||
|
||||
dispatch = subprocess.run(
|
||||
["drone", "@ai_mail", "dispatch", "@drone",
|
||||
"Watchdog ping test", "Reply with OK then exit."],
|
||||
capture_output=True, text=True, timeout=60,
|
||||
)
|
||||
assert dispatch.returncode == 0, f"dispatch failed: {dispatch.stderr}"
|
||||
|
||||
result = agent_handler.watch_agent("@drone", timeout_seconds=300, poll_interval=2.0)
|
||||
assert result["woke"] is True
|
||||
assert result["agent_state"] in ("completed", "crashed")
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
@pytest.mark.skipif(
|
||||
os.environ.get("WATCHDOG_INTEGRATION") != "1",
|
||||
reason="Set WATCHDOG_INTEGRATION=1 to run live dispatch tests",
|
||||
)
|
||||
def test_watch_agent_live_dispatch_timeout_path():
|
||||
"""Dispatch a real agent with a very short watchdog timeout; verify timeout state."""
|
||||
import subprocess
|
||||
|
||||
subprocess.run(
|
||||
["drone", "@ai_mail", "dispatch", "@drone",
|
||||
"Long watchdog test", "Wait at least 30 seconds then reply."],
|
||||
capture_output=True, text=True, timeout=60,
|
||||
)
|
||||
|
||||
result = agent_handler.watch_agent("@drone", timeout_seconds=2, poll_interval=0.5)
|
||||
assert result["agent_state"] == "timeout"
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
def test_watch_agent_crash_path_skipped():
|
||||
"""Crash-path integration test — skipped.
|
||||
|
||||
Cheaply triggering a real agent crash mid-task would require either
|
||||
crafting a malformed dispatch (risk: corrupting the ai_mail flow) or
|
||||
SIGKILLing a live monitor (risk: leaving stale locks). The unit test
|
||||
test_watch_agent_crashed_via_bounce_file already covers the bounce-file
|
||||
branch via the same code path the monitor uses.
|
||||
"""
|
||||
pytest.skip("Crash path covered by unit test test_watch_agent_crashed_via_bounce_file")
|
||||
@@ -0,0 +1,279 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: test_watchdog_module.py
|
||||
# Description: Tests for the watchdog module router
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
"""Tests for the watchdog module router (Phase 1, FPLAN-0186)."""
|
||||
|
||||
import sys
|
||||
from unittest.mock import patch
|
||||
|
||||
import pytest
|
||||
|
||||
from aipass.devpulse.apps.modules import watchdog as wd_mod
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _bypass_caller_guard():
|
||||
"""Force _guard_caller to always pass so tests don't depend on cwd."""
|
||||
with patch.object(wd_mod, "_guard_caller", return_value=True):
|
||||
yield
|
||||
|
||||
|
||||
def test_handle_command_rejects_unrelated_command():
|
||||
"""Router returns False for commands that aren't 'watchdog'."""
|
||||
assert wd_mod.handle_command("feedback", []) is False
|
||||
|
||||
|
||||
def test_handle_command_no_args_shows_introspection(capsys):
|
||||
"""No args -> module introspection (mentions 'watchdog')."""
|
||||
result = wd_mod.handle_command("watchdog", [])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "watchdog" in combined.lower()
|
||||
|
||||
|
||||
def test_handle_command_help_flag(capsys):
|
||||
"""--help prints HELP_TEXT."""
|
||||
result = wd_mod.handle_command("watchdog", ["--help"])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "Usage" in combined or "usage" in combined.lower()
|
||||
|
||||
|
||||
def test_handle_command_unknown_subcommand(capsys):
|
||||
"""Unknown subcommand returns clean error message."""
|
||||
result = wd_mod.handle_command("watchdog", ["bogus"])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "bogus" in combined.lower() or "unknown" in combined.lower()
|
||||
|
||||
|
||||
def _fake_registry_module(active=None, kill_result=None, kill_all_result=None):
|
||||
"""Build a fake watchdog.registry module for router-level tests."""
|
||||
fake = type(sys)("fake_registry_mod")
|
||||
fake.calls = []
|
||||
|
||||
def list_active(storage_path=None, prune_stale=True):
|
||||
"""Fake registry list_active — returns the preset active list."""
|
||||
fake.calls.append(("list_active", prune_stale))
|
||||
return list(active or [])
|
||||
|
||||
def kill_watch(handle, storage_path=None):
|
||||
"""Fake registry kill_watch — returns the preset result."""
|
||||
fake.calls.append(("kill_watch", handle))
|
||||
return kill_result or {
|
||||
"handle": handle,
|
||||
"killed": True,
|
||||
"was_alive": True,
|
||||
"reason": "fake kill",
|
||||
}
|
||||
|
||||
def kill_all(storage_path=None):
|
||||
"""Fake registry kill_all — returns the preset result list."""
|
||||
fake.calls.append(("kill_all",))
|
||||
return list(kill_all_result or [])
|
||||
|
||||
fake.list_active = list_active
|
||||
fake.kill_watch = kill_watch
|
||||
fake.kill_all = kill_all
|
||||
return fake
|
||||
|
||||
|
||||
def _fake_timer_module_with_format():
|
||||
"""Minimal fake timer module exposing ``format_human``."""
|
||||
fake = type(sys)("fake_timer_mod_fmt")
|
||||
fake.format_human = lambda seconds: f"{seconds}s"
|
||||
return fake
|
||||
|
||||
|
||||
def _patch_registry_imports(fake_registry, fake_timer=None):
|
||||
"""Patch importlib.import_module to return the right fake per module path."""
|
||||
timer = fake_timer or _fake_timer_module_with_format()
|
||||
|
||||
def fake_import(name):
|
||||
"""Patched ``importlib.import_module`` — routes to fake registry/timer."""
|
||||
if name.endswith(".registry"):
|
||||
return fake_registry
|
||||
if name.endswith(".timer"):
|
||||
return timer
|
||||
raise ImportError(f"unexpected import in test: {name}")
|
||||
|
||||
return patch("importlib.import_module", side_effect=fake_import)
|
||||
|
||||
|
||||
def test_cancel_requires_handle(capsys):
|
||||
"""`cancel` with no args prints usage."""
|
||||
result = wd_mod.handle_command("watchdog", ["cancel"])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "usage" in combined.lower() or "cancel" in combined.lower()
|
||||
|
||||
|
||||
def test_cancel_handle_routes_to_registry(capsys):
|
||||
"""`cancel <handle>` calls registry.kill_watch and prints the result."""
|
||||
fake = _fake_registry_module(kill_result={
|
||||
"handle": "agent-abc123",
|
||||
"killed": True,
|
||||
"was_alive": True,
|
||||
"reason": "SIGTERM — pid 1234 exited in 0.1s",
|
||||
})
|
||||
with _patch_registry_imports(fake):
|
||||
result = wd_mod.handle_command("watchdog", ["cancel", "agent-abc123"])
|
||||
assert result is True
|
||||
assert ("kill_watch", "agent-abc123") in fake.calls
|
||||
combined = capsys.readouterr().out
|
||||
assert "agent-abc123" in combined
|
||||
assert "KILLED" in combined
|
||||
|
||||
|
||||
def test_cancel_all_routes_to_registry(capsys):
|
||||
"""`cancel --all` calls registry.kill_all and prints every result line."""
|
||||
fake = _fake_registry_module(kill_all_result=[
|
||||
{"handle": "timer-111111", "killed": True, "was_alive": True, "reason": "ok"},
|
||||
{"handle": "schedule-222222", "killed": True, "was_alive": True, "reason": "ok"},
|
||||
])
|
||||
with _patch_registry_imports(fake):
|
||||
result = wd_mod.handle_command("watchdog", ["cancel", "--all"])
|
||||
assert result is True
|
||||
assert ("kill_all",) in fake.calls
|
||||
out = capsys.readouterr().out
|
||||
assert "timer-111111" in out
|
||||
assert "schedule-222222" in out
|
||||
|
||||
|
||||
def test_cancel_all_empty(capsys):
|
||||
"""`cancel --all` with nothing active reports 'no active watches to cancel'."""
|
||||
fake = _fake_registry_module(kill_all_result=[])
|
||||
with _patch_registry_imports(fake):
|
||||
wd_mod.handle_command("watchdog", ["cancel", "--all"])
|
||||
out = capsys.readouterr().out.lower()
|
||||
assert "no active watches" in out
|
||||
|
||||
|
||||
def test_status_reports_no_active_watches(capsys):
|
||||
"""status reports 'no active watches' when the registry is empty."""
|
||||
fake = _fake_registry_module(active=[])
|
||||
with _patch_registry_imports(fake):
|
||||
result = wd_mod.handle_command("watchdog", ["status"])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "no active watches" in combined.lower()
|
||||
|
||||
|
||||
def test_status_prints_active_watches(capsys):
|
||||
"""status prints every active watch with handle + type + pid."""
|
||||
active = [
|
||||
{
|
||||
"handle": "agent-abc123",
|
||||
"type": "agent",
|
||||
"pid": 4321,
|
||||
"started_epoch": 1000.0,
|
||||
"elapsed_seconds": 134,
|
||||
"metadata": {"agent_id": "@drone", "timeout_seconds": 1800},
|
||||
},
|
||||
{
|
||||
"handle": "schedule-def456",
|
||||
"type": "schedule",
|
||||
"pid": 4322,
|
||||
"started_epoch": 2000.0,
|
||||
"elapsed_seconds": 45,
|
||||
"metadata": {"scheduled_for": "02:00:00", "command": "drone @git status"},
|
||||
},
|
||||
]
|
||||
fake = _fake_registry_module(active=active)
|
||||
with _patch_registry_imports(fake):
|
||||
result = wd_mod.handle_command("watchdog", ["status"])
|
||||
assert result is True
|
||||
out = capsys.readouterr().out
|
||||
assert "agent-abc123" in out
|
||||
assert "schedule-def456" in out
|
||||
assert "@drone" in out
|
||||
assert "2 active" in out or "2 active watch" in out
|
||||
|
||||
|
||||
def test_list_routes_to_status(capsys):
|
||||
"""`list` is an alias — same output as `status`."""
|
||||
fake = _fake_registry_module(active=[])
|
||||
with _patch_registry_imports(fake):
|
||||
result = wd_mod.handle_command("watchdog", ["list"])
|
||||
assert result is True
|
||||
out = capsys.readouterr().out.lower()
|
||||
assert "watchdog status" in out or "no active watches" in out
|
||||
|
||||
|
||||
def test_agent_subcommand_requires_id(capsys):
|
||||
"""`agent` with no id prints usage."""
|
||||
result = wd_mod.handle_command("watchdog", ["agent"])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "usage" in combined.lower() or "watchdog agent" in combined.lower()
|
||||
|
||||
|
||||
def test_agent_subcommand_invokes_handler(capsys):
|
||||
"""`agent <id>` lazily imports and invokes watch_agent."""
|
||||
fake_result = {
|
||||
"woke": True,
|
||||
"reason": "fake clean exit",
|
||||
"elapsed": 5,
|
||||
"agent_state": "completed",
|
||||
"exit_code": 0,
|
||||
"agent_id": "@drone",
|
||||
}
|
||||
fake_module = type(sys)("fake_agent_mod")
|
||||
fake_module.watch_agent = lambda agent_id, timeout_seconds=1800: fake_result
|
||||
|
||||
with patch("importlib.import_module", return_value=fake_module):
|
||||
result = wd_mod.handle_command("watchdog", ["agent", "@drone"])
|
||||
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "completed" in combined
|
||||
assert "@drone" in combined
|
||||
|
||||
|
||||
def test_agent_subcommand_parses_timeout_flag():
|
||||
"""--timeout flag is parsed and passed to the handler."""
|
||||
captured_args = {}
|
||||
|
||||
def fake_watch_agent(agent_id, timeout_seconds=1800):
|
||||
"""Fake agent watcher that records its arguments."""
|
||||
captured_args["agent_id"] = agent_id
|
||||
captured_args["timeout"] = timeout_seconds
|
||||
return {
|
||||
"woke": True,
|
||||
"reason": "fake",
|
||||
"elapsed": 1,
|
||||
"agent_state": "completed",
|
||||
"exit_code": 0,
|
||||
"agent_id": agent_id,
|
||||
}
|
||||
|
||||
fake_module = type(sys)("fake_agent_mod")
|
||||
fake_module.watch_agent = fake_watch_agent
|
||||
|
||||
with patch("importlib.import_module", return_value=fake_module):
|
||||
wd_mod.handle_command("watchdog", ["agent", "@flow", "--timeout", "60"])
|
||||
|
||||
assert captured_args == {"agent_id": "@flow", "timeout": 60}
|
||||
|
||||
|
||||
def test_agent_subcommand_invalid_timeout(capsys):
|
||||
"""Invalid --timeout value reports a clean error."""
|
||||
result = wd_mod.handle_command(
|
||||
"watchdog", ["agent", "@flow", "--timeout", "notanumber"]
|
||||
)
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "invalid" in combined.lower() or "--timeout" in combined.lower()
|
||||
@@ -0,0 +1,423 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: test_watchdog_registry.py
|
||||
# Description: Tests for the watchdog watch registry (Phase 4, FPLAN-0186)
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
"""Tests for watchdog registry — register/deregister/list/kill (Phase 4)."""
|
||||
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
import pytest
|
||||
|
||||
from aipass.devpulse.apps.handlers.watchdog import registry as watch_registry
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def store_path(tmp_path):
|
||||
"""Fresh, isolated registry file per test — never touches real .trinity/."""
|
||||
return tmp_path / "watchdog_active.json"
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# register / deregister
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_register_creates_entry_and_returns_handle(store_path):
|
||||
handle = watch_registry.register(
|
||||
"agent",
|
||||
metadata={"agent_id": "@drone", "timeout_seconds": 1800},
|
||||
storage_path=store_path,
|
||||
)
|
||||
assert handle.startswith("agent-")
|
||||
assert len(handle) > len("agent-")
|
||||
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert raw["version"] == 1
|
||||
assert len(raw["watches"]) == 1
|
||||
entry = raw["watches"][0]
|
||||
assert entry["handle"] == handle
|
||||
assert entry["type"] == "agent"
|
||||
assert entry["pid"] == os.getpid()
|
||||
assert entry["metadata"]["agent_id"] == "@drone"
|
||||
assert "started_at" in entry
|
||||
assert "started_epoch" in entry
|
||||
|
||||
|
||||
def test_register_multiple_watches(store_path):
|
||||
h1 = watch_registry.register("agent", {"agent_id": "@drone"}, storage_path=store_path)
|
||||
h2 = watch_registry.register("timer", {"duration": "5m"}, storage_path=store_path)
|
||||
h3 = watch_registry.register("schedule", {"scheduled_for": "02:00"}, storage_path=store_path)
|
||||
|
||||
handles = {h1, h2, h3}
|
||||
assert len(handles) == 3
|
||||
assert h1.startswith("agent-")
|
||||
assert h2.startswith("timer-")
|
||||
assert h3.startswith("schedule-")
|
||||
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert len(raw["watches"]) == 3
|
||||
stored = {w["handle"] for w in raw["watches"]}
|
||||
assert stored == handles
|
||||
|
||||
|
||||
def test_register_rejects_empty_watch_type(store_path):
|
||||
with pytest.raises(ValueError):
|
||||
watch_registry.register("", storage_path=store_path)
|
||||
with pytest.raises(ValueError):
|
||||
watch_registry.register(" ", storage_path=store_path)
|
||||
|
||||
|
||||
def test_deregister_removes_entry(store_path):
|
||||
handle = watch_registry.register("timer", {"duration": "1s"}, storage_path=store_path)
|
||||
removed = watch_registry.deregister(handle, storage_path=store_path)
|
||||
assert removed is True
|
||||
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert raw["watches"] == []
|
||||
|
||||
|
||||
def test_deregister_nonexistent_returns_false(store_path):
|
||||
# Empty store
|
||||
assert watch_registry.deregister("ghost-abcdef", storage_path=store_path) is False
|
||||
|
||||
# Non-empty store, wrong handle
|
||||
watch_registry.register("timer", {"duration": "1s"}, storage_path=store_path)
|
||||
assert watch_registry.deregister("ghost-abcdef", storage_path=store_path) is False
|
||||
|
||||
|
||||
def test_deregister_only_removes_target(store_path):
|
||||
h1 = watch_registry.register("agent", {}, storage_path=store_path)
|
||||
h2 = watch_registry.register("timer", {}, storage_path=store_path)
|
||||
|
||||
watch_registry.deregister(h1, storage_path=store_path)
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert [w["handle"] for w in raw["watches"]] == [h2]
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# list_active
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_list_active_empty(store_path):
|
||||
assert watch_registry.list_active(storage_path=store_path) == []
|
||||
|
||||
|
||||
def test_list_active_returns_all_with_elapsed(store_path):
|
||||
watch_registry.register("agent", {"agent_id": "@drone"}, storage_path=store_path)
|
||||
watch_registry.register("timer", {"duration": "5m"}, storage_path=store_path)
|
||||
time.sleep(0.05)
|
||||
|
||||
active = watch_registry.list_active(storage_path=store_path, prune_stale=False)
|
||||
assert len(active) == 2
|
||||
for entry in active:
|
||||
assert "elapsed_seconds" in entry
|
||||
assert entry["elapsed_seconds"] >= 0
|
||||
|
||||
|
||||
def test_list_active_prunes_stale_by_default(store_path):
|
||||
watch_registry.register("timer", {}, storage_path=store_path)
|
||||
watch_registry.register("schedule", {}, storage_path=store_path)
|
||||
|
||||
# Force every pid to look dead.
|
||||
with patch.object(watch_registry, "is_pid_alive", return_value=False):
|
||||
active = watch_registry.list_active(storage_path=store_path, prune_stale=True)
|
||||
|
||||
assert active == []
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert raw["watches"] == [] # pruned from disk too
|
||||
|
||||
|
||||
def test_list_active_keeps_stale_when_prune_false(store_path):
|
||||
watch_registry.register("timer", {}, storage_path=store_path)
|
||||
|
||||
with patch.object(watch_registry, "is_pid_alive", return_value=False):
|
||||
active = watch_registry.list_active(storage_path=store_path, prune_stale=False)
|
||||
|
||||
assert len(active) == 1
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert len(raw["watches"]) == 1 # still on disk
|
||||
|
||||
|
||||
def test_list_active_selective_prune(store_path):
|
||||
"""Only entries with dead pids should be pruned."""
|
||||
h_alive = watch_registry.register("agent", {"label": "alive"}, storage_path=store_path)
|
||||
h_dead = watch_registry.register("agent", {"label": "dead"}, storage_path=store_path)
|
||||
|
||||
# Patch the dead entry's pid on disk to something definitely unused.
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
for watch in raw["watches"]:
|
||||
if watch["handle"] == h_dead:
|
||||
watch["pid"] = 999999
|
||||
store_path.write_text(json.dumps(raw, indent=2), encoding='utf-8')
|
||||
|
||||
active = watch_registry.list_active(storage_path=store_path, prune_stale=True)
|
||||
surviving_handles = {a["handle"] for a in active}
|
||||
assert h_alive in surviving_handles
|
||||
assert h_dead not in surviving_handles
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# is_pid_alive
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_is_pid_alive_current_process():
|
||||
assert watch_registry.is_pid_alive(os.getpid()) is True
|
||||
|
||||
|
||||
def test_is_pid_alive_impossible_pid():
|
||||
# 999999 is well above typical kernel.pid_max default — unlikely to exist.
|
||||
assert watch_registry.is_pid_alive(999999) is False
|
||||
|
||||
|
||||
def test_is_pid_alive_rejects_invalid_input():
|
||||
assert watch_registry.is_pid_alive(-1) is False
|
||||
assert watch_registry.is_pid_alive(0) is False
|
||||
# Non-int input can't be used by os.kill — handled defensively.
|
||||
assert watch_registry.is_pid_alive("1234") is False # type: ignore[arg-type]
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# kill_watch / kill_all
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_kill_watch_handle_not_found(store_path):
|
||||
result = watch_registry.kill_watch("ghost-abcdef", storage_path=store_path)
|
||||
assert result == {
|
||||
"handle": "ghost-abcdef",
|
||||
"killed": False,
|
||||
"was_alive": False,
|
||||
"reason": "handle not found",
|
||||
}
|
||||
|
||||
|
||||
def test_kill_watch_already_dead_pid(store_path):
|
||||
"""Handle for a dead pid should still be deregistered cleanly."""
|
||||
handle = watch_registry.register("timer", {}, storage_path=store_path)
|
||||
# Point the entry at a dead pid without touching the current pid of this test.
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
raw["watches"][0]["pid"] = 999999
|
||||
store_path.write_text(json.dumps(raw, indent=2), encoding='utf-8')
|
||||
|
||||
result = watch_registry.kill_watch(handle, storage_path=store_path)
|
||||
assert result["killed"] is True
|
||||
assert result["was_alive"] is False
|
||||
|
||||
# Deregistered
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert raw["watches"] == []
|
||||
|
||||
|
||||
def test_kill_watch_happy_path(store_path):
|
||||
"""Spawn a sleep subprocess, register it, kill via registry — verify dead."""
|
||||
proc = subprocess.Popen([sys.executable, "-c", "import time; time.sleep(30)"])
|
||||
try:
|
||||
# Manually install an entry pointing at the subprocess pid so kill_watch
|
||||
# targets it instead of this test's own pid.
|
||||
handle = watch_registry.register("timer", {"duration": "30s"}, storage_path=store_path)
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
for watch in raw["watches"]:
|
||||
if watch["handle"] == handle:
|
||||
watch["pid"] = proc.pid
|
||||
store_path.write_text(json.dumps(raw, indent=2), encoding='utf-8')
|
||||
|
||||
result = watch_registry.kill_watch(handle, storage_path=store_path)
|
||||
assert result["handle"] == handle
|
||||
assert result["was_alive"] is True
|
||||
assert result["killed"] is True
|
||||
|
||||
# Subprocess should have exited (wait briefly in case SIGTERM is still
|
||||
# in flight on a slow CI box).
|
||||
proc.wait(timeout=5)
|
||||
assert proc.returncode is not None
|
||||
finally:
|
||||
if proc.poll() is None:
|
||||
proc.kill()
|
||||
proc.wait(timeout=5)
|
||||
|
||||
|
||||
def test_kill_all_multiple_watches(store_path):
|
||||
h1 = watch_registry.register("timer", {}, storage_path=store_path)
|
||||
h2 = watch_registry.register("schedule", {}, storage_path=store_path)
|
||||
# Point both at dead pids so kill_all completes fast and doesn't touch real processes.
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
for watch in raw["watches"]:
|
||||
watch["pid"] = 999999
|
||||
store_path.write_text(json.dumps(raw, indent=2), encoding='utf-8')
|
||||
|
||||
results = watch_registry.kill_all(storage_path=store_path)
|
||||
handles = {r["handle"] for r in results}
|
||||
assert handles == {h1, h2}
|
||||
assert all(r["killed"] for r in results)
|
||||
|
||||
raw_after = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert raw_after["watches"] == []
|
||||
|
||||
|
||||
def test_kill_all_empty(store_path):
|
||||
assert watch_registry.kill_all(storage_path=store_path) == []
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Atomic write + concurrency sanity
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_atomic_write_leaves_no_tmp(store_path):
|
||||
watch_registry.register("timer", {}, storage_path=store_path)
|
||||
tmp_sibling = store_path.with_suffix(store_path.suffix + ".tmp")
|
||||
assert not tmp_sibling.exists()
|
||||
|
||||
|
||||
def test_sequential_register_deregister_preserves_entries(store_path):
|
||||
"""Sanity test for read-modify-write: many ops in a row don't lose data."""
|
||||
handles = [
|
||||
watch_registry.register("timer", {"i": i}, storage_path=store_path)
|
||||
for i in range(10)
|
||||
]
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert len(raw["watches"]) == 10
|
||||
|
||||
# Remove every other one.
|
||||
for handle in handles[::2]:
|
||||
assert watch_registry.deregister(handle, storage_path=store_path) is True
|
||||
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert len(raw["watches"]) == 5
|
||||
surviving = {w["handle"] for w in raw["watches"]}
|
||||
assert surviving == set(handles[1::2])
|
||||
|
||||
|
||||
def test_full_cycle_integration(store_path):
|
||||
"""register -> list -> deregister end to end."""
|
||||
h = watch_registry.register("agent", {"agent_id": "@drone"}, storage_path=store_path)
|
||||
|
||||
active = watch_registry.list_active(storage_path=store_path, prune_stale=False)
|
||||
assert len(active) == 1
|
||||
assert active[0]["handle"] == h
|
||||
assert active[0]["metadata"]["agent_id"] == "@drone"
|
||||
|
||||
assert watch_registry.deregister(h, storage_path=store_path) is True
|
||||
assert watch_registry.list_active(storage_path=store_path, prune_stale=False) == []
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Handler integration — each handler registers at start, deregisters on exit
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_agent_handler_registers_and_deregisters(store_path, monkeypatch):
|
||||
"""watch_agent registers on entry and deregisters in finally (even on early return)."""
|
||||
from aipass.devpulse.apps.handlers.watchdog import agent as agent_handler
|
||||
|
||||
# Force the handler's default storage path to our tmp file so register
|
||||
# lands in the right place.
|
||||
monkeypatch.setattr(
|
||||
watch_registry, "_default_storage_path", lambda: store_path
|
||||
)
|
||||
|
||||
# @__definitely_not_a_branch__ hits the early "not found" return —
|
||||
# exercises the deregister-in-finally path with minimal work.
|
||||
result = agent_handler.watch_agent("@__definitely_not_a_branch__", timeout_seconds=1)
|
||||
assert "handle" in result
|
||||
assert result["handle"].startswith("agent-")
|
||||
|
||||
# Registry should be empty after the call.
|
||||
assert watch_registry.list_active(storage_path=store_path, prune_stale=False) == []
|
||||
|
||||
|
||||
def test_timer_wake_in_registers_and_deregisters(store_path, monkeypatch):
|
||||
"""wake_in with a short duration registers then deregisters."""
|
||||
from aipass.devpulse.apps.handlers.watchdog import timer as timer_handler
|
||||
|
||||
monkeypatch.setattr(
|
||||
watch_registry, "_default_storage_path", lambda: store_path
|
||||
)
|
||||
|
||||
# Take a peek mid-flight by patching time.sleep to snapshot the registry.
|
||||
snapshots: list[list] = []
|
||||
real_sleep = time.sleep
|
||||
|
||||
def spy_sleep(duration):
|
||||
"""Capture the registry state while the timer is mid-wait."""
|
||||
snapshots.append(
|
||||
watch_registry.list_active(storage_path=store_path, prune_stale=False)
|
||||
)
|
||||
real_sleep(duration)
|
||||
|
||||
with patch("aipass.devpulse.apps.handlers.watchdog.timer.time.sleep", spy_sleep):
|
||||
result = timer_handler.wake_in("1s")
|
||||
|
||||
assert result["state"] == "woke"
|
||||
assert "handle" in result
|
||||
|
||||
# Mid-flight snapshot must have seen the entry.
|
||||
assert any(
|
||||
any(w["handle"].startswith("timer-") for w in snap)
|
||||
for snap in snapshots
|
||||
), "timer handler never registered mid-wait"
|
||||
|
||||
# After wake_in returns, the registry must be empty.
|
||||
assert watch_registry.list_active(storage_path=store_path, prune_stale=False) == []
|
||||
|
||||
|
||||
def test_schedule_wake_at_registers_and_deregisters(store_path, monkeypatch):
|
||||
"""wake_at with a tiny relative delay registers then deregisters."""
|
||||
from aipass.devpulse.apps.handlers.watchdog import schedule as schedule_handler
|
||||
|
||||
monkeypatch.setattr(
|
||||
watch_registry, "_default_storage_path", lambda: store_path
|
||||
)
|
||||
|
||||
# Fast-forward clock so wake_at returns immediately without real waiting.
|
||||
from datetime import datetime, timedelta
|
||||
start = datetime(2026, 4, 14, 12, 0, 0)
|
||||
calls = {"n": 0}
|
||||
|
||||
def fake_now():
|
||||
"""Returns ``start`` once then jumps 10s forward to trip the break."""
|
||||
calls["n"] += 1
|
||||
if calls["n"] == 1:
|
||||
return start
|
||||
return start + timedelta(seconds=10)
|
||||
|
||||
result = schedule_handler.wake_at("+5s", now_fn=fake_now)
|
||||
|
||||
assert result["state"] == "woke"
|
||||
assert "handle" in result
|
||||
|
||||
# Registry cleaned up.
|
||||
assert watch_registry.list_active(storage_path=store_path, prune_stale=False) == []
|
||||
|
||||
|
||||
def test_handler_deregisters_on_exception(store_path, monkeypatch):
|
||||
"""If a handler raises mid-wait, the finally block must still deregister."""
|
||||
from aipass.devpulse.apps.handlers.watchdog import timer as timer_handler
|
||||
|
||||
monkeypatch.setattr(
|
||||
watch_registry, "_default_storage_path", lambda: store_path
|
||||
)
|
||||
|
||||
# Make time.sleep raise after the register call.
|
||||
def exploding_sleep(duration):
|
||||
"""Simulate a KeyboardInterrupt / cancellation mid-sleep."""
|
||||
raise RuntimeError("boom")
|
||||
|
||||
with patch("aipass.devpulse.apps.handlers.watchdog.timer.time.sleep", exploding_sleep):
|
||||
with pytest.raises(RuntimeError, match="boom"):
|
||||
timer_handler.wake_in("5s")
|
||||
|
||||
# Even though wake_in raised, the finally block must have deregistered.
|
||||
assert watch_registry.list_active(storage_path=store_path, prune_stale=False) == []
|
||||
@@ -0,0 +1,358 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: test_watchdog_schedule.py
|
||||
# Description: Tests for the watchdog schedule handler + router integration
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
"""Tests for watchdog schedule handler (Phase 3, FPLAN-0186)."""
|
||||
|
||||
import sys
|
||||
from datetime import datetime, timedelta
|
||||
from unittest.mock import patch
|
||||
|
||||
import pytest
|
||||
|
||||
from aipass.devpulse.apps.handlers.watchdog import schedule as schedule_handler
|
||||
from aipass.devpulse.apps.modules import watchdog as wd_mod
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# parse_schedule — wall-clock
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def fixed_now():
|
||||
# 2026-04-14 10:00:00 — any 02:00 target is "past", any 14:30 is "future".
|
||||
return datetime(2026, 4, 14, 10, 0, 0)
|
||||
|
||||
|
||||
def test_parse_schedule_future_same_day(fixed_now):
|
||||
target = schedule_handler.parse_schedule("14:30", now=fixed_now)
|
||||
assert target == datetime(2026, 4, 14, 14, 30, 0)
|
||||
|
||||
|
||||
def test_parse_schedule_past_rolls_to_tomorrow(fixed_now):
|
||||
target = schedule_handler.parse_schedule("02:00", now=fixed_now)
|
||||
assert target == datetime(2026, 4, 15, 2, 0, 0)
|
||||
|
||||
|
||||
def test_parse_schedule_with_seconds(fixed_now):
|
||||
target = schedule_handler.parse_schedule("14:30:15", now=fixed_now)
|
||||
assert target == datetime(2026, 4, 14, 14, 30, 15)
|
||||
|
||||
|
||||
def test_parse_schedule_midnight_rolls_tomorrow(fixed_now):
|
||||
# 00:00 is strictly before 10:00 today -> tomorrow.
|
||||
target = schedule_handler.parse_schedule("00:00", now=fixed_now)
|
||||
assert target == datetime(2026, 4, 15, 0, 0, 0)
|
||||
|
||||
|
||||
def test_parse_schedule_end_of_day_same_day(fixed_now):
|
||||
target = schedule_handler.parse_schedule("23:59", now=fixed_now)
|
||||
assert target == datetime(2026, 4, 14, 23, 59, 0)
|
||||
|
||||
|
||||
def test_parse_schedule_equal_to_now_rolls_tomorrow(fixed_now):
|
||||
# Target equal to now is "past" by policy — rolls forward so a caller
|
||||
# who says "wake at 10:00" at 10:00 sharp gets tomorrow, not "now".
|
||||
target = schedule_handler.parse_schedule("10:00", now=fixed_now)
|
||||
assert target == datetime(2026, 4, 15, 10, 0, 0)
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# parse_schedule — relative
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.mark.parametrize("text,delta_seconds", [
|
||||
("+30m", 1800),
|
||||
("+1h", 3600),
|
||||
("+45s", 45),
|
||||
("+1h30m", 5400),
|
||||
("+2h", 7200),
|
||||
])
|
||||
def test_parse_schedule_relative(fixed_now, text, delta_seconds):
|
||||
target = schedule_handler.parse_schedule(text, now=fixed_now)
|
||||
assert target == fixed_now + timedelta(seconds=delta_seconds)
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# parse_schedule — invalid
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.mark.parametrize("text", [
|
||||
"",
|
||||
" ",
|
||||
"abc",
|
||||
"25:00",
|
||||
"12:60",
|
||||
"-5m",
|
||||
"+xyz",
|
||||
"+",
|
||||
"14:30:99",
|
||||
])
|
||||
def test_parse_schedule_invalid(text):
|
||||
with pytest.raises(ValueError):
|
||||
schedule_handler.parse_schedule(text, now=datetime(2026, 4, 14, 10, 0, 0))
|
||||
|
||||
|
||||
def test_parse_schedule_none_raises():
|
||||
with pytest.raises(ValueError):
|
||||
schedule_handler.parse_schedule(None) # type: ignore[arg-type]
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# format_wait
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_format_wait_hours_and_minutes():
|
||||
now = datetime(2026, 4, 14, 10, 0, 0)
|
||||
target = now + timedelta(hours=2, minutes=15)
|
||||
assert schedule_handler.format_wait(target, now) == "in 2h 15m"
|
||||
|
||||
|
||||
def test_format_wait_minutes_only():
|
||||
now = datetime(2026, 4, 14, 10, 0, 0)
|
||||
target = now + timedelta(minutes=5)
|
||||
assert schedule_handler.format_wait(target, now) == "in 5m"
|
||||
|
||||
|
||||
def test_format_wait_seconds():
|
||||
now = datetime(2026, 4, 14, 10, 0, 0)
|
||||
target = now + timedelta(seconds=45)
|
||||
assert schedule_handler.format_wait(target, now) == "in 45s"
|
||||
|
||||
|
||||
def test_format_wait_overdue():
|
||||
now = datetime(2026, 4, 14, 10, 0, 0)
|
||||
target = now - timedelta(minutes=3)
|
||||
assert schedule_handler.format_wait(target, now) == "overdue by 3m"
|
||||
|
||||
|
||||
def test_format_wait_zero():
|
||||
now = datetime(2026, 4, 14, 10, 0, 0)
|
||||
assert schedule_handler.format_wait(now, now) == "in 0s"
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# wake_at — injected clock
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
class _FakeClock:
|
||||
"""Clock that advances by a fixed delta on every call."""
|
||||
|
||||
def __init__(self, start: datetime, step_seconds: float = 2.0):
|
||||
self.current = start
|
||||
self.step = timedelta(seconds=step_seconds)
|
||||
|
||||
def __call__(self) -> datetime:
|
||||
value = self.current
|
||||
self.current = self.current + self.step
|
||||
return value
|
||||
|
||||
|
||||
def test_wake_at_relative_with_injected_clock():
|
||||
start = datetime(2026, 4, 14, 10, 0, 0)
|
||||
clock = _FakeClock(start, step_seconds=10.0)
|
||||
|
||||
with patch.object(schedule_handler.time, "sleep", return_value=None):
|
||||
result = schedule_handler.wake_at("+30s", now_fn=clock)
|
||||
|
||||
assert result["woke"] is True
|
||||
assert result["state"] == "woke"
|
||||
assert result["reason"] == "schedule fired"
|
||||
assert result["command"] is None
|
||||
assert result["command_exit_code"] is None
|
||||
assert result["command_stdout"] is None
|
||||
assert result["command_stderr"] is None
|
||||
assert result["scheduled_for"] == (start + timedelta(seconds=30)).isoformat()
|
||||
assert result["elapsed"] >= 30
|
||||
|
||||
|
||||
def test_wake_at_wall_clock_with_injected_clock():
|
||||
start = datetime(2026, 4, 14, 10, 0, 0)
|
||||
clock = _FakeClock(start, step_seconds=60.0)
|
||||
|
||||
with patch.object(schedule_handler.time, "sleep", return_value=None):
|
||||
result = schedule_handler.wake_at("10:05", now_fn=clock)
|
||||
|
||||
assert result["scheduled_for"] == datetime(2026, 4, 14, 10, 5, 0).isoformat()
|
||||
assert result["elapsed"] >= 300
|
||||
|
||||
|
||||
def test_wake_at_real_short_relative():
|
||||
"""Use a real 1s sleep to verify wake_at actually returns without mocking."""
|
||||
result = schedule_handler.wake_at("+1s")
|
||||
assert result["woke"] is True
|
||||
assert result["elapsed"] >= 1
|
||||
assert result["command"] is None
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# wake_at — command execution
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_wake_at_runs_echo_command():
|
||||
start = datetime(2026, 4, 14, 10, 0, 0)
|
||||
clock = _FakeClock(start, step_seconds=5.0)
|
||||
|
||||
with patch.object(schedule_handler.time, "sleep", return_value=None):
|
||||
result = schedule_handler.wake_at("+1s", command="echo hello", now_fn=clock)
|
||||
|
||||
assert result["command"] == "echo hello"
|
||||
assert result["command_exit_code"] == 0
|
||||
assert "hello" in (result["command_stdout"] or "")
|
||||
|
||||
|
||||
def test_wake_at_failing_command_no_exception():
|
||||
start = datetime(2026, 4, 14, 10, 0, 0)
|
||||
clock = _FakeClock(start, step_seconds=5.0)
|
||||
|
||||
with patch.object(schedule_handler.time, "sleep", return_value=None):
|
||||
result = schedule_handler.wake_at("+1s", command="false", now_fn=clock)
|
||||
|
||||
assert result["command"] == "false"
|
||||
assert result["command_exit_code"] != 0
|
||||
|
||||
|
||||
def test_wake_at_nonexistent_command():
|
||||
start = datetime(2026, 4, 14, 10, 0, 0)
|
||||
clock = _FakeClock(start, step_seconds=5.0)
|
||||
|
||||
with patch.object(schedule_handler.time, "sleep", return_value=None):
|
||||
result = schedule_handler.wake_at(
|
||||
"+1s",
|
||||
command="this-command-does-not-exist-xyz-123",
|
||||
now_fn=clock,
|
||||
)
|
||||
|
||||
# shell=True routes through /bin/sh which reports 127 for not-found.
|
||||
assert result["command_exit_code"] != 0
|
||||
assert result["command"] == "this-command-does-not-exist-xyz-123"
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Chunked sleep
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_wake_at_sleeps_in_chunks():
|
||||
"""Long waits should call time.sleep repeatedly, never once with a big value."""
|
||||
start = datetime(2026, 4, 14, 10, 0, 0)
|
||||
# Real-time clock — no advancement per call — so wake_at relies on sleep.
|
||||
# We fake sleep to advance our synthetic clock instead.
|
||||
state = {"now": start}
|
||||
|
||||
def fake_clock():
|
||||
return state["now"]
|
||||
|
||||
def fake_sleep(seconds):
|
||||
state["now"] = state["now"] + timedelta(seconds=seconds)
|
||||
|
||||
with patch.object(schedule_handler.time, "sleep", side_effect=fake_sleep) as sleep_mock:
|
||||
schedule_handler.wake_at("+30s", now_fn=fake_clock)
|
||||
|
||||
# Each chunk is capped at _SLEEP_CHUNK_SECONDS (5.0) so 30s -> >= 6 calls.
|
||||
max_chunk = schedule_handler._SLEEP_CHUNK_SECONDS
|
||||
assert sleep_mock.call_count >= 6
|
||||
for call in sleep_mock.call_args_list:
|
||||
(value,) = call.args
|
||||
assert value <= max_chunk + 0.0001
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Router integration
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def _bypass_caller_guard():
|
||||
with patch.object(wd_mod, "_guard_caller", return_value=True):
|
||||
yield
|
||||
|
||||
|
||||
def _fake_schedule_module(**overrides):
|
||||
"""Build a fake schedule module with the public surface the router calls."""
|
||||
fake = type(sys)("fake_schedule_mod")
|
||||
fake.calls = []
|
||||
|
||||
def wake_at(time_str, command=None, now_fn=None):
|
||||
fake.calls.append(("wake_at", time_str, command))
|
||||
return overrides.get("wake_at", {
|
||||
"woke": True,
|
||||
"reason": "schedule fired",
|
||||
"elapsed": 0,
|
||||
"scheduled_for": "2026-04-14T10:00:00",
|
||||
"state": "woke",
|
||||
"command": command,
|
||||
"command_exit_code": 0 if command else None,
|
||||
"command_stdout": "hi\n" if command else None,
|
||||
"command_stderr": "" if command else None,
|
||||
})
|
||||
|
||||
fake.wake_at = wake_at
|
||||
return fake
|
||||
|
||||
|
||||
def test_router_schedule_wall_clock(_bypass_caller_guard):
|
||||
fake = _fake_schedule_module()
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
result = wd_mod.handle_command("watchdog", ["schedule", "02:00"])
|
||||
assert result is True
|
||||
assert ("wake_at", "02:00", None) in fake.calls
|
||||
|
||||
|
||||
def test_router_schedule_relative(_bypass_caller_guard):
|
||||
fake = _fake_schedule_module()
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
wd_mod.handle_command("watchdog", ["schedule", "+30m"])
|
||||
assert ("wake_at", "+30m", None) in fake.calls
|
||||
|
||||
|
||||
def test_router_schedule_with_command(_bypass_caller_guard):
|
||||
fake = _fake_schedule_module()
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
wd_mod.handle_command("watchdog", ["schedule", "02:00", "drone @git status"])
|
||||
assert ("wake_at", "02:00", "drone @git status") in fake.calls
|
||||
|
||||
|
||||
def test_router_schedule_help(_bypass_caller_guard, capsys):
|
||||
result = wd_mod.handle_command("watchdog", ["schedule", "--help"])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "schedule" in combined.lower()
|
||||
|
||||
|
||||
def test_router_schedule_empty_shows_help(_bypass_caller_guard, capsys):
|
||||
# Without a positional time, the subhandler prints help rather than erroring.
|
||||
result = wd_mod.handle_command("watchdog", ["schedule"])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "schedule" in combined.lower()
|
||||
|
||||
|
||||
def test_router_schedule_invalid_time(_bypass_caller_guard, capsys):
|
||||
"""ValueError from wake_at surfaces as a clean router error."""
|
||||
fake = type(sys)("fake_schedule_mod")
|
||||
|
||||
def wake_at(time_str, command=None, now_fn=None):
|
||||
raise ValueError(f"bad schedule: {time_str}")
|
||||
|
||||
fake.wake_at = wake_at
|
||||
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
result = wd_mod.handle_command("watchdog", ["schedule", "notatime"])
|
||||
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "invalid" in combined.lower()
|
||||
@@ -0,0 +1,366 @@
|
||||
# =================== AIPass ====================
|
||||
# Name: test_watchdog_timer.py
|
||||
# Description: Tests for the watchdog timer handler + router integration
|
||||
# Version: 1.0.0
|
||||
# Created: 2026-04-14
|
||||
# Modified: 2026-04-14
|
||||
# =============================================
|
||||
|
||||
"""Tests for watchdog timer handler (Phase 2, FPLAN-0186)."""
|
||||
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
import pytest
|
||||
|
||||
from aipass.devpulse.apps.handlers.watchdog import timer as timer_handler
|
||||
from aipass.devpulse.apps.modules import watchdog as wd_mod
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# parse_duration
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.mark.parametrize("text,expected", [
|
||||
("30s", 30),
|
||||
("5m", 300),
|
||||
("2h", 7200),
|
||||
("1h30m", 5400),
|
||||
("45", 45),
|
||||
("0s", 0),
|
||||
("120", 120),
|
||||
("1h1m1s", 3661),
|
||||
])
|
||||
def test_parse_duration_valid(text, expected):
|
||||
assert timer_handler.parse_duration(text) == expected
|
||||
|
||||
|
||||
@pytest.mark.parametrize("text", [
|
||||
"abc",
|
||||
"",
|
||||
" ",
|
||||
"-5m",
|
||||
"5x",
|
||||
"5m5",
|
||||
"hm",
|
||||
])
|
||||
def test_parse_duration_invalid(text):
|
||||
with pytest.raises(ValueError):
|
||||
timer_handler.parse_duration(text)
|
||||
|
||||
|
||||
def test_parse_duration_none_raises():
|
||||
with pytest.raises(ValueError):
|
||||
timer_handler.parse_duration(None) # type: ignore[arg-type]
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# format_human
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.mark.parametrize("seconds,expected", [
|
||||
(1, "1s"),
|
||||
(59, "59s"),
|
||||
(60, "1m 00s"),
|
||||
(125, "2m 05s"),
|
||||
(3600, "1h 0m 00s"),
|
||||
(5400, "1h 30m 00s"),
|
||||
(3725, "1h 2m 05s"),
|
||||
(0, "0s"),
|
||||
])
|
||||
def test_format_human(seconds, expected):
|
||||
assert timer_handler.format_human(seconds) == expected
|
||||
|
||||
|
||||
def test_format_human_rejects_negative():
|
||||
with pytest.raises(ValueError):
|
||||
timer_handler.format_human(-1)
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# wake_in
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_wake_in_short_duration():
|
||||
"""Very short wake-in returns the expected shape and elapsed is in the ballpark."""
|
||||
started = time.monotonic()
|
||||
result = timer_handler.wake_in("1s")
|
||||
elapsed_real = time.monotonic() - started
|
||||
|
||||
assert result["woke"] is True
|
||||
assert result["state"] == "woke"
|
||||
assert result["reason"] == "timer fired"
|
||||
assert result["duration"] == "1s"
|
||||
assert 1 <= result["elapsed"] <= 3
|
||||
assert elapsed_real < 3
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# timer_start / timer_stop
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def store_path(tmp_path):
|
||||
return tmp_path / "watchdog_timers.json"
|
||||
|
||||
|
||||
def test_timer_start_then_stop(store_path):
|
||||
start = timer_handler.timer_start("phase-a", storage_path=store_path)
|
||||
assert start["state"] == "started"
|
||||
assert start["name"] == "phase-a"
|
||||
assert "started_at" in start
|
||||
|
||||
time.sleep(1.1)
|
||||
|
||||
stop = timer_handler.timer_stop("phase-a", storage_path=store_path)
|
||||
assert stop["state"] == "stopped"
|
||||
assert stop["name"] == "phase-a"
|
||||
assert stop["elapsed_seconds"] >= 1
|
||||
assert stop["human"].endswith("s")
|
||||
assert "stopped_at" in stop
|
||||
|
||||
raw = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert "phase-a" not in raw["active"]
|
||||
assert len(raw["history"]) == 1
|
||||
assert raw["history"][0]["name"] == "phase-a"
|
||||
|
||||
|
||||
def test_timer_start_duplicate_returns_error(store_path):
|
||||
timer_handler.timer_start("dup", storage_path=store_path)
|
||||
second = timer_handler.timer_start("dup", storage_path=store_path)
|
||||
assert second["state"] == "error"
|
||||
assert "already running" in second["reason"]
|
||||
|
||||
|
||||
def test_timer_stop_without_start_returns_error(store_path):
|
||||
result = timer_handler.timer_stop("ghost", storage_path=store_path)
|
||||
assert result["state"] == "error"
|
||||
assert "not running" in result["reason"]
|
||||
|
||||
|
||||
def test_timer_start_empty_name(store_path):
|
||||
result = timer_handler.timer_start(" ", storage_path=store_path)
|
||||
assert result["state"] == "error"
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# timer_list
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_timer_list_mixes_active_and_history(store_path):
|
||||
timer_handler.timer_start("alpha", storage_path=store_path)
|
||||
timer_handler.timer_start("beta", storage_path=store_path)
|
||||
time.sleep(1.1)
|
||||
timer_handler.timer_stop("beta", storage_path=store_path)
|
||||
|
||||
snapshot = timer_handler.timer_list(storage_path=store_path)
|
||||
active_names = [a["name"] for a in snapshot["active"]]
|
||||
history_names = [h["name"] for h in snapshot["history"]]
|
||||
|
||||
assert "alpha" in active_names
|
||||
assert "beta" not in active_names
|
||||
assert "beta" in history_names
|
||||
|
||||
alpha_entry = next(a for a in snapshot["active"] if a["name"] == "alpha")
|
||||
assert alpha_entry["elapsed_so_far_seconds"] >= 0
|
||||
assert "human" in alpha_entry
|
||||
|
||||
|
||||
def test_timer_list_empty_store(store_path):
|
||||
snapshot = timer_handler.timer_list(storage_path=store_path)
|
||||
assert snapshot == {"active": [], "history": []}
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# timer_report
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_timer_report_contains_sections(store_path):
|
||||
timer_handler.timer_start("reporting", storage_path=store_path)
|
||||
time.sleep(1.1)
|
||||
timer_handler.timer_stop("reporting", storage_path=store_path)
|
||||
timer_handler.timer_start("still-active", storage_path=store_path)
|
||||
|
||||
report = timer_handler.timer_report(storage_path=store_path)
|
||||
assert "Watchdog Timer Report" in report
|
||||
assert "Active:" in report
|
||||
assert "History" in report
|
||||
assert "reporting" in report
|
||||
assert "still-active" in report
|
||||
assert "Total tracked" in report
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Persistence + atomic writes
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_persistence_across_reloads(store_path):
|
||||
timer_handler.timer_start("persistent", storage_path=store_path)
|
||||
|
||||
raw_after_start = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert "persistent" in raw_after_start["active"]
|
||||
|
||||
time.sleep(1.1)
|
||||
stop_result = timer_handler.timer_stop("persistent", storage_path=store_path)
|
||||
assert stop_result["state"] == "stopped"
|
||||
|
||||
raw_after_stop = json.loads(store_path.read_text(encoding='utf-8'))
|
||||
assert "persistent" not in raw_after_stop["active"]
|
||||
assert any(h["name"] == "persistent" for h in raw_after_stop["history"])
|
||||
|
||||
|
||||
def test_atomic_write_cleans_up_tmp(store_path):
|
||||
timer_handler.timer_start("atomic", storage_path=store_path)
|
||||
tmp_sibling = store_path.with_suffix(store_path.suffix + ".tmp")
|
||||
assert not tmp_sibling.exists()
|
||||
timer_handler.timer_stop("atomic", storage_path=store_path)
|
||||
assert not tmp_sibling.exists()
|
||||
|
||||
|
||||
def test_concurrent_timers_independent(store_path):
|
||||
timer_handler.timer_start("t1", storage_path=store_path)
|
||||
timer_handler.timer_start("t2", storage_path=store_path)
|
||||
timer_handler.timer_stop("t1", storage_path=store_path)
|
||||
|
||||
snapshot = timer_handler.timer_list(storage_path=store_path)
|
||||
active_names = [a["name"] for a in snapshot["active"]]
|
||||
assert active_names == ["t2"]
|
||||
assert [h["name"] for h in snapshot["history"]] == ["t1"]
|
||||
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Router integration
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def _bypass_caller_guard():
|
||||
"""Force _guard_caller pass so router tests don't depend on cwd."""
|
||||
with patch.object(wd_mod, "_guard_caller", return_value=True):
|
||||
yield
|
||||
|
||||
|
||||
def _fake_timer_module(**overrides):
|
||||
"""Build a fake handler module with the public surface the router calls."""
|
||||
fake = type(sys)("fake_timer_mod")
|
||||
fake.calls = []
|
||||
|
||||
def wake_in(duration):
|
||||
fake.calls.append(("wake_in", duration))
|
||||
return overrides.get("wake_in", {
|
||||
"woke": True,
|
||||
"reason": "timer fired",
|
||||
"elapsed": 1,
|
||||
"duration": duration,
|
||||
"state": "woke",
|
||||
})
|
||||
|
||||
def timer_start(name, storage_path=None):
|
||||
fake.calls.append(("timer_start", name))
|
||||
return overrides.get("timer_start", {
|
||||
"name": name,
|
||||
"started_at": "now",
|
||||
"state": "started",
|
||||
})
|
||||
|
||||
def timer_stop(name, storage_path=None):
|
||||
fake.calls.append(("timer_stop", name))
|
||||
return overrides.get("timer_stop", {
|
||||
"name": name,
|
||||
"elapsed_seconds": 12,
|
||||
"human": "12s",
|
||||
"state": "stopped",
|
||||
})
|
||||
|
||||
def timer_list(storage_path=None):
|
||||
fake.calls.append(("timer_list",))
|
||||
return overrides.get("timer_list", {"active": [], "history": []})
|
||||
|
||||
def timer_report(storage_path=None):
|
||||
fake.calls.append(("timer_report",))
|
||||
return overrides.get("timer_report", "report body")
|
||||
|
||||
fake.wake_in = wake_in
|
||||
fake.timer_start = timer_start
|
||||
fake.timer_stop = timer_stop
|
||||
fake.timer_list = timer_list
|
||||
fake.timer_report = timer_report
|
||||
return fake
|
||||
|
||||
|
||||
def test_router_timer_wake_in(_bypass_caller_guard, capsys):
|
||||
fake = _fake_timer_module()
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
result = wd_mod.handle_command("watchdog", ["timer", "1s"])
|
||||
assert result is True
|
||||
assert ("wake_in", "1s") in fake.calls
|
||||
|
||||
|
||||
def test_router_timer_start(_bypass_caller_guard):
|
||||
fake = _fake_timer_module()
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
wd_mod.handle_command("watchdog", ["timer", "start", "build-phase-3"])
|
||||
assert ("timer_start", "build-phase-3") in fake.calls
|
||||
|
||||
|
||||
def test_router_timer_stop(_bypass_caller_guard):
|
||||
fake = _fake_timer_module()
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
wd_mod.handle_command("watchdog", ["timer", "stop", "build-phase-3"])
|
||||
assert ("timer_stop", "build-phase-3") in fake.calls
|
||||
|
||||
|
||||
def test_router_timer_list(_bypass_caller_guard):
|
||||
fake = _fake_timer_module()
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
wd_mod.handle_command("watchdog", ["timer", "list"])
|
||||
assert ("timer_list",) in fake.calls
|
||||
|
||||
|
||||
def test_router_timer_report(_bypass_caller_guard, capsys):
|
||||
fake = _fake_timer_module()
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
wd_mod.handle_command("watchdog", ["timer", "report"])
|
||||
assert ("timer_report",) in fake.calls
|
||||
captured = capsys.readouterr()
|
||||
assert "report body" in (captured.out + captured.err)
|
||||
|
||||
|
||||
def test_router_timer_help(_bypass_caller_guard, capsys):
|
||||
result = wd_mod.handle_command("watchdog", ["timer", "--help"])
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "timer" in combined.lower()
|
||||
|
||||
|
||||
def test_router_timer_invalid_duration(_bypass_caller_guard, capsys):
|
||||
"""Invalid duration from wake_in surfaces as a clean error via the router."""
|
||||
fake = type(sys)("fake_timer_mod")
|
||||
|
||||
def wake_in(duration):
|
||||
raise ValueError(f"bad duration: {duration}")
|
||||
|
||||
fake.wake_in = wake_in
|
||||
fake.timer_start = lambda *a, **kw: {}
|
||||
fake.timer_stop = lambda *a, **kw: {}
|
||||
fake.timer_list = lambda *a, **kw: {}
|
||||
fake.timer_report = lambda *a, **kw: ""
|
||||
|
||||
with patch("importlib.import_module", return_value=fake):
|
||||
result = wd_mod.handle_command("watchdog", ["timer", "notaduration"])
|
||||
|
||||
assert result is True
|
||||
captured = capsys.readouterr()
|
||||
combined = captured.out + captured.err
|
||||
assert "invalid" in combined.lower()
|
||||
@@ -100,7 +100,6 @@ BRANCH_COLORS = {
|
||||
'AIPASS': 'bold white', # All internal AIPass agent sessions
|
||||
'AIPL': 'bright_blue', # AIPL external project
|
||||
'VERA-STUDIO': 'bright_cyan', # Vera Studio external project
|
||||
'COMPASS': 'bright_green', # Compass external project
|
||||
}
|
||||
|
||||
|
||||
|
||||
@@ -536,5 +536,5 @@ from aipass.seedgo.apps.handlers.json import json_handler
|
||||
#@comments:2025-11-29:claude: Added explicit "Drone: CLI Router Pattern" section with FORBIDDEN import examples
|
||||
#@comments:2025-11-29:claude: Clarified that branches should NEVER import from aipass.drone.apps.modules - Drone resolves @ before calling them
|
||||
#@comments:2026-03-07:claude: Cleaned old references - updated to AIPass namespace imports, removed sys.path/AIPASS_ROOT patterns, fixed hardcoded home paths, old->seedgo naming
|
||||
#@comments:2026-03-08:claude: Resolved patrick's "whole file needs updating" flag - fixed old branch names (3 occurrences), removed stale AIPASS_ROOT reference, corrected handler independence summary to reflect that handlers CAN import cross-branch services, bumped to Draft v2
|
||||
#@comments:2026-03-08:claude: Resolved the user's "whole file needs updating" flag - fixed old branch names (3 occurrences), removed stale AIPASS_ROOT reference, corrected handler independence summary to reflect that handlers CAN import cross-branch services, bumped to Draft v2
|
||||
#@comments:2026-03-31:claude: Cleaned stale references in comments and docs throughout seedgo standards
|
||||
|
||||
@@ -25,7 +25,7 @@
|
||||
- FINAL: 176 tests, all passing. Test coverage 6%→93%. Overall audit 99%.
|
||||
- Ready to PR
|
||||
|
||||
## Questions for Patrick
|
||||
## Questions for the AIPass Developer
|
||||
(none yet)
|
||||
|
||||
## Notes
|
||||
|
||||
@@ -63,7 +63,7 @@ drone @spawn regenerate-registry --all # Regenerate all template class re
|
||||
**External project support:**
|
||||
```bash
|
||||
# Creating inside an existing AIPass project auto-detects the project registry
|
||||
drone @spawn create ~/Projects/compass/navigator # Registers in COMPASS_REGISTRY.json
|
||||
drone @spawn create ~/Projects/MyProject/agent_name # Registers in MYPROJECT_REGISTRY.json
|
||||
```
|
||||
|
||||
**Python API:**
|
||||
|
||||
Reference in New Issue
Block a user