#668 — poll loop re-drained a rate-limited backlog in a flood loop. The offset advanced AFTER process_update, so a rate-limited/rejected/erroring update never advanced it and the same backlog was re-fetched. Fix: advance the offset BEFORE process_update, so a consumed update never pins it (base_bot.py run loop).
#669 — three fixes: (1) systemd unit gets KillMode=process so a Restart is not killed by the old instance's cgroup teardown (the suicide-loop); (2) create_bot_via_botfather now RAISES RuntimeError with an actionable message (names the set-secret command) instead of silently returning None when telethon config is missing/unready — fail-honestly (botfather_client.py); (3) stale config-mechanism docstrings corrected (bot_factory/bot_operations).
Bonus (unbriefed but correct + beneficial): @skills also Windows-hardened _is_pid_alive (OpenProcess+GetExitCodeProcess on win32, os.kill moved into the POSIX branch) + refactored _check_lock to use it, and switched TEMP_DIR to tempfile.gettempdir(). Side effect: base_bot.py os.kill is now platform-guarded.
Built by @skills, verified by devpulse: 653 telegram tests green (incl lock/pid tests exercising the refactor); #668 offset-before-process verified by inspection; #669.2 raise covered by test_botfather_client. Note: @skills dispatch bounced on a usage-limit retry AFTER completing the work — verified the on-disk result independently.
Rides PR#659 (issue-clearing, no main-merge). Source: devpulse todos #41/#52.
Two hardening items surfaced during #664 that @memory could not touch (cross-branch edit gate blocked it).
ITEM 1 — _find_repo_root fail-loud (lifecycle/rollover.py). The PreCompact rollover hook's _find_repo_root() returned None SILENTLY when AIPASS_HOME/cwd was wrong -> rollover no-ops invisibly (the exact silent-skip that hid #664 for months). Now logs a logger.error with the AIPASS_HOME value + cwd before returning None (still degrades, just visibly).
ITEM 2 — edit_gate soft entry-count guard (security/edit_gate.py). edit_gate enforced per-entry CHARACTER caps but not entry COUNTS, so a branch could drift past its count cap between rollovers. New _check_section_counts warns (NEVER blocks) when a rolling section exceeds its cap, reading the SAME memory.config.json rollover caps @memory uses (config_loader.section('rollover') -> per_branch/defaults -> count); wrapped so a config-import failure degrades silently.
Built by @hooks, verified by devpulse: 70 tests green (+14 incl never-blocks guarantee, boundary cases, per-branch override, import-failure resilience); LIVE repro proves item1 logs the error on a bad root and item2 warns over-cap (20/15) without blocking; config structure confirmed to match memory's real caps (not inert).
Rides PR#659 (issue-clearing, no main-merge). Source: #664 verify (S292).
registry.is_owner (apps/handlers/registry.py:382) @-normalized the email but never lowercased, so a mixed-case branch name (registry names are mixed-case: DEVPULSE vs devpulse) returned False against the seated owner while the lowercase form returned True. Harmless today — the only live caller (@ai_mail dispatch_monitor._wake_sender) lowercases first — but the frozen TDPLAN-0012 contract promises a normalized email, and PART-4 owner-gating of watchdog/feedback may pass a raw branch name.
Fix: lowercase BOTH sides of the comparison (passed-in email AND registry owner email), @-strip preserved. +1 case-insensitivity test. Built by @spawn, verified by devpulse: LIVE repro — every case variant of the owner (DEVPULSE/@DEVPULSE/DevPulse) resolves True, non-owners (seedgo/@SEEDGO) and empty stay False; 316 spawn tests green (+1), seedgo 100%.
Rides PR#659 (issue-clearing, no main-merge). Source: #678/TDPLAN-0012 verify.
Two rough edges on the JSONL stall detector, both hardened in one pass on my own module (apps/handlers/watchdog/agent.py).
PART 1 (false-positive): _has_jsonl_activity inferred liveness purely from JSONL file-size growth over the 120s window. An agent doing ONE genuinely long operation (big Read, long Bash, heavy compute) writes no new JSONL lines for that span -> read as idle -> STALLED fires WHILE the agent is actively working. Fix: watch_agent now also treats an in-flight tool_use as activity. While a tool runs, the assistant's tool_use is the last transcript entry; new _last_entry_is_inflight_tool() tail-reads the newest .jsonl and detects it (fully defensive -> False on any parse/shape drift, degrading to size-based). LIVE-PROVEN against real Claude Code transcripts: sampled my own session across a 10s in-flight bash -> tool_use line is written at tool START and persists the whole call (the sub-second flush lag is irrelevant at the 120s horizon).
PART 2 (invisible stall): the stall only hit _stderr()+logger. The Monitor tool that arms the watchdog turns each STDOUT line into a live event but only captures stderr to a file (never surfaced) -> devpulse never saw the stall until the 600s timeout. Fix: new _stdout_event() emits the stall (+ a long-running-tool advisory for a possibly-hung tool, + a resumed signal) to stdout so Monitor relays it live; the verbose trail stays on stderr+logger.
Stall logic extracted into a StallTracker class (kills deep-nesting). +9 tests (unit + full-loop stdout proofs + real-transcript schema check); 142 watchdog tests green, seedgo audit 100%, no type errors.
Rides PR#659 (issue-clearing campaign, no main-merge).
@spawn: owner + registry_id written into the SEALED registry entries (authority lives in registry, not the self-editable passport). ensure_project_has_owner() now keys off citizen_class=manager (was earliest-created, which mislabeled @aipass) and writes the registry entry. get_owner()/is_owner() resolvers added. 315 tests, seedgo 100%.
@hooks: new registry_gate PreToolUse handler seals *_REGISTRY.json — blocks raw writes/tee/sed/rm + Edit/Write/MultiEdit, redirects to drone @spawn; per-clause bypass defeats compound-command smuggling; reads allowed. 82 tests, seedgo 100%.
@ai_mail: wake-back reslope — SKIP_SENDERS blocklist replaced by an is_owner allowlist. Only the project owner is woken when their dispatched agent completes; all other guards intact (depth cap, lock, occupancy, honest messaging, dispatch_wake.log). seedgo 100% on the changed file.
devpulse cross-part verify (REAL unmocked resolver): is_owner resolves devpulse-only; gate 13/13 incl compound-smuggle blocked + reads/drone-@spawn allowed; wake-back wakes owner / skips non-owner / respects depth-cap; 195 new-suite tests green together. Note: AIPASS_REGISTRY.json is gitignored — this ships the CODE; owner data regenerates per-install via ensure_project_has_owner. Still open (PART 4): gate watchdog+feedback on is_owner; portability of owner-only privileges across projects.
- New stdlib-only bash launcher at repo root: pre-setup only 'install' works (delegates to setup.sh, full flag pass-through); post-setup execs the venv aipass binary transparently
- 13 launcher tests (tests/test_launcher.py), bypass.json architecture entry for the test file
- README Quick Start leads with ./aipass install; CHANGELOG entry
CI flagged @aipass at 99%: install.py had no no-args introspection gate and a direct mkdir. Both are sanctioned patterns (binary-invoked 'aipass install' bare-runs; bootstrap dir-prep before services exist) matching doctor/init_flow/profile precedent. @aipass authored the bypass entries; landing them. Local audit now 100%, pyright clean.
Follow-up to 4105a7e. CI Windows Test surfaced 3 Windows-only test failures where
the sweep cross-platformed the code but tests still asserted POSIX behavior, plus
a fleet-wide architecture ripple from the template scaffold test:
- ai_mail: 3 Popen detach sites now use creationflags=CREATE_NEW_PROCESS_GROUP on
win32 / start_new_session on POSIX (real win32 detachment); test asserts the
platform-correct kwargs.
- drone: test_rm.py /tmp assertions guarded to POSIX-only (win32 uses
tempfile.gettempdir()); the swept code already dropped hardcoded /tmp on win32.
- seedgo: architecture checker exempts template scaffold test files
(test_scaffold.py via TEMPLATE_IGNORE_PATTERNS) from branch conformance -- a
template example test is not a structural requirement in every branch. Restores
all 17 branches to 100%.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013uzDhtcZ6wT1T9e2AHPQig
The 3 branches the template checker correctly flagged had never had their
.aipass/aipass_local_prompt.md filled in — they booted with a NEEDS CONFIGURATION
placeholder and no branch-specific identity. Each branch wrote its own real prompt
(identity, key commands, architecture, critical rules, integration points;
~63-67 lines, PROMPT_STYLE.md format).
Dispatched @cli/@drone/@prax (each owns its identity); verified independently —
0 stub markers, all three Template 100%, real coherent content.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013uzDhtcZ6wT1T9e2AHPQig
Pass 3 / final. The definitive-marker scan still matched {{BRANCH}} inside markdown
inline code — spawn's README documents 'Replace `{{BRANCH}}` in...', which is
scaffolding docs, not an un-rendered stub (spawn scored 66%). For .md files, fenced
+ inline code is now stripped once up front before BOTH the definitive and
single-curly scans; passport.json (JSON) still scans raw. Safe because real stubs
carry markers in prose/headings (the '## Status: NEEDS CONFIGURATION' line), never
exclusively in code.
Verified system-wide: Template avg 80%→94%; spawn + seedgo cleared to 100%; only
the three genuine unconfigured prompt stubs (cli/drone/prax) still flag. +3 tests
(24/24), full suite green (1132). Completes the checker-solid work begun in 26893fb.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013uzDhtcZ6wT1T9e2AHPQig
The advisory 'template' stale-checker matched marker strings anywhere in a file,
firing on documentation ABOUT templates rather than un-rendered stubs. Two root
causes fixed:
1. Scanned .trinity/*.json (all memory) — local.json/observations.json accumulate
marker mentions (seedgo's own note about the checker, prax's template_pusher
note). Now scans passport.json only, the sole spawn-templated trinity file.
2. Single-curly {…} regex ran on every .md, matching inline JSON/f-strings/code
paths in READMEs. Now single-curly detection runs on the branch prompt only
(README template has no single-curly placeholders) and strips fenced + inline
code first.
Definitive-marker detection unchanged — real stubs (cli/drone/prax prompts) still
flag. Verified live: seedgo 100%, drone/prax flag only the real prompt stub.
+4 tests (21/21). Dispatched to @seedgo (owner), verified independently.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013uzDhtcZ6wT1T9e2AHPQig
_advance_pending kept the pending file with a frozen processing_message_id, so
any Stop firing without a fresh placeholder (remote/mirror input or multi-Stop
turns) re-edited the same Telegram message instead of posting a new one. Clear
processing_message_id after the first delivery so subsequent Stops fall through
to send a new message. Live-proven on the devpulse bot; +2 regression tests
in TestAdvancePending (114/114).
devpulse-tier (owner) verb: fetch origin, VERSION GUARD (tag X.Y.Z must match origin/main pyproject + __init__), EXISTS GUARD (refuse if tag exists), tag origin/main + push -> fires publish.yml. 'tag --list' is global tier. Removes the last user-input step from releases (Patrick request, S274). Merge playbook (merge.md) updated to use it. Verb proven live: cut v2.6.1.
PATCH bump riding into PR#646 so main's merge commit carries the release version. pyproject + __init__ = 2.6.1 (must match the v2.6.1 tag). CHANGELOG [2026-07-02] leads with the release rollup + all 6 CI-stabilization fixes.
test_partial_line_not_consumed asserted +1 byte for the newline, but write_text() text mode translates \n->\r\n on Windows (2 bytes) -> off-by-one, failing windows-setup only. Switched both transcript write sites to write_bytes() for deterministic LF cross-platform. Production _tail_transcript_bytes is already CRLF-safe (reads rb, splits b'\n', strips \r) — test-only fix.
.gitignore exceptions still pointed at templates/builder/ after the TDPLAN-0010 rename (13463c0), so DASHBOARD.local.json + 10 other template dirs/files under templates/aipass_framework/ were silently gitignored — on disk (dirty tree passed) but absent in clean clones/CI. Result: spawn produced no DASHBOARD.local.json and test_full_spawn failed only in a clean checkout. Fixed all 23 .gitignore exception paths + tracked the now-visible template files (all placeholder/seed content: {{BRANCHNAME}}/{{DATE}}/{{CITIZEN_NUMBER}}).
Root-cause fixes for PR#646 red (dev broke after DPLAN-0226/FPLAN-0289/TDPLAN-0010 batch):
- seedgo: branch_audit honors ADVISORY (template_check no longer averaged into gate) + presence_gate added to hooks-snapshot fixture (4 tests)
- hooks: cc_sessions README entry + seedgo modules bypass (reads external ~/.claude, not branch data)
- spawn: retire passport(disabled).py/passport_ops(disabled).py to .archive/ (disabled suffix kept broken cross-import visible to type checker)
- ai_mail: broker-fd test gives testbranch a real .trinity/passport.json for the new marker-walk resolution (f914ab6)
Navmap (@hooks): one bullet in the Memory section — caps are hook-enforced,
the live cap is rendered in each file's *_meta line; read it before writing,
draft to ~80%, one-pass rewrite if rejected. Devpulse branch prompt: same
behavior, explicitly notes caps are NOT listed (single source =
memory.config.json → entry_limits, auto-rendered by @memory's tab_renderer).
Change the config → enforcement + in-file docs follow mechanically; prompts
never go stale.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q3mZT61WsKVN3srCwVDBiW
New 'drone @backup share <file_path> [--public]': uploads a single file to Drive
(AIPass Backups/Shared), sets a read permission (default: restricted to the
authenticated user; --public: anyone-with-link), returns the webViewLink
(webContentLink fallback). Reuses upload_single_file + DriveClient; idempotent
via _find_existing_file; fail-loud on every path. 21 new tests, all Drive API
mocked (zero live calls). Existing commands untouched. DPLAN-0230 v1.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q3mZT61WsKVN3srCwVDBiW
Repurpose the heartbeat into a ~2s transcript-tail loop that edits the "Processing" message in place (block-level: thinking/tool/text), plain text, coalesced, no-op-skipped, 429 retry_after aware, with 4096 rollover. Opt-in per-bot "stream" flag, default OFF; batch path byte-for-byte unchanged. @hooks reviewed: no change needed (already edits processing_message_id for the single-chunk final). Race hardened: re-check delivered before each edit. 37/37 streaming + 653/653 TG tests green; live-proven on the devpulse bot.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q3mZT61WsKVN3srCwVDBiW