Root cause: get_module_logs_dir() last fallback unconditionally created
src/aipass/{name}/logs/ for any unknown module, causing external projects
(AIPL polyglot) running with AIPASS_HOME set to pollute the AIPass src tree
with src/aipass/unknown_branch/logs/*.
Fix (already committed in 3c5ac29):
- Check AIPASS_CALLER_CWD env var (set by drone during cross-project dispatch,
DPLAN-0121) and walk up to the caller's project root (.git/pyproject.toml)
- Final fallback: system_logs/external/{module_name} — never create unknown
dirs in the AIPass source tree
- Added inspect.stack() auto-detection when module_name is not provided
- Added _warn_routing() helper for lazy prax logger access (avoids circular
imports — logger.py imports load.py at module level)
This PR:
- Regression tests: test_unknown_module_routes_to_system_logs_external and
test_aipass_caller_cwd_routes_to_caller_project (both green)
- Cleanup: moved 27 leaked AIPL polyglot logs from src/aipass/unknown_branch/
to /tmp/aipl_leaked_logs/ for AIPass Developer review; removed directory
- bypass.json: architecture + documentation exemptions for tests/test_config.py
CC @polyglot: AIPL-side logger config may need AIPASS_CALLER_CWD set during
cross-project dispatch so logs route to ~/Projects/AIPL/ correctly. The fix
is transparent if drone sets AIPASS_CALLER_CWD; no AIPL changes required for
the basic fix, but explicit env var support improves log placement accuracy.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
TRIGGER_ROOT.parent.parent resolves to src/, not the project root.
Fix to TRIGGER_ROOT.parent.parent.parent so all three SYSTEM_LOGS_DIR
constants resolve to /home/patrick/Projects/AIPass/system_logs.
Fixes recurring 'Failed to start log watcher' on every service start.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Removed 10 orphaned functions/constants left after DPLAN-0112 Option C stripped
dispatch: _find_repo_root, _is_medic_enabled, _is_branch_muted,
_get_registered_emails, _is_rate_limited, _record_dispatch, _log_suppression,
_build_notification_message, and associated constants. Handler now only uses
_log_warning + json_handler. 370 tests passing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Extracted _restore_fingerprint_tracking() helper and flattened early-return
guards to reduce nesting depth from 5 to 3. Seedgo deep_nesting now passes.
370 tests passing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_fire_to_handlers now tracks consecutive failures per handler. After 5
consecutive failures, the handler is auto-disabled (skipped) and a CRITICAL
log is emitted. Success resets the failure count. Disabled handlers re-enable
on process restart (in-memory tracking). Prevents broken handlers from
flooding logs forever. 370 tests passing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Moved CB state from trigger_config.json to dedicated trigger_cb_state.json.
Now persists full state: recent_errors timestamps, half_open_allow, summary_sent,
and per-fingerprint dispatch tracking (last_dispatch + count). Restored on
startup. record_dispatch() now triggers persist. Uses atomic_write_json +
json_file_lock. Module-level dict init moved before CB load so fingerprint
data restores correctly. 370 tests passing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Added json_file_lock() context manager to config.py using fcntl.flock
with .lock sidecar files. Wrapped all read-modify-write cycles:
- error_registry.py: report(), _save/_clear_circuit_breaker_state()
- medic_state.py: set_enabled(), mute_branch(), unmute_branch()
- log_watcher.py: _save_seen_hashes(), _save_log_positions()
Combined with existing atomic_write_json for both concurrency and crash safety.
370 tests passing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
error_logged.py handler stripped to monitor-only (log event, no email/dispatch).
Centralized watcher (watchers/log_watcher.py) now calls registry_report() +
fires error_detected for ERROR-level lines, routing through full Medic v2
pipeline (count threshold, circuit breaker, fingerprint backoff). Falls back
to error_logged (monitor-only) if registry unavailable. 370 tests passing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
report_error() and log_watcher primary path both had a gate that only fired
error_detected at count==1 (new) and count==2 (exact match). After count
passed 2, the event never fired again, so the handler's backoff logic never
got a chance to re-dispatch. Fix: always fire the event and let the
error_detected handler's Medic v2 gating (circuit breaker, backoff, rate
limiting) decide. 370 tests passing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The fallback path in log_watcher._process_log_line() used _is_duplicate_error()
to skip repeated errors entirely. This prevented registry count from incrementing
past 1, so dispatch (which requires count>=2) never fired. Fix: retry lazy
registry import in fallback path; if truly unavailable, track count locally
and fire event with count so error_detected handler can apply threshold.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
After delivering error notification email, now calls wake_branch() to spawn
an agent in the target branch immediately. Changes reply_to from @trigger
to @devpulse so resolution reports go to devpulse. 370 tests passing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>