fix(devpulse): watchdog stall threshold 120s to 300s (false stalls on long tool calls, verified live S330) + branch settings.local carries the 350k autoCompactWindow dial (Patrick ruling S326)

This commit is contained in:
AIOSAI
2026-07-21 18:04:51 -07:00
parent 57285740f2
commit 76ce13830d
4 changed files with 18 additions and 4 deletions
+7
View File
@@ -11,6 +11,13 @@ PyPI version — not the changelog header.
## [2026-07-21]
**fix(devpulse)** — watchdog stall threshold 120s → 300s: the 120s
no-JSONL-activity heuristic fired `[watchdog.stall]` on healthy agents doing
long tool calls; 300s matches observed real-stall behavior (verified live
S330). Branch `.claude/settings.local.json` carries the devpulse
`autoCompactWindow: 350000` dial (Patrick ruling S326 — devpulse compacts
~292k, dispatched agents stay pinned at 200k).
**fix(seedgo)** — checker accuracy arc (S330): AST-based import analysis
lands in the checkers (dead_code, encapsulation, handlers, readme,
test_quality, unused_function) — 13 false positives eliminated fleet-wide,
@@ -1,4 +1,5 @@
{
"autoCompactWindow": 350000,
"permissions": {
"allow": [],
"deny": [
@@ -423,7 +423,12 @@ class StallTracker:
(#634 part 2). The verbose trail stays on ``_stderr`` + logger.
"""
STALL_THRESHOLD = 120.0
# Quiet spans with zero JSONL output are routine, not stalls: a model composing
# a long response writes nothing until the turn completes, and compaction is one
# multi-minute silent API call. 120s false-fired on every stall of S329 (8/8);
# 300s clears normal turns and most compactions while a wedged agent still
# surfaces well inside a long watch.
STALL_THRESHOLD = 300.0
# A single tool call held in-flight this long is surfaced as a soft advisory
# (heavy op or a hung tool). Below the 600s default timeout so long watches
# get a mid-flight heads-up instead of waiting on the timeout.
@@ -516,7 +521,7 @@ def watch_agent(
poll_interval: Seconds between checks. Default 5.0 — the per-tick work (lock
stat, PID liveness, one-dir JSONL size scan) is cheap, so a tight cadence
just burns CPU. 5s keeps completion latency invisible on multi-minute
dispatches while the 120s stall threshold has ample resolution.
dispatches while the 300s stall threshold has ample resolution.
Returns:
dict with keys: woke, reason, elapsed, agent_state, exit_code, agent_id.
@@ -312,7 +312,7 @@ def test_stalltracker_inflight_tool_prevents_stall(monkeypatch, capsys):
monkeypatch.setattr(agent_handler, "_last_entry_is_inflight_tool", lambda *a, **kw: True)
t = agent_handler.StallTracker("@x", Path("/nope"), {}, now=0.0, pid=123)
for now in (60.0, 120.0, 180.0, 240.0):
for now in (120.0, 240.0, 360.0, 480.0):
t.observe(now=now)
out = capsys.readouterr().out
@@ -391,7 +391,8 @@ def test_watch_agent_surfaces_stall_to_stdout(monkeypatch, tmp_path, capsys):
monkeypatch.setattr(agent_handler, "_pid_alive", lambda pid: True)
monkeypatch.setattr(agent_handler, "_has_jsonl_activity", lambda *a, **kw: False)
monkeypatch.setattr(agent_handler, "_last_entry_is_inflight_tool", lambda *a, **kw: False)
_fake_clock_sleep(agent_handler, monkeypatch, lock_file)
# Lock must outlive STALL_THRESHOLD (300s) or the watch completes stall-free.
_fake_clock_sleep(agent_handler, monkeypatch, lock_file, unlink_at=400.0)
result = agent_handler.watch_agent("@fakebranch", timeout_seconds=100000, poll_interval=0.01)
out = capsys.readouterr().out