30 β PhaseRunner¶
Scope: Complete implementation reference for
jobs/phase_runner.pyβ the strict-order (sequenced) phase execution subsystem. Every state-machine step, watchdog rule, and timing-capture path here is derived directly from the source. A developer should be able to re-implement the strict-order runner from this document alone.Ownership note: this subsystem was extracted from
core/manager.pyin the manager re-bundle. The orchestration now lives entirely injobs/phase_runner.py(class PhaseRunner). What stays on the manager is a single 1-line delegator (maybe_advance_phase) plus the brand-tunable watchdog timing (_phase_timing+ the_PHASE_*module constants). If you find a doc that attributes the phase machine tocore/manager.py, this page is the correct owner.
1. Overview¶
A path-optimizing brand (the Roborock S6 is the live example) re-routes a multi-room batch into whatever order its firmware thinks is fastest, ignoring the order you dispatched. To honor a strict cleaning order on such a brand, the framework switches to the sequenced job model: one room per phase, dispatched in queue order, each waited on before the next is sent.
PhaseRunner owns the two halves of making that work reliably:
- The watchdog β a per-phase state machine (
_run_advanced_phase) that handles the device's nasty post-dock behavior: the S6 finishes a room, docks, starts charging, and ignores anapp_segment_cleansent at that instant. So each next room is dispatched from a background task that settles, dispatches, verifies the robot actually started THIS room, and retries if not. - Per-phase timing capture β
_capture_finishing_phase_timingsnapshots each finishing phase's room timing from its own counter slice before advance resets the queue, so finalization can reconstruct per-phase timings instead of mis-attributing the whole run to the last phase's room.
Two phase kinds. A phase is either a
room_groupclean phase (the single-room clean described above,phase_typeabsent or"room_group") or a stop phase that docks and holds between clean groups:phase_type == "charge_wait"(dock + poll battery to a target %, driven by_run_charge_wait_phase) or"wait"(dock + hold for N minutes, driven by_run_wait_phase). Stop phases carry no rooms; they exist so a native run profile can express vacuum β charge β mop as one learned job. They are introduced only by a run profile's orderedsteps(Β§3.1); a plain room dispatch never produces one. Everything in Β§4βΒ§7 branches on the phase kind.
Module: custom_components/eufy_vacuum/jobs/phase_runner.py
Class: PhaseRunner (constructed once with the core manager, the
bundled-subsystem pattern). The manager builds it as self.phase_runner and
reads/writes the same manager.data["active_jobs"] store.
Strict order is opt-in and brand-gated. A clean run becomes sequenced (one room per phase) when
strict_order=Trueis requested and the brand does not natively honor order (honors_clean_orderis False). Eufy, which honors order, folds a clean group back to a single atomic phase and never runs the intra-group watchdog. A run profile with charge/wait stops forcesstrict_order=True(_build_effective_start_plan, Β§3.1) so a path-optimizing brand runs each group in the shown order; the stop phases themselves still split the job into multiple phases on either brand (a stepped Eufy run is multi-phase but each clean group stays one atomic phase). See 06-job-lifecycle Β§2 and 29-roborock-adapter.
2. What it owns (and what it doesn't)¶
| Concern | Owner | Symbol |
|---|---|---|
| Public advance entry point | PhaseRunner (manager delegates) | maybe_advance_phase |
| Per-phase watchdog state machine | PhaseRunner | _run_advanced_phase |
| Spawn a stop-phase poller, guarded against a double-spawn | PhaseRunner | _spawn_dock_poller (+ the _dock_poller_active set) |
| Re-arm a dock phase whose in-memory poller was lost (pause+resume / HA restart) | PhaseRunner | rearm_dock_phase_if_needed |
Drive a charge_wait stop phase (dock + poll battery) |
PhaseRunner | _run_charge_wait_phase |
Drive a wait stop phase (dock + hold minutes) |
PhaseRunner | _run_wait_phase |
| Dispatch a phase's segment + per-room fan + per-phase global pre-calls | PhaseRunner | _dispatch_active_phase |
| Verify the device started THIS room | PhaseRunner | _await_phase_started |
| Release the dispatch-pending guard | PhaseRunner | _clear_phase_dispatch_pending |
| Coarse "is it cleaning at all" fallback | PhaseRunner | _vacuum_started_cleaning |
| Per-phase timing/area snapshot | PhaseRunner | _capture_finishing_phase_timing β _phase_room_timing / _wall_seconds / _learned_room_area_m2 |
| Build the per-room phase list | queue engine | GenericRoomIdsEngine.build_phases |
Materialize a run profile's steps into a phase list (clean groups + stops) |
run plan | RunPlanManager._build_steps_phases |
| Advance the stored job to the next phase | queue engine | advance_active_job_phase |
| Concatenate per-phase timings at finalize | learning | history_store finalization |
| Watchdog TIMING (defaults + brand overrides) | core/manager.py |
_phase_timing + _PHASE_* constants |
The timing split is deliberate: the orchestration is the subsystem's job, but the
brand-tuned defaults are a core-level surface so an adapter can override any
subset via dispatch.phase_timing. _run_advanced_phase /
_await_phase_started read the resolved timing back through
self._manager._phase_timing(...) β see Β§6.
3. The sequenced job model (upstream of PhaseRunner)¶
PhaseRunner consumes a job that was already built as phased by the queue engine. The relevant upstream contract:
GenericRoomIdsEngine.build_phases(strict_order=True)(inqueue/dispatch_engines.py) emits one single-segment phase per resolved room in queue order instead of one batch phase. Each phase carries its own room'spasses, so per-room passes is honored too. Withoutstrict_orderit returns a single batch phase ([self.build_payload(...)]). Atomic engines ignore the flag entirely.RoborockSegmentEngineis a naming subclass ofGenericRoomIdsEngineand inherits this behavior.build_active_job_state(..., phases=...)(inqueue/queue_engine.py) stores the ordered phase list pluscurrent_phase_index/phase_count. Whenphases is None(atomic β a plain room dispatch, and any run without strict-order or stops) those keys are omitted, so the stored job is byte-identical to pre-sequencing.advance_active_job_phase(active_job)returns a new job dict swapped to the next phase (nextresolved_rooms/payload/room_count/queue_*, per-phase progress reset,current_phase_indexincremented). It returnsNonefor an atomic job (nophases, orlen(phases) < 2) or when already on the final phase β in both cases the caller finalizes exactly as before.
Two job-record keys that advance_active_job_phase sets are central to the
watchdog handoff:
has_observed_active_lifecycle = Falseβ the next phase must be observed active again before it can finalize, so the just-ended phase's stale completion signal can't immediately re-finalize the new phase._phase_dispatch_pendingβ set True until the watchdog confirms the device actually started THIS room. While pending, the completion listener must not finalize/advance on the previous room's lingering dock signals (see Β§7). The stop-phase drivers (_run_charge_wait_phase/_run_wait_phase) rely on this same flag: they inherit it from the advance and never clear it, so the intentional charge/idle dock is not read as a completion (Β§6.6).
3.1 Stepped runs β the run-profile step builder¶
A run profile can carry an ordered steps list (set via the
set_run_profile_steps service) so a single learned job runs clean group β
stop β clean group. When such a profile starts, its steps are stashed under
data["_pending_run_steps"][vacuum_entity_id][map_id] and materialized 1:1 into
active_job["phases"] by RunPlanManager._build_steps_phases
(planning/run_plan.py) β a separate phase builder from
build_phases, feeding the same stored-phase-list contract above:
- a
room_groupstep β its brand dispatch phase(s), scoped to the group's included (enabled + not-blocked) room ids, with per-group settings: each group room's own fields (clean_mode/fan_speed/water_level/ β¦) overlay the global effective-room view, so the same room can vacuum in one group and mop in the next; - a
charge_waitstep β a phase{phase_type: "charge_wait", target_battery_percent}; awaitstep β{phase_type: "wait", wait_minutes}; both carry no rooms. - a
zonestep β a phase{phase_type: "zone", zone_ids: [...]}β a real clean phase (unlike a stop), carrying no rooms but a saved-zone id list. It dispatches viadispatch_zone_clean(the ids resolve to bbox rects at dispatch time) and advances on clean-complete through the same room-group watchdog hook as a clean phase β no new completion mechanism. Azonemay sit at the tail (after the last room group); leading stops don't apply to it. - Leading/trailing stops are dropped and consecutive same-type stops
collapse (two charges β the last target) so phase 0 is always a real clean
the initial dispatch can send; if no stop survives, it falls back to one atomic
clean. (A
zoneis a clean phase, so it is never trimmed as a "stop.")
_build_effective_start_plan forces strict_order=True whenever the run's steps
contain any charge_wait/wait stop, so a path-optimizing brand runs each
group's rooms in the shown order (a multi-room group then splits into per-room
phases); it is a no-op for an order-honoring brand (Eufy folds it back to
False). The stashed steps are popped only on the real dispatch; preflight
callers peek so they can't consume the stash early.
4. Entry point β maybe_advance_phase¶
The manager method is a 1-line delegator to
self.phase_runner.maybe_advance_phase(...). It is kept on the manager only
because production (listeners/lifecycle.py) and the listener tests reference
manager.maybe_advance_phase.
The completion hook in listeners/lifecycle.py calls it right before it would
finalize a completed phase. Behavior of PhaseRunner.maybe_advance_phase:
- Capture the finishing phase's timing FIRST β
_capture_finishing_phase_timing(...)snapshots the just-completed phase's room timing from its own counter slice (Β§5). This must run beforeadvance_active_job_phaseresets the queue, and it also runs for the final phase (advance returnsNoneimmediately after, but the capture already ran). - Call
advance_active_job_phase(active_job). - Returns
Noneβ an atomic job or the final phase β returnFalse(the caller finalizes as today). - Returns the advanced dict β stamp
current_room_started_at, persist it todata["active_jobs"][vacuum_entity_id][map_id], then spawn the driver for the new phase's kind as a background task (not awaited, so the completion listener returns promptly), routing onphases[next_idx]["phase_type"]: a stop phase ("charge_wait"/"wait") is spawned through_spawn_dock_poller(...)(Β§6.6.1, the double-spawn-guarded entry that picks_run_charge_wait_phasevs_run_wait_phase); a clean phase (aroom_groupor azone) is spawned as_run_advanced_phase(...)directly. ReturnTrue(the caller skips finalization).
maybe_advance_phase is also re-entrant from the stop-phase drivers: once a
charge_wait/wait phase's condition is met, its driver calls
maybe_advance_phase itself to move on to the next phase (Β§6.7 / Β§6.8).
Return contract: True = job advanced to a next phase, caller must skip
finalization; False = atomic job or final phase, caller finalizes normally.
5. Per-phase timing capture¶
Why this exists: the whole-run counter stream cannot be segmented across the
per-room dock trips. The CV/transit segmenter's transit capture breaks on the
intermediate dockings, so running the whole stream against the (single) last-phase
queue would either yield empty room_timings or credit the entire run's
battery/area to one room. Instead, each phase is segmented alone at the moment
it finishes, while its own counter slice is still intact.
5.1 _capture_finishing_phase_timing¶
- Atomic jobs (no
phaseslist) β no-op. Out-of-range / already-captured phases β no-op (idempotent; the_timing_end_tmarker is set even for an empty capture so it can't re-run). - Stop phases (
phase_typeincharge_wait/wait) record an EMPTY timing. Their counter slice is flat (the robot charges or idles on the dock), so finalize reads them as a dock/hold interval, not a phantom zero-metric room. The_timing_end_tmarker is still set so the next clean phase's slice starts after the stop. - A
zonephase captures azone_timing, not a room timing. It is a real clean, but it belongs to a saved zone (zone_id), not a room β so_capture_finishing_phase_timingroutes it to_capture_zone_phase_timing(Β§5.5), which writesphases[idx]["zone_timing"]and leavesroom_timingempty (no room to credit). - Slice selection: the phase's slice is the
counter_sampleswhose timestamp is strictly after the previous phase's recorded_timing_end_t(or from the run start for phase 0). All timestamps come from_iso_now()so a lexical string compare slices correctly. - One room per phase: the target is
queue_room_ids[0]; its slug comes fromresolved_rooms. - Empty-phase guard: a real per-room delta needs β₯ 2 samples carrying a
counter (
cleaning_timeorcleaning_areanotNone). A phase that never cleaned (the watchdog gave up, or a stale completion signal) records an empty timing rather than a phantom room β so finalize reads the run as not-fully-captured instead of fabricating a room with a learned area. - Result is stashed on the phase as
phases[idx]["room_timing"]plusphases[idx]["_timing_end_t"].
5.2 _phase_room_timing β within-slice deltas¶
One strict-order phase = one room, so no segmenter is needed. cleaning_time
and cleaning_area are cumulative across the run, so the per-phase figure is
last β first within the slice. (Using the segmenter here would report a later
phase's cumulative total, not its own area.) The returned shape:
{
"room_id": room_id,
"slug": slug,
"cleaning_start": <first slice ts>,
"cleaning_end": <last slice ts>,
"cleaning_seconds": <cleaning_time delta, int>,
"cleaning_wall_seconds": <wall-clock between first/last ts, or cleaning_seconds>,
"area_m2": <cleaning_area delta, rounded 3dp>,
"battery_delta": <first β last battery, or None>,
"boundary": "phase",
}
_wall_seconds(t0, t1) is a static helper: whole seconds between two ISO
timestamps, best-effort 0 on parse failure.
5.3 Area fallback β _learned_room_area_m2¶
A within-phase cleaning_area delta of β 0 (a stale/flat sensor through the
phase) is unusable. When _phase_room_timing produces area_m2 <= 0, the capture
substitutes the room's learned area from the map registry
(maps[...][...][rooms][rid], keys learned_area_m2 then area_m2) and tags it
area_source: "learned_fallback". Returns None when no learned area is known.
5.4 Finalization concat (downstream, in learning)¶
At finalize, learning/history_store.py checks for active_job_state["phases"].
If present, it concatenates each phase's room_timing in phase order rather
than running the whole counter stream through _build_transit_blocks:
transitions = []β inter-phase gaps are dock overhead, not room-to-room transit.transit_capture_validis True only when every phase contributed a non-empty timing (_every_phase_captured); a single empty phase marks the run as not-fully-captured (excluded from learning). A stop phase's timing is intentionally empty (Β§5.1), so a stepped run that includes acharge_wait/waitstop is currently reported not-fully-captured for transit-learning purposes β the per-roomroom_timingentries of its clean phases are still concatenated.- Atomic jobs leave
phasesabsent β the legacy_build_transit_blockspath.
See 10-learning-system for how finalized
room_timings feed per-room learned stats.
5.5 Zone-phase timing (_capture_zone_phase_timing)¶
A zone phase is a real clean but is owned by a saved zone, so it captures a
zone_timing snapshot instead of a room_timing. _capture_finishing_phase_timing
detects phase_type == "zone" and delegates to _capture_zone_phase_timing, which
writes on the phase:
phases[idx]["zone_timing"] = {
"zone_ids": [<saved-zone id>, ...],
"zone_names": [<name at run time>, ...], # snapshotted, survives a later rename/delete
"clean_mode": "mop" | "vacuum", # from job_metadata.has_mop_mode
"wall_seconds": <phase dispatch β completion, whole seconds>,
"area_m2": <summed saved-zone area, or None>,
"cleaning_start": <ISO>,
"cleaning_end": <ISO>,
}
The learned quantity is wall-clock (_wall_seconds(started_at, now)), not a
counter slice: for a small mop zone the clock is dominated by dock-prep and pad
wash, so wall time is the honest cost while area is not (see
10-learning-system Β§2.5).
zone_names and area_m2 come from _saved_zone_names / _saved_zone_area_sum
over the map's saved_zones store; the names are snapshotted so the record stays
readable even if the zone is later renamed or deleted (learning still keys on
zone_id). At finalize, learning/history_store.py harvests every zone phase's
zone_timing into the record's zone_timings + zone_count, and
learning.manager folds each single-zone phase into the learned_zones store.
6. The watchdog state machine β _run_advanced_phase¶
This is the heart of the subsystem. It runs as a background task β spawned by
maybe_advance_phase for an advanced phase, and by start_selected_rooms (in
core/manager.py) with initial=True for phase 0.
The state machine, per phase: settle β (dispatch) β verify β retry.
6.1 Settle¶
- Resolve timing via
pt = self._manager._phase_timing(vacuum_entity_id). - Advanced phase (
initial=False):await asyncio.sleep(settle). If the phase's target room is the dock room (_phase_target_is_dock_roomreads the map's per-roomis_dock_roomflag), the settle is extended todock_settle_secondsβ the robot is parked + charging right on the target, so its post-dock ignore-transient is the longest. - Initial phase (
initial=True): skip the settle β phase 0 was already dispatched bystart_selected_rooms, so the first iteration is verify-only.
6.2 The attempt loop (1 .. max_attempts)¶
Each iteration first re-reads the live job and bails if the phase is no longer
ours β any of: job gone, status != "started", current_phase_index != phase_index,
or _cancel_in_flight set. (async_cancel sets _cancel_in_flight before the
status flips, so the watchdog stops before it can re-dispatch during the
return-to-base window.)
Then:
- Dispatch via
_dispatch_active_phase(...)β except on the initial phase's first attempt (its first send already happened at job start; attempt 1 is verify-only, and a re-dispatch only happens on a retry if the device ignored it). - Verify via
await _await_phase_started(...): - True β the device started THIS room β call
_clear_phase_dispatch_pending(...)and return (phase under way). - False β the device ignored the dispatch / is still docked / is still finishing the previous room β loop to the next attempt (re-dispatch).
After max_attempts with no start, the loop logs a warning and leaves the run
stalled β recoverable by the user via Cancel Run, rather than silently hung.
max_attempts is the per-phase watchdog.
6.3 _dispatch_active_phase¶
- Run this phase's global pre-calls per phase
(
manager._run_global_pre_calls(...)) from THIS phase's own resolved rooms. A brand that exposes fan/water only as global device settings (Roborock mop intensity via aselect) sets them here, so a vacuum group then a mop group each apply their own water setting β phase 0's pre-call fires instart_selected_rooms, phases 1+ fire here. No-op when the adapter declares nodispatch.global_pre_calls(Eufy, S6). - Apply this phase's per-room live settings (fan) before the dispatch
(
manager.active_job.apply_per_room_live_settings_awaited(...)) so the room starts at its own fan value with nocurrent_roompoll lag. - Live-resolve the segment ids by slug per phase
(
manager._resolve_live_dispatch_payload(...)) so a re-segment between rooms can't clean the wrong room (no-op withoutdispatch.resolve_live_ids_by_slug). - Dispatch via
manager._dispatch_clean_payload(...). - Log at INFO (so a strict-order run is diagnosable without debug): the exact re-dispatched payload + attempt number. An empty segment list = the next room was skipped (e.g. a slug that didn't resolve to a live id).
6.4 _await_phase_started β the verification logic¶
This is where an earlier naive check (the job-active binary / "is it cleaning at
all") false-passed: the device's inCleaning flag stays on across the whole
job, so a clean it ignored at the dock looked like success and the watchdog
never retried (only ~1 room in 4 actually fired). The strong signal is the
brand's native current-room matching the phase's target while actually
cleaning, sustained.
- Native path (adapter declares
entities.active_cleaning_target): - Accumulates
cleaning_in_targetseconds when all hold:vacuum.state == "cleaning"(rules out docked-in-the-target-room), the native signal (manager.active_job.native_current_room_target_id(...)) is present and== job["current_room_id"]. - Confirms (True) once
cleaning_in_target >= confirm_seconds. The signal dips in and out, so a dip just doesn't add to the tally β we don't require strict continuity. - Bounds an attempt by no-progress (idle) time, not a fixed overall window: a long cross-room transit merely delays when the tally starts, so it can't falsely fail a device genuinely on its way; a device that never reaches the room accrues idle and retries promptly.
- Idle-exit weak confirm: once
idle >= verify_seconds, returncleaning_in_target > 0. If we observed genuine cleaning of the target at any point (the tally only accrues while cleaning AND on-target), the device started the room and has since finished/docked β a small room that completed in underconfirm_seconds. Treat that as confirmed, because re-dispatching an already-cleaned room is ignored by the device and would stall the phase forever. Only a true no-show (never cleaned the target) returns False to retry. - No-native fallback: brands with no native current-room signal fall back to
the coarse
_vacuum_started_cleaning(...)(immediate, unchanged). - Early-out (True): if the phase advanced / finalized / paused or
_cancel_in_flightis set mid-poll, return True (nothing to retry).
6.5 _vacuum_started_cleaning (coarse fallback)¶
True when the vacuum entity reports cleaning, or the adapter's job-active
binary (entities.job_active β the device inCleaning flag) is on. Used only by
the no-native fallback in _await_phase_started.
6.6 _clear_phase_dispatch_pending¶
Once the watchdog confirms the phase's room actually started, this clears
_phase_dispatch_pending so the room's real completion can finalize/advance
normally. It only clears when the job is still on this exact phase (a later
advance owns its own pending flag), then best-effort persists via
manager._async_save_logged().
6.6.1 _spawn_dock_poller + rearm_dock_phase_if_needed (stop-phase poller lifecycle)¶
A stop phase's only driver is an in-memory asyncio task (_run_charge_wait_phase
/ _run_wait_phase); unlike a room_group phase, no room-finish event can advance
it. Two paths would otherwise spawn a second concurrent poller for the same
phase, so both route through one guarded entry.
_spawn_dock_poller(*, vacuum_entity_id, map_id, phase_type, phase_index) -> boolβ the single spawn point. It keys a(vacuum_entity_id, map_id, phase_index)tuple into the in-memoryself._dock_poller_activeset before creating the task; if the key is already present it returnsFalsewithout spawning. It picks_run_charge_wait_phasevs_run_wait_phaseonphase_type. The chosen driver'sfinally(in its thin wrapper)discards the same key, so a later re-arm can spawn a fresh poller once this one has exited. The key includesphase_index(not just vac/map) so a poller advancing to the next dock phase can spawn that phase's poller before its ownfinallyclears the old key. Bothmaybe_advance_phase(Β§4) andrearm_dock_phase_if_neededcall it.rearm_dock_phase_if_needed(*, vacuum_entity_id, map_id) -> boolβ recovers a dock phase whose in-memory poller was lost. A pause+resume (status flips back to"started"but re-arms nothing) or an HA restart (the task is gone andasync_initializeforce-clears_phase_dispatch_pending) leaves acharge_wait/waitphase with no live driver, so the run wedges in"started"forever. When the active job is"started"and its current phase (current_phase_index) is a stop phase, it re-asserts_phase_dispatch_pending(the restart cleared it, so the intentional dock isn't read as a completion mid- re-arm) and re-spawns the matching poller via_spawn_dock_poller. It is a no-op for atomic /room_group/ finalized jobs, and the double-spawn guard makes it idempotent against a live poller. The re-spawned poller is idempotent by its own already-at-target / deadline /_still_ourslogic; thewaitpoller recomputes its deadline from the persistedwait_started_at, so a restart mid-wait does not restart the full timer. Called from two sites:core/manager.py::async_initialize(on load) β loops every"started"job onceself.phase_runnerexists.jobs/active_job.py::async_resume_active_job(on resume) β right afterresume_active_jobflips the status back to"started".
6.7 _run_charge_wait_phase (the charge stop driver)¶
A charge_wait phase has no rooms β its "work" is parking on the dock and
charging, so no room-finish event can advance it; this coroutine owns its whole
lifecycle. It is spawned by maybe_advance_phase when the advanced phase's
phase_type == "charge_wait".
- Reads
target_battery_percent(default100) andcharge_wait_timeout_minutes(default180) off the phase dict. - Already at/above target on entry β advance immediately via
maybe_advance_phase, no charge. - Otherwise it stamps
charge_from_battery/charge_started_aton the phase (Wave-4 observability for the live snapshot), sends the robot home (vacuum.return_to_base, blocking β a no-op if already docked+charging), then polls battery every ~30 s (get_battery_levelfromcore/charging.py). - On reaching the target it stamps
charge_to_battery/charge_ended_atand callsmaybe_advance_phaseto continue. - Timeout β
async_cancel_active_job(cancel_reason="charge_timeout", forced_lifecycle_state="charge_timeout")so the un-cleaned remaining rooms are finalized as missed rather than left hung. - A genuine user Cancel (
_cancel_in_flight), pause, or advance (the_still_oursguard: job present,status == "started",current_phase_index == phase_index) β bail; the phase is no longer ours. - It keeps
_phase_dispatch_pendingset (inherited from the advance), so the completion gate (Β§7) never reads the intentional charge-dock as a completion.
The ETA the live snapshot shows (charge_eta_minutes / charge_eta_source)
comes from battery/manager.py::compute_time_to_target_pct (a CC/CV split at
80 %, falling back to wall-clock rather than a fabricated estimate) β computed at
snapshot time, not inside this driver.
6.8 _run_wait_phase (the timed stop driver)¶
The time-based twin of Β§6.7 (e.g. a mop-dry pause between a vacuum pass and a mop
pass), spawned when the advanced phase's phase_type == "wait". It reads
wait_minutes (default 5, floored at 1), stamps wait_started_at, parks on
the dock (vacuum.return_to_base, blocking), then polls every ~15 s until
the deadline; on elapse it stamps wait_ended_at and calls maybe_advance_phase.
Like the charge driver it has no rooms, keeps _phase_dispatch_pending set, and
bails on the same _still_ours (cancel/pause/advance) check.
7. The completion-gate handoff¶
The watchdog cooperates with the completion listener
(listeners/lifecycle.py) through _phase_dispatch_pending:
phase completes
β should_finalize_completed = True
β if active_job["_phase_dispatch_pending"]: should_finalize_completed = False # (A)
β if not should_finalize_completed: continue
β if await manager.maybe_advance_phase(...): continue # (B) advanced β skip finalize
β finalize_learning_for_active_job(...) # (C) atomic / final phase
- (A) A Roborock sits docked + charging between phases β precisely its
completion signal. While the next phase's dispatch is still pending (the
watchdog hasn't confirmed it started), the prior room's lingering dock signal
must not finalize the next room before it ever starts. This same guard also
protects a stop phase's intentional dock:
_run_charge_wait_phase/_run_wait_phaseinherit_phase_dispatch_pendingand never clear it, so the charge/idle dock is not read as a completion. No-op for non-sequenced jobs (the flag is only set on a phase advance). - (B) A completed non-final phase advances + re-dispatches instead of
finalizing (this is
maybe_advance_phasereturning True). - (C) Atomic jobs β a plain room dispatch β and the final phase fall through to normal finalization.
See 06-job-lifecycle Β§6βΒ§7 for the full end-of-job flow.
8. Watchdog timing β the manager seam¶
The watchdog timing is deliberately not owned by this subsystem. It lives on
core/manager.py so it is a brand-tunable surface:
- In-core defaults are the
_PHASE_*module constants incore/manager.py:
| Constant | Default | Meaning |
|---|---|---|
_PHASE_SETTLE_SECONDS |
10 |
Post-dock settle before the first dispatch of an advanced phase. |
_PHASE_DOCK_SETTLE_SECONDS |
45 |
Longer settle when the target room is the dock room (worst-case ignore-transient). |
_PHASE_VERIFY_SECONDS |
90 |
No-progress (idle) budget before an attempt gives up and re-dispatches. |
_PHASE_CONFIRM_SECONDS |
45 |
Cumulative cleaning-of-target seconds required to release the pending guard. |
_PHASE_POLL_SECONDS |
5 |
How often the verify loop re-checks. |
_PHASE_MAX_ATTEMPTS |
3 |
Per-phase watchdog cap; after this the run is left stalled. |
- The resolver is
core/manager.py::_phase_timing(vacuum_entity_id): it merges the adapter'sdispatch.phase_timingoverrides over the_PHASE_*defaults, casting each declared value to the default's type. A brand whose post-dock transient differs declares its own; anything omitted falls back to the default (byte-identical). Defaults are read live per call so the tests' module-constant monkeypatching still applies.
PhaseRunner reads this single resolver β never the constants directly β via
self._manager._phase_timing(...) inside _run_advanced_phase and
_await_phase_started.
The Roborock S6 declares its own dispatch.phase_timing
(adapters/roborock/adapter.py), lowering confirm_seconds to 15 (a sub-15s
room is rare on the S6; the idle-exit weak-confirm backstops anything faster) and
otherwise matching the defaults. Eufy declares nothing β fully default β inert.
See
22-adapter-config-reference Β§dispatch.phase_timing.
9. Integration points¶
| Caller | Method | When |
|---|---|---|
listeners/lifecycle.py (completion hook) |
manager.maybe_advance_phase() β PhaseRunner.maybe_advance_phase() |
A phase's room set finished; decide advance vs finalize |
core/manager.py::start_selected_rooms |
phase_runner._run_advanced_phase(..., initial=True) |
A sequenced job starts (verify phase 0) |
PhaseRunner.maybe_advance_phase |
phase_runner._run_advanced_phase(...) directly (clean phase) / _spawn_dock_poller(...) (stop phase) |
The next phase is dispatched, routed on its phase_type |
core/manager.py::async_initialize (on load) |
phase_runner.rearm_dock_phase_if_needed(...) |
Re-arm a stop phase whose in-memory poller the restart lost |
jobs/active_job.py::async_resume_active_job (on resume) |
phase_runner.rearm_dock_phase_if_needed(...) |
Re-arm a stop phase after a pause+resume re-armed nothing |
queue/dispatch_engines.py |
build_phases(strict_order=True) |
Build the one-room-per-phase list (upstream) |
planning/run_plan.py |
_build_steps_phases(...) |
Materialize a run profile's steps into clean-group + stop phases (upstream) |
queue/queue_engine.py |
advance_active_job_phase() |
Swap the stored job to the next phase |
learning/history_store.py |
concat phases[*]["room_timing"] |
Finalize per-phase timings into the run record |
core/manager.py::_phase_timing |
(read by the watchdog) | Resolve brand timing overrides |
10. Edge cases & invariants¶
- Atomic jobs are fully inert. No
phaseskey βmaybe_advance_phasereturns False,_capture_finishing_phase_timingno-ops, the watchdog is never spawned. A plain (non-stepped, non-strict) run never enters this machine. - Stop phases carry no rooms. A
charge_wait/waitphase is advanced only by its own driver (_run_charge_wait_phase/_run_wait_phase) reaching its condition, never by a room-finish event; it records an empty per-phase timing and keeps_phase_dispatch_pendingset so its dock is not read as a completion. - A stop phase's poller survives pause+resume and HA restart. Its driver is an
in-memory task with no persistent scheduler, so
rearm_dock_phase_if_needed(Β§6.6.1) re-spawns it on resume (async_resume_active_job) and on load (async_initialize), re-asserting_phase_dispatch_pending. The_dock_poller_activeguard (via_spawn_dock_poller) makes a re-arm idempotent against a live poller, so the phase is never double-driven. Without it a charge/wait run wedges in"started"forever after a pause+resume or restart. - Stops force strict order. Any run profile whose
stepsinclude a stop is planned withstrict_order=True; on an order-honoring brand (Eufy) that folds back to no intra-group reordering but the multi-phase (stop-bracketed) structure still applies. - Cancellation is race-safe.
async_cancelsets_cancel_in_flightbefore the status flips; both_run_advanced_phaseand_await_phase_startedcheck it each iteration and bail, so the watchdog can't re-dispatch during return-to-base. - Idempotent capture.
_capture_finishing_phase_timingwrites_timing_end_teven for an empty capture, so it never double-counts or re-runs a phase. - A stalled run is recoverable, never silently hung. After
max_attemptsthe watchdog stops and logs; the user recovers via Cancel Run. - The final phase's timing is still captured.
maybe_advance_phasecaptures beforeadvance_active_job_phasereturnsNone, so the last room is recorded.
Cross-links¶
- 06-job-lifecycle β full job start β finalize flow; the
completion gate that calls
maybe_advance_phase. - 07-queue-engine β
build_phases/build_active_job_state/advance_active_job_phaseand the dispatch engines. - 29-roborock-adapter β the live path-optimizing brand;
honors_clean_order, native current-room signal,per_room_live_settings. - 10-learning-system β how concatenated per-phase
room_timingsfeed learned per-room stats. - 22-adapter-config-reference Β§dispatch.phase_timing β the brand override block for the watchdog timing.