summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
authorhathach <[email protected]>2026-08-20 16:45:32 +0700
committerhathach <[email protected]>2026-08-20 16:45:32 +0700
commitbe86281b81b70adc1ce3746d620c912bf032fac2 (patch)
treee751548c35e8b3f09d55a5303e7e7d7acf3c6a6b
parent6905639b07c69fec68e9ebc77f7d27ac2775ee41 (diff)
hil-pool-check: document the probe power-cycle escalation; name the env script directly
A probe whose firmware has wedged reports flash-failed with the probe present and "probe toggle unconfirmed". The tool's own recovery cannot fix that: an authorized toggle re-enumerates but never removes power, and hil_pool_check.py:304 already notes that ST-Link, WCH-Link, CP210x and picoprobe keep their sysfs kobject across one. So the check correctly gives up, and the operator was left to invent the next rung. Write it down, as what this rig's hardware actually does rather than what uhubctl advertises: the Renesas cards list their root hubs as ppps-capable but do not implement it (owner-confirmed - VBUS never drops, only D+/D-), so a root-port cycle is a harder forced re-enumeration that a wedged probe can ride out, worth exactly one attempt; and the AMD 0000:02:00.0, where the WCH-Links live, has no port-power switching at all - nothing to cycle, straight to a physical replug. Which card a probe hangs off decides which case applies, so the procedure starts from readlink. The ordering rules encode the shared-rig protocol: let the full run finish (a bounce re-enumerates siblings and corrupts checks still in flight), hold --all with this host's --config before the cycle (hil_lock.py hold validates nothing against the roster and nothing maps a sysfs busport to a board name, so a narrower hand-listed hold reserves nothing while reporting success - and --all defaults to tinyusb.json, which on the tusb rig would reserve 27 boards that do not exist there), release BEFORE the re-check (hil_pool_check.py self-locks every board it checks, so a hold still in place makes the verification report locked against your own hold and verify nothing), and drive the cycle through usb_recover.sh root-cycle by its full in-repo path - it is on no PATH and sudo's secure_path excludes the checkout. Never a bare `uhubctl -a cycle`: without -S it writes sysfs disable, whose disable_store takes the root hub's lock uninterruptibly and then usb_disconnect()s the wedged child - the one input that turns a probe wedge into a bus-wide wedge. Give the script the wedged probe's own busport, not the hub path: the serial guard and the success check both read the path you pass, and the hub's inode always changes when its own port cycles. Reporting asks for both passes: a final table showing every board healthy hides that a probe needed power-cycling to get there, which is the signal that it will recur. Also: the ESP-IDF env hints name `. $HOME/code/esp-idf/export.sh` instead of the `get-idf` alias, which lives only in interactive shells and fails from scripts.
-rw-r--r--.claude/skills/hil-pool-check/SKILL.md75
-rw-r--r--test/hil/helper/hil_pool_check.py6
2 files changed, 78 insertions, 3 deletions
diff --git a/.claude/skills/hil-pool-check/SKILL.md b/.claude/skills/hil-pool-check/SKILL.md
index 6a8f66087..3c056c16b 100644
--- a/.claude/skills/hil-pool-check/SKILL.md
+++ b/.claude/skills/hil-pool-check/SKILL.md
@@ -41,7 +41,10 @@ ssh ci.lan 'bash -lc "cd ~/code/tinyusb && python3 test/hil/helper/hil_pool_chec
Missing firmware is **built on the spot** — never skipped (`--no-build` opts out; those boards
then report `flash-failed`). Builds need the family env, exported on the rig in
`~/.profile`/`~/.bashrc`: `PICO_SDK_PATH` for rp2040/rp2350 (`~/code/pico/pico-sdk`), the
-ESP-IDF env (`get-idf`) for espressif — which also needs `esptool` on PATH (pip's
+ESP-IDF env for espressif — source it as `. $HOME/code/esp-idf/export.sh`, NOT as `get-idf`:
+that is an interactive shell alias (`~/.bashrc`), and aliases are not expanded in non-interactive
+shells, so scripts and agents get `get-idf: command not found` even under `bash -lc`. It also
+needs `esptool` on PATH (pip's
`~/.local/bin/esptool`; a non-login shell may lack it — run via `bash -lc`). An explicit `-B` is
searched exclusively for *existing* firmware; builds still land in `cmake-build/` and are noted
`built <example>`. Espressif boards park too when the IDF env is present. A first run on an
@@ -57,8 +60,78 @@ are *unverified*, not healthy — read the footer, not just `$?`. A `⚠ pid …
means stale firmware or a silent flash no-op (J-Link lore); a device off the bus entirely needs
the usb-kernel-recover skill or a physical replug.
+## When the tool's probe recovery fails
+
+`flash-failed` with the probe ✅ present and a `probe toggle unconfirmed` note means the probe's
+own firmware is wedged, not the board. The tool's recovery is an `authorized` toggle, which is a
+USB re-enumeration and never removes power, so probes that keep their sysfs kobject across it
+(ST-Link, WCH-Link, CP210x, picoprobe) survive the toggle still wedged. Confirm with the flasher's
+own list — `STM32_Programmer_CLI -l st-link`, or `JLinkExe -CommandFile <script>` with
+`ShowEmuList` in it: a probe that enumerates but reports a blank serial/firmware is answering the
+kernel and not the tool, which is a host-to-probe fault. A dead target reports the opposite: the
+probe identifies itself normally and then fails to connect.
+
+The next rung is a root-port bounce, and what it buys depends on which card the probe hangs off
+(`readlink -f /sys/bus/usb/devices/usb<bus>` gives the PCI address):
+
+- **Renesas** (five cards here): `uhubctl` lists their root hubs as `ppps`-capable, but the cards
+ do not implement it — VBUS never drops, only D+/D− (see usb-kernel-recover). A cycle is therefore
+ a harder forced re-enumeration, **not** a power cycle: worth one attempt, but a probe that rode
+ out the `authorized` toggle can ride this out too. Do not read `ppps` here as power control.
+- **AMD `0000:02:00.0`** (where the WCH-Links live): no port-power switching at all — `uhubctl`
+ does not list it. There is nothing to cycle; go straight to a physical replug.
+
+The leaf hubs are ganged, so a bounce hits every device under that root port. Escalate by hand, in
+this order:
+
+1. **Let the full run finish first.** Never cycle mid-run: the bounce re-enumerates siblings and
+ would corrupt the checks still in flight for other boards.
+2. Identify the subtree and its blast radius, so the report can name what was disturbed:
+ ```bash
+ ls -d /sys/bus/usb/devices/<bus>-<rootport>.* # siblings that will be bounced
+ ```
+3. **Hold `--all` for the cycle, and release before the re-check.** `hil_lock.py status` only
+ observes; CI can take a board a second later and flash straight into the bounce. The bounce
+ hits every board under the root port and nothing maps a sysfs busport to a board name, so
+ `--all` is the only reservation that actually covers them:
+ ```bash
+ python3 test/hil/helper/hil_lock.py hold --all --config test/hil/tinyusb.json --reason "probe power cycle"
+ ```
+ It is all-or-nothing: a refusal naming `hil_test.py` means a CI job is mid-test — wait, do
+ not force, and do not substitute a partial hold. Release before step 5: `hil_pool_check.py`
+ self-locks every board it checks and reports 🔒 locked for any it cannot take, so a hold
+ still in place makes the whole verification pass report locked against you and verify nothing.
+4. Cycle the ROOT port through the recovery script — rung 2 of usb-kernel-recover, which owns this
+ invocation:
+ ```bash
+ sudo .claude/skills/usb-kernel-recover/scripts/usb_recover.sh \
+ root-cycle <the wedged probe's own busport> [expected-serial]
+ # e.g. 13-1.6, NOT the 13-1 hub path from step 2 — the script derives the root port
+ # itself. Give the full path: the script ships inside the checkout, is on no PATH, and
+ # sudo's secure_path excludes the repo, so a bare `usb_recover.sh` is command-not-found
+ ```
+ Give it the device, not the hub: the expected-serial guard and the success check both read the
+ path you pass, so handing it the hub compares the hub's serial and watches the hub's inode,
+ which always changes when its own root port is cycled — it prints success while the probe is
+ still dead. **Never a bare `uhubctl -a cycle` here.** Without `-S` it writes sysfs `disable`,
+ whose `disable_store` takes the root hub's lock uninterruptibly and then calls
+ `usb_disconnect()` on the child — against the wedged probe you are trying to clear, that blocks
+ while holding the root hub's lock and poisons the whole bus. The script passes `-S`.
+5. Release the locks, then re-check the affected boards:
+ `python3 test/hil/helper/hil_pool_check.py -b BOARD [-b …]`. Include the bounced siblings — a
+ cycle that fixes one probe can leave another unenumerated.
+
+If the second pass still fails, the probe needs a physical replug: no software rung on this rig
+removes VBUS, so there is nothing further to try.
+
## Reporting
The user-facing answer to a pool check IS the tool's summary table: paste the complete per-board
table (and footer counts) verbatim — never truncate rows or reduce it to a prose digest like
"27/27 healthy"; at most one line of commentary below it.
+
+When an escalation above was needed, add a short note under the table naming: which boards needed
+it, which root port was cycled (or that a replug was needed instead), which siblings bounced, and
+the second-pass result for each.
+Report BOTH passes — a final table showing every board ok hides the fact that a probe had to be
+power-cycled to get there, which is exactly the signal that predicts it recurring.
diff --git a/test/hil/helper/hil_pool_check.py b/test/hil/helper/hil_pool_check.py
index 371aff1e1..8ad68bc6d 100644
--- a/test/hil/helper/hil_pool_check.py
+++ b/test/hil/helper/hil_pool_check.py
@@ -328,7 +328,8 @@ def flash(board: dict, fw, allow_recovery: bool, probe_port: str, note: list) ->
if rc == 0:
return True
if rc == 127: # flasher binary missing: retries/probe recovery can't fix env
- note.append(f'flasher tool missing ({err}) — esptool needs the ESP-IDF env (get-idf)'
+ note.append(f'flasher tool missing ({err}) — esptool needs the ESP-IDF env '
+ f'(. $HOME/code/esp-idf/export.sh)'
if board['flasher']['name'].lower() == 'esptool' else
f'flasher tool missing: {err}')
return False
@@ -487,7 +488,8 @@ def ensure_fw(board: dict, variant: str, example: str, note: list):
rc = build_example(board, variant, example)
if rc == 127 and board['flasher']['name'].lower() == 'esptool':
_builds[key] = (None, 'no-env')
- note.append(f'cannot build {base}: ESP-IDF env missing (get-idf)')
+ note.append(f'cannot build {base}: ESP-IDF env missing '
+ f'(source $HOME/code/esp-idf/export.sh)')
return None
if rc == 124: # hung build: a deps/cache retry cannot cure it, don't double the stall
_builds[key] = (None, 'timeout')