summaryrefslogtreecommitdiff
path: root/docs
diff options
context:
space:
mode:
Diffstat (limited to 'docs')
-rw-r--r--docs/getting_started.rst2
-rw-r--r--docs/reference/boards.rst9
-rw-r--r--docs/reference/device_issues.rst56
-rw-r--r--docs/superpowers/followup/pr3803-flasher-recover.md280
-rw-r--r--docs/superpowers/followup/pr3803-hil-blindness-reporting.md185
-rw-r--r--docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md118
-rw-r--r--docs/superpowers/followup/pr3803-pci-rebind-stranding.md157
-rw-r--r--docs/superpowers/followup/pr3803-usbtest-recovery-reserve.md175
-rw-r--r--docs/superpowers/plans/2026-07-14-usb-target-debug-handoff.md125
-rw-r--r--docs/superpowers/plans/2026-07-23-esp-target-debug-skill.md74
-rw-r--r--docs/superpowers/plans/2026-07-23-target-debug-skill-enhancement.md570
-rw-r--r--docs/superpowers/plans/2026-07-24-etm-trace-agent-integration.md299
-rw-r--r--docs/superpowers/plans/2026-07-27-openocd-unified-fork.md602
-rw-r--r--docs/superpowers/plans/2026-07-28-hil-test-split.md355
-rw-r--r--docs/superpowers/plans/2026-07-29-hil-select.md856
-rw-r--r--docs/superpowers/plans/2026-08-15-ci-hs-reset-edges.md782
-rw-r--r--docs/superpowers/plans/2026-08-16-drop-ep0-prime-verify.md314
-rw-r--r--docs/superpowers/specs/2026-07-09-claude-agents-workflows-design.md15
-rw-r--r--docs/superpowers/specs/2026-07-23-esp-target-debug-design.md102
-rw-r--r--docs/superpowers/specs/2026-07-28-hil-test-refactor-design.md141
-rw-r--r--docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md179
-rw-r--r--docs/superpowers/specs/2026-07-30-hil-usbtest-fleet-wedge-design.md236
-rw-r--r--docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md162
-rw-r--r--docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md90
24 files changed, 5877 insertions, 7 deletions
diff --git a/docs/getting_started.rst b/docs/getting_started.rst
index 7fcc2f5d1..bce028ccb 100644
--- a/docs/getting_started.rst
+++ b/docs/getting_started.rst
@@ -181,7 +181,7 @@ Some examples require udev permissions to access USB devices:
.. code-block:: bash
- $ cp `examples/device/99-tinyusb.rules <https://github.com/hathach/tinyusb/tree/master/examples/device/99-tinyusb.rules>`_ /etc/udev/rules.d/
+ $ sudo cp examples/device/99-tinyusb-examples.rules /etc/udev/rules.d/
$ sudo udevadm control --reload-rules && sudo udevadm trigger
Next Steps
diff --git a/docs/reference/boards.rst b/docs/reference/boards.rst
index 148175671..8b0f798ba 100644
--- a/docs/reference/boards.rst
+++ b/docs/reference/boards.rst
@@ -81,6 +81,15 @@ espressif_s3_devkitm Espresif S3 DevKitM espressif https://do
espressif_saola_1 Espresif S2 Saola 1 espressif https://docs.espressif.com/projects/esp-dev-kits/en/latest/esp32s2/esp32-s2-saola-1/index.html
========================= ============================== ========= ======================================================================================================== ======
+Geehy
+-----
+
+==================== ==================== ========= ============================================== ======
+Board Name Family URL Note
+==================== ==================== ========= ============================================== ======
+apm32f072_dev_board APM32F072 Dev Board apm32f0xx
+==================== ==================== ========= ============================================== ======
+
GigaDevice
----------
diff --git a/docs/reference/device_issues.rst b/docs/reference/device_issues.rst
index ae9cd55f1..b95a3fc1e 100644
--- a/docs/reference/device_issues.rst
+++ b/docs/reference/device_issues.rst
@@ -20,6 +20,51 @@ Most severe issues are:
- USB.5: In USB full-speed host mode, linked list on done queue is broken.
- USB.15: USB high-speed device in endpoint TX data corruption
+NXP i.MX RT1015/RT1020/RT1024/RT1050/RT1060/RT1064
+-----------------------------------------------------
+**Severity: High** when an isochronous IN endpoint is used behind a hub
+
+Reference: ERR050101 "USB: Endpoint conflict issue in device mode", listed in the errata sheet of
+every part above - `IMXRT1015CE`_, `IMXRT1020CE`_, `IMXRT1024CE`_, `IMXRT1050CE`_, `IMXRT1060CE`_
+and `IMXRT1064CE`_. On RT1060 and RT1064 it applies to rev A silicon only and is fixed in rev B; on
+RT1015, RT1020, RT1024 and RT1050 it is marked *no fix scheduled*, so all silicon is affected.
+RT1010, RT116x, RT117x and RT118x do not list it.
+
+.. _IMXRT1015CE: https://www.nxp.com/docs/en/errata/IMXRT1015CE.pdf
+.. _IMXRT1020CE: https://www.nxp.com/docs/en/errata/IMXRT1020CE.pdf
+.. _IMXRT1024CE: https://www.nxp.com/docs/en/errata/IMXRT1024CE.pdf
+.. _IMXRT1050CE: https://www.nxp.com/docs/en/errata/IMXRT1050CE.pdf
+.. _IMXRT1060CE: https://www.nxp.com/docs/en/errata/IMXRT1060CE.pdf
+.. _IMXRT1064CE: https://www.nxp.com/docs/en/errata/IMXRT1064CE.pdf
+
+While an isochronous IN endpoint is active, an IN token addressed to *that same endpoint number on
+another device sharing the host* can silently unprime one of this device's OUT endpoints - control,
+bulk, interrupt or isochronous alike. NXP states the unpriming cannot be detected by software and
+raises no interrupt, so the endpoint simply stops answering OUT tokens and the transfer never
+completes. Typically seen when the device is behind a hub with other devices attached.
+
+Workaround: give isochronous IN endpoints a number that no other device on the same host uses for
+any IN endpoint - endpoints 1-3 are used by nearly every composite device, so choose a high number
+(``examples/device/usbtest`` uses endpoint 7 on this family for that reason). Devices without an
+isochronous IN endpoint are unaffected.
+
+NXP LPC55S2x/LPC552x
+---------------------------------
+**Severity: Low** (both need specific conditions)
+
+Reference: `LPC55S2x Errata Sheet`_ USB.3, USB.5
+
+.. _LPC55S2x Errata Sheet: https://www.nxp.com/docs/en/errata/ES_LPC55S2x.pdf
+
+USB.3: As a high-speed device behind certain full-speed hubs, the device does not correctly detect
+the host's KJ chirp sequence and can behave erratically due to wrong speed detection. The documented
+workaround is to set the FORCE_FS bit in DEVCMDSTAT on bus reset when the reported link speed is
+full speed. TinyUSB does not implement this workaround.
+
+USB.5: An isochronous IN endpoint sending a 1024-byte maximum-packet-size packet raises no endpoint
+interrupt and its command/status entry is not updated. Workaround: cap the isochronous IN maximum
+packet size at 1023 bytes in the descriptor.
+
WCH CH32F20x/CH32V20x/CH32V30x
---------------------------------
**Severity: Medium**
@@ -33,3 +78,14 @@ Reference: `CH32V30X Reference Manual`_ USBFS/USBHS controller chapter
Data corruption may occur on isochronous endpoints. Due to the lacking of FIFO for interrupt status registers, later completed transfer will overwrite `INT_ST` and `RX_LEN` register if previous transfer processing is not completed.
Other types of transfers are not affected.
+
+Puya PY32F071/072
+---------------------------------
+**Severity: Very Low**
+
+Reference: `PY32F07x Reference Manual` USBD chapter
+
+The USB device controller (MUSB-like) has 5 application endpoints EP1-EP5 with fixed FIFO sizes
+shared between IN and OUT of the same endpoint number: EP1 = 512 B, EP2-4 = 128 B, EP5 = 64 B.
+This is much lower than the max ISO ep size of 1024 for high EP numbers.
+Place large isochronous endpoints on EP1 and size descriptors accordingly.
diff --git a/docs/superpowers/followup/pr3803-flasher-recover.md b/docs/superpowers/followup/pr3803-flasher-recover.md
new file mode 100644
index 000000000..e9fff7480
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-flasher-recover.md
@@ -0,0 +1,280 @@
+# `flasher_recover` Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Give the 15 HIL boards whose flasher cannot reach its probe past a poisoned usbfs
+node a second, convoy-safe flasher used only for recovery.
+
+**Architecture:** An optional roster key `flasher_recover` beside `flasher`.
+`hil_flash.recover_flasher(board)` picks it when present; `hil_test` substitutes it into the
+`--recover-board` JSON so `usbtest.py` never learns a second entry exists. Delivery over
+openocd's jlink driver is convoy-safe by construction, but the flash command form must
+differ from the one `flash_openocd` uses, so the recovery gets its own flasher name.
+
+**Tech Stack:** Python 3.13 stdlib, openocd 0.12.0+dev (build 0ce743125 on ci.lan),
+libjaylink, J-Link probes.
+
+## Global Constraints
+
+- Roster JSON: `test/hil/tinyusb.json`. `flasher_recover` is OPTIONAL; absent means today's
+ behaviour (`recover_flasher` returns the primary).
+- Never change the shape of `board['flasher']` — it is read as a dict in `hil_flash`,
+ `hil_test`, `usbtest`, `hil_pool_check`, `hil_select` and the roster lint, and is shipped
+ as JSON to a subprocess.
+- Flasher dispatch is by name: `getattr(hil_flash, f'flash_{name}')` / `reset_{name}`.
+- `RECOVER_FLASH_TIMEOUT = 90`, `RECOVER_RESET_TIMEOUT = 30` (`usbtest.py`). Any board whose
+ flash cannot finish inside 90 s is not a candidate.
+- Tests run offline: `cd test/hil && python3 test/test_hil_select.py`.
+
+## What is already established
+
+**Landed on PR #3803 and inert without roster entries:** `hil_flash.recover_flasher()`,
+`convoy_safe()` accepting openocd-over-jlink, `hil_test` substituting the recovery flasher
+into `--recover-board`, and `test_hil_select.FlasherRecoverEntry` (4 tests).
+
+**Verified in source:**
+- openocd's jlink driver ignores `adapter usb vid_pid` — `jlink.c` never reads
+ `adapter_usb_get_vids/pids`; selection is `adapter serial` / USB address / usb location.
+ Do NOT lint a jlink recovery entry for `vid_pid`.
+- It is convoy-safe anyway: libjaylink `discovery_usb.c` returns early unless
+ `idVendor == 0x1366` and the PID is in its table, and only THEN calls `libusb_open`. A
+ wedged `cafe:4010` DUT is never opened.
+- CMSIS-DAP stays pin-gated: `cmsis_dap_usb_bulk.c:107` skips before `libusb_open`, and
+ `id_filter` is only `vids[0] || pids[0]`.
+
+**Measured on ci.lan 2026-08-17**, base args
+`-f interface/jlink.cfg -c "transport select swd" -c "adapter speed 4000" -f target/<cfg>`:
+
+| Board | target cfg | flash | reset |
+|--------------------------|--------------|-------|-------|
+| stm32f407disco | stm32f4x | OK | OK |
+| stm32f072disco | stm32f0x | OK | OK |
+| stm32f723disco | stm32f7x | OK | OK |
+| stm32l476disco | stm32l4x | OK | OK |
+| feather_nrf52840_express | nrf52 | OK | OK |
+| metro_m4_express | atsame5x | OK | OK |
+| frdm_k64f | k60 | OK | OK |
+
+`frdm_k64f` is host-only (`tests.device == false`) — verify its reset over UART
+(`/dev/serial/by-id/usb-SEGGER_J-Link_000621000000-if00`), never by USB disconnect.
+
+**Excluded, with reasons:** `lpcxpresso11u37` — 118 s for 24 KB at 1 MHz with a verify
+mismatch, versus 0.277 s via JLinkExe; cannot fit `RECOVER_FLASH_TIMEOUT`.
+`mimxrt1064_evk`, `ra4m1_ek`, `nrf54lm20dk` — no target config exists in this openocd
+build, so they cannot be covered at all. **The board that wedges most (mimxrt1064_evk) is
+therefore still uncovered by this work.**
+
+**The blocker this plan solves:** `flash_openocd` issues `program <fw> verify reset exit`,
+which fails over the jlink transport on BOTH families tried (`stm32f4x`, `stm32f0x`) with
+`Examination failed` → `auto_probe failed`, with or without a preceding `init; reset halt`.
+Every successful flash above used the explicit sequence in Task 1.
+
+**Why this is a separate PR:** it adds a roster capability and a new flasher backend, which
+is a different scope from containing a wedge; and it needs bench time on seven boards.
+
+## File Structure
+
+- `test/hil/hil_flash.py` — add `flash_openocd_seq` / `reset_openocd_seq`; extend
+ `convoy_safe` to accept the new name. This is the only file that learns the command form.
+- `test/hil/tinyusb.json` — seven `flasher_recover` entries.
+- `test/hil/test/test_hil_select.py` — extend `FlasherRecoverEntry`; add a roster lint.
+
+---
+
+### Task 1: `openocd_seq` flasher backend
+
+**Files:**
+- Modify: `test/hil/hil_flash.py` (beside `flash_openocd`, ~line 100)
+- Test: `test/hil/test/test_hil_select.py`
+
+**Interfaces:**
+- Consumes: `_openocd_cmd_base(flasher)`, `hil_util.run_cmd`.
+- Produces: `flash_openocd_seq(board, firmware, timeout=None)`,
+ `reset_openocd_seq(board, timeout=None)`, both returning
+ `subprocess.CompletedProcess`; `convoy_safe()` returns True for
+ `{'name': 'openocd_seq', 'args': '...interface/jlink.cfg...'}`.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_openocd_seq_is_convoy_safe_over_jlink(self):
+ self.assertTrue(hil_flash.convoy_safe(
+ {'name': 'openocd_seq', 'args': '-f interface/jlink.cfg -f target/stm32f4x.cfg'}))
+
+ def test_openocd_seq_uses_explicit_flash_commands_not_program(self):
+ """`program` fails over the jlink transport: Examination failed -> auto_probe
+ failed, measured on stm32f4x and stm32f0x."""
+ seen = {}
+ real = hil_util.run_cmd
+ hil_util.run_cmd = lambda cmd, **k: seen.setdefault('cmd', cmd) or real('true')
+ try:
+ hil_flash.flash_openocd_seq(
+ {'flasher': {'name': 'openocd_seq', 'uid': 'X', 'args': '-f interface/jlink.cfg'}},
+ '/tmp/fw.elf', timeout=5)
+ finally:
+ hil_util.run_cmd = real
+ self.assertIn('flash write_image erase /tmp/fw.elf', seen['cmd'])
+ self.assertIn('verify_image /tmp/fw.elf', seen['cmd'])
+ self.assertNotIn('program ', seen['cmd'])
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_select.py FlasherRecoverEntry -v`
+Expected: FAIL — `module 'hil_flash' has no attribute 'flash_openocd_seq'`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def flash_openocd_seq(board, firmware, timeout=None):
+ # Explicit commands, NOT `program`: over the jlink transport `program` fails at the
+ # flash bank probe ("Examination failed" -> "auto_probe failed"), measured on
+ # stm32f4x and stm32f0x, with or without a preceding reset halt. This sequence
+ # succeeded on all seven candidate boards.
+ flasher = board['flasher']
+ verify = f' -c "verify_image {firmware}"' if flasher.get('verify', True) else ''
+ return hil_util.run_cmd(
+ f'{_openocd_cmd_base(flasher)} -c "init" -c "reset halt" '
+ f'-c "flash write_image erase {firmware}"{verify} -c "reset run" -c "shutdown"',
+ timeout=timeout)
+
+
+def reset_openocd_seq(board, timeout=None):
+ flasher = board['flasher']
+ return hil_util.run_cmd(
+ f'{_openocd_cmd_base(flasher)} -c "init" -c "reset run" -c "shutdown"',
+ timeout=timeout)
+```
+
+In `convoy_safe`, replace `if name != 'openocd':` with:
+
+```python
+ if name not in ('openocd', 'openocd_seq'):
+ return False
+```
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_select.py FlasherRecoverEntry -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/hil_flash.py test/hil/test/test_hil_select.py
+git commit -m "hil: add openocd_seq flasher for convoy-safe recovery delivery"
+```
+
+---
+
+### Task 2: Roster entries for the seven validated boards
+
+**Files:**
+- Modify: `test/hil/tinyusb.json`
+- Test: `test/hil/test/test_hil_select.py`
+
+**Interfaces:**
+- Consumes: `flash_openocd_seq` / `reset_openocd_seq` from Task 1.
+- Produces: seven boards for which `hil_flash.convoy_safe(hil_flash.recover_flasher(b))`
+ is True.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_roster_recover_entries_are_convoy_safe_and_named_openocd_seq(self):
+ import json, pathlib
+ roster = json.loads((pathlib.Path(__file__).parent.parent / 'tinyusb.json').read_text())
+ recover = [b for b in roster['boards'] if 'flasher_recover' in b]
+ self.assertGreaterEqual(len(recover), 7)
+ for b in recover:
+ f = b['flasher_recover']
+ self.assertEqual(f['name'], 'openocd_seq', b['name'])
+ self.assertIn('interface/jlink.cfg', f['args'], b['name'])
+ self.assertIn('adapter speed', f['args'], b['name']) # required; see below
+ self.assertTrue(hil_flash.convoy_safe(f), b['name'])
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_select.py FlasherRecoverEntry -v`
+Expected: FAIL — `0 >= 7`
+
+- [ ] **Step 3: Add the entries**
+
+`adapter speed` is REQUIRED: without it examination fails outright on the jlink driver.
+Add to each board below, using the SAME `uid` as its primary jlink entry:
+
+```json
+"flasher_recover": {
+ "name": "openocd_seq",
+ "uid": "<same probe serial as flasher.uid>",
+ "args": "-f interface/jlink.cfg -c \"transport select swd\" -c \"adapter speed 4000\" -f target/<cfg>.cfg"
+}
+```
+
+| Board | `uid` | `<cfg>` |
+|--------------------------|----------------|-----------|
+| stm32f407disco | 000773661813 | stm32f4x |
+| stm32f072disco | 779541626 | stm32f0x |
+| stm32f723disco | 000776606156 | stm32f7x |
+| stm32l476disco | 777632258 | stm32l4x |
+| feather_nrf52840_express | 681295394 | nrf52 |
+| metro_m4_express | 123456 | atsame5x |
+| frdm_k64f | 000621000000 | k60 |
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_select.py -v`
+Expected: PASS, and no other selector test regresses.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/tinyusb.json test/hil/test/test_hil_select.py
+git commit -m "hil: give seven J-Link boards a convoy-safe recovery flasher"
+```
+
+---
+
+### Task 3: Bench validation on the rig
+
+**Files:** none — this task produces evidence, not code.
+
+- [ ] **Step 1: Confirm the rig is idle and take the locks**
+
+```bash
+ssh [email protected] 'if pgrep -f "[h]il_test.py" >/dev/null; then echo BUSY; exit 1; fi'
+ssh [email protected] 'cd ~/actions-runner/_work/tinyusb/tinyusb && \
+ nohup timeout 900 python3 test/hil/helper/hil_lock.py hold <boards...> --reason "flasher_recover validation" &'
+```
+
+Guard with `if`, never `cmd && echo || echo` — that form only gates the echo and will take
+locks during a live CI run.
+
+- [ ] **Step 2: For each board, flash then reset through the recovery entry**
+
+```bash
+python3 test/hil/hil_test.py -b <board> test/hil/tinyusb.json # normal path still works
+```
+
+Then force the recovery path by running usbtest with the recovery flags and a firmware that
+hangs a case, or drive `hil_flash.flash_openocd_seq` / `reset_openocd_seq` directly.
+
+- [ ] **Step 3: Verify**
+
+Device boards: `sudo dmesg` shows `USB disconnect` then a fresh enumeration.
+`frdm_k64f`: UART shows the boot banner (see above).
+Every flash must finish well inside `RECOVER_FLASH_TIMEOUT` (90 s).
+
+- [ ] **Step 4: Release locks and record the results in the PR body**
+
+---
+
+## Out of scope, and why
+
+- **`mimxrt1064_evk`** needs an i.MX RT target config that this openocd build does not
+ have. Sourcing or writing one is its own investigation; until then the board with the
+ most wedges has no automated recovery.
+- **Changing `flash_openocd`** to the explicit form would cover these boards without a new
+ name, but `program` is what nine pinned CMSIS-DAP boards use in CI daily and no CMSIS-DAP
+ image could be built in the originating worktree (no pico-sdk) to re-validate it.
diff --git a/docs/superpowers/followup/pr3803-hil-blindness-reporting.md b/docs/superpowers/followup/pr3803-hil-blindness-reporting.md
new file mode 100644
index 000000000..69ff939b0
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-hil-blindness-reporting.md
@@ -0,0 +1,185 @@
+# Blindness Reporting Gaps Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Make a HIL worker's sysfs blindness reach the report in the two cases where it
+currently does not — an untested producer, and a board that raises.
+
+**Architecture:** A worker returns `hil_util.sysfs_blind()` as the last field of its result
+tuple; `_blind_note()` turns that into a report banner. Two holes: nothing tests the
+producer, and a board that raises returns no tuple at all, so its blindness is lost.
+
+**Tech Stack:** Python 3.13 stdlib, multiprocessing Pool with `maxtasksperchild=1`.
+
+## Global Constraints
+
+- A blind worker answers `SYSFS_UNKNOWN` for every attribute, so its "device not found"
+ means "could not tell". The report must say so or a red cell reads as a broken board.
+- `maxtasksperchild=1`: one worker per board, so the flag is per-board and must not be
+ smeared across boards.
+- Tests: `cd test/hil && python3 test/test_hil_bounded.py`.
+
+## What is already established
+
+- `hil_test.test_board` returns `(..., hil_util.sysfs_blind(), stray)`; `_blind_note(mret)`
+ renders the banner; wired into all three report paths.
+- **The producer is provably untested**: replacing `hil_util.sysfs_blind()` with `False` in
+ the return leaves all tests green. Nothing drives `test_board` — it needs a board dict, a
+ real flock, a flasher and `test_example` per test.
+- Blindness fired for real on ci.lan: four workers went blind in one run, and cells failed
+ *because* of it (`Printer device not found ... (this worker is blind)`).
+
+**Why this is a separate PR:** closing it means making `test_board` testable, which is a
+refactor of the harness's orchestration layer — a different scope from the containment
+work, and the reason the gap was accepted rather than papered over.
+
+## File Structure
+
+- `test/hil/hil_test.py` — extract the result-tuple assembly from `test_board` so it can be
+ built and asserted without running a board; carry blindness out of the raise path.
+- `test/hil/test/test_hil_bounded.py` — tests for both.
+
+---
+
+### Task 1: Make the result tuple assembly testable
+
+**Files:**
+- Modify: `test/hil/hil_test.py` (`test_board`, the `return (name, err_count, ...)` at the
+ end of the try block)
+- Test: `test/hil/test/test_hil_bounded.py`
+
+**Interfaces:**
+- Produces: `_board_result(name, err_count, failed_tests, rows, t_total, board_wide_fail)`
+ returning the 7-tuple `(name, err_count, failed, rows, t_total, blind, stray)`, reading
+ `hil_util.sysfs_blind()` and `hil_health.kill_own_children()` itself.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+class BoardResultCarriesBlindness(unittest.TestCase):
+ def test_a_blind_worker_reports_it(self):
+ from helper import hil_util, hil_health
+ self.addCleanup(setattr, hil_util, 'sysfs_blind', hil_util.sysfs_blind)
+ self.addCleanup(setattr, hil_health, 'kill_own_children', hil_health.kill_own_children)
+ hil_util.sysfs_blind = lambda: True
+ hil_health.kill_own_children = lambda: 0
+ row = hil_test._board_result('b', 0, [], [], 1.0, False)
+ self.assertTrue(row[5], 'blindness did not reach the result tuple')
+ self.assertIn('b', hil_test._blind_note([row]))
+
+ def test_a_sighted_worker_does_not(self):
+ from helper import hil_util, hil_health
+ self.addCleanup(setattr, hil_util, 'sysfs_blind', hil_util.sysfs_blind)
+ self.addCleanup(setattr, hil_health, 'kill_own_children', hil_health.kill_own_children)
+ hil_util.sysfs_blind = lambda: False
+ hil_health.kill_own_children = lambda: 0
+ row = hil_test._board_result('b', 0, [], [], 1.0, False)
+ self.assertFalse(row[5])
+ self.assertEqual(hil_test._blind_note([row]), '')
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py BoardResultCarriesBlindness -v`
+Expected: FAIL — `module 'hil_test' has no attribute '_board_result'`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def _board_result(name, err_count, failed_tests, rows, t_total, board_wide_fail):
+ """Assemble a worker's result tuple. Separate from test_board so the two fields only
+ the WORKER can answer -- its process-global blindness latch and what it could not kill
+ -- are testable without running a board."""
+ stray = hil_health.kill_own_children()
+ return (name, err_count, [] if board_wide_fail else sorted(set(failed_tests)),
+ rows, t_total, hil_util.sysfs_blind(), stray)
+```
+
+Replace the tail of `test_board` with:
+
+```python
+ return _board_result(name, err_count, failed_tests, rows, t_total, board_wide_fail)
+```
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py -v`
+Expected: PASS, and the existing `BlindWorkerReachesTheReport` tests still pass.
+
+- [ ] **Step 5: Verify the mutation is now caught**
+
+Replace `hil_util.sysfs_blind()` with `False` inside `_board_result` and re-run; the suite
+MUST fail. Restore it.
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add test/hil/hil_test.py test/hil/test/test_hil_bounded.py
+git commit -m "test/hil: make the worker result tuple testable, covering blindness"
+```
+
+---
+
+### Task 2: Carry blindness out of the worker-raise path
+
+**Files:**
+- Modify: `test/hil/hil_test.py` (`test_board`'s except/finally, and `main`'s worker-raise
+ handler that builds synthetic rows)
+- Test: `test/hil/test/test_hil_bounded.py`
+
+**Interfaces:**
+- Consumes: `_board_result` from Task 1.
+- Produces: a board that raises still contributes a row whose blindness field is accurate.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_a_board_that_raises_still_reports_blindness(self):
+ """The result tuple is returned inside a try whose finally only releases the lock,
+ so a board that dies by exception contributed nothing -- and its blindness, the
+ thing that most explains its failure, was lost with it."""
+ from helper import hil_util
+ self.addCleanup(setattr, hil_util, 'sysfs_blind', hil_util.sysfs_blind)
+ hil_util.sysfs_blind = lambda: True
+ row = hil_test._board_result_on_error('b', RuntimeError('boom'))
+ self.assertTrue(row[5])
+ self.assertIn('b', hil_test._blind_note([row]))
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py BoardResultCarriesBlindness -v`
+Expected: FAIL — no `_board_result_on_error`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def _board_result_on_error(name, exc):
+ """A row for a board that died by exception. err_count 1, no per-test detail, but the
+ blindness and stray fields are still accurate -- they explain the failure more often
+ than the exception text does."""
+ rows = [(name, {BOUNDARY_CELL: f'{REPORT_CELL["fail"]} {type(exc).__name__}'}, None)]
+ return _board_result(name, 1, [], rows, 0.0, True)
+```
+
+Wrap the body of `test_board` so the exception path returns it instead of propagating.
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/hil_test.py test/hil/test/test_hil_bounded.py
+git commit -m "test/hil: keep a raising board's blindness in the report"
+```
+
+---
+
+## Caution
+
+`test_board`'s `finally` releases the board flock. Any restructuring MUST keep that
+release on every path, including the new error path — a leaked flock locks the board until
+the host reboots.
diff --git a/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md b/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md
new file mode 100644
index 000000000..fe377f741
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md
@@ -0,0 +1,118 @@
+# IAR HIL Leg Re-run Spec Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Let the `hil-hfp-iar` CI leg re-run only its failed boards, as the other two HIL
+legs already do.
+
+**Architecture:** `hil_test.py` writes a `<config>.failed` spec into `HIL_REPORT_DIR`; a
+workflow step reads it on the next attempt and passes the boards back as arguments. The IAR
+leg passes `--retry 1` like the others but sets no `HIL_REPORT_DIR` and has no read-back
+step, so its spec is written into the workspace and never read.
+
+**Tech Stack:** GitHub Actions YAML, self-hosted runner.
+
+## Global Constraints
+
+- `.github/workflows/build.yml`. The two working legs are `hil-tinyusb` (matrix) — see its
+ `Set HIL report dir (per run+job; persists across run attempts)` and `Get re-run spec from
+ previous attempt` steps — and they are the pattern to copy.
+- The report dir must be keyed by run id AND job so a matrix leg does not collide with
+ another, and must survive across run attempts (that is the whole point).
+- The IAR leg is the only HIL job that BUILDS inline; its `Build` step is bounded at
+ `timeout-minutes: 30` under a 120-minute job ceiling. Do not disturb that.
+
+## What is already established
+
+- Verified by reading the workflow: `hil-hfp-iar` has neither `HIL_REPORT_DIR` nor a
+ `Get re-run spec` step, while passing `--retry 1`.
+- Consequence: a GitHub re-run of that job re-tests its whole matrix. **This is not a
+ regression** — that leg never had the mechanism — and the unread spec costs only a file.
+- The report artifact upload for that leg is named `hil-report-hfp-iar`.
+
+**Why this is a separate PR:** it is CI plumbing with no code change, it needs a real
+re-run on the self-hosted runner to prove, and it duplicates ~15 lines of workflow that
+would be better factored — a decision worth making on its own.
+
+## File Structure
+
+- `.github/workflows/build.yml` — the `hil-hfp-iar` job only.
+
+---
+
+### Task 1: Give the IAR leg a persistent report dir and a re-run spec
+
+**Files:**
+- Modify: `.github/workflows/build.yml` (job `hil-hfp-iar`)
+
+**Interfaces:**
+- Consumes: `hil_test.py`'s existing `--report-dir` / `.failed` behaviour — no code change.
+- Produces: `env.HIL_REPORT_DIR` for the job, and `$RERUN_ARGS` for the test step.
+
+- [ ] **Step 1: Copy the two steps from `hil-tinyusb`, before the Build step**
+
+```yaml
+ - name: Set HIL report dir (per run+job; persists across run attempts)
+ run: |
+ BASE=$HOME/hil-reports
+ echo "HIL_REPORT_DIR=$BASE/${GITHUB_RUN_ID}-hfp-iar" >> "$GITHUB_ENV"
+
+ - name: Get re-run spec from previous attempt
+ run: |
+ SPEC="$HIL_REPORT_DIR/hfp.json.failed"
+ if [ -f "$SPEC" ]; then
+ echo "RERUN_ARGS=$(cat "$SPEC")" >> "$GITHUB_ENV"
+ echo "re-running only: $(cat "$SPEC")"
+ fi
+```
+
+Match the exact spec filename `hil_test.py` writes for this leg's config — read
+`_write_failed_spec` and the `failed_fname` construction rather than assuming.
+
+- [ ] **Step 2: Pass the spec to the test step**
+
+```yaml
+ python3 test/hil/hil_test.py --retry 1 $SEL_ARGS hfp.json $RERUN_ARGS
+```
+
+`--retry 1` stays FIRST so argparse's last-wins keeps any explicit override working.
+
+- [ ] **Step 3: Point the artifact upload at the report dir**
+
+```yaml
+ path: ${{ env.HIL_REPORT_DIR }}/hil_report.md
+```
+
+- [ ] **Step 4: Validate the YAML**
+
+Run: `python3 -c "import yaml,sys; d=yaml.safe_load(open('.github/workflows/build.yml')); j=d['jobs']['hil-hfp-iar']; print(j['timeout-minutes'], [s.get('name') for s in j['steps']])"`
+Expected: the ceiling is still 120, the Build step still carries `timeout-minutes: 30`, and
+the two new steps appear before Build.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add .github/workflows/build.yml
+git commit -m "ci: let the IAR HIL leg re-run only its failed boards"
+```
+
+---
+
+### Task 2: Prove it on a real re-run
+
+**Files:** none — evidence only.
+
+- [ ] **Step 1:** Push and let `hil-hfp-iar` run to a failure (or force one).
+- [ ] **Step 2:** Confirm `$HIL_REPORT_DIR/hfp.json.failed` exists on the runner after the
+ job.
+- [ ] **Step 3:** Use GitHub's "Re-run failed jobs" and confirm the log line
+ `re-running only: ...` and that only those boards are tested.
+- [ ] **Step 4:** Record the run URL in the PR body.
+
+---
+
+## Consider first
+
+Three jobs would then carry the same ~15 lines. Factoring them into a composite action, or
+computing the report dir inside `hil_test.py` from `GITHUB_RUN_ID`, may be the better
+change — decide that before copying the block a third time.
diff --git a/docs/superpowers/followup/pr3803-pci-rebind-stranding.md b/docs/superpowers/followup/pr3803-pci-rebind-stranding.md
new file mode 100644
index 000000000..de1f7163b
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-pci-rebind-stranding.md
@@ -0,0 +1,157 @@
+# `pci-rebind` Stranding Investigation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Settle when a PCI unbind/rebind of an xHCI controller strands it driverless, so
+the `usb-kernel-recover` skill can state a rule instead of a hypothesis.
+
+**Architecture:** No product code. This is a controlled reproduction against the rig's
+kernel, ending in a documentation change and — if the boundary turns out to be
+detectable — a guard in `usb_recover.sh`.
+
+**Tech Stack:** Linux 6.12.96 (ci.lan), Renesas uPD720201 xHCI, `usb_recover.sh`.
+
+## Global Constraints
+
+- ci.lan is a live CI rig. Take every affected board's lock first
+ (`hil_lock.py hold --all --reason ...`) and confirm no `hil_test.py` is running, with an
+ `if`, not an `&&` chain.
+- A stranded controller takes every fixture on it offline; recovery is
+ `usb_recover.sh pci-bind <addr>` or, failing that, a PVE **host** power cycle — an
+ operator action. Do not start this without being able to reach the host.
+- The rig has two Renesas controllers plus an AMD one; pick the controller with the fewest
+ fixtures for the experiment.
+
+## What is already established
+
+**The skill claimed, unconditionally, that `pci-rebind`'s re-bind hangs on the D-state URB
+and leaves the controller with no driver.** That claim was generalised from ONE observation
+and was used to delete `pci-rebind` and `pci-bind` from `usb_recover.sh` entirely.
+
+**It was refuted in the field on 2026-08-17.** After `hub-cycle 17-2.7` failed to clear a
+wedge, `pci-rebind 0000:05:00.0` recovered the controller in about one second:
+
+```
+02:34:41 remove, state 4 / USB bus 18 deregistered
+02:34:41 remove, state 1 / USB bus 17 deregistered
+02:34:42 xHCI Host Controller / new USB bus registered, assigned bus number 1
+02:34:42 new USB bus registered, assigned bus number 2
+```
+
+Both actions were restored, with the guidance scoped to failure mode: **dead controller →
+use it; device-lock convoy → do not**. Buses renumbered 17/18 → 1/2, which is why rig-wide
+operations need every board's lock.
+
+**What is NOT known:** why the earlier attempt stranded and this one did not. The leading
+hypothesis is that it turns on whether a live D-state URB exists **on that controller** at
+the moment of the re-bind — but in the 02:34 incident the wedged board (17-2.7) was on that
+very controller, which weakens it. An alternative is that `hub-cycle` had already cleared
+the holder, leaving only a dead controller.
+
+**Why this is a separate PR:** it is an experiment that risks taking the rig offline, and
+its output is a documentation change plus possibly a guard — a different scope from any
+code change.
+
+## File Structure
+
+- `.claude/skills/usb-kernel-recover/SKILL.md` — replace the hypothesis in section 3b and
+ the Common-mistakes entry with whatever the experiment establishes.
+- `.claude/skills/usb-kernel-recover/scripts/usb_recover.sh` — only if the boundary is
+ detectable from userspace.
+
+---
+
+### Task 1: Reproduce a controller-scoped D-state wedge
+
+**Files:** none.
+
+- [ ] **Step 1: Establish the safety net**
+
+```bash
+ssh [email protected] 'if pgrep -f "[h]il_test.py" >/dev/null; then echo BUSY; exit 1; fi'
+# hold ALL boards on the target controller
+```
+
+Confirm host access to pve.lan before continuing.
+
+- [ ] **Step 2: Create a wedge deliberately**
+
+Run `usbtest.py` against a board known to hang (`mimxrt1064_evk` has wedged eight times,
+TEST 9/10/24/27), or drive `testusb` directly until a case does not return.
+
+- [ ] **Step 3: Confirm the holder and its controller**
+
+```bash
+ps -eo pid,stat,etimes,wchan:22,args | awk '$2 ~ /D/'
+sudo cat /proc/<pid>/stack # usbdev_ioctl + [usbtest] = the owner
+readlink -f /sys/bus/usb/devices/usb<N> # bus -> PCI addr
+```
+
+Record whether the holder is on the SAME controller you will rebind.
+
+---
+
+### Task 2: Rebind and record the outcome
+
+**Files:** none.
+
+- [ ] **Step 1: Rebind, with a bounded observer**
+
+```bash
+timeout 120 sudo usb_recover.sh pci-rebind <addr>; echo "rc=$?"
+```
+
+- [ ] **Step 2: Record which of the three outcomes occurred**
+
+1. Re-bind completes, controller recovers (as on 2026-08-17).
+2. Re-bind hangs; `/sys/bus/pci/devices/<addr>/driver` is gone → **stranded**.
+3. Re-bind completes but the wedge persists.
+
+Capture `sudo journalctl -k --since ...` around the attempt either way.
+
+- [ ] **Step 3: If stranded, recover**
+
+```bash
+sudo usb_recover.sh pci-bind <addr>
+```
+
+If that hangs too, the only remaining step is a PVE host power cycle — an operator action.
+
+- [ ] **Step 4: Repeat at least three times**
+
+One observation is what produced the wrong rule in the first place. Vary whether a D-state
+holder is live on that controller at rebind time; that is the hypothesis under test.
+
+---
+
+### Task 3: Write down what was learned
+
+**Files:**
+- Modify: `.claude/skills/usb-kernel-recover/SKILL.md`
+
+- [ ] **Step 1: Replace section 3b's scoping with the measured rule**
+
+State the condition under which stranding occurs, with the journal lines. If the experiment
+does NOT reproduce stranding, say that too, with the attempt count — "not reproduced in N
+attempts" is a better record than an unexplained warning.
+
+- [ ] **Step 2: If the boundary is detectable, guard the script**
+
+For example, refuse `pci-rebind` when a D-state holder exists on that controller, since the
+holder is enumerable from `/proc` and the controller from `readlink`. Only add this if the
+experiment shows it predicts the outcome.
+
+- [ ] **Step 3: Commit**
+
+```bash
+git add .claude/skills/usb-kernel-recover/
+git commit -m "skills: replace the pci-rebind stranding hypothesis with measurement"
+```
+
+---
+
+## Abort criteria
+
+Stop and hand back to the operator if: a rebind strands the controller and `pci-bind` does
+not recover it; `uhubctl` starts hanging (the convoy has spread to the hub locks); or a CI
+run starts while the rig is in a broken state.
diff --git a/docs/superpowers/followup/pr3803-usbtest-recovery-reserve.md b/docs/superpowers/followup/pr3803-usbtest-recovery-reserve.md
new file mode 100644
index 000000000..eb8959520
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-usbtest-recovery-reserve.md
@@ -0,0 +1,175 @@
+# usbtest Recovery Reserve Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Make the post-hang recovery reserve a derived, asserted property instead of an
+accident of four independently-set constants.
+
+**Architecture:** `hil_test` passes `--budget` and `--outer-timeout` to `usbtest.py`, which
+decides at runtime whether a recovery still fits. Today the reserve survives only because
+the four numbers happen to line up; nothing ties them together or fails when they stop.
+
+**Tech Stack:** Python 3.13 stdlib.
+
+## Global Constraints
+
+- `usbtest.py`: `RECOVER_FLASH_TIMEOUT = 90`, `RECOVER_RESET_TIMEOUT = 30`.
+- `hil_test.py`: `USBTEST_BATTERY_BUDGET = 260`, `USBTEST_RECOVERY_BUDGET = 250`,
+ `USBTEST_OVERSHOOT = 120`; `outer = BATTERY_BUDGET + (RECOVERY_BUDGET if recovery else
+ OVERSHOOT)`, used for both the child's `--outer-timeout` and the parent's `run_cmd` bound.
+- All five are env-overridable via `hil_util.pos_int_env`, so a rig can change them.
+- Tests: `cd test/hil && python3 test/test_hil_health.py` and `test_hil_bounded.py`.
+
+## What is already established
+
+The reserve holds at the shipped values, checked by hand:
+
+- The battery checks its budget BEFORE dispatching a case, so it can overshoot by one
+ case — worst case `260 + 60 + 5 = 325 s`.
+- Recovery is gated on `_time_left() >= RECOVER_RESET_TIMEOUT`, where
+ `_time_left() = outer_timeout - elapsed - 35`; with `outer = 510` that allows recovery
+ until `elapsed = 445 s`, and the reflash until `385 s`.
+- So ~60 s of margin survives, and recovery does fire.
+
+**The defect is structural, not arithmetic:** lower `--outer-timeout`, raise `--timeout`, or
+raise `USBTEST_BATTERY_BUDGET` via the env and the reserve silently disappears. The failure
+mode is a skipped reflash that leaves the D-state holder for the next job — the exact thing
+the containment exists to prevent — with no error anywhere.
+
+**Why this is a separate PR:** it changes the timing contract between `hil_test` and
+`usbtest.py`, which affects every board's run duration, so it wants its own review and a
+full rig run.
+
+## File Structure
+
+- `test/hil/usbtest.py` — a `reserve_ok()` predicate plus a startup assertion.
+- `test/hil/hil_test.py` — derive the battery budget from the outer bound rather than
+ setting both independently.
+- `test/hil/test/test_hil_health.py` — tests.
+
+---
+
+### Task 1: Assert the reserve at startup
+
+**Files:**
+- Modify: `test/hil/usbtest.py` (constants block, and `main()` after argparse)
+- Test: `test/hil/test/test_hil_health.py`
+
+**Interfaces:**
+- Produces: `usbtest.reserve_ok(budget, outer, case_timeout)` returning bool.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+class RecoveryReserveIsChecked(unittest.TestCase):
+ """The battery may overshoot its budget by ONE already-started case, so the outer bound
+ must leave room for that overshoot AND a bounded recovery afterwards."""
+
+ def setUp(self):
+ import usbtest
+ self.u = usbtest
+
+ def test_the_shipped_numbers_leave_room(self):
+ self.assertTrue(self.u.reserve_ok(budget=260, outer=510, case_timeout=60))
+
+ def test_a_tighter_outer_bound_is_rejected(self):
+ self.assertFalse(self.u.reserve_ok(budget=260, outer=380, case_timeout=60))
+
+ def test_a_longer_case_timeout_is_rejected(self):
+ self.assertFalse(self.u.reserve_ok(budget=260, outer=510, case_timeout=200))
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_health.py RecoveryReserveIsChecked -v`
+Expected: FAIL — `module 'usbtest' has no attribute 'reserve_ok'`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def reserve_ok(budget: int, outer: int, case_timeout: int) -> bool:
+ """Does `outer` leave room for the battery's worst case AND a bounded recovery?
+
+ The budget is checked BEFORE dispatch, so the battery can run to
+ `budget + case_timeout + 5` (the +5 is run_case's reap). _time_left() subtracts a
+ further 35 s of fixed tail. A reflash needs RECOVER_FLASH_TIMEOUT beyond that.
+ """
+ worst_case_end = budget + case_timeout + 5
+ return outer - worst_case_end - 35 >= RECOVER_FLASH_TIMEOUT
+```
+
+In `main()`, after parsing args:
+
+```python
+ if args.budget and args.outer_timeout and not reserve_ok(
+ args.budget, args.outer_timeout, args.timeout):
+ print(f'warning: --outer-timeout {args.outer_timeout} leaves no room for a bounded '
+ f'recovery after a --budget {args.budget} battery with --timeout '
+ f'{args.timeout} cases; a HUNG board will be left wedged', file=sys.stderr)
+```
+
+Warn, do not exit: a caller that deliberately runs without recovery is legitimate.
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_health.py RecoveryReserveIsChecked -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/usbtest.py test/hil/test/test_hil_health.py
+git commit -m "usbtest: check the recovery reserve instead of assuming it"
+```
+
+---
+
+### Task 2: Derive the outer bound from one place
+
+**Files:**
+- Modify: `test/hil/hil_test.py` (constants block ~line 227, and `test_device_usbtest`)
+- Test: `test/hil/test/test_hil_bounded.py`
+
+**Interfaces:**
+- Consumes: `usbtest.reserve_ok` semantics (duplicate the arithmetic, do not import
+ usbtest — `hil_test` must not import it).
+- Produces: an assertion at module import that the shipped constants satisfy the reserve.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_the_shipped_constants_satisfy_the_reserve(self):
+ """Whatever the env overrides, the pair hil_test computes must leave recovery room:
+ outer - (budget + case_timeout + 5) - 35 >= 90."""
+ outer = hil_test.USBTEST_BATTERY_BUDGET + hil_test.USBTEST_RECOVERY_BUDGET
+ self.assertGreaterEqual(outer - (hil_test.USBTEST_BATTERY_BUDGET + 60 + 5) - 35, 90)
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Temporarily set `HIL_USBTEST_RECOVERY_BUDGET=100` and run; expect FAIL. Unset.
+
+- [ ] **Step 3: Add the guard**
+
+```python
+# The recovery reserve is a PROPERTY of these two, not a coincidence: the battery may
+# overshoot its budget by one already-started case (checked before dispatch), and a bounded
+# reflash needs 90 s after a 35 s fixed tail. Env overrides make this checkable at import
+# rather than discoverable when a wedge is left unrecovered.
+if USBTEST_RECOVERY_BUDGET - 60 - 5 - 35 < 90:
+ print(f'warning: HIL_USBTEST_RECOVERY_BUDGET={USBTEST_RECOVERY_BUDGET} leaves no room '
+ f'for a bounded reflash after a one-case overshoot; HUNG boards will stay wedged',
+ file=sys.stderr)
+```
+
+- [ ] **Step 4: Run tests to verify they pass**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/hil_test.py test/hil/test/test_hil_bounded.py
+git commit -m "hil: warn when the timeout constants leave no recovery reserve"
+```
diff --git a/docs/superpowers/plans/2026-07-14-usb-target-debug-handoff.md b/docs/superpowers/plans/2026-07-14-usb-target-debug-handoff.md
new file mode 100644
index 000000000..b56d035d9
--- /dev/null
+++ b/docs/superpowers/plans/2026-07-14-usb-target-debug-handoff.md
@@ -0,0 +1,125 @@
+# Hand-off: `usb-target-debug` skill + `target-debugger` agent
+
+**Status: agreed but NOT started.** Design discussion happened 2026-07-13 in session
+`c31a4617-43b1-491d-9865-3e35f393996b` (post-merge of the agents/workflows harness,
+PR #3762 / `ac595bc5c`). This document is the implementation brief for a fresh session.
+
+**Agreed sequencing: skill first → dogfood on 1-2 real HIL failures → then the agent
+as its own small PR.** Do not build both at once — the agent charter's hard parts are
+exactly what dogfooding the skill answers.
+
+## The gap being filled
+
+When HIL fails today, *what failed* is covered (hil-validate workflow, hil-operator
+agent) but the deep *why* loop — instrument the target, capture on both sides,
+correlate — has no skill and no agent. Every hard case so far (musb babble, rusb2
+FRDY wedge, ch32v307 Heisenbug) fell back to interactive main-session work.
+
+Why no existing agent can do it:
+
+- **hil-operator** (sonnet) is deliberately mechanical: lock → flash → `hil_test.py`
+ → recover. It never edits source, so it cannot inject instrumentation.
+- **port-dev** can edit source but its charter is scoped changes verified by a
+ *build*; it has no hardware mandate.
+- The host-side capture knowledge lives in skills (`usbmon`, `usb-debug`); the
+ device-side half exists only as CLAUDE.md recipes plus session memory.
+
+The skill completes the debugging trio:
+
+| Skill | Answers | Status |
+|---|---|---|
+| `usbmon` | what the host actually exchanged (URBs) | on master |
+| `usb-debug` | why the host acted (dmesg / dynamic debug) | ships in PR #3758 (untracked copy in tree) |
+| `usb-target-debug` | what the device did | **this hand-off** |
+
+## Part 1 — `usb-target-debug` skill (do this first)
+
+Create `.claude/skills/usb-target-debug/SKILL.md`. Match the style of
+`.claude/skills/usbmon/SKILL.md` and `usb-debug/SKILL.md`: frontmatter `name` +
+`description` where the description states concretely *when* to reach for it
+(HIL test fails and host-side capture can't explain it; device silently NAKs,
+wedges, or misbehaves; need TU_LOG/device-state evidence from real hardware).
+
+Playbook to codify — all techniques already proven on this rig:
+
+1. **TU_LOG capture** — build with `LOG=2` (add `LOGGER=rtt` for RTT); UART capture
+ from the board's debug serial; RTT via `JLinkGDBServer -RTTTelnetPort 19021` +
+ `JLinkRTTClient` (non-interactive: `timeout 20s JLinkRTTClient > rtt.log`).
+ Note which log level perturbs timing (see warning #6).
+2. **GDB recipes per probe family** — J-Link, OpenOCD (ST-Link / CMSIS-DAP /
+ WCH-Link). Base connect/load recipes already exist in CLAUDE.md "GDB Debugging";
+ the skill adds the debug-loop specifics: breakpoints in ISR context, dumping
+ endpoint/FIFO registers, watchpoints on driver state variables.
+3. **RAM ring-buffer trace pattern** (used to crack the musb babble): instrument
+ the dcd/hcd with a small RAM ring of event records instead of TU_LOG when
+ printing perturbs timing; let the failure happen; halt and dump the ring via
+ GDB. Include a minimal C snippet (fixed-size struct ring, no allocation,
+ ISR-safe single-writer).
+4. **J-Link PC-sampling** (nailed the rusb2 FRDY wedge): statistically sample PC
+ without halting to find where the core spins — the non-intrusive option when
+ halting or logging masks the bug.
+5. **Dual-side capture**: usbmon on the host + RTT/ring-buffer on the target,
+ simultaneously; correlate host URBs against device events on one timeline.
+ This is the default posture for enumeration/transfer bugs, not an escalation.
+6. **Warnings**: observation can mask the bug (the ch32v307 case changed behavior
+ under logging/debug — prefer ring-buffer over TU_LOG, PC-sampling over halting,
+ and say so explicitly); a J-Link core reset does NOT drop a DWC2 soft-connect
+ pullup, so a wedged DUT stays wedged on the host side (cross-ref
+ `usb-recover/SKILL.md`).
+7. **Rig discipline**: hold the board lock for the whole manual session —
+ `python3 test/hil/board_lock.py hold <board> --reason "target debug: <bug>"`
+ … work … `release <board>`. Never stop the actions-runner. Board → probe
+ mapping via `test/hil/tinyusb.json`; `JLINK_DEVICE`/`OPENOCD_OPTION` via
+ `hw/bsp/*/boards/*/board.cmake` or `board.mk`.
+
+**Where to ship**: its own small PR (usb-recover/usb-debug already belong to
+PR #3758 — don't grow that one), or fold into #3758 if it is still open and being
+rebased anyway. User's call at the time.
+
+## Part 2 — `target-debugger` agent (later, after dogfooding)
+
+Create `.claude/agents/target-debugger.md` as its own PR once the skill has been
+through at least one real debug session.
+
+Agreed charter outline:
+
+- **Frontmatter**: `model: opus`; omit `tools:` (= all tools — it must edit source
+ AND drive hardware). Note the registry supports no `effort` field — the agreed
+ opus/**xhigh** tier is requested per `agent()` call by whichever workflow or
+ session spawns it.
+- **Loop**: instrument → build → flash under one held board lock → dual-side
+ capture (host usbmon + target RTT/ring-buffer/GDB) → correlate → refine
+ hypothesis → repeat. Deliberately serial: no fan-out win; the value is
+ backgrounding a long debug session and the codified playbook.
+- **Strictly one instance**, holds the board lock for the entire session — its work
+ is exactly the "hardware work outside hil_test.py" case in the lock protocol.
+- **Skills are its source of truth** (mirror hil-operator's pattern): read
+ `usb-target-debug`, `usbmon`, `usb-debug`, `usb-recover`, `hil` SKILL.md files
+ before acting.
+- **Hard rule — instrumentation is temporary**: the instrumentation diff must be
+ reverted (or explicitly listed in the hand-back report) at session end; the *fix*
+ itself goes to port-dev. Keeps charters clean: this agent produces a diagnosis
+ and evidence, not a merged patch.
+
+Questions dogfooding must answer before the charter is written (do NOT guess these
+now — that was the whole reason for skill-first):
+
+1. When to stop instrumenting and report a partial diagnosis vs keep digging.
+2. Maximum board-lock hold time / check-in cadence for a backgrounded session.
+3. What "revert instrumentation" means when a partial fix emerged mid-debug
+ (revert + attach diff? keep on a branch?).
+
+## Conventions and references for the implementing session
+
+- Skill style exemplars: `.claude/skills/usbmon/SKILL.md`, `usb-debug/SKILL.md`,
+ `usb-recover/SKILL.md` (the latter two are #3758's copies, present untracked).
+- Agent style exemplars: `.claude/agents/hil-operator.md` (lock discipline,
+ skills-as-source-of-truth), `port-dev.md` (source-edit + verify charter).
+- When the agent lands, update the harness spec's agent roster:
+ `docs/superpowers/specs/2026-07-09-claude-agents-workflows-design.md`
+ (convention: spec evolves in-repo; plans like this file are per-effort records).
+- Agents register from `.claude/agents/*.md` at session start — a new agent file
+ is only visible to sessions launched after it exists.
+- Past cases to mine for the skill's examples: musb babble (ring-buffer trace),
+ rusb2 FRDY wedge (J-Link PC-sampling), ch32v307 Heisenbug (observation
+ sensitivity) — details in session memory and the referenced session transcript.
diff --git a/docs/superpowers/plans/2026-07-23-esp-target-debug-skill.md b/docs/superpowers/plans/2026-07-23-esp-target-debug-skill.md
new file mode 100644
index 000000000..d08902111
--- /dev/null
+++ b/docs/superpowers/plans/2026-07-23-esp-target-debug-skill.md
@@ -0,0 +1,74 @@
+# esp-target-debug Skill Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Create `.claude/skills/esp-target-debug/SKILL.md` (Espressif built-in USB-Serial-JTAG debug backend) with every recipe verified on the rig's P4, the S3 PHY boundary verified both ways, plus pointer edits in `target-debug` and the `target-debugger` agent.
+
+**Architecture:** Per spec `docs/superpowers/specs/2026-07-23-esp-target-debug-design.md`. Verification-first: hardware gates 1–6 run before the skill text lands, so only proven content ships unmarked. One lock session per board.
+
+**Tech Stack:** ESP-IDF at `$HOME/code/esp-idf` (`export.sh` → `openocd-esp32`, `riscv32-esp-elf-gdb`, `xtensa-esp32s3-elf-gdb`, `esptool.py`), rig boards `espressif_p4_function_ev` (uid 6055F9F98715), `espressif_s3_devkitm` (uid 84F703C084E4).
+
+## Global Constraints
+
+- Worktree `/home/hathach/code/tinyusb/.claude/worktrees/improve-debug-skill-agent`, branch `claude/improve-debug-skill-agent`.
+- Board-lock discipline per `hil` skill; reflash pristine firmware before release; evidence (command + output snippet) in commit message bodies.
+- Formatting: aligned table columns, skill-name-only cross-references.
+- Unverified content ships tagged `(untested)` or not at all.
+- Espressif anything requires `. $HOME/code/esp-idf/export.sh` in that shell first.
+
+---
+
+### Task 1: P4 recon + coexistence gate (spec gates 1)
+
+- [x] **Step 1: Environment + firmware recon**
+
+```bash
+ls $HOME/code/esp-idf/export.sh && source $HOME/code/esp-idf/export.sh && which openocd riscv32-esp-elf-gdb
+ls /home/hathach/code/tinyusb/examples/cmake-build-espressif_p4_function_ev 2>/dev/null || echo "no prebuilt"
+lsusb -d 303a:1001 # USB-SJ devices present
+```
+If no prebuilt firmware: build `device/cdc_msc_freertos` for the P4 (`idf.py -DBOARD=espressif_p4_function_ev build` in that example, per CLAUDE.md), else use the prebuilt binary. Identify the ELF path for gdb symbolization.
+
+- [x] **Step 2: Lock P4, ensure known firmware, confirm DUT traffic**
+
+```bash
+python3 test/hil/board_lock.py hold espressif_p4_function_ev --reason "esp-target-debug verify: coexistence"
+# flash known build (esptool/idf.py flash -p <port-by-uid>), settle, then confirm enumeration:
+lsusb | grep -i cafe # TinyUSB VID on the DUT port
+# generate traffic: echo > /dev/ttyACM<N> of the cdc, or timeout 5s cat
+```
+
+- [x] **Step 3: Attach openocd over USB-SJ while the device runs**
+
+```bash
+openocd -f board/esp32p4-builtin.cfg -c 'adapter serial 60:55:F9:F9:87:15' & # gdb :3333 — USB-SJ iSerial = MAC with colons
+riscv32-esp-elf-gdb -batch -ex 'target extended-remote :3333' -ex 'monitor halt' \
+ -ex bt -ex 'monitor resume' <p4 elf>
+```
+Expected: backtrace with symbols; after resume the CDC device still answers (re-run the traffic check). Record: does the DUT drop off the bus during halt (host URB timeouts — expected per target-debug) and does it recover on resume without re-enumeration?
+
+- [x] **Step 4: Release-or-continue checkpoint** — keep the lock for Task 2 (same session). No commit yet; evidence to `/tmp/esp_evidence.txt`.
+
+### Task 2: P4 budget, watchpoint, threads, console (spec gates 2–4)
+
+- [x] **Step 1: Breakpoint/watchpoint budget** — RISC-V trigger count: in gdb `monitor riscv info` or set watchpoints until rejection; verify a hardware watchpoint on a TinyUSB variable (e.g. `watch -l` on a usbd counter) reports and hits.
+- [x] **Step 2: FreeRTOS threads** — `info threads` after halt; expect ESP-IDF tasks incl. the USB task; note whether it works at attach or needs run→stop (mirror the ARM finding).
+- [x] **Step 3: Console during traffic** — OUTCOME: stock builds route the console to UART0 (the CP2102 flasher tty — boot log captured there); the USB-SJ CDC carries no log without sdkconfig `ESP_CONSOLE_USB_SERIAL_JTAG`, which stays (untested) in the skill.
+- [x] **Step 4: Reflash pristine, release P4 lock.** Evidence appended to `/tmp/esp_evidence.txt`.
+
+### Task 3: P4 apptrace spike — GATED (spec gate 5)
+
+Budget 30 min. `openocd -c 'esp apptrace start ...'` against a firmware built with apptrace enabled? Stock HIL firmware has no apptrace init — if a code change would be required, that's the gate answer: land apptrace as `(untested — needs CONFIG_APPTRACE + firmware init)` with the recipe sketch. Only a working capture lands unmarked.
+
+### Task 4: S3 boundary (spec gate 6)
+
+- [x] **Step 1: Lock S3, flash `board_test`** (no TinyUSB → PHY free). Attach `openocd -f board/esp32s3-builtin.cfg -c 'adapter serial 84F703C084E4'` + `xtensa-esp32s3-elf-gdb`: halt + bt works.
+- [x] **Step 2: Flash a USB device example** — record the exact failure: does 303a:1001 vanish from lsusb (PHY switched), does openocd fail to attach or die mid-session? Capture verbatim error.
+- [x] **Step 3: Reflash pristine (a USB example — that is the CI-expected state), release.**
+
+### Task 5: Write the skill + integration edits + commit
+
+- [x] **Step 1: Write `.claude/skills/esp-target-debug/SKILL.md`** per spec section order (role/defer, PHY map with verified boundary symptoms, toolchain+attach with the real commands from Tasks 1–4, technique mapping table with verified annotations, rig deltas, external-JTAG TODO). Aligned tables.
+- [x] **Step 2: `target-debug` pointer** (2 lines, after probe-mapping bullets) + `target-debugger` agent table row.
+- [x] **Step 3: pre-commit, single commit** with evidence summary from `/tmp/esp_evidence.txt`.
+- [x] **Step 4: Retrieval sanity** — one fresh-subagent scenario: "debug a TinyUSB hang on the rig's P4" routes to esp-target-debug (not JLink recipes); "same on S3 while cdc_msc runs" routes to the PHY boundary + external-JTAG TODO.
diff --git a/docs/superpowers/plans/2026-07-23-target-debug-skill-enhancement.md b/docs/superpowers/plans/2026-07-23-target-debug-skill-enhancement.md
new file mode 100644
index 000000000..36a3144c2
--- /dev/null
+++ b/docs/superpowers/plans/2026-07-23-target-debug-skill-enhancement.md
@@ -0,0 +1,570 @@
+# target-debug Skill & target-debugger Agent Enhancement Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Extend `.claude/skills/target-debug/SKILL.md` (and its agent) with the full debugger facility arsenal from the J-Link, OpenOCD, and GDB manuals — breakpoint/watchpoint depth, OpenOCD RTT, vector catch + fault autopsy, SWO/ITM trace, flash verification — each recipe hardware-verified on the ci rig before it lands unmarked.
+
+**Architecture:** The skill's organizing spine is its intrusiveness table ("pick the least intrusive technique that can answer the question"); every new facility slots into that model with an honest cost row. Recipes keep the existing dense, copy-paste style. The skill's value is that its recipes are *proven on this rig* — so each task pairs drafting with a bounded hardware verification, and anything unverifiable lands tagged `(untested)` or is dropped.
+
+**Tech Stack:** arm-none-eabi-gdb 15.2, OpenOCD 0.12.0+dev, SEGGER J-Link V7.94b (`JLinkExe`, `JLinkGDBServer`, `JLinkSWOViewerCLExe`), ci rig boards from `test/hil/tinyusb.json` (10 jlink / 6 openocd / 1 stlink probes).
+
+**Reference manual:** "Debugging with GDB", **Tenth Edition** (for GDB 18.0.50) — prefer the calibre-library copy via the `read-doc` skill, but **verify the edition on the title page first**: the library also holds an outdated Ninth Edition (2002, GDB 5.1.1, txt) that predates `dprintf`/`watch -l` — do not use it. Fallback fetch: `curl -sL -o /tmp/gdb.pdf https://sourceware.org/gdb/current/onlinedocs/gdb.pdf` (HTML pages block fetchers; the PDF does not). Sections used by this plan: §5.1.2 Setting Watchpoints, §5.1.6 Break Conditions, §5.1.7 Breakpoint Command Lists, §5.1.8 Dynamic Printf (PDF page = book page + 18). NOTE: the manual documents GDB 18; the rig runs 15.2 — the installed `arm-none-eabi-gdb`'s `help <cmd>` is authoritative for feature availability.
+
+## Global Constraints
+
+- Worktree: `/home/hathach/code/tinyusb/.claude/worktrees/improve-debug-skill-agent`, branch `claude/improve-debug-skill-agent`. All paths below are relative to it.
+- J-Link User Guide link must be exactly `https://kb.segger.com/UM08001_J-Link_/_J-Trace_User_Guide` (user-specified, verified live 2026-07-23).
+- **Hardware-verify before landing**: a recipe is committed unmarked only with captured evidence from a rig board; otherwise tag it `(untested)` inline or drop it. Record evidence (command + output snippet) in the task's commit message body.
+- Rig discipline (from `hil` + `target-debug` skills): `python3 test/hil/board_lock.py hold <board> --reason "skill-enhance verify: <what>"` before touching hardware, `release` after; reflash pristine firmware before release; NEVER stop the actions-runner; one J-Link client per probe at a time; we are ON host `ci` (config `test/hil/tinyusb.json`).
+- Hardware tasks are strictly serial (one board session at a time). Bash timeouts ≥ 10 min for flash+debug cycles.
+- Style: match the skill's existing voice — dense, recipe-first, caveats inline. Skill word budget after all tasks: ≤ 2 700 words (`wc -w`, currently 1 763).
+- Run `pre-commit run --files <changed>` before every commit. No Co-Authored-By trailers.
+- Board selection is runtime data (boards come/go, locks): resolve with the exact python snippet in Task 2 Step 2 and reuse `$JB` (jlink board) / `$OB` (openocd board) thereafter.
+
+---
+
+### Task 1: Manuals reference block
+
+**Files:**
+- Modify: `.claude/skills/target-debug/SKILL.md` (insert new `## Manuals` section immediately before `## Warnings`)
+
+**Interfaces:**
+- Produces: `## Manuals` section that later tasks' text may reference as "see Manuals".
+
+- [x] **Step 1: Insert the Manuals section**
+
+In `.claude/skills/target-debug/SKILL.md`, find the line `## Warnings` and insert immediately before it:
+
+```markdown
+## Manuals
+
+- J-Link / J-Trace User Guide (UM08001): <https://kb.segger.com/UM08001_J-Link_/_J-Trace_User_Guide> — flash breakpoints, RTT, SWO, monitor mode, Commander commands.
+- OpenOCD User's Guide: <https://openocd.org/doc/html/index.html> — `rtt`, `bp`/`wp`, `cortex_m vector_catch` / `maskisr`, `itm`/`tpiu`.
+- "Debugging with GDB" (the official manual; §5.1 covers break/watch/dprintf):
+ calibre library first (`read-doc` skill) — use the **Tenth Edition (GDB 18)**
+ copy, not the 2002 Ninth-Edition txt also present; fallback
+ `curl -sL -o /tmp/gdb.pdf https://sourceware.org/gdb/current/onlinedocs/gdb.pdf`
+ (the HTML mirror blocks fetchers; the PDF works). The installed
+ `arm-none-eabi-gdb`'s `help <cmd>` is authoritative for what this rig runs.
+
+```
+
+- [x] **Step 2: Verify formatting and word count**
+
+Run: `cd /home/hathach/code/tinyusb/.claude/worktrees/improve-debug-skill-agent && grep -A5 '^## Manuals' .claude/skills/target-debug/SKILL.md && wc -w .claude/skills/target-debug/SKILL.md`
+Expected: section present before `## Warnings`; word count ≤ 1 830.
+
+- [x] **Step 3: Commit**
+
+```bash
+cd /home/hathach/code/tinyusb/.claude/worktrees/improve-debug-skill-agent
+pre-commit run --files .claude/skills/target-debug/SKILL.md
+git add .claude/skills/target-debug/SKILL.md
+git commit -m "docs(target-debug): link J-Link UM08001, OpenOCD and GDB manuals"
+```
+
+---
+
+### Task 2: Breakpoint & watchpoint arsenal (GDB + OpenOCD)
+
+**Files:**
+- Modify: `.claude/skills/target-debug/SKILL.md` — extend the `## GDB — state autopsy and watchpoints` section
+- Read-only reference: `test/hil/tinyusb.json` (board resolution)
+
+**Interfaces:**
+- Consumes: nothing from other tasks.
+- Produces: board env vars `$JB`, `$OB` resolution snippet (reused by Tasks 3-6); the "halt-per-hit cost model" wording that Task 7's table row cites.
+
+- [x] **Step 1: Draft the section extension**
+
+In `.claude/skills/target-debug/SKILL.md`, the GDB section currently ends with the paragraph beginning `While halted the device answers **nothing**`. Insert immediately BEFORE that paragraph:
+
+```markdown
+**Hardware budget — read it off the chip, not from memory** (counts differ
+per core: M0+ typically 4 bp/2 wp, M3/M4 6/4, M7 8/4):
+
+```gdb
+p ((*(unsigned*)0xE0002000)>>4) & 0xF # FPB NUM_CODE = hw breakpoints (M7 adds bits[14:12])
+p (*(unsigned*)0xE0001000)>>28 # DWT_CTRL NUMCOMP = watchpoint comparators
+```
+
+- `hbreak`/`thbreak` force a hardware breakpoint (code in flash can't take a
+ software break unless the probe does flash breakpoints — J-Link does,
+ OpenOCD needs `bp <addr> 2 hw`); `tbreak` = one-shot.
+- `watch -l <expr>` watches the *address* the expression evaluates to once —
+ cheap and what you almost always want; `rwatch`/`awatch` trap reads/any
+ access (hardware-only — they error rather than fall back). OpenOCD (telnet
+ :4444) adds a data-VALUE match GDB cannot express: `wp <addr> 4 w <value>
+ [mask]` — fires only when the written value matches (e.g. catch who writes
+ 0 into a busy flag, ignoring writes of 1).
+- **Demand the word "Hardware" in the confirmation.** `watch` silently falls
+ back to a SOFTWARE watchpoint when no DWT comparator fits (expression too
+ wide/complex, budget exhausted): GDB then single-steps the whole program —
+ hundreds of times slower, certain USB death. `Watchpoint 2:` without
+ "Hardware" = delete it; `set can-use-hw-watchpoints 1` is the default but
+ narrowing the expression (`watch -l`, cast to a 4-byte int) is the real fix.
+- Conditional breaks/watches (`break dcd_edpt_xfer if ep_addr==0x81`) are
+ evaluated by GDB on the HOST with our stubs — neither JLinkGDBServer nor
+ OpenOCD supports target-side agent expressions on Cortex-M — so every hit
+ is a halt+resume (~ms) whether the condition matches or not: fine
+ post-wedge or on cold paths, wrong under live USB traffic.
+- `commands <bpnum> ... end` auto-runs GDB commands at each hit (start with
+ `silent`, end with `continue` for hands-free evidence collection) — same
+ halt-per-hit cost.
+- `dprintf <loc>,"fmt",args` = printf without recompiling. Stay on the
+ default `dprintf-style gdb` (host prints): the `call` style runs the
+ target's own printf mid-halt and `agent` needs stub support — neither is
+ viable on these probes. Same cost model as conditional breaks; for
+ ISR-rate events use the RAM ring buffer instead.
+- Stepping while the USB ISR fires between every step is chaos: OpenOCD
+ `cortex_m maskisr steponly` masks interrupts during single-steps only.
+ The bus keeps running either way — the host may still reset a device that
+ stops responding mid-step.
+- While halted you can poke state to test a hypothesis (`set var
+ _usbd_dev.ep_status[2][1].busy = 0`) — but that invalidates the snapshot
+ as post-mortem evidence; dump first, poke after.
+```
+
+- [x] **Step 2: Resolve verification boards (runtime data)**
+
+```bash
+cd /home/hathach/code/tinyusb
+python3 - <<'EOF'
+import json
+cfg = json.load(open('test/hil/tinyusb.json'))
+jl = [b['name'] for b in cfg['boards'] if b['flasher']['name']=='jlink']
+oo = [b['name'] for b in cfg['boards'] if b['flasher']['name']=='openocd']
+print('JLINK candidates:', jl)
+print('OPENOCD candidates:', oo)
+EOF
+```
+Pick the first candidate of each that `python3 test/hil/board_lock.py status` shows unlocked; export as `JB=<jlink board>` `OB=<openocd board>`. Look up `flasher.uid` for each in `test/hil/tinyusb.json` (`JB_UID`, `OB_UID`) and `JLINK_DEVICE`/`OPENOCD_OPTION` from `hw/bsp/*/boards/$JB/board.cmake` (family via `ls -d hw/bsp/*/boards/$JB`).
+
+- [x] **Step 3: Hardware-verify the budget reads on both probe families**
+
+```bash
+python3 test/hil/board_lock.py hold $JB --reason "skill-enhance verify: bp/wp budget"
+printf 'mem32 E0002000, 1\nmem32 E0001000, 1\nqc\n' | \
+ JLinkExe -device $JLINK_DEVICE -SelectEmuBySN $JB_UID -if swd -speed 4000 -autoconnect 1 -nogui 1
+python3 test/hil/board_lock.py release $JB
+```
+Expected: two register values; decode NUM_CODE and NUMCOMP by hand and check they are plausible (2-8 range). Repeat for `$OB` via `openocd $OPENOCD_OPTION -c init -c 'mdw 0xE0002000' -c 'mdw 0xE0001000' -c shutdown` under its own lock.
+If a register reads 0 on one board, note which core and adjust the skill text's example counts if contradicted.
+
+- [x] **Step 4: Hardware-verify dprintf + commands round-trip on $JB**
+
+With the board lock held and an already-flashed example (any; do not reflash), a JLinkGDBServer on :2331 (per CLAUDE.md GDB Debugging), run bounded — `commands` blocks cannot be passed via `-ex`, so use a command file:
+
+```bash
+cat > /tmp/bpcmd.gdb <<'EOF'
+target remote :2331
+set var $count=0
+watch -l *(unsigned*)&_usbd_dev
+delete
+dprintf tud_task_ext,"tick\n"
+break tud_task_ext
+commands 3
+silent
+set var $count=$count+1
+continue
+end
+continue&
+shell sleep 3
+interrupt
+print $count
+EOF
+timeout 120 arm-none-eabi-gdb -batch -x /tmp/bpcmd.gdb \
+ $(find examples/cmake-build-$JB -name 'cdc_msc.elf' | head -1)
+```
+Expected: the `watch` line answers `Hardware watchpoint 1:` (the word
+"Hardware" present — this is the skill's software-fallback check, then
+deleted), "tick" lines printed, and `$count > 0`. (`tud_task_ext` is the real
+symbol — `tud_task` is an inline wrapper; the breakpoint is number 3 after
+the watchpoint and dprintf.) Kill the GDB server, reflash pristine
+(`ninja`-flash target or `hil_test.py` flash path), release the lock.
+
+- [x] **Step 5: Apply the Step-1 text, run pre-commit, commit**
+
+```bash
+cd /home/hathach/code/tinyusb/.claude/worktrees/improve-debug-skill-agent
+pre-commit run --files .claude/skills/target-debug/SKILL.md
+git add .claude/skills/target-debug/SKILL.md
+git commit -m "docs(target-debug): breakpoint/watchpoint arsenal with halt-per-hit cost model
+
+Verified on <JB> (J-Link) + <OB> (OpenOCD): FPB/DWT budget reads, dprintf,
+breakpoint command lists. <paste the two register values here>"
+```
+
+---
+
+### Task 3: OpenOCD RTT — RTT is not J-Link-only
+
+**Files:**
+- Modify: `.claude/skills/target-debug/SKILL.md` — `## TU_LOG capture` section
+
+**Interfaces:**
+- Consumes: `$OB`, `$OB_UID`, `$OPENOCD_OPTION` from Task 2 Step 2.
+- Produces: the corrected claim "RTT works on any OpenOCD-driven probe" that Task 7's agent text repeats.
+
+- [x] **Step 1: Replace the J-Link-only claim**
+
+In the `## TU_LOG capture` section, replace:
+
+```markdown
+Build with `LOG=2` (`LOG=3` adds per-transfer noise and much more timing skew).
+`LOGGER=rtt` routes it over the debug probe (J-Link only) — no UART wiring:
+```
+
+with:
+
+```markdown
+Build with `LOG=2` (`LOG=3` adds per-transfer noise and much more timing skew).
+`LOGGER=rtt` routes it over the debug probe — no UART wiring. SEGGER's host
+tools need a J-Link, but OpenOCD serves the same RTT buffer on ST-Link /
+CMSIS-DAP / WCH-Link boards:
+```
+
+- [x] **Step 2: Add the OpenOCD RTT recipe**
+
+Immediately after the existing J-Link/UART capture code block (ends with `... | tee /tmp/uart.log`), add:
+
+```markdown
+```bash
+# OpenOCD RTT (any probe OpenOCD drives) — in telnet :4444 (or -c equivalents):
+rtt setup 0x20000000 0x8000 "SEGGER RTT" # search range = RAM ORIGIN + LENGTH (from the .ld / map file)
+rtt start # after firmware booted; rerun after each reflash
+rtt server start 19021 0
+# then: timeout 20s nc localhost 19021 > /tmp/rtt.log
+```
+
+OpenOCD polls the buffer (default 10 ms): bursty logs can drop lines a J-Link
+would keep — prefer J-Link where both exist; the drain-model warning below
+applies unchanged.
+```
+
+- [x] **Step 3: Hardware-verify on $OB**
+
+```bash
+python3 test/hil/board_lock.py hold $OB --reason "skill-enhance verify: openocd rtt"
+cd examples/device/cdc_msc && cmake -B build-rtt -DBOARD=$OB -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel -DLOG=2 -DLOGGER=rtt && cmake --build build-rtt
+# flash it (ninja -C build-rtt cdc_msc-openocd), then:
+openocd $OPENOCD_OPTION & # gdb :3333, telnet :4444
+{ echo 'rtt setup 0x20000000 0x8000 "SEGGER RTT"'; echo 'rtt start'; echo 'rtt server start 19021 0'; sleep 1; } | nc -q1 localhost 4444
+timeout 10s nc localhost 19021 > /tmp/ob_rtt.log; head /tmp/ob_rtt.log
+```
+Expected: TinyUSB boot banner / log lines in `/tmp/ob_rtt.log`. Adjust the search range from the board's linker script if the control block isn't found ("rtt: No control block found") and mirror any correction into the Step-2 text. Kill openocd, reflash pristine cdc_msc (no LOG), release lock, delete `build-rtt`.
+
+- [x] **Step 4: Commit**
+
+```bash
+pre-commit run --files .claude/skills/target-debug/SKILL.md
+git add .claude/skills/target-debug/SKILL.md
+git commit -m "docs(target-debug): RTT via OpenOCD on non-J-Link probes
+
+Verified on <OB>: rtt setup/start/server + nc capture of boot log.
+<paste first captured log line>"
+```
+
+---
+
+### Task 4: Vector catch + fault autopsy
+
+**Files:**
+- Modify: `.claude/skills/target-debug/SKILL.md` — new section after `## GDB — state autopsy and watchpoints`
+
+**Interfaces:**
+- Consumes: `$JB` from Task 2. (Corrected during execution: $OB/rp2040 is ARMv6-M — no CFSR/BFAR and only VC_HARDERR, so the full autopsy verify needs the ARMv7-M $JB; the payload is a bad LOAD because stores fault imprecisely with BFAR invalid.)
+- Produces: section title `## Vector catch + fault autopsy` cited by Task 7's table row.
+
+- [x] **Step 1: Insert the new section**
+
+After the GDB section (i.e. before `## RAM ring-buffer trace`), insert:
+
+```markdown
+## Vector catch + fault autopsy — catch the crash, not the wedge
+
+A "wedge" that is really a fault (HardFault loop, lockup) autopsies best AT
+the faulting instruction, not minutes later. Arm before reproducing:
+
+```gdb
+# tool-agnostic (any probe, incl. J-Link): DEMCR trap bits — halt on fault
+set *(unsigned*)0xE000EDFC |= (1<<10)|(1<<9)|(1<<8)|(1<<7)|(1<<6)|(1<<5)|(1<<4)
+# = VC_HARDERR|INTERR|BUSERR|STATERR|CHKERR|NOCPERR|MMERR; bit0 VC_CORERESET halts at reset
+```
+
+OpenOCD native form: `cortex_m vector_catch hard_err bus_err state_err chk_err mm_err`.
+When it fires the core halts at the fault; decode:
+
+```gdb
+p/x *(unsigned*)0xE000ED28 # CFSR — low byte MemManage, byte1 BusFault, top half UsageFault
+p/x *(unsigned*)0xE000ED2C # HFSR — bit30 FORCED = an escalated lower-priority fault
+p/x *(unsigned*)0xE000ED38 # BFAR — faulting address (valid if CFSR bit15 BFARVALID)
+x/8wx $msp # stacked frame: r0 r1 r2 r3 r12 lr pc xpsr — pc = culprit
+```
+
+`arm-none-eabi-addr2line -e <elf> <stacked pc>` names the line. Caveats: a
+vector-catch halt is still a halt (host-side URB timeouts apply); the bits
+persist until power-cycle — clear them (`... &= ~0x7F1`) before handing the
+board back; RISC-V ports have no DEMCR — use a breakpoint on the trap handler.
+```
+
+- [x] **Step 2: Hardware-verify with a deliberate fault on $JB (ARMv7-M)**
+
+Create the fault build (NOT committed):
+
+```bash
+python3 test/hil/board_lock.py hold $JB --reason "skill-enhance verify: vector catch"
+cd examples/device/cdc_msc # executed on $JB (stm32f407disco, ARMv7-M) via JLinkExe — see commit evidence
+# temporary patch — revert after: fault 5 s after boot
+python3 - <<'EOF'
+import pathlib
+p = pathlib.Path('src/main.c'); s = p.read_text()
+import re
+s = re.sub(r'\\nint main\\(void\\)',
+ '\\nstatic void _fault_after_5s(void){ static uint32_t t0=0; if(!t0) t0=tusb_time_millis_api();'
+ ' if(tusb_time_millis_api()-t0>5000) (void)*(volatile uint32_t*)0xCF000000u; }\\n\\nint main(void)', s, count=1) # board_millis is gone; helper must sit after the includes
+s = s.replace('led_blinking_task();', 'led_blinking_task(); _fault_after_5s();', 1)
+p.write_text(s)
+EOF
+grep -n '_fault_after_5s' src/main.c # expect 3 hits: definition + call + (none in decl block)
+cmake -B build-fault -DBOARD=$JB -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel && cmake --build build-fault
+```
+(If `app_led_task`/`board_millis` anchors differ in the current `main.c`, place the same 3-line helper on whatever per-loop task function exists — the fault line `*(volatile uint32_t*)0xCF000000u = 0;` is the payload.)
+Flash `build-fault`, then:
+
+```bash
+# executed variant: DEMCR armed + autopsy via JLinkExe command file on $JB (see commit c1d2d305f evidence); OpenOCD-native form:
+openocd $OPENOCD_OPTION -c init -c 'cortex_m vector_catch hard_err bus_err' &
+timeout 60 arm-none-eabi-gdb -batch -ex 'target remote :3333' -ex 'monitor reset run' \
+ -ex 'shell sleep 8' -ex 'interrupt' \
+ -ex 'p/x *(unsigned*)0xE000ED28' -ex 'p/x *(unsigned*)0xE000ED38' -ex 'x/8wx $msp' \
+ build-fault/cdc_msc.elf
+```
+Expected: halted in the fault path, CFSR BusFault bits set, **BFAR = 0xCF000000**, stacked pc addr2lines to `_fault_after_5s`. If the write is silently ignored on this core (some buses RAZ/WI), switch payload to a NULL-function call `((void(*)(void))0x1)();` and note UsageFault/INVSTATE instead.
+
+- [x] **Step 3: Clean up hardware state**
+
+`git checkout -- src/main.c`, delete `build-fault/`, clear DEMCR bits (`set *(unsigned*)0xE000EDFC &= ~0x7F1` via a final gdb attach or power-cycle note), reflash pristine cdc_msc, `board_lock.py release $JB`.
+
+- [x] **Step 4: Commit**
+
+```bash
+pre-commit run --files .claude/skills/target-debug/SKILL.md
+git add .claude/skills/target-debug/SKILL.md
+git commit -m "docs(target-debug): vector catch + Cortex-M fault autopsy recipe
+
+Verified on <OB>: deliberate bad-address write halted via vector_catch,
+CFSR=<val> BFAR=0xCF000000, stacked pc resolved by addr2line."
+```
+
+---
+
+### Task 5: SWO/ITM experiment — exception trace & hardware PC sampling
+
+This is an EXPERIMENT task with an explicit gate: the section lands **unmarked only if packets are actually captured** on a rig board; otherwise it lands tagged `(untested — SWO wiring unconfirmed on this rig)`. Budget: 30 min of hardware time, then decide.
+
+**Files:**
+- Modify: `.claude/skills/target-debug/SKILL.md` — new subsection inside the PC-sampling section (after the OpenOCD variant paragraph)
+
+**Interfaces:**
+- Consumes: `$JB`, `$JB_UID`, `$JLINK_DEVICE` from Task 2.
+- Produces: verified-or-tagged status consumed by Task 7's table row for SWO.
+
+- [x] **Step 1: Probe for SWO output (gate experiment)**
+
+```bash
+python3 test/hil/board_lock.py hold $JB --reason "skill-enhance verify: SWO"
+# arm DWT sources while the fw runs (background mem write, no halt):
+printf 'w4 E0001000, 0x00011401\nqc\n' | JLinkExe -device $JLINK_DEVICE -SelectEmuBySN $JB_UID -if swd -speed 4000 -autoconnect 1 -nogui 1
+# EXCTRCENA(16)|PCSAMPLENA(12)|SYNCTAP(10)|CYCCNTENA(0); tune POSTPRESET[4:1] if PC samples flood — then hand the probe to the viewer:
+timeout 20s JLinkSWOViewerCLExe -device $JLINK_DEVICE -usb $JB_UID -swofreq 4000000 -itmmask 0xFFFFFFFF | head -40
+```
+Gate: ANY decoded output (stimulus, PC samples, exception packets) = SWO wired on `$JB` → land unmarked with the observed invocation. No output → try one more J-Link board, then land tagged. Either way `release $JB` after reflashing nothing (this experiment flashes nothing).
+
+- [x] **Step 2: Insert the section (wording per gate outcome)**
+
+Append to the `## PC-sampling` section:
+
+```markdown
+### SWO/ITM — hardware-timed trace on one pin (J-Link)
+
+If the board routes SWO (TRACESWO), DWT emits packets with ZERO code change:
+**exception trace** (`DWT_CTRL` bit16 EXCTRCENA) — every IRQ enter/exit,
+timestamped, the ISR-ordering evidence the ring buffer needs code for — and
+**hardware PC sampling** (bit12 PCSAMPLENA), better histograms than DWT_PCSR
+polling. Arm the bits, then give the probe to the viewer (one client rule):
+
+```bash
+printf 'w4 E0001000, 0x00011401\nqc\n' | JLinkExe -device $JLINK_DEVICE -SelectEmuBySN <uid> ...
+timeout 20s JLinkSWOViewerCLExe -device $JLINK_DEVICE -usb <uid> -swofreq 4000000 -itmmask 0xFFFFFFFF
+```
+
+SWO needs the pin physically wired to the probe — many rig boards route only
+SWDIO/SWCLK. If the viewer shows nothing, that is the wiring, not the recipe.
+```
+
+If the gate FAILED on both boards, append ` (untested — SWO wiring unconfirmed on this rig)` to the subsection heading and keep the text.
+
+- [x] **Step 3: Commit**
+
+```bash
+pre-commit run --files .claude/skills/target-debug/SKILL.md
+git add .claude/skills/target-debug/SKILL.md
+git commit -m "docs(target-debug): SWO exception-trace / hw PC-sampling recipe
+
+Gate result on <JB>: <captured packet types | no SWO output — tagged untested>."
+```
+
+---
+
+### Task 6: Flash verification, FreeRTOS thread awareness, semihosting & monitor-mode notes
+
+**Files:**
+- Modify: `.claude/skills/target-debug/SKILL.md` — `## Warnings` section + GDB section tail
+
+**Interfaces:**
+- Consumes: `$JB`, `$JB_UID`, `$JLINK_DEVICE` from Task 2.
+- Produces: warning-list entries cited in Task 7's retrieval test scenarios.
+
+- [x] **Step 1: Add flash-content verification to Warnings**
+
+In `## Warnings`, after the "A marginal link can fake a deterministic firmware bug" bullet, add:
+
+```markdown
+- **"Flash OK" can lie** (silent no-op: old firmware keeps running after a
+ green flash). When behavior contradicts the code you think is flashed,
+ verify flash against the build:
+ `arm-none-eabi-objcopy -O binary fw.elf /tmp/fw.bin`, then J-Link
+ `verifybin /tmp/fw.bin,<flash-base>` (Commander) or OpenOCD
+ `verify_image /tmp/fw.bin <flash-base>` — a mismatch means reflash with
+ verification before debugging another minute.
+```
+
+- [x] **Step 2: Add FreeRTOS + semihosting + monitor-mode notes to the GDB section**
+
+Append to the end of the `## GDB — state autopsy and watchpoints` section (after the Task-2 additions):
+
+```markdown
+FreeRTOS examples (`*_freertos`): add `-rtos GDBServer/RTOSPlugin_FreeRTOS`
+to JLinkGDBServer (OpenOCD: `-rtos FreeRTOS` on the target) and `info
+threads` / `thread <n>` shows every task's stack — a USB task blocked on a
+queue vs. spinning is one `bt` away. Semihosting is never the answer here:
+each call traps and halts the core — RTT does the same job without stopping.
+**Monitor-mode debugging** (J-Link, M3+) can keep the USB ISR serviced while
+you sit at a breakpoint — needs SEGGER's `JLINK_MONITOR.c`/ISR files compiled
+in + `SetMonModeDebug=1`; not set up in this repo, reach for it when a bug
+truly needs live breakpoints without killing the bus:
+<https://kb.segger.com/Monitor_Mode_Debugging> (untested).
+```
+
+- [x] **Step 3: Hardware-verify verifybin + FreeRTOS awareness on $JB**
+
+```bash
+python3 test/hil/board_lock.py hold $JB --reason "skill-enhance verify: verifybin+rtos"
+# (a) verifybin positive path against whatever is flashed — first reflash a known build:
+# flash examples/cmake-build-$JB/device/cdc_msc, then:
+arm-none-eabi-objcopy -O binary examples/cmake-build-$JB/device/cdc_msc/cdc_msc.elf /tmp/fw.bin
+printf 'verifybin /tmp/fw.bin,<flash-base from board .ld>\nqc\n' | \
+ JLinkExe -device $JLINK_DEVICE -SelectEmuBySN $JB_UID -if swd -speed 4000 -autoconnect 1 -nogui 1
+# (b) rtos plugin: flash cdc_msc_freertos for $JB (build if missing), start
+JLinkGDBServer -device $JLINK_DEVICE -select usb=$JB_UID -if swd -speed 4000 -port 2331 -nogui -rtos GDBServer/RTOSPlugin_FreeRTOS &
+timeout 60 arm-none-eabi-gdb -batch -ex 'target remote :2331' -ex 'monitor halt' -ex 'info threads' \
+ <path to cdc_msc_freertos.elf>
+python3 test/hil/board_lock.py release $JB # after pristine reflash
+```
+Expected: (a) `Verify successful.` (b) `info threads` lists FreeRTOS tasks (`usbd`, `IDLE`, ...). If the plugin errors ("Could not load RTOS plugin"), drop the JLinkGDBServer variant from the Step-2 text and keep only the OpenOCD `-rtos FreeRTOS` form tagged `(untested)`.
+
+- [x] **Step 4: Commit**
+
+```bash
+pre-commit run --files .claude/skills/target-debug/SKILL.md
+git add .claude/skills/target-debug/SKILL.md
+git commit -m "docs(target-debug): flash verifybin, FreeRTOS thread awareness, monitor-mode pointer
+
+Verified on <JB>: verifybin 'Verify successful.'; info threads listed <n> tasks."
+```
+
+---
+
+### Task 7: Intrusiveness table integration, agent update, retrieval test
+
+**Files:**
+- Modify: `.claude/skills/target-debug/SKILL.md` — the technique/intrusiveness table
+- Modify: `.claude/agents/target-debugger.md` — primary-playbook bullet
+
+**Interfaces:**
+- Consumes: verified/untested status of every technique from Tasks 2-6.
+
+- [x] **Step 1: Extend the intrusiveness table**
+
+The table under `## Pick the least intrusive technique that can answer the question` currently has 5 rows (PC-sampling → GDB halt). Replace it with (keep the header row and any wording the earlier tasks did not contradict):
+
+```markdown
+| Technique | Intrusiveness | Reach for it when |
+|---|---|---|
+| PC-sampling | none — no halt, no code change | core wedged/spinning somewhere unknown (rusb2 FRDY) |
+| SWO exception trace / hw PC-sample | none — needs SWO pin wired | ISR ordering/timing with zero code change |
+| Vector catch | none until a fault fires | crash-shaped wedges — autopsy AT the faulting pc |
+| RAM ring-buffer | ~tens of cycles per event | ISR ordering/timing bugs (musb babble) |
+| TU_LOG (RTT) | µs per line | logic bugs that survive logging (J-Link or OpenOCD rtt) |
+| TU_LOG (UART) | ms per line — blocking write | same, when no debug-probe RTT path |
+| dprintf / conditional breakpoint | halt+resume per hit (~ms) | low-rate probes post-wedge; never ISR-rate events |
+| GDB halt / breakpoints | stops USB service entirely | post-mortem state autopsy once wedged |
+```
+
+If Task 5's gate failed, keep the SWO row but append ` (untested)` in its "Reach for it" cell.
+
+- [x] **Step 2: Update the agent's playbook bullet**
+
+In `.claude/agents/target-debugger.md`, replace:
+
+```markdown
+- `.claude/skills/target-debug/SKILL.md` — your primary playbook: technique
+ choice by intrusiveness, channel choice by link topology, capture recipes,
+ GDB autopsy, all rig warnings.
+```
+
+with:
+
+```markdown
+- `.claude/skills/target-debug/SKILL.md` — your primary playbook: technique
+ choice by intrusiveness, channel choice by link topology, capture recipes,
+ breakpoint/watchpoint budget and cost model, vector catch + fault autopsy,
+ GDB autopsy, all rig warnings.
+```
+
+- [x] **Step 3: Word-count and stale-reference check**
+
+Run: `wc -w .claude/skills/target-debug/SKILL.md` — expected ≤ 2 700. If over, trim prose (not recipes) until under.
+Run: `grep -n 'J-Link only' .claude/skills/target-debug/SKILL.md` — expected: no output (Task 3 removed the claim).
+
+- [x] **Step 4: Retrieval test (skill-TDD GREEN gate)**
+
+Dispatch a fresh read-only subagent (Explore) that reads ONLY the updated `.claude/skills/target-debug/SKILL.md` and answers:
+
+1. "A CH32 board's firmware wedges; you suspect a HardFault loop. Least-intrusive next step?" — expected: vector catch (with the RISC-V caveat noted: CH32 is RISC-V → breakpoint on trap handler).
+2. "You need RTT logs on an ST-Link-only board." — expected: OpenOCD `rtt setup/start/server`, NOT "impossible/J-Link only".
+3. "Who is writing 0 into a busy flag, under live traffic?" — expected: OpenOCD value-match watchpoint `wp <addr> 4 w 0`, NOT a GDB conditional watch (halt-per-hit cost).
+4. "Flash reported OK but behavior matches last week's build." — expected: verifybin/verify_image.
+5. "You set `watch xfer_status[2][1]` and GDB answered `Watchpoint 2:` (no 'Hardware'). Proceed?" — expected: NO — software-watchpoint fallback single-steps the program; delete and narrow the expression.
+
+All five must route correctly; a miss = fix the text (usually the table row or a heading), re-test.
+
+- [x] **Step 5: Final commit**
+
+```bash
+pre-commit run --files .claude/skills/target-debug/SKILL.md .claude/agents/target-debugger.md
+git add .claude/skills/target-debug/SKILL.md .claude/agents/target-debugger.md
+git commit -m "docs(target-debug): integrate new techniques into intrusiveness table; agent playbook bullet
+
+Retrieval test: 4/4 scenarios routed correctly."
+```
+
+---
+
+## Deferred / out of scope (deliberate)
+
+- **ETM / J-Trace instruction trace** — no J-Trace hardware on the rig; UM08001 "Trace" chapter is linked for the day one arrives.
+- **Monitor-mode debugging as a working recipe** — needs SEGGER monitor files compiled into firmware (a firmware feature, not a doc change); landed as a pointer + `(untested)` in Task 6.
+- **ITM stimulus-port logging backend for TU_LOG** — would be a `lib/` + `LOGGER=itm` firmware feature; out of scope for a skill-doc plan.
+- **GDB tracepoints (`trace`/`tfind`)** — need a tracing-capable stub; neither JLinkGDBServer nor OpenOCD implements them for Cortex-M.
diff --git a/docs/superpowers/plans/2026-07-24-etm-trace-agent-integration.md b/docs/superpowers/plans/2026-07-24-etm-trace-agent-integration.md
new file mode 100644
index 000000000..ebcc089fd
--- /dev/null
+++ b/docs/superpowers/plans/2026-07-24-etm-trace-agent-integration.md
@@ -0,0 +1,299 @@
+# etm-trace Skill Tightening + target-debugger Agent Integration Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Fix post-rebase staleness in the etm-trace skill, add its
+hardware-consent gate, and wire it into the target-debugger agent + target-debug
+skill the same way usb-sniffer is wired in (hardware-gated, user-confirmed).
+
+**Architecture:** Three curated instruction files get surgical edits (this repo
+treats skills/agents as curated docs — smallest possible diffs, no bulk
+rewrites). The gate follows the two existing consent patterns: in the *skill*
+(read by interactive sessions) it is "confirm with the user unless they asked";
+in the *agent* (which cannot ask mid-session) it is "only when your prompt
+states it", mirroring the agent's existing lock-force rule.
+
+**Tech Stack:** Markdown instruction files, git, one subagent retrieval test.
+
+## Global Constraints
+
+- Work in the existing worktree `/home/hathach/code/tinyusb/.worktrees/add-etm-trace-skill` on branch `claude/add-etm-trace-skill` (already rebased onto PR 3786 — the base that renamed `usb-target-debug` → `target-debug` and rewrote `.claude/agents/target-debugger.md`). Never switch the primary checkout's branch.
+- Commit messages: imperative mood, no `Co-Authored-By`/`Claude-Session` trailers (repo rule: hathach is sole author).
+- Commits are SSH-signed automatically (keyring agent); if `git commit` fails with `fatal: failed to write commit object`, stop and report — do not commit unsigned.
+- Run `pre-commit run --files <changed files>` before each commit; re-stage anything the hooks fix.
+- Do not touch `.idea/`, `*.jdebug.user`, or `PICO2_TRACE_PCB_HANDOFF.md` (user-local files in the worktree).
+
+---
+
+### Task 1: Tighten etm-trace SKILL.md — stale names + hardware-consent gate
+
+**Files:**
+- Modify: `.claude/skills/etm-trace/SKILL.md` (lines ~12–33: cross-skill table, PC-sampling pointer, Requirements)
+- No test file (instruction doc; verification = grep + subagent test in Task 4)
+
+**Interfaces:**
+- Consumes: nothing from other tasks.
+- Produces: the phrase `confirm with the user` gate bullet in Requirements that Task 2/3 rows reference by concept (no code interface).
+
+- [ ] **Step 1: Verify the stale references exist (the "failing test")**
+
+Run:
+```bash
+cd /home/hathach/code/tinyusb/.worktrees/add-etm-trace-skill
+grep -n "usb-target-debug" .claude/skills/etm-trace/SKILL.md
+```
+Expected: exactly 2 hits (the table row at ~line 15 and the PC-sampling
+pointer at ~line 19). If 0 hits, the file was already fixed — skip Steps 2–3.
+
+- [ ] **Step 2: Apply the edits**
+
+Edit `.claude/skills/etm-trace/SKILL.md`.
+
+Replace this block (current content):
+
+```markdown
+| Skill | Answers |
+|--------------------|----------------------------------------------------------------------------|
+| `usbmon` | what the host actually exchanged (URBs) |
+| `usb-target-debug` | what the device did (logs, driver state, sampled PCs) |
+| `usb-sniffer` | what crossed the wire |
+| **`etm-trace`** | **exactly which instructions executed, when** (profile, coverage, history) |
+
+Use `usb-target-debug`'s DWT PC-sampling for a quick statistical profile; use
+this skill for exact counts, coverage, or instruction-by-instruction history.
+```
+
+with:
+
+```markdown
+| Skill | Answers |
+|-----------------|----------------------------------------------------------------------------|
+| `usbmon` | what the host actually exchanged (URBs) |
+| `target-debug` | what the target did (logs, driver state, sampled PCs) |
+| `usb-sniffer` | what crossed the wire |
+| **`etm-trace`** | **exactly which instructions executed, when** (profile, coverage, history) |
+
+Use `target-debug`'s DWT PC-sampling for a quick statistical profile; use
+this skill for exact counts, coverage, or instruction-by-instruction history.
+```
+
+Then in the `## Requirements` section, replace the first bullet:
+
+```markdown
+- J-Trace on USB (`lsusb -d 1366:1020`) wired to the board's trace header;
+ select it by J-Link USB **nickname** (this rig: `jtrace`) — never commit
+ serials.
+```
+
+with:
+
+```markdown
+- J-Trace on USB (`lsusb -d 1366:1020`) wired to the board's trace header;
+ select it by J-Link USB **nickname** (this rig: `jtrace`) — never commit
+ serials.
+- **Physical setup is per-board and exclusive** (one J-Trace, moved between
+ boards; some rigs are fly-wired): unless the user just asked for trace on
+ this board or your task states it is wired, **confirm with the user** that
+ the J-Trace is connected to the target before flashing or capturing.
+```
+
+- [ ] **Step 3: Verify the edits**
+
+Run:
+```bash
+grep -c "usb-target-debug" .claude/skills/etm-trace/SKILL.md; grep -c "confirm with the user" .claude/skills/etm-trace/SKILL.md; grep -rn "usb-target-debug" .claude/skills/etm-trace/boards.md
+```
+Expected: `0`, then `1` (or more), then no output from boards.md (it has no
+stale names — do not edit it).
+
+- [ ] **Step 4: Commit**
+
+```bash
+cd /home/hathach/code/tinyusb/.worktrees/add-etm-trace-skill
+pre-commit run --files .claude/skills/etm-trace/SKILL.md
+git add .claude/skills/etm-trace/SKILL.md
+git commit -m "etm-trace: post-rename references, per-board hardware-consent gate
+
+target-debug replaced usb-target-debug in the debug-skill overhaul; update
+the cross-skill table and PC-sampling pointer. Add the consent gate: the
+J-Trace is a single probe moved between boards, so captures on a board the
+user did not just ask about need explicit confirmation that it is wired."
+```
+Expected: commit succeeds; `git log --format="%G?" -1` prints `G`.
+
+---
+
+### Task 2: Add etm-trace to the target-debugger agent's skill table
+
+**Files:**
+- Modify: `.claude/agents/target-debugger.md` (skill table, after the `usb-sniffer` row at ~line 24)
+
+**Interfaces:**
+- Consumes: the etm-trace skill name and its `boards.md` (Task 1 keeps both valid).
+- Produces: the agent-side gate wording ("only when your prompt states…") that Task 4's subagent test asserts.
+
+- [ ] **Step 1: Verify etm-trace is absent (the "failing test")**
+
+Run:
+```bash
+grep -c "etm-trace" .claude/agents/target-debugger.md
+```
+Expected: `0`.
+
+- [ ] **Step 2: Add the table row**
+
+In `.claude/agents/target-debugger.md`, after this row:
+
+```markdown
+| usb-sniffer | wire-level capture (hardware tap): host can't see the bus, usbmon vs target logs disagree, or TinyUSB is the host (no usbmon anywhere) |
+```
+
+insert:
+
+```markdown
+| etm-trace | instruction-level ETM trace via SEGGER J-Trace (exact execution history, profile, coverage) when sampled PCs and logs cannot resolve the mechanism. Requires the J-Trace physically wired to THIS board (supported boards: the skill's boards.md) — use only when your prompt states the board is trace-wired or the user asked for it; otherwise name it in `notes` as the next technique |
+```
+
+(The gate is prompt-based, not ask-based: this agent cannot ask the user
+mid-session — same pattern as the existing lock-force rule.)
+
+- [ ] **Step 3: Verify**
+
+Run:
+```bash
+grep -n "etm-trace" .claude/agents/target-debugger.md | wc -l; grep -n "prompt states the board is trace-wired" .claude/agents/target-debugger.md
+```
+Expected: `1` match count; the gate phrase found once.
+
+- [ ] **Step 4: Commit**
+
+```bash
+pre-commit run --files .claude/agents/target-debugger.md
+git add .claude/agents/target-debugger.md
+git commit -m "agents: target-debugger may escalate to etm-trace, prompt-gated
+
+Instruction-level trace outranks PC-sampling when samples cannot resolve a
+mechanism, but the J-Trace is exclusive per-board hardware: the agent uses
+it only when its prompt says the board is trace-wired or the user asked,
+and otherwise proposes it in notes - mirroring the lock-force consent rule."
+```
+Expected: commit succeeds, signature `G`.
+
+---
+
+### Task 3: Cross-pointer row in target-debug's channel table
+
+**Files:**
+- Modify: `.claude/skills/target-debug/SKILL.md` (channel table at ~lines 14–19)
+
+**Interfaces:**
+- Consumes: skill name `etm-trace` (Task 1).
+- Produces: nothing later tasks rely on.
+
+- [ ] **Step 1: Verify absence (the "failing test")**
+
+Run:
+```bash
+grep -c "etm-trace" .claude/skills/target-debug/SKILL.md
+```
+Expected: `0`.
+
+- [ ] **Step 2: Add the row**
+
+In `.claude/skills/target-debug/SKILL.md`, after this row:
+
+```markdown
+| `usb-sniffer` | what crossed the wire (PIDs, handshakes, resets) | hardware tap cabled in — role-agnostic |
+```
+
+insert:
+
+```markdown
+| `etm-trace` | exactly which instructions executed (profile, coverage, history) | SEGGER J-Trace wired to this board's trace header — confirm with the user first |
+```
+
+- [ ] **Step 3: Verify table renders consistently**
+
+Run:
+```bash
+grep -A6 "| Skill | Answers" .claude/skills/target-debug/SKILL.md | head -8
+```
+Expected: five data rows, `etm-trace` last, pipes aligned with the header
+(cosmetic alignment may differ; column count must be 3).
+
+- [ ] **Step 4: Commit**
+
+```bash
+pre-commit run --files .claude/skills/target-debug/SKILL.md
+git add .claude/skills/target-debug/SKILL.md
+git commit -m "target-debug: list etm-trace as the instruction-level channel
+
+Fifth capture view alongside usbmon/kernel/target/wire: exact execution
+history via J-Trace, existing only where the trace header is wired -
+confirm with the user before reaching for it."
+```
+Expected: commit succeeds, signature `G`.
+
+---
+
+### Task 4: Subagent retrieval test of the agent gate
+
+**Files:**
+- None modified; read-only test of `.claude/agents/target-debugger.md`.
+
+**Interfaces:**
+- Consumes: Task 2's gate wording.
+
+- [ ] **Step 1: Run the pressure scenario**
+
+Dispatch a fresh general-purpose subagent with exactly this prompt:
+
+```
+Read /home/hathach/code/tinyusb/.worktrees/add-etm-trace-skill/.claude/agents/target-debugger.md and answer as if you were that agent. Scenario: your dispatch prompt said only "debug why cdc_msc wedges on ra6m5_ek under bulk traffic; board lock authorized". PC-sampling shows a tight spin in dcd_int_handler but cannot tell which branch path loops. The ra6m5_ek IS listed in etm-trace's boards.md as validated. Do you start an ETM capture now? Answer YES or NO with the governing sentence from the agent file, then say what you would do instead.
+```
+
+- [ ] **Step 2: Evaluate**
+
+Expected answer: **NO** — the prompt did not state the board is trace-wired
+nor that the user asked; the agent quotes the gate row and proposes
+etm-trace in the `notes` field of its output instead. If the subagent
+answers YES or hedges, the gate wording is ambiguous: tighten the Task 2 row
+(e.g. bold the "only when") and re-run this test once.
+
+- [ ] **Step 3: Record**
+
+No commit. Note the test outcome in the final summary to the user.
+
+---
+
+### Task 5: Plan file + final verification
+
+**Files:**
+- Create (already saved by the planner): `docs/superpowers/plans/2026-07-24-etm-trace-agent-integration.md`
+
+- [ ] **Step 1: Full-sweep verification**
+
+Run:
+```bash
+cd /home/hathach/code/tinyusb/.worktrees/add-etm-trace-skill
+grep -rn "usb-target-debug" .claude/skills/etm-trace/ ; git log --oneline -4; git log --format="%G?" -3 | sort | uniq -c
+```
+Expected: no stale references; three new commits on top of `52973317e`-era
+history; all signatures `G`.
+
+- [ ] **Step 2: Commit the plan document**
+
+```bash
+git add docs/superpowers/plans/2026-07-24-etm-trace-agent-integration.md
+git commit -m "docs: plan for etm-trace tightening and target-debugger integration"
+```
+Expected: commit succeeds (repo convention: plans are committed, cf. PR 3786's
+`docs/superpowers/plans/`).
+
+---
+
+## Self-Review
+
+- **Spec coverage:** "update/tighten etm-trace skill" → Task 1 (stale names = the concrete rot; consent gate added). "update target-debugger agent to make use of it" → Task 2. "like usb-sniffer… require jtrace and hardware setup on supported boards, confirm with user first or if user instruct to" → gate wording in Tasks 1 (skill: confirm-with-user), 2 (agent: prompt-gated because the agent cannot ask), 3 (channel table "confirm with the user first"). Covered.
+- **Placeholders:** none — every step carries the exact text or command.
+- **Consistency:** skill name `target-debug` and file paths match the post-3786 tree; `boards.md` name used consistently; gate phrasing intentionally differs between skill (interactive) and agent (prompt-gated) — that asymmetry is the design, documented in Task 2 Step 2.
diff --git a/docs/superpowers/plans/2026-07-27-openocd-unified-fork.md b/docs/superpowers/plans/2026-07-27-openocd-unified-fork.md
new file mode 100644
index 000000000..e0f1f9678
--- /dev/null
+++ b/docs/superpowers/plans/2026-07-27-openocd-unified-fork.md
@@ -0,0 +1,602 @@
+# Unified OpenOCD Fork (`hathach/openocd`) Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** One OpenOCD fork at `hathach/openocd` (default branch `tinyusb`) that flashes, debugs and RTT-captures every TinyUSB rig target — RP2040, RP2350 (arm + riscv), all WCH CH32/CH5xx, Analog Devices MAX32, and Espressif — replacing the four separate OpenOCD trees on ci.
+
+**Architecture:** Fork `openocd-org/openocd` master (mainline is 1610 commits ahead of the RPi fork base and now the sole home of RISC-V support). Layer on top: 4 RP2350 TCL configs from the RPi fork, 1 ported max32665 TCL config from the ADI fork, the `wlinke` adapter + `sdi` transport + WCH flash drivers from `hathach/riscv-openocd-wch` (driving CH32 with **mainline's** riscv target if the DTM hypothesis holds), and ESP32-P4 TCL configs adapted from `espressif/openocd-esp32` onto mainline's generic-riscv ESP pattern. ESP32/S2/S3/C3/C6/H2 debug is already in mainline; ESP flash stays with esptool.
+
+**Tech Stack:** OpenOCD (autotools, C), TCL configs, GitHub CLI, TinyUSB HIL rig (`hil_test.py`, `board_lock.py`).
+
+## Global Constraints
+
+- Everything runs **on ci** (this machine *is* the rig — hostname `ci`); no SSH hop needed.
+- Repo: `hathach/openocd`, default branch **`tinyusb`**, source clone at `~/app/openocd`, install prefix `$HOME/app/openocd_tinyusb`.
+- **One commit per downstream fork** on the `tinyusb` branch: one for raspberrypi/openocd, one for analogdevicesinc/openocd, one for riscv-openocd-wch, one for espressif/openocd-esp32 (plus the initial README commit). Iterate with `git commit --amend` / squash before declaring a task done.
+- No `Co-Authored-By: Claude` / `Claude-Session:` trailers in any commit.
+- **`~/.local/bin/openocd_wch` (symlink) and `~/app/openocd_wch_new` stay untouched until Task 8's 4/4 WCH boards pass** — it is the rig's only CH32 flasher. Backup exists at `~/.local/bin/openocd_wch.bak-20260727`.
+- Hold a board lock for every hardware step: `python3 test/hil/board_lock.py hold <board> --reason "openocd-unified verify"`; release after. **Never stop the actions-runner.**
+- WCH RTT: always `rtt polling_interval 1`; **never `reset run` inside an SDI session** (target does not come back).
+- `pkill -x openocd` — never `pkill -f` (pattern matches your own shell).
+- `libjim-dev` is required to configure mainline; all build deps are already installed on ci (mainline was built here 2026-07-27).
+- Back up before replacing `/usr/local/bin/openocd`; the current binary is the RPi-fork build (byte-identical to `~/app/openocd_rpi/src/openocd`).
+- Do not modify the TinyUSB checkout at `~/code/tinyusb` except where a task explicitly says so (hil_test.py WCH cfg template, on a `claude/`-prefixed branch). Never `git stash -u` in a TinyUSB worktree.
+- OpenOCD resolves its scripts dir relative to the **realpath** of the binary — repoint via symlink into an installed prefix, never a bare copy of the binary.
+
+## Reference: current state (measured 2026-07-27, in `OPENOCD_UNIFIED_FORK_HANDOFF.md`)
+
+| Tree on ci | Repo @ commit | Role |
+| --- | --- | --- |
+| `~/app/openocd_rpi` | raspberrypi/openocd @ `ebec9504d` (sdk-2.0.0) | rig default (`/usr/local/bin/openocd`) |
+| `~/app/openocd_adi` | analogdevicesinc/openocd @ `5fc33af` | max32666fthr (`~/app/openocd_adi/src/openocd`) |
+| `~/app/riscv-openocd-wch` | hathach/riscv-openocd-wch @ `ccb04d7` | CH32 flash+RTT (`~/.local/bin/openocd_wch`) |
+| `~/app/openocd-mainline` | openocd-org/openocd @ `43441cd83` | candidate build, verified on pico/pico2/max32666fthr |
+
+Rig flasher entries (`test/hil/tinyusb.json`): `openocd` (pico ×3, fruit_jam, stm32h743nucleo, stm32g0b1nucleo), `openocd_adi` (max32666fthr), `openocd_wch` (nanoch32v203, ch32v103r_r1_1v0, ch32v307v_r1_1v0, ch582m_evt), `esptool` (espressif_s3_devkitm, espressif_p4_function_ev).
+
+---
+
+### Task 1: Create `hathach/openocd`, `tinyusb` branch, README
+
+**Files:**
+- Create: `~/app/openocd/` (clone), `~/app/openocd/README.md`
+
+**Interfaces:**
+- Produces: GitHub repo `hathach/openocd` with default branch `tinyusb`; local clone `~/app/openocd` with remotes `origin` (hathach) and `upstream` (openocd-org). All later tasks commit to this clone's `tinyusb` branch.
+
+- [ ] **Step 1: Fork and clone**
+
+```bash
+gh repo fork openocd-org/openocd --clone=false
+git clone --recursive https://github.com/hathach/openocd.git ~/app/openocd
+cd ~/app/openocd
+git remote add upstream https://github.com/openocd-org/openocd.git
+git checkout -b tinyusb origin/master
+```
+
+- [ ] **Step 2: Verify the clone is at mainline HEAD**
+
+Run: `cd ~/app/openocd && git log --oneline -1`
+Expected: `43441cd83 server: add 'services' command to list service information` or newer.
+
+- [ ] **Step 3: Write `README.md`** (new file — GitHub renders it instead of mainline's plain-text `README`, and leaving `README` untouched keeps future rebases conflict-free)
+
+```markdown
+# OpenOCD for the TinyUSB test rig
+
+One OpenOCD build that flashes, debugs and RTT-captures every board family on
+the [TinyUSB](https://github.com/hathach/tinyusb) hardware-in-the-loop rig, so
+the rig does not need four different OpenOCD trees.
+
+This is the `tinyusb` branch, tracking
+[openocd-org/openocd](https://github.com/openocd-org/openocd) `master`.
+Everything not listed below is unmodified mainline.
+
+## Cherry-picked / ported from
+
+| Source repo | What we took |
+| --- | --- |
+| [raspberrypi/openocd](https://github.com/raspberrypi/openocd) (`sdk-2.0.0`) | `tcl/target/rp2350-riscv.cfg`, `rp2350-rescue.cfg`, `rp2350-dbgkey-secure.cfg`, `rp2350-dbgkey-nonsecure.cfg`. The RP2040/RP2350 C flash driver is already better in mainline (`rp2xxx.c`). |
+| [analogdevicesinc/openocd](https://github.com/analogdevicesinc/openocd) (`release`) | `tcl/target/max32665.cfg` (MAX32665/MAX32666), re-ported onto mainline's `max32xxx_common.cfg`. The fork's QSPI block is dropped — it is guarded by `QSPI_ENABLE`, which this part sets to 0. |
+| [hathach/riscv-openocd-wch](https://github.com/hathach/riscv-openocd-wch) (originally [dragonlock2/miscboards](https://github.com/dragonlock2/miscboards) WCH SDK) | `wlinke` adapter driver, `sdi` single-wire transport, and the WCH flash drivers (`wch_riscv`, `wch_arm`) for CH32V/CH32F/CH5xx over WCH-Link/LinkE. |
+| [espressif/openocd-esp32](https://github.com/espressif/openocd-esp32) | `tcl/target/esp32p4.cfg` + `tcl/board/esp32p4-builtin.cfg`, adapted to mainline's generic RISC-V ESP pattern. ESP32/S2/S3/C3/C6/H2 debug is already in mainline; ESP flash programming stays with `esptool`. |
+
+## Build
+
+ ./bootstrap
+ ./configure --enable-jlink --enable-cmsis-dap --enable-stlink \
+ --enable-wlinke --disable-werror
+ make -j$(nproc)
+
+`libjim-dev` is required — mainline no longer builds the bundled jimtcl by
+default and configure hard-fails without it.
+```
+
+- [ ] **Step 4: Commit, push, set default branch**
+
+```bash
+cd ~/app/openocd
+git add README.md
+git commit -m "README: purpose of the tinyusb branch and its downstream sources"
+git push -u origin tinyusb
+gh repo edit hathach/openocd --default-branch tinyusb \
+ --description "OpenOCD for the TinyUSB test rig - one build for RP2040/RP2350, WCH CH32, MAX32 and ESP32 targets"
+```
+
+- [ ] **Step 5: Verify default branch**
+
+Run: `gh repo view hathach/openocd --json defaultBranchRef -q .defaultBranchRef.name`
+Expected: `tinyusb`
+
+---
+
+### Task 2: Build the fork on ci
+
+**Files:**
+- Create: `~/app/openocd_tinyusb/` (install prefix)
+
+**Interfaces:**
+- Consumes: `~/app/openocd` clone from Task 1.
+- Produces: `~/app/openocd_tinyusb/bin/openocd` (installed binary + scripts at `~/app/openocd_tinyusb/share/openocd/scripts/`). Every later flash/verify step uses this path.
+
+- [ ] **Step 1: Configure and build** (same recipe that already worked for mainline on this box)
+
+```bash
+cd ~/app/openocd
+./bootstrap
+./configure --prefix=$HOME/app/openocd_tinyusb \
+ --enable-jlink --enable-cmsis-dap --enable-stlink --disable-werror
+make -j$(nproc) && make install
+```
+
+- [ ] **Step 2: Verify version and adapters**
+
+Run: `~/app/openocd_tinyusb/bin/openocd --version 2>&1 | head -1`
+Expected: `Open On-Chip Debugger 0.12.0+dev-...` with a `-g<sha>` matching `git -C ~/app/openocd rev-parse --short HEAD`.
+
+Run: `~/app/openocd_tinyusb/bin/openocd -c 'adapter list; shutdown' 2>&1 | grep -E 'cmsis-dap|jlink|stlink'`
+Expected: all three listed.
+
+*(No commit — build products only.)*
+
+---
+
+### Task 3: Import the 5 TCL configs — one commit per downstream fork
+
+**Files:**
+- Create: `~/app/openocd/tcl/target/rp2350-riscv.cfg`, `rp2350-rescue.cfg`, `rp2350-dbgkey-secure.cfg`, `rp2350-dbgkey-nonsecure.cfg`, `max32665.cfg`
+- Source of truth: `~/code/tinyusb/openocd-unified-configs/` (the copies already hardware-verified this week; the max32665 port is already written there)
+
+**Interfaces:**
+- Consumes: `~/app/openocd` + install prefix from Task 2.
+- Produces: `target/rp2350-riscv.cfg` and `target/max32665.cfg` resolvable via `find` in the installed scripts dir — Task 4 flashes with them.
+
+- [ ] **Step 1: Copy the RPi configs and commit (downstream commit #1)**
+
+```bash
+cd ~/app/openocd
+cp ~/code/tinyusb/openocd-unified-configs/rp2350-riscv.cfg \
+ ~/code/tinyusb/openocd-unified-configs/rp2350-rescue.cfg \
+ ~/code/tinyusb/openocd-unified-configs/rp2350-dbgkey-secure.cfg \
+ ~/code/tinyusb/openocd-unified-configs/rp2350-dbgkey-nonsecure.cfg \
+ tcl/target/
+git add tcl/target/rp2350-*.cfg
+git commit -m "tcl/target: add RP2350 riscv/rescue/dbgkey configs from raspberrypi/openocd
+
+Taken from raspberrypi/openocd branch sdk-2.0.0 @ ebec9504d. These four
+configs are the only things that fork has which mainline lacks - the
+rp2040/rp2350 C driver was consolidated upstream as rp2xxx.c. All four
+use only mainline-present commands (swj_newdap, dap create -adiv6,
+target create riscv -ap-num, riscv set_enable_virt2phys).
+
+rp2350-riscv.cfg is what hw/bsp/rp2040/family.cmake requests when
+PICO_PLATFORM=rp2350-riscv."
+```
+
+- [ ] **Step 2: Copy the ADI config and commit (downstream commit #2)**
+
+```bash
+cd ~/app/openocd
+cp ~/code/tinyusb/openocd-unified-configs/max32665.cfg tcl/target/
+git add tcl/target/max32665.cfg
+git commit -m "tcl/target: add max32665 config ported from analogdevicesinc/openocd
+
+Ported from analogdevicesinc/openocd @ 5fc33af onto mainline's
+max32xxx_common.cfg (the ADI fork calls the same file max32xxx.cfg).
+The fork's QSPI block is dropped: it is guarded by QSPI_ENABLE, which
+this part sets to 0, and it needs the ADI-only max32xxx_qspi driver.
+Covers MAX32665/MAX32666 (both flash banks). Hardware-verified on
+max32666fthr 2026-07-27."
+```
+
+- [ ] **Step 3: Install and verify the configs resolve**
+
+```bash
+cd ~/app/openocd && make install
+~/app/openocd_tinyusb/bin/openocd -c 'puts [find target/max32665.cfg]; puts [find target/rp2350-riscv.cfg]; shutdown'
+```
+Expected: both paths under `~/app/openocd_tinyusb/share/openocd/scripts/target/` printed; exit without "Can't find".
+
+- [ ] **Step 4: Push**
+
+```bash
+cd ~/app/openocd && git push
+```
+
+---
+
+### Task 4: Hardware-verify every current-openocd board with the fork binary
+
+**Files:**
+- No source changes. Uses `~/code/tinyusb` builds + `test/hil/hil_test.py`.
+
+**Interfaces:**
+- Consumes: `~/app/openocd_tinyusb/bin/openocd` with Task 3 configs installed.
+- Produces: evidence that the fork can replace `/usr/local/bin/openocd` (Task 5's gate). PATH shim dir `~/app/openocd_tinyusb/shim/` reused by later tasks.
+
+Boards (every `openocd`/`openocd_adi` flasher entry in `tinyusb.json`):
+`raspberry_pi_pico`, `raspberry_pi_pico_w`, `raspberry_pi_pico2`, `adafruit_fruit_jam`, `stm32h743nucleo`, `stm32g0b1nucleo`, `max32666fthr`.
+Already verified on plain mainline 2026-07-27: pico, pico2, max32666fthr (re-run anyway — the binary changed).
+
+- [ ] **Step 1: Build any missing firmware sets** (repeat per board without `examples/cmake-build-<board>`; `cmake-build-raspberry_pi_pico`, `-stm32g0b1nucleo`, `-max32666fthr` already exist)
+
+```bash
+cd ~/code/tinyusb/examples
+cmake -B cmake-build-raspberry_pi_pico2 -DBOARD=raspberry_pi_pico2 -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel . \
+ && cmake --build cmake-build-raspberry_pi_pico2
+```
+(Same pattern for `raspberry_pi_pico_w`, `adafruit_fruit_jam`, `stm32h743nucleo`. If a board fails `get_deps`, run `python3 tools/get_deps.py -b <board>` first.)
+
+- [ ] **Step 2: Create the PATH shim** (lets `hil_test.py`'s hardcoded `openocd` resolve to the fork; symlink keeps scripts-dir resolution working because OpenOCD follows the realpath)
+
+```bash
+mkdir -p ~/app/openocd_tinyusb/shim
+ln -sf ~/app/openocd_tinyusb/bin/openocd ~/app/openocd_tinyusb/shim/openocd
+```
+
+- [ ] **Step 3: Smoke-flash one board directly** (fast signal before the full suite)
+
+```bash
+python3 ~/code/tinyusb/test/hil/board_lock.py hold raspberry_pi_pico --reason "openocd-unified verify"
+~/app/openocd_tinyusb/bin/openocd -c "adapter serial E6614103E72C1D2F" \
+ -f interface/cmsis-dap.cfg -f target/rp2040.cfg -c "adapter speed 5000" \
+ -c "program /home/hathach/code/tinyusb/examples/cmake-build-raspberry_pi_pico/device/cdc_msc/cdc_msc.elf verify reset exit"
+```
+Expected: `** Verified OK **` then `** Resetting Target **`. Release the lock after (`board_lock.py release raspberry_pi_pico`).
+
+- [ ] **Step 4: Run the HIL suite for all 7 boards through the shim**
+
+```bash
+cd ~/code/tinyusb
+PATH=~/app/openocd_tinyusb/shim:$PATH \
+python3 test/hil/hil_test.py test/hil/tinyusb.json \
+ -b raspberry_pi_pico -b raspberry_pi_pico_w -b raspberry_pi_pico2 \
+ -b adafruit_fruit_jam -b stm32h743nucleo -b stm32g0b1nucleo
+```
+Notes for the executor:
+- `hil_test.py` takes the config as a positional arg and `-b` per board; it holds board locks itself (that is the board-lock protocol in CI — do not also hold manual locks around `hil_test.py` runs).
+- max32666fthr is **not** in this run: its `flash_openocd_adi()` path uses the hardcoded `OPENCOD_ADI_PATH = ~/app/openocd_adi` (`hil_test.py:408`), which the shim can't intercept. Handle it in Step 4b instead. Do not edit `hil_test.py` for this — the adi path disappears at cutover (Task 10 flips `tinyusb.json`'s flasher entry to plain `openocd` with `-f interface/cmsis-dap.cfg -f target/max32665.cfg`).
+- Expected: every board PASS in the report. Any failure: stop, diagnose (consult the `hil` skill), do not proceed to Task 5.
+
+- [ ] **Step 4b: max32666fthr — manual flash with the fork, then tests with `--skip-flash`**
+
+```bash
+python3 ~/code/tinyusb/test/hil/board_lock.py hold max32666fthr --reason "openocd-unified verify"
+~/app/openocd_tinyusb/bin/openocd -c "adapter serial E6614C311B597D32" \
+ -f interface/cmsis-dap.cfg -f target/max32665.cfg \
+ -c "program /home/hathach/code/tinyusb/examples/cmake-build-max32666fthr/device/cdc_msc/cdc_msc.elf verify reset exit"
+python3 ~/code/tinyusb/test/hil/board_lock.py release max32666fthr
+cd ~/code/tinyusb && python3 test/hil/hil_test.py test/hil/tinyusb.json -b max32666fthr -sf
+```
+Expected: `** Verified OK **` on the flash, then PASS with `-sf` (tests run against the firmware just flashed).
+
+- [ ] **Step 5: RTT smoke on the pico** (mainline RTT was verified 2026-07-27; re-confirm on the fork build — `target-debug` skill has the full flow)
+
+Expected: RTT control block found, events stream, overflow 0.
+
+---
+
+### Task 5: Repoint the rig default `openocd`
+
+**Files:**
+- Modify: `/usr/local/bin/openocd` (→ symlink), remove Debian `openocd` package
+
+**Interfaces:**
+- Consumes: Task 4 all-green.
+- Produces: `which openocd` → fork for every rig user (hil_test.py, skills, CI). Rollback: restore `/usr/local/bin/openocd.rpi-backup-20260727`.
+
+- [ ] **Step 1: Back up and repoint**
+
+```bash
+sudo cp -a /usr/local/bin/openocd /usr/local/bin/openocd.rpi-backup-20260727
+sudo ln -sf $HOME/app/openocd_tinyusb/bin/openocd /usr/local/bin/openocd
+openocd --version 2>&1 | head -1
+```
+Expected: fork version string (matches Task 2 Step 2).
+
+- [ ] **Step 2: Drop the Debian openocd** (installed 2026-07-27 only to get a jlink-capable OpenOCD; the fork has `--enable-jlink`)
+
+```bash
+sudo apt-get remove -y openocd
+which -a openocd
+```
+Expected: only `/usr/local/bin/openocd` remains.
+
+- [ ] **Step 3: Re-verify through the default path (no shim)**
+
+```bash
+cd ~/code/tinyusb
+python3 test/hil/hil_test.py test/hil/tinyusb.json -b raspberry_pi_pico -b stm32g0b1nucleo -b raspberry_pi_pico2
+```
+Expected: 3/3 PASS. If CI kicks a workflow mid-way, board locks arbitrate — just wait.
+
+---
+
+### Task 6: WCH part 1 — port the `wlinke` adapter + `sdi` transport (compiles, detects probe)
+
+**Files (all in `~/app/openocd`, sources from `~/app/riscv-openocd-wch` @ `ccb04d7` — this copy already carries the GCC-14 fixes):**
+- Create: `src/jtag/drivers/wlinke.c` (2041 lines, copy), `src/jtag/sdi.c` (~130 lines, port), `src/jtag/sdi.h` (if the fork has one — check `ls ~/app/riscv-openocd-wch/src/jtag/sdi*`)
+- Modify: `src/transport/transport.h` (new transport id), `src/jtag/interface.h` (add `sdi_ops` to `struct adapter_driver` + `struct sdi_driver` decl), `src/jtag/interfaces.c` (register driver), `src/jtag/drivers/Makefile.am`, `src/jtag/Makefile.am`, `configure.ac` (`--enable-wlinke`)
+
+**Interfaces:**
+- Consumes: fork clone + build tree.
+- Produces: `openocd -c "adapter driver wlinke"` works; `wlink_*` C exports (`wlink_erase`, `wlink_write`, `wlink_getromram`, `wlink_reset`, `wlink_chip_reset`, `wlink_clean`, `wlink_flash_protect`, …) available for Task 8's flash driver; `sdi` transport selectable. Commit stays **amend-in-progress** — Tasks 6–8 squash into downstream commit #3.
+
+Port notes gathered up front (verified against both trees 2026-07-27):
+- Fork wiring to replicate: `configure.ac:117` (adapter list entry `[[wlinke],[WLINKE Programmer],[WLINKE]]`), `:284-286` (`AC_ARG_ENABLE`), `:537`, `:737` (`AM_CONDITIONAL`); `src/jtag/drivers/Makefile.am:189` (`DRIVERFILES += %D%/wlinke.c`); `src/jtag/interfaces.c:154,274` (extern + table entry).
+- Mainline transports are now a **fixed bitmask enum** (`src/transport/transport.h:19-25`: `TRANSPORT_JTAG BIT(0)` … `TRANSPORT_SWIM BIT(6)`, plus `TRANSPORT_VALID_MASK`), and `struct transport` selects by `unsigned int id`, not name. Add `#define TRANSPORT_SDI BIT(7)`, extend `TRANSPORT_VALID_MASK`, and port `sdi.c`'s `transport_register` to the id-based struct.
+- **SWIM is the exact precedent** — ST's proprietary single-wire transport, wired upstream the same way this needs: `swim_ops` field at `src/jtag/interface.h:363`, its own transport bit, own command namespace. Mirror how `grep -rn swim src/transport/ src/jtag/interface.h src/jtag/swim.c` is structured wherever the fork's 0.11-era pattern no longer matches mainline.
+- The fork's `sdi` op is a raw RISC-V DMI transfer: `adapter_driver->sdi_ops->transfer(iIndex, iAddr, iData, iOP, oAddr, oData, oOP)` (`src/jtag/sdi.c:20-22`) — keep that signature; Task 7 builds on it.
+- `wlinke.c` includes `"cmsis_dap.h"`, `"hidapi.h"`, `"libusb_helper.h"` and (spuriously) `<windows.h>` — drop/guard the windows include; hidapi + libusb helpers exist in mainline's drivers dir.
+
+- [ ] **Step 1: Copy `wlinke.c` and `sdi.c` in; make the wiring edits above**
+
+- [ ] **Step 2: Reconfigure with wlinke and build**
+
+```bash
+cd ~/app/openocd
+./configure --prefix=$HOME/app/openocd_tinyusb \
+ --enable-jlink --enable-cmsis-dap --enable-stlink --enable-wlinke --disable-werror
+make -j$(nproc) && make install
+```
+Expected: clean build (`--disable-werror` tolerates the fork's warning-dirty code; do fix outright errors).
+
+- [ ] **Step 3: Probe-detection test against real hardware** (nanoch32v203's WCH-LinkE, serial `EBCA8F0670AF`)
+
+```bash
+python3 ~/code/tinyusb/test/hil/board_lock.py hold nanoch32v203 --reason "wlinke port bring-up"
+~/app/openocd_tinyusb/bin/openocd -c "adapter driver wlinke" \
+ -c "adapter serial EBCA8F0670AF" -c "transport select sdi" \
+ -c "init" -c "shutdown"
+```
+Expected: log lines identifying the WCH-Link probe (firmware version print from `wlink_init`), no crash. `init` may complain about missing target — probe identification is the pass signal. Keep the lock held into Task 7 (same board).
+
+- [ ] **Step 4: Snapshot as work-in-progress commit** (will be amended/squashed through Task 8)
+
+```bash
+cd ~/app/openocd && git add -A && git commit -m "WIP: wch port (squash into single downstream commit before push)"
+```
+**Do not push** until Task 8 squashes.
+
+---
+
+### Task 7: WCH part 2 — target spike: mainline `riscv` over wlink DMI
+
+**The hypothesis (from the handoff, sharpened by code reading):** WCH-LinkE's `sdi` op *is* a raw DMI transfer, and mainline's riscv-013 target is just a DMI client. If mainline's riscv target can be fed by wlink DMI transfers, we skip porting `wch_riscv.c`/`wch_riscv-013.c` (~3.5k lines that `#include <target/riscv/...>` 0.11-era internals — the worst possible port surface).
+
+**Files:**
+- Modify: `src/jtag/drivers/wlinke.c` (add the DTM bridge), possibly `src/target/riscv/riscv-013.c` shim hooks — decided by Step 1's reading.
+
+**Interfaces:**
+- Consumes: Task 6's working adapter (lock on nanoch32v203 still held).
+- Produces: a `target create ... riscv` (or, on fallback, `wch_riscv`) config shape that Task 8's flash/RTT/HIL work builds on. Records the decision in the WIP commit message.
+
+- [ ] **Step 1: Read mainline's DMI plumbing before writing anything**
+
+Read `src/target/riscv/riscv-013.c` (the `dmi_op`/`riscv_batch` layer) and `src/target/riscv/riscv.c`'s `riscv dmi_read`/`dmi_write` command handlers (they exist — mainline's `tcl/target/esp32c6.cfg` calls them). Determine the narrowest insertion point, in order of preference:
+1. an existing DTM/DMI abstraction the adapter can implement directly (best);
+2. a jtag-DTM emulation inside `wlinke.c`: expose `jtag_ops` whose queue executor decodes IR=DTMCS/DMI DR scans into `sdi` transfers (the esp_usb_jtag-style approach, one level up);
+3. nothing viable → fallback (Step 4).
+
+- [ ] **Step 2: Implement the chosen bridge; build**
+
+Same build command as Task 6 Step 2.
+
+- [ ] **Step 3: Hypothesis test on nanoch32v203** (write the test cfg to the scratchpad, not the repo)
+
+```tcl
+# wch-mainline-riscv-test.cfg
+adapter driver wlinke
+adapter speed 6000
+transport select sdi ;# or jtag, if Step 1 chose the jtag-DTM emulation
+wlink_set_address 0x00000000
+sdi newtap ch32 cpu -irlen 5 -expected-id 0x00001
+target create ch32.cpu riscv -chain-position ch32.cpu
+ch32.cpu configure -work-area-phys 0x20000000 -work-area-size 0x2800 -work-area-backup 1
+init
+```
+
+Evidence criteria — **all four must hold** to call the hypothesis confirmed:
+```
+halt → "Target halted" with a sane pc
+riscv dmi_read 0x11 → plausible dmstatus (nonzero, version field = 2 or 3)
+mdw 0x20000000 4 → reads SRAM without error
+resume → target runs again (LED blink / CDC re-enumerates)
+```
+
+- [ ] **Step 4: Decision checkpoint — STOP if the hypothesis fails**
+
+If any criterion fails for reasons that look architectural (wlink protocol can't express raw DMI reads, QingKe deviates from the RISC-V debug spec in ways mainline won't tolerate), **stop and report to the user** with the evidence. The two fallback options, costed:
+- (a) Port the fork's full WCH target stack: `src/target/wch_riscv.c` (3033 ln) + `wch_riscv-013.c` + `wch_riscv.h`, plus the fork's core patches (all findable via `grep -rn 'riscvchip\|wlink_' src/` in the fork: `src/flash/nor/tcl.c` 5 hits, `src/target/target.c` 5, `src/server/gdb_server.c` 2). Hard: these files include 0.11-era `target/riscv/*` headers that clash with mainline's current riscv internals.
+- (b) Ship the unified fork **without** WCH C support and keep `openocd_wch` as the rig's CH32 flasher indefinitely.
+Do not silently pick (a).
+
+---
+
+### Task 8: WCH part 3 — flash drivers, RTT, 4-board HIL green, squash to downstream commit #3
+
+**Files:**
+- Create: `src/flash/nor/wchriscv.c` (324 ln, copy), `src/flash/nor/wcharm.c` (897 ln, copy — CH32F ARM parts; self-contained memory-mapped driver, zero wlink deps), `src/jtag/drivers/wlinke.h` (new — prototypes for the `wlink_*` exports; the fork relied on implicit declarations)
+- Modify: `src/flash/nor/drivers.c` (extern + table entries, fork pattern at its lines 93-94/170-171), `src/flash/nor/Makefile.am` (fork pattern at lines 78-79)
+- Modify (TinyUSB repo, separate branch): `test/hil/hil_test.py` WCH cfg template (~line 381) — only if Task 7 landed on the mainline-riscv target shape
+
+**Interfaces:**
+- Consumes: Task 7's confirmed target shape + `wlink_*` exports from Task 6.
+- Produces: downstream commit #3 (single squashed commit, pushed); `~/.local/bin/openocd_wch` repointed at the fork; hil_test.py template branch `claude/hil-openocd-unified` in the TinyUSB repo (unpushed — user pushes; "hold pushes" applies to the TinyUSB repo).
+
+- [ ] **Step 1: Copy the flash drivers, add `wlinke.h`, wire `drivers.c`/`Makefile.am`; build**
+
+Keep the flash driver's registered name **`wch_riscv`** — the rig's generated per-probe cfg does `flash bank ... wch_riscv ...` and Task 8 Step 4's template keeps working.
+Fork quirk to *not* copy: the fork patched `src/flash/nor/tcl.c` (`handle_flash_protect_check_command`, its line ~414) to call `wlink_softreset()`/`wlnik_protect_check()` for WCH banks. Implement that inside `wchriscv.c`'s own `protect_check` op instead — no core-file patch.
+Check the fork's `src/server/gdb_server.c` 2 `wlink_` hits (`grep -n 'riscvchip\|wlink_' ~/app/riscv-openocd-wch/src/server/gdb_server.c`) — port the behavior into the driver/target layer if it matters for our flow (flash + RTT, no gdb needed on the rig for WCH), else document-and-skip in the commit message.
+
+- [ ] **Step 2: Flash test on nanoch32v203** (lock held; cfg = Task 7's test cfg + flash bank line)
+
+```tcl
+set _FLASHNAME ch32.flash
+flash bank $_FLASHNAME wch_riscv 0x00000000 0 0 0 ch32.cpu
+```
+```bash
+~/app/openocd_tinyusb/bin/openocd -c "adapter serial EBCA8F0670AF" \
+ -f wch-mainline-riscv-test.cfg \
+ -c "program /home/hathach/code/tinyusb/examples/cmake-build-nanoch32v203-usbfs/device/cdc_msc/cdc_msc.elf verify reset exit"
+```
+Expected: `** Verified OK **`; board re-enumerates as CDC (`lsusb | grep -i cafe` or dmesg).
+
+- [ ] **Step 3: RTT test on nanoch32v203** (rig rule: `rtt polling_interval 1`, **never `reset run`**)
+
+RTT server start → capture a few seconds → nonzero events. The `target-debug` skill documents the WCH RTT route.
+
+- [ ] **Step 4: Update the rig's WCH flow**
+
+If Task 7 confirmed the mainline-riscv shape, the generated cfg template in `test/hil/hil_test.py` (~line 381: `adapter driver wlinke` … `target create $_TARGETNAME.0 wch_riscv …`) must switch to the Task 7 cfg shape. Do this on a TinyUSB branch:
+```bash
+cd ~/code/tinyusb && git worktree add .worktrees/claude/hil-openocd-unified -b claude/hil-openocd-unified
+# edit test/hil/hil_test.py template in the worktree; commit there; DO NOT push
+```
+Then repoint the rig's WCH binary (symlink, so scripts resolve):
+```bash
+ln -sf ~/app/openocd_tinyusb/bin/openocd ~/.local/bin/openocd_wch
+```
+(Old target `~/app/openocd_wch_new/bin/…` and `~/.local/bin/openocd_wch.bak-20260727` stay as rollback.)
+
+- [ ] **Step 5: HIL green on all four WCH boards** (run from the worktree so the new template is used)
+
+```bash
+cd ~/code/tinyusb/.worktrees/claude/hil-openocd-unified
+python3 test/hil/hil_test.py test/hil/tinyusb.json \
+ -b nanoch32v203 -b ch32v103r_r1_1v0 -b ch32v307v_r1_1v0 -b ch582m_evt
+```
+Expected: 4/4 PASS. Firmware for missing `cmake-build-<board>` sets: build first (nanoch32v203 sets exist; ch32v103/307/ch582m may need `tools/get_deps.py -b <board>` + the examples build). Known flake: ch32v103r throughput is ~40% flaky historically — retry before blaming the port. If ch582m misbehaves specifically, note it and check `wlinke.c`'s riscvchip dispatch for CH58x.
+
+- [ ] **Step 6: Squash Tasks 6–8 into downstream commit #3 and push**
+
+```bash
+cd ~/app/openocd
+git reset --soft $(git log --grep='WIP: wch port' --format=%H | tail -1)^
+git commit -m "jtag, flash: add WCH-LinkE adapter, sdi transport and CH32 flash drivers
+
+Ported from hathach/riscv-openocd-wch @ ccb04d7 (originally
+dragonlock2/miscboards WCH SDK, base openocd 0.11.0):
+- src/jtag/drivers/wlinke.c: WCH-Link/LinkE USB adapter (GCC-14 fixes included)
+- src/jtag/sdi.c: WCH single-wire debug transport, re-worked onto
+ mainline's id-based transport API (TRANSPORT_SDI)
+- src/flash/nor/wchriscv.c, wcharm.c: CH32V/CH5xx (wlink protocol) and
+ CH32F (memory-mapped) flash drivers
+CH32 cores are driven by mainline's riscv target over wlink DMI
+transfers; the fork's wch_riscv target stack is not needed.
+The fork's core patches (flash/nor/tcl.c protect-check hack) moved into
+the wch_riscv flash driver's protect_check op.
+
+Verified on ci rig: nanoch32v203, ch32v103r_r1_1v0, ch32v307v_r1_1v0,
+ch582m_evt - flash + verify + HIL suite + RTT (nanoch32v203)."
+git push
+```
+(Amend the target-stack paragraph if the fallback path was taken instead.)
+Release the nanoch32v203 lock if still held.
+
+---
+
+### Task 9: Espressif — ESP32-P4 configs, S3 attach verification, downstream commit #4
+
+Mainline already has: `src/target/espressif/` (esp32/s2/s3 xtensa targets + apptrace/semihosting), the `esp_usb_jtag` adapter driver, and builtin cfgs for c2/c3/c6/h2/s3. Missing vs the rig: anything ESP32-P4. Flash stays esptool (rig flashes ESP via `idf.py`/esptool; the espressif fork's flash-stub stack is explicitly out of scope).
+
+**Files:**
+- Create: `~/app/openocd/tcl/target/esp32p4.cfg`, `~/app/openocd/tcl/board/esp32p4-builtin.cfg`
+
+**Interfaces:**
+- Consumes: install prefix; espressif fork cfgs fetched from GitHub.
+- Produces: downstream commit #4; P4 + S3 debug-attach evidence.
+
+- [ ] **Step 1: Verify S3 attach with pure mainline inheritance** (no new files; proves the "espressif support" baseline)
+
+```bash
+python3 ~/code/tinyusb/test/hil/board_lock.py hold espressif_s3_devkitm --reason "openocd-unified esp verify"
+~/app/openocd_tinyusb/bin/openocd -f board/esp32s3-builtin.cfg -c "init; halt"
+```
+Expected: both xtensa cores detected over USB-Serial-JTAG (303a:1001), `Target halted`. Then `resume; shutdown`, release lock. Gotchas live in the `esp-target-debug` skill (S3's debug port can be occupied when TinyUSB firmware owns the USB peripheral — use the same recovery steps as that skill).
+
+- [ ] **Step 2: Fetch and adapt the P4 configs (write both files)**
+
+```bash
+curl -fsSL https://raw.githubusercontent.com/espressif/openocd-esp32/master/tcl/target/esp32p4.cfg -o /tmp/claude-1000/-home-hathach-code-tinyusb/7dee5f9e-874b-4680-bb09-01a5d13fbd37/scratchpad/esp32p4-espressif.cfg
+curl -fsSL https://raw.githubusercontent.com/espressif/openocd-esp32/master/tcl/board/esp32p4-builtin.cfg -o /tmp/claude-1000/-home-hathach-code-tinyusb/7dee5f9e-874b-4680-bb09-01a5d13fbd37/scratchpad/esp32p4-builtin-espressif.cfg
+```
+Espressif's cfg creates an `esp32p4`-type target (their `esp_riscv` C stack — not in mainline). Rewrite `tcl/target/esp32p4.cfg` following **mainline's own ESP RISC-V pattern** — `tcl/target/esp32c6.cfg` + `esp_common.cfg` (generic `riscv` target create, chip quirks via `riscv dmi_write` with the `_RISCV_*` register constants from `esp_common.cfg`) — carrying over from Espressif's file: `_CPUTAPID`, memory map/workarea, the dual-core SMP topology (P4 is 2× RV32 — model on how mainline handles SMP, and on Espressif's `_ESP_SMP_TARGET`), and the `_ESP_EFUSE_MAC_ADDR_REG` value. `tcl/board/esp32p4-builtin.cfg` = `esp_usb_jtag` adapter + `transport select jtag` + source the target cfg (mirror `board/esp32c6-builtin.cfg`, adjusting `ESP_USB_JTAG_*` ids to Espressif's P4 values).
+Also check `src/jtag/drivers/esp_usb_jtag.c` accepts the P4 (VID/PID 303a:1001 is shared; verify any chip-id gating).
+
+- [ ] **Step 3: P4 attach test**
+
+```bash
+python3 ~/code/tinyusb/test/hil/board_lock.py hold espressif_p4_function_ev --reason "openocd-unified esp verify"
+cd ~/app/openocd && make install
+~/app/openocd_tinyusb/bin/openocd -f board/esp32p4-builtin.cfg -c "init; halt"
+```
+Evidence criteria: both HP cores halt, `mdw 0x4ff00000 4` (P4 HP TCM/SRAM — cross-check the address against Espressif's cfg memory map before running) reads, `resume` works. Known nuance from prior sessions: P4 attach can need the reset-into-attach dance — the `esp-target-debug` skill documents it; an attach that only works with that dance still counts as pass (note it in the commit).
+**Decision checkpoint:** if the generic-riscv shape cannot attach P4 for architectural reasons (needs Espressif's C-level `esp_riscv` assist), stop and report — options are cherry-picking their `esp_riscv` stack (large) vs shipping P4 as esptool-flash-only with debug via ESP-IDF's openocd as today. Do not silently pick either.
+
+- [ ] **Step 4: Commit (downstream commit #4) and push**
+
+```bash
+cd ~/app/openocd
+git add tcl/target/esp32p4.cfg tcl/board/esp32p4-builtin.cfg
+git commit -m "tcl: add ESP32-P4 target/board configs adapted from espressif/openocd-esp32
+
+Adapted from espressif/openocd-esp32 master onto mainline's generic
+RISC-V ESP pattern (tcl/target/esp32c6.cfg + esp_common.cfg): generic
+riscv targets over esp_usb_jtag instead of the fork's esp_riscv C
+stack. Flash programming stays with esptool, matching how the rig
+flashes all Espressif boards. ESP32/S2/S3/C3/C6/H2 were already
+supported by mainline.
+
+Verified on ci rig: espressif_p4_function_ev and espressif_s3_devkitm
+attach/halt/resume over built-in USB-Serial-JTAG."
+git push
+```
+Release both ESP board locks.
+
+---
+
+### Task 10: Final sweep, README truth-up, rig config flip
+
+**Files:**
+- Modify: `~/app/openocd/README.md` (only if scope shifted in Tasks 7–9)
+- Modify (TinyUSB worktree from Task 8): `test/hil/tinyusb.json` — max32666fthr flasher `openocd_adi` → `openocd` with args `-f interface/cmsis-dap.cfg -f target/max32665.cfg` (plain openocd now serves it)
+
+**Interfaces:**
+- Consumes: everything green from Tasks 4–9.
+- Produces: the finished fork; TinyUSB branch `claude/hil-openocd-unified` with hil_test.py + tinyusb.json changes, committed, **unpushed** (user pushes per standing instruction).
+
+- [ ] **Step 1: Full HIL regression across every openocd-family board**
+
+```bash
+cd ~/code/tinyusb/.worktrees/claude/hil-openocd-unified
+python3 test/hil/hil_test.py test/hil/tinyusb.json \
+ -b raspberry_pi_pico -b raspberry_pi_pico_w -b raspberry_pi_pico2 \
+ -b adafruit_fruit_jam -b stm32h743nucleo -b stm32g0b1nucleo -b max32666fthr \
+ -b nanoch32v203 -b ch32v103r_r1_1v0 -b ch32v307v_r1_1v0 -b ch582m_evt
+```
+Expected: 11/11 PASS (ch32v103r throughput may need its usual retries).
+
+- [ ] **Step 2: README truth-up**
+
+Re-read `README.md` against what actually landed (WCH target route, P4 outcome). Fix any row that no longer matches; amend into the README commit or add
+`git commit -m "README: reflect verified scope"`. Push.
+
+- [ ] **Step 3: Verify the one-commit-per-fork shape**
+
+Run: `git -C ~/app/openocd log --oneline upstream/master..tinyusb`
+Expected: exactly 5 commits (or 6 with a README truth-up): README, RPi configs, ADI config, WCH port, ESP32-P4 configs. If not, interactive-free cleanup: `git rebase --onto` / `reset --soft` re-squash, then `git push --force-with-lease` (fork branch, ours alone — safe).
+
+- [ ] **Step 4: Commit the TinyUSB-side changes in the worktree (do not push)**
+
+```bash
+cd ~/code/tinyusb/.worktrees/claude/hil-openocd-unified
+git add test/hil/hil_test.py test/hil/tinyusb.json
+git commit -m "test(hil): drive WCH boards and max32666fthr through the unified openocd"
+```
+Leave for the user to push/PR.
+
+- [ ] **Step 5: Leftovers report** (no deletions now)
+
+Write a short status into `OPENOCD_UNIFIED_FORK_HANDOFF.md` (append a "2026-07-XX outcome" section): what was repointed, rollback paths (`/usr/local/bin/openocd.rpi-backup-20260727`, `~/.local/bin/openocd_wch.bak-20260727`), and that `~/app/openocd_rpi`, `~/app/openocd_adi`, `~/app/openocd-mainline`, `~/app/openocd_mainline`, `~/app/openocd_wch_new`, `~/app/riscv-openocd-wch` can be retired **after a week of green CI** — not now.
diff --git a/docs/superpowers/plans/2026-07-28-hil-test-split.md b/docs/superpowers/plans/2026-07-28-hil-test-split.md
new file mode 100644
index 000000000..6b8528973
--- /dev/null
+++ b/docs/superpowers/plans/2026-07-28-hil-test-split.md
@@ -0,0 +1,355 @@
+# hil_test.py Split Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Split `test/hil/hil_test.py` (2370 ln) into a test-focused core plus `hil_lock.py` (board locks + controller permits + operator CLI, superseding `board_lock.py`) and `hil_flash.py` (run_cmd + flash backends + firmware/serial lookup), with no behavior change.
+
+**Architecture:** Pure code motion per `docs/superpowers/specs/2026-07-28-hil-test-refactor-design.md`. Import graph: `hil_test` → {`hil_lock`, `hil_flash`}; helpers import nothing local. Call sites use module-qualified names (`hil_lock.flash_permit(...)`), never wildcard mirroring.
+
+**Tech Stack:** Python 3.11+ (existing `TypedDict`/`NotRequired` usage), stdlib only in the helpers (fcntl, json, glob, multiprocessing objects passed in).
+
+## Global Constraints
+
+- Work in worktree `.claude/worktrees/hil-test-split` (branch `claude/hil-test-split`); never touch the primary checkout.
+- Behavior-preserving: `hil_test.py` CLI args, log lines, report format, lock/permit semantics, flash behavior all byte-identical. The ONLY user-visible change is the CLI filename `board_lock.py` → `hil_lock.py`.
+- Moved functions are moved **verbatim** — no reformatting, no comment editing, no "improvements". A diff of a moved function's body against its old self must be empty.
+- Commit messages: imperative, scoped, no Co-Authored-By/Claude-Session trailers.
+- Every commit leaves the tree working: `python3 -m py_compile` clean on all touched modules, and `python3 .claude/skills/hil/pool_check.py --scan-only` exits 0 (safe on the rig: scan-only takes no locks, flashes nothing).
+- Hardware steps (Task 4) run on the `ci` rig only, from this worktree, and rely on the tools' own board flocks — never pre-hold boards you are about to run `hil_test.py`/`pool_check.py` on.
+
+---
+
+### Task 1: Create hil_flash.py; repoint hil_test + pool_check flash call sites
+
+**Files:**
+- Create: `test/hil/hil_flash.py`
+- Modify: `test/hil/hil_test.py` (delete moved code; add import; qualify call sites)
+- Modify: `.claude/skills/hil/pool_check.py` (flash-related imports)
+- Modify: `test/hil/hil_ci.sh` (scp list)
+
+**Interfaces:**
+- Produces (used by Tasks 2-4): module `hil_flash` with `CMD_TIMEOUT`, `run_cmd(cmd, cwd=None, timeout=CMD_TIMEOUT)`, `cmd_stdout_text(out)`, `OPENCOD_ADI_PATH`, `TINYUSB_ROOT`, `flash_jlink/reset_jlink`, `flash_stlink/reset_stlink`, `flash_stflash/reset_stflash`, `flash_openocd/reset_openocd`, `flash_openocd_wch/reset_openocd_wch`, `flash_openocd_adi/reset_openocd_adi`, `flash_wlink_rs/reset_wlink_rs`, `flash_esptool/reset_esptool`, `flash_uniflash/reset_uniflash`, `flash_lm4flash/reset_lm4flash`, `find_firmware(variant, example)`, `get_serial_dev(id, vendor_str, product_str, ifnum)`, module globals `build_dir = 'cmake-build'`, `verbose = False`.
+
+- [ ] **Step 1: Create `test/hil/hil_flash.py`**
+
+Header (new code), then the moved blocks verbatim:
+
+```python
+#!/usr/bin/env python3
+# SPDX-License-Identifier: MIT
+# Firmware flashing for the TinyUSB HIL rig: run_cmd, one flash_*/reset_* pair per
+# flasher type (dispatched by config name via getattr), find_firmware, and the
+# fixture serial-port resolver get_serial_dev (here, not hil_test: flash_esptool
+# needs it and helpers must not import hil_test).
+# Callers set module globals `build_dir` and `verbose` (hil_test.main from argparse,
+# pool_check directly) exactly as they set hil_test's globals today.
+
+import glob
+import json
+import os
+import signal
+import subprocess
+import sys
+from pathlib import Path
+
+verbose = False
+build_dir = 'cmake-build'
+```
+
+Then MOVE (cut from `hil_test.py`, paste unchanged, in this order):
+1. `CMD_TIMEOUT = int(os.getenv('HIL_CMD_TIMEOUT', '180'))` (from the constants block; leave `POOL_TIMEOUT`/`SERIAL_*_TIMEOUT` in hil_test)
+2. `def cmd_stdout_text(out)`
+3. `OPENCOD_ADI_PATH = Path.home() / 'app' / 'openocd_adi'` and `TINYUSB_ROOT = Path(__file__).resolve().parents[2]`
+4. `def get_serial_dev(id, vendor_str, product_str, ifnum)`
+5. `def run_cmd(cmd, cwd=None, timeout=CMD_TIMEOUT)`
+6. All ten `flash_*`/`reset_*` pairs listed in Interfaces, in current file order
+7. `def find_firmware(variant, example)`
+
+- [ ] **Step 2: Delete the moved code from `hil_test.py` and qualify call sites**
+
+In `hil_test.py`: add `import hil_flash` under the existing imports; delete the moved definitions and the `build_dir = 'cmake-build'` global (line ~165) plus `global build_dir` in `main`. Repoint every use, all module-qualified:
+- `globals()[f'flash_{...}']` → `getattr(hil_flash, f'flash_{...}')` (1 site, in `test_example`)
+- `globals()[f'reset_{...}']` → `getattr(hil_flash, f'reset_{...}')` (3 sites: `test_host_device_info`, `test_host_cdc_msc_hid`, `test_host_msc_file_explorer`)
+- bare `run_cmd(` → `hil_flash.run_cmd(` ; `cmd_stdout_text(` → `hil_flash.cmd_stdout_text(` ; `find_firmware(` → `hil_flash.find_firmware(` ; `get_serial_dev(` → `hil_flash.get_serial_dev(` ; `TINYUSB_ROOT` → `hil_flash.TINYUSB_ROOT` (in `build_board`, `CONTROLLER_CACHE` stays hil_test-local)
+- In `main()`: `build_dir = args.build_dir` → `hil_flash.build_dir = args.build_dir`; where `verbose` is set, add `hil_flash.verbose = args.verbose` (hil_test keeps its own `verbose` for test-side prints)
+- `run_cmd`'s `elif verbose:` branch now reads `hil_flash.verbose` (it moved with the function — verify it references the module-local name, not hil_test's)
+
+Find every remaining call site mechanically:
+
+Run: `grep -nE 'run_cmd|cmd_stdout_text|find_firmware|get_serial_dev|flash_[a-z]|reset_[a-z]|TINYUSB_ROOT|OPENCOD' test/hil/hil_test.py | grep -v hil_flash`
+Expected: only hits inside comments/strings and the `reset_{flasher}` dispatch f-strings already qualified.
+
+- [ ] **Step 3: Repoint pool_check's flash imports**
+
+In `.claude/skills/hil/pool_check.py`: add `import hil_flash` next to `import hil_test`; replace `hil_test.find_firmware` → `hil_flash.find_firmware` (3 sites), `hil_test.cmd_stdout_text` → `hil_flash.cmd_stdout_text`, `hil_test.get_serial_dev` → `hil_flash.get_serial_dev`, `hil_test.TINYUSB_ROOT` → `hil_flash.TINYUSB_ROOT`, `hil_test.build_dir` → `hil_flash.build_dir` (2 sites incl. `main`'s assignment), `hil_test.verbose = args.verbose` → `hil_flash.verbose = args.verbose`, `getattr(hil_test, f'flash_...')`/`getattr(hil_test, f'reset_...')` → `getattr(hil_flash, ...)` (4 sites). Keep `import hil_test` and the pymtp shim for now (locks still live there; removed in Task 2).
+
+- [ ] **Step 4: Add hil_flash.py to the hil_ci.sh scp list**
+
+```bash
+scp -q "$ROOT_DIR/test/hil/hil_test.py" \
+ "$ROOT_DIR/test/hil/hil_flash.py" \
+ "$ROOT_DIR/test/hil/pymtp.py" \
+ "$CONFIG" \
+ "$REMOTE:$REMOTE_DIR/test/hil/"
+```
+
+- [ ] **Step 5: Verify**
+
+Run: `python3 -m py_compile test/hil/hil_flash.py test/hil/hil_test.py .claude/skills/hil/pool_check.py && python3 test/hil/hil_test.py --help >/dev/null && python3 .claude/skills/hil/pool_check.py --scan-only`
+Expected: compiles; help prints nothing to stderr; scan-only prints the table and exits 0.
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add test/hil/hil_flash.py test/hil/hil_test.py test/hil/hil_ci.sh .claude/skills/hil/pool_check.py
+git commit -m "hil: extract flashing into hil_flash.py"
+```
+
+---
+
+### Task 2: Create hil_lock.py core (flock protocol + controller permits); repoint hil_test + pool_check
+
+**Files:**
+- Create: `test/hil/hil_lock.py`
+- Modify: `test/hil/hil_test.py`
+- Modify: `.claude/skills/hil/pool_check.py`
+- Modify: `test/hil/hil_ci.sh`
+
+**Interfaces:**
+- Produces: module `hil_lock` with `BOARD_LOCK_DIR`, `CI_REASON = 'hil_test.py'`, `lock_path(board)`, `flock_nb(board)`, `write_record(fh, reason)`, `clear_record(fh)`, `read_record(board)`, `acquire_board_lock(board, reason=CI_REASON)`, `FLASH_PARALLEL`, `USBTEST_PARALLEL`, `CONTROLLER_SLOTS`, `controller_of(uid)`, `controller_slot(pci)`, `controller_permit`, `flash_permit(uid)`, `usbtest_permit(uid)`, `init_scheduling(b_sems, f_sems, cmap, cmeta, hints, log_fn=None)`.
+
+- [ ] **Step 1: Create `test/hil/hil_lock.py` with the flock core**
+
+New code (the protocol, factored from today's three copies — `board_lock.py` `cmd_hold`/`read_info`, `hil_test.acquire_board_lock`, pool_check `lock_board`; behavior identical to `hil_test.acquire_board_lock` for the acquire path):
+
+```python
+#!/usr/bin/env python3
+# SPDX-License-Identifier: MIT
+"""Board locks + controller permits for the TinyUSB HIL rig.
+
+Board locks are kernel flocks in BOARD_LOCK_DIR arbitrating hardware access
+between dev sessions and CI's hil_test.py (never stop the actions-runner).
+Controller permits are in-process semaphores budgeting flashes and usbtest
+batteries per host controller; they have no CLI meaning. The CLI below
+(hold/release/status) manages board locks only; it supersedes board_lock.py.
+"""
+import argparse
+import fcntl
+import glob
+import json
+import os
+import re
+import select
+import signal
+import sys
+import time
+
+BOARD_LOCK_DIR = '/tmp/tinyusb-hil-locks'
+CI_REASON = 'hil_test.py' # release-protected holder tag (release refuses to kill it)
+PROFILE = os.environ.get('HIL_PROFILE') == '1'
+
+
+def lock_path(board: str) -> str:
+ return os.path.join(BOARD_LOCK_DIR, f'{board}.lock')
+
+
+def flock_nb(board: str):
+ """Open-or-create the lock file WITHOUT truncating (a losing racer must not
+ wipe the winner's record) and take LOCK_EX|LOCK_NB. Returns the open handle;
+ raises OSError when the flock is held elsewhere (handle already closed)."""
+ fd = os.open(lock_path(board), os.O_RDWR | os.O_CREAT, 0o666)
+ fh = os.fdopen(fd, 'r+')
+ try:
+ fcntl.flock(fh, fcntl.LOCK_EX | fcntl.LOCK_NB)
+ except OSError:
+ fh.close()
+ raise
+ return fh
+
+
+def write_record(fh, reason: str) -> None:
+ """Best-effort holder record; the flock itself is already held."""
+ try:
+ fh.truncate(0)
+ fh.seek(0)
+ json.dump({'pid': os.getpid(), 'reason': reason,
+ 'since': time.strftime('%Y-%m-%dT%H:%M:%S%z')}, fh)
+ fh.flush()
+ except OSError:
+ pass
+
+
+def clear_record(fh) -> None:
+ """Clear our record before dropping the flock so records stay truthful."""
+ try:
+ fh.truncate(0)
+ except OSError:
+ pass
+
+
+def read_record(board: str):
+ try:
+ with open(lock_path(board)) as f:
+ return json.load(f)
+ except (OSError, ValueError):
+ return None
+```
+
+Then MOVE `acquire_board_lock` from `hil_test.py` verbatim, with exactly two mechanical edits: signature becomes `def acquire_board_lock(board_name, reason=CI_REASON):` and the record-write dict's `'reason': 'hil_test.py'` becomes `'reason': reason`. Do NOT rewrite its body in terms of `flock_nb` — on conflict it reads holder info from the still-open handle before closing, which `flock_nb` (closes on conflict) cannot provide; the fail-open warning text and RuntimeError message must survive character-for-character.
+
+- [ ] **Step 2: Move the controller-permit block into `hil_lock.py`**
+
+MOVE verbatim from `hil_test.py`: the scheduling comment block + `FLASH_PARALLEL`, `USBTEST_PARALLEL`, `CONTROLLER_SLOTS`, the five module globals (`usbtest_sems`, `flash_sems`, `controller_map`, `controller_meta`, `controller_hints`), `controller_of`, `controller_slot`, `controller_permit`, `flash_permit`, `usbtest_permit`. Two mechanical adaptations:
+- add at module scope `log = print` and a setter, replacing the two `log_line(...)` calls inside `controller_of`/`controller_permit` with `log(...)`:
+
+```python
+log = print # hil_test.init_worker points this at log_line via init_scheduling
+
+
+def init_scheduling(b_sems, f_sems, cmap, cmeta, hints, log_fn=None):
+ """Install per-worker scheduling state (called from hil_test.init_worker)."""
+ global usbtest_sems, flash_sems, controller_map, controller_meta, controller_hints, log
+ usbtest_sems, flash_sems = b_sems, f_sems
+ controller_map, controller_meta, controller_hints = cmap, cmeta, hints
+ if log_fn is not None:
+ log = log_fn
+```
+
+- `PROFILE` inside `controller_permit` now resolves to hil_lock's own module constant (defined in Step 1).
+
+- [ ] **Step 3: Repoint `hil_test.py`**
+
+Add `import hil_lock`. Delete the moved lock + permit code and the five globals. `init_worker` keeps its exact signature and initargs; its body sets the hil_test globals it still owns (`print_lock`, `shuffle_seed`) and forwards the rest:
+
+```python
+def init_worker(lock, seed, b_mutexes, f_sems, cmap, cmeta, hints_by_uid):
+ global print_lock, shuffle_seed
+ print_lock = lock
+ shuffle_seed = seed
+ hil_lock.init_scheduling(b_mutexes, f_sems, cmap, cmeta, hints_by_uid, log_fn=log_line)
+```
+
+Qualify remaining uses: `acquire_board_lock(name)` → `hil_lock.acquire_board_lock(name)` (in `test_board`), `flash_permit(` → `hil_lock.flash_permit(`, `usbtest_permit(` → `hil_lock.usbtest_permit(`, and `main()`'s startup log line + Semaphore construction read `hil_lock.FLASH_PARALLEL`/`hil_lock.USBTEST_PARALLEL`/`hil_lock.CONTROLLER_SLOTS`. `controller_map` reads in the hint-persistence block of `main` use the Manager dict it already holds locally (`cmap`) — no hil_lock global access there; verify.
+
+- [ ] **Step 4: Repoint pool_check to hil_lock and drop its private copies + hil_test import**
+
+In `pool_check.py`: replace `lock_board`/`unlock_board` bodies with the shared core —
+
+```python
+import hil_lock
+
+def lock_board(name: str):
+ try:
+ fh = hil_lock.flock_nb(name)
+ except OSError:
+ info = hil_lock.read_record(name)
+ return json.dumps(info) if info else 'unknown holder'
+ hil_lock.write_record(fh, 'pool_check')
+ return fh
+
+
+def unlock_board(fh) -> None:
+ hil_lock.clear_record(fh)
+ fh.close()
+```
+
+(Behavior note: `lock_board` currently returns the raw record text; JSON-dumping the parsed record is equivalent for display. `hil_lock.BOARD_LOCK_DIR` replaces `hil_test.BOARD_LOCK_DIR`; `os.makedirs(...)` call stays, now on `hil_lock.BOARD_LOCK_DIR`.) Then delete `import hil_test` and the pymtp stub block (`try: import pymtp ... sys.modules['pymtp'] = ...`) — pool_check now imports only `hil_lock` + `hil_flash`.
+
+Run: `grep -n 'hil_test' .claude/skills/hil/pool_check.py`
+Expected: only the docstring mention of the protocol/history, no code references (update the docstring's "imports test/hil/hil_test.py" line to name hil_lock/hil_flash).
+
+- [ ] **Step 5: Add hil_lock.py to the hil_ci.sh scp list** (same block as Task 1 Step 4, one more line: `"$ROOT_DIR/test/hil/hil_lock.py" \`)
+
+- [ ] **Step 6: Verify**
+
+Run: `python3 -m py_compile test/hil/hil_lock.py test/hil/hil_test.py .claude/skills/hil/pool_check.py && python3 test/hil/hil_test.py --help >/dev/null && python3 .claude/skills/hil/pool_check.py --scan-only`
+Expected: clean compile, working scan table, exit 0.
+
+- [ ] **Step 7: Commit**
+
+```bash
+git add test/hil/hil_lock.py test/hil/hil_test.py test/hil/hil_ci.sh .claude/skills/hil/pool_check.py
+git commit -m "hil: extract board locks and controller permits into hil_lock.py"
+```
+
+---
+
+### Task 3: Absorb board_lock.py CLI into hil_lock.py; delete board_lock.py; rename in docs
+
+**Files:**
+- Modify: `test/hil/hil_lock.py` (append CLI)
+- Delete: `test/hil/board_lock.py`
+- Modify: `.claude/skills/hil/SKILL.md`, `.claude/agents/hil-operator.md`, `.claude/agents/target-debugger.md`, `.claude/skills/etm-trace/SKILL.md`, `.claude/skills/usb-kernel-recover/SKILL.md`, `.claude/skills/target-debug/SKILL.md`
+
+**Interfaces:**
+- Produces: `python3 test/hil/hil_lock.py hold|release|status` — identical subcommands, flags, output, and exit codes to today's `board_lock.py`.
+
+- [ ] **Step 1: Move the CLI from `board_lock.py` into `hil_lock.py`**
+
+MOVE verbatim to the end of `hil_lock.py`: `boards_from_config`, `is_locked`, `cmd_hold`, `cmd_release`, `cmd_status`, `main()`, and the `if __name__ == '__main__':` guard. Mechanical adaptations only:
+- `LOCK_DIR` → `BOARD_LOCK_DIR` (all sites), `lock_path` already exists (delete the duplicate), `read_info` → `read_record` (all sites; delete the duplicate definition)
+- `cmd_hold`'s holder loop body (the open/flock/json.dump block) becomes `fh = flock_nb(b)` + `write_record(fh, reason)` inside the existing try/except OSError
+- `_bow_out`'s per-handle truncate loop becomes `clear_record(h)` per handle
+- `cmd_release`'s probe uses `flock_nb(b)` in a try/except OSError (held → existing record/victim logic, with the literal `'hil_test.py'` comparison becoming `CI_REASON`); the free-path truncate becomes `clear_record(fh)`
+- `main()`'s module docstring reference for `--help` text: keep the usage lines, updating the tool name to `hil_lock.py`
+
+Then delete `test/hil/board_lock.py` (`git rm test/hil/board_lock.py`).
+
+- [ ] **Step 2: Rename `board_lock.py` → `hil_lock.py` in the six live docs**
+
+Run: `cd <worktree> && sed -i 's/board_lock\.py/hil_lock.py/g' .claude/skills/hil/SKILL.md .claude/agents/hil-operator.md .claude/agents/target-debugger.md .claude/skills/etm-trace/SKILL.md .claude/skills/usb-kernel-recover/SKILL.md .claude/skills/target-debug/SKILL.md`
+Then: `grep -rn 'board_lock' .claude/ test/ --include='*.md' --include='*.py' --include='*.sh'`
+Expected: zero hits outside `docs/superpowers/` history (which stays untouched).
+
+- [ ] **Step 3: Verify CLI behavior end-to-end**
+
+```bash
+python3 test/hil/hil_lock.py status # expect: no locks (or current holders)
+python3 test/hil/hil_lock.py hold stm32f072disco --reason "split test" &
+sleep 1
+python3 test/hil/hil_lock.py status # expect: stm32f072disco: {... 'reason': 'split test' ...}
+python3 test/hil/hil_lock.py hold stm32f072disco --reason "rival" || echo "conflict OK" # expect: ERROR ... locked + conflict OK
+python3 test/hil/hil_lock.py release stm32f072disco # expect: released holder pid NNN
+python3 test/hil/hil_lock.py status # expect: no locks
+```
+
+Also verify CI-holder protection: create a fake record `echo '{"pid": 1, "reason": "hil_test.py"}' > /tmp/tinyusb-hil-locks/faketest.lock` — since pid 1 holds no flock, `release faketest` must clear the stale record without printing the mid-test error; then `rm -f /tmp/tinyusb-hil-locks/faketest.lock`.
+
+- [ ] **Step 4: Commit**
+
+```bash
+git add -A test/hil .claude
+git commit -m "hil: fold board_lock CLI into hil_lock.py, retire board_lock.py"
+```
+
+---
+
+### Task 4: Rig verification + pre-commit
+
+**Files:** none new (fixes only if verification fails)
+
+- [ ] **Step 1: pool_check flash path on one board**
+
+Run: `python3 .claude/skills/hil/pool_check.py -b stm32f407disco`
+Expected: `✅ dfu_runtime ✅ cafe:...`, exit 0.
+
+- [ ] **Step 2: Capture a pre-refactor baseline report**
+
+Run: `cd /home/hathach/code/tinyusb && python3 test/hil/hil_test.py -b stm32f407disco -B examples test/hil/tinyusb.json && cp hil_report.md /tmp/claude-1000/-home-hathach-code-tinyusb/*/scratchpad/hil_report_master.md`
+(Primary checkout = pre-refactor code but same rig/config; its working tree already carries the new probe uids.)
+
+- [ ] **Step 3: Run the same board from the worktree and diff the report shape**
+
+Run: `cd .claude/worktrees/hil-test-split && python3 test/hil/hil_test.py -b stm32f407disco -B /home/hathach/code/tinyusb/examples test/hil/tinyusb.json && diff <(sed 's/[0-9.]*s//g;s/[0-9.]* [kMG]B\/s//g' hil_report.md) <(sed 's/[0-9.]*s//g;s/[0-9.]* [kMG]B\/s//g' /tmp/claude-1000/-home-hathach-code-tinyusb/*/scratchpad/hil_report_master.md)`
+Expected: empty diff after stripping timings/speeds. Note: `-B` accepts the absolute path so the worktree run reuses the primary checkout's built firmware; `find_firmware` resolves `TINYUSB_ROOT/<build_dir>` and an absolute `-B` overrides relative rooting — if it does not (Path join semantics), instead symlink `ln -s /home/hathach/code/tinyusb/examples/cmake-build-stm32f407disco examples/cmake-build-stm32f407disco` in the worktree and use `-B examples`.
+
+- [ ] **Step 4: pre-commit + final grep hygiene**
+
+Run: `pre-commit run --files test/hil/hil_test.py test/hil/hil_lock.py test/hil/hil_flash.py test/hil/hil_ci.sh .claude/skills/hil/pool_check.py $(git diff --name-only HEAD~3 -- '*.md')`
+Expected: all hooks pass.
+
+- [ ] **Step 5: Commit any verification fixes**
+
+```bash
+git add -A && git commit -m "hil: post-split verification fixes" # only if Steps 1-4 required changes
+```
diff --git a/docs/superpowers/plans/2026-07-29-hil-select.md b/docs/superpowers/plans/2026-07-29-hil-select.md
new file mode 100644
index 000000000..a8abf9887
--- /dev/null
+++ b/docs/superpowers/plans/2026-07-29-hil-select.md
@@ -0,0 +1,856 @@
+# PR-Scoped HIL Selection Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** A diff→(boards, tests) selector (`test/hil/hil_select.py`) that scopes CI's HIL build+test jobs on pull requests and is reusable locally, per `docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md`.
+
+**Architecture:** Pure-stdlib classification engine (changed files → per-board test selection, fail-open to full) + thin CLI emitting JSON with per-rig `hil_test.py` arg strings; consumed by `hil_ci_set_matrix.py --select` (prunes hil-build) and shell steps in the three HIL jobs (prunes rig runs). Test lists shared via new `hil_examples.py`.
+
+**Tech Stack:** Python 3.11 stdlib only (`re`, `json`, `glob`, `subprocess` for git), `unittest` for tests, GitHub Actions YAML.
+
+## Global Constraints
+
+- Work in worktree `.claude/worktrees/hil-select` (branch `claude/hil-select`); never touch the primary checkout.
+- `hil_select.py`, `hil_examples.py`, `test_hil_select.py` import NOTHING outside the stdlib and each other — in particular never `hil_test`/`hil_flash`/`hil_lock` (GitHub's bare runner has no pyserial/pymtp).
+- Fail-open: any changed file matching no classification rule ⇒ `full: true`. Scoping applies to `pull_request` events only; push/scheduled runs stay full.
+- Behavior-preserving for existing tools: `hil_test.py` runtime behavior unchanged (only its test-list constants move to `hil_examples.py`); `hil_ci_set_matrix.py` without `--select` emits byte-identical output to today.
+- The selector only ever emits board names present in the given roster (`config['boards']`); `boards-skip` is invisible to it.
+- Commit messages: imperative, scoped, NO Co-Authored-By/Claude-Session trailers.
+- Every commit: `python3 -m py_compile` clean on touched python files, `python3 test/hil/test_hil_select.py` green (once it exists), `pre-commit run --files <touched>` clean.
+
+---
+
+### Task 1: hil_examples.py + selection engine with unit tests
+
+**Files:**
+- Create: `test/hil/hil_examples.py`
+- Create: `test/hil/hil_select.py` (engine only; CLI comes in Task 2)
+- Create: `test/hil/test_hil_select.py`
+- Modify: `test/hil/hil_test.py` (import test lists from hil_examples)
+- Modify: `test/hil/hil_ci.sh` (scp list gains `hil_examples.py`)
+
+**Interfaces:**
+- Produces `hil_examples.py`: `device_tests: list[str]`, `dual_tests: list[str]`, `host_test: list[str]` — the three lists moved VERBATIM (incl. comments) from `hil_test.py`.
+- Produces `hil_select.py` engine API used by Task 2:
+ - `classify(changed_files: list[str], repo_root: str, rosters: list[tuple[str, list[dict]]]) -> dict`
+ returning `{'full': bool, 'boards': {board_name: 'all' | sorted list[str]}, 'reasons': list[str]}`
+ where `rosters` = `[(config_path, config['boards']), ...]`.
+ - `board_roles(board: dict) -> set[str]` — subset of `{'device', 'host'}` from the roster
+ entry's `tests` flags (`device`/`host`/`dual` booleans; an `only` list contributes the
+ roles of its entries' path prefixes; `dual` implies both roles).
+ - `board_family(board_name: str, repo_root: str) -> str | None` — the `<family>` for which
+ `hw/bsp/<family>/boards/<board_name>` exists.
+ - `port_families(port_dir: str, repo_root: str) -> set[str]` — directories of
+ `hw/bsp/*/family.cmake` and `hw/bsp/*/family.mk` whose text contains `port_dir`
+ (e.g. `raspberrypi/rp2040`).
+ - `class_examples(class_dir: str, role: str, repo_root: str) -> set[str]` — tests from
+ `hil_examples` lists whose example `tusb_config.h` enables the class for that role (regex
+ `#define\s+CFG_TUD_<C>\s+\(?\s*0*[1-9]` / `CFG_TUH_<C>`; exceptions per spec:
+ `dfu_rt_device.*`→`CFG_TUD_DFU_RUNTIME`, `dfu_device.*`→`CFG_TUD_DFU`, class dir `net`
+ → `CFG_TUD_ECM_RNDIS|CFG_TUD_NCM`). Test path `device/x` ⇒ config at
+ `examples/device/x/src/tusb_config.h`; same pattern for `host/` and `dual/`.
+
+- [ ] **Step 1: Move the test lists into `hil_examples.py`**
+
+Create `test/hil/hil_examples.py`:
+
+```python
+#!/usr/bin/env python3
+# SPDX-License-Identifier: MIT
+# HIL example test lists, shared by hil_test.py (runner) and hil_select.py
+# (PR-diff selector). Stdlib-only: hil_select runs on bare CI runners.
+```
+
+then MOVE the `device_tests`, `dual_tests`, `host_test` list definitions (and their preceding
+comment block "The per-board run order is shuffled...") VERBATIM from `hil_test.py` into it.
+In `hil_test.py`, add `from hil_examples import device_tests, dual_tests, host_test` where the
+lists were (a `from`-import of data constants is fine here — they are read-only lists used by
+name throughout `test_board`). Add `"$ROOT_DIR/test/hil/hil_examples.py" \` to the
+`hil_ci.sh` scp list after the `hil_lock.py` line.
+
+- [ ] **Step 2: Verify the move broke nothing**
+
+Run: `cd /home/hathach/code/tinyusb/.claude/worktrees/hil-select && python3 -m py_compile test/hil/hil_examples.py test/hil/hil_test.py && python3 test/hil/hil_test.py --help >/dev/null && echo ok`
+Expected: `ok`
+
+- [ ] **Step 3: Write the failing unit tests (spec acceptance cases)**
+
+Create `test/hil/test_hil_select.py`. ROSTER is a trimmed but real-shaped fixture; tests call
+the engine API directly (no git, no CLI):
+
+```python
+#!/usr/bin/env python3
+# SPDX-License-Identifier: MIT
+# Unit tests for hil_select.py — pure logic, no hardware, no git. Run directly:
+# python3 test/hil/test_hil_select.py
+import os
+import sys
+import unittest
+
+sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
+import hil_select
+from hil_examples import device_tests, dual_tests, host_test
+
+REPO = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
+
+ROSTER = [
+ # device-only, rp2040 family
+ {'name': 'raspberry_pi_pico', 'uid': 'u1',
+ 'tests': {'device': True, 'host': True, 'dual': True}},
+ # device-only, stm32f4 family
+ {'name': 'stm32f407disco', 'uid': 'u2',
+ 'tests': {'device': True, 'host': False, 'dual': False}},
+ # host-only board
+ {'name': 'raspberry_pi_pico2', 'uid': 'u3',
+ 'tests': {'device': False, 'host': True, 'dual': False}},
+ # only-list board (espressif-style)
+ {'name': 'espressif_s3_devkitm', 'uid': 'u4',
+ 'tests': {'only': ['device/cdc_msc_freertos', 'host/device_info']}},
+]
+ROSTERS = [('test/hil/tinyusb.json', ROSTER)]
+
+
+def sel(files):
+ return hil_select.classify(files, REPO, ROSTERS)
+
+
+class TestPortRule(unittest.TestCase):
+ def test_dcd_rp2040_selects_pico_family_only(self):
+ s = sel(['src/portable/raspberrypi/rp2040/dcd_rp2040.c'])
+ self.assertFalse(s['full'])
+ self.assertIn('raspberry_pi_pico', s['boards'])
+ self.assertNotIn('stm32f407disco', s['boards'])
+ self.assertNotIn('espressif_s3_devkitm', s['boards'])
+ # device role: no host tests in pico's list
+ self.assertTrue(all(not t.startswith('host/') for t in s['boards']['raspberry_pi_pico']))
+ # host-only boards drop out entirely on a device-role change
+ self.assertNotIn('raspberry_pi_pico2', s['boards'])
+
+ def test_shared_port_file_is_both_roles(self):
+ s = sel(['src/portable/synopsys/dwc2/dwc2_common.c'])
+ self.assertFalse(s['full'])
+ self.assertNotIn('raspberry_pi_pico', s['boards']) # rp2040 is not a dwc2 family
+ self.assertIn('stm32f407disco', s['boards']) # stm32f4 is
+
+
+class TestCoreRoleRule(unittest.TestCase):
+ def test_usbd_selects_all_device_tests_everywhere(self):
+ s = sel(['src/device/usbd.c'])
+ self.assertFalse(s['full'])
+ self.assertNotIn('raspberry_pi_pico2', s['boards']) # host-only board dropped
+ pico = s['boards']['raspberry_pi_pico']
+ self.assertTrue(set(device_tests).issubset(set(pico)))
+ self.assertTrue(set(dual_tests).issubset(set(pico))) # dual survives device role
+ self.assertTrue(all(not t.startswith('host/') for t in pico))
+ # only-list board: selection intersects its only-list
+ esp = s['boards']['espressif_s3_devkitm']
+ self.assertEqual(esp, ['device/cdc_msc_freertos'])
+
+ def test_host_change_drops_device(self):
+ s = sel(['src/host/usbh.c'])
+ self.assertFalse(s['full'])
+ self.assertIn('raspberry_pi_pico2', s['boards'])
+ self.assertNotIn('stm32f407disco', s['boards']) # device-only board dropped
+
+
+class TestClassRule(unittest.TestCase):
+ def test_cdc_device_selects_cdc_examples_only(self):
+ s = sel(['src/class/cdc/cdc_device.c'])
+ self.assertFalse(s['full'])
+ pico = s['boards']['raspberry_pi_pico']
+ self.assertIn('device/cdc_msc', pico)
+ self.assertIn('device/cdc_dual_ports', pico)
+ self.assertNotIn('device/msc_dual_lun', pico) # CFG_TUD_CDC 0 there
+ self.assertNotIn('device/usbtest', pico) # CFG_TUD_CDC 0 there
+ self.assertTrue(all(not t.startswith('host/') for t in pico))
+
+ def test_msc_host_selects_host_side(self):
+ s = sel(['src/class/msc/msc_host.c'])
+ self.assertFalse(s['full'])
+ self.assertNotIn('stm32f407disco', s['boards']) # device-only board
+ pico2 = s['boards']['raspberry_pi_pico2']
+ self.assertIn('host/msc_file_explorer', pico2)
+ self.assertTrue(all(not t.startswith('device/') for t in pico2))
+
+
+class TestFallbackRules(unittest.TestCase):
+ def test_unknown_tool_is_full(self):
+ s = sel(['tools/random_new_script.py'])
+ self.assertTrue(s['full'])
+
+ def test_docs_only_is_empty_not_full(self):
+ s = sel(['docs/info/contributing.rst', 'README.rst'])
+ self.assertFalse(s['full'])
+ self.assertEqual(s['boards'], {})
+
+ def test_bsp_family_selects_family_boards(self):
+ s = sel(['hw/bsp/rp2040/family.cmake'])
+ self.assertFalse(s['full'])
+ self.assertIn('raspberry_pi_pico', s['boards'])
+ self.assertEqual(s['boards']['raspberry_pi_pico'], 'all')
+ self.assertNotIn('stm32f407disco', s['boards'])
+
+ def test_bsp_board_narrows_to_board(self):
+ s = sel(['hw/bsp/rp2040/boards/raspberry_pi_pico/board.h'])
+ self.assertFalse(s['full'])
+ self.assertEqual(list(s['boards'].keys()), ['raspberry_pi_pico'])
+
+ def test_example_change_selects_that_example(self):
+ s = sel(['examples/device/cdc_msc/src/main.c'])
+ self.assertFalse(s['full'])
+ self.assertEqual(s['boards']['raspberry_pi_pico'], ['device/cdc_msc'])
+
+ def test_core_common_is_full(self):
+ for f in ['src/tusb.c', 'src/common/tusb_fifo.c', 'src/osal/osal_freertos.h']:
+ self.assertTrue(sel([f])['full'], f)
+
+ def test_harness_is_full(self):
+ for f in ['test/hil/hil_test.py', '.github/workflows/build.yml', 'hw/mcu/nxp/x.c', 'lib/foo/x.c']:
+ self.assertTrue(sel([f])['full'], f)
+
+ def test_mixed_roles_no_pruning(self):
+ s = sel(['src/device/usbd.c', 'src/host/usbh.c'])
+ self.assertFalse(s['full'])
+ self.assertIn('raspberry_pi_pico2', s['boards'])
+ self.assertIn('stm32f407disco', s['boards'])
+
+
+if __name__ == '__main__':
+ unittest.main(verbosity=1)
+```
+
+- [ ] **Step 4: Run tests to verify they fail**
+
+Run: `python3 test/hil/test_hil_select.py 2>&1 | tail -2`
+Expected: `ModuleNotFoundError: No module named 'hil_select'` (or import error).
+
+- [ ] **Step 5: Implement the engine**
+
+Create `test/hil/hil_select.py`:
+
+```python
+#!/usr/bin/env python3
+# SPDX-License-Identifier: MIT
+"""PR-diff -> HIL selection: which rig boards and which tests a change can affect.
+
+Stdlib-only (runs on bare CI runners; never imports hil_test/hil_flash/hil_lock).
+Fail-open: any file no rule classifies forces the full matrix. See
+docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md.
+"""
+import argparse
+import glob
+import json
+import os
+import re
+import subprocess
+import sys
+
+from hil_examples import device_tests, dual_tests, host_test
+
+ALL_TESTS = {'device': device_tests, 'dual': dual_tests, 'host': host_test}
+
+# class dir -> config macro suffix exceptions (rule 3); dfu is per-file, handled inline
+NET_MACROS = ('ECM_RNDIS', 'NCM')
+
+_NONCODE_RE = re.compile(
+ r'^(docs/|\.claude/|.*\.(md|rst|txt)$|LICENSE)')
+_FULL_RE = re.compile(
+ r'^(src/common/|src/osal/|src/tusb\.c$|src/tusb\.h$|src/tusb_option\.h$|'
+ r'test/hil/|\.github/workflows/build.*\.yml$|\.github/actions/|'
+ r'tools/build\.py$|tools/get_deps\.py$|tools/cmake/|hw/mcu/|lib/|'
+ r'hw/bsp/(family_support\.cmake|board_api\.h|board\.c|ansi_escape\.h)$|'
+ r'examples/build_system/|examples/CMakeLists\.txt$)')
+
+
+def test_role(test: str) -> str:
+ return test.split('/', 1)[0] # 'device' | 'dual' | 'host'
+
+
+def board_roles(board: dict) -> set:
+ t = board.get('tests', {})
+ roles = set()
+ if t.get('device'):
+ roles.add('device')
+ if t.get('host'):
+ roles.add('host')
+ if t.get('dual'):
+ roles.update(('device', 'host'))
+ for only in t.get('only', []):
+ r = test_role(only)
+ roles.update(('device', 'host') if r == 'dual' else (r,))
+ return roles
+
+
+def board_tests(board: dict) -> list:
+ """Every test this board would run today (mirrors hil_test.test_board's default)."""
+ t = board.get('tests', {})
+ if 'only' in t:
+ run = list(t['only'])
+ else:
+ run = []
+ if t.get('device'):
+ run += device_tests
+ if t.get('dual'):
+ run += dual_tests
+ if t.get('host'):
+ run += host_test
+ return [x for x in run if x not in t.get('skip', [])]
+
+
+def board_family(board_name: str, repo_root: str):
+ hits = glob.glob(os.path.join(repo_root, 'hw/bsp/*/boards', board_name))
+ return os.path.basename(os.path.dirname(os.path.dirname(hits[0]))) if hits else None
+
+
+def port_families(port_dir: str, repo_root: str) -> set:
+ fams = set()
+ for f in glob.glob(os.path.join(repo_root, 'hw/bsp/*/family.cmake')) + \
+ glob.glob(os.path.join(repo_root, 'hw/bsp/*/family.mk')):
+ try:
+ if port_dir in open(f).read():
+ fams.add(os.path.basename(os.path.dirname(f)))
+ except OSError:
+ pass
+ return fams
+
+
+def _config_enables(cfg_path: str, macros) -> bool:
+ try:
+ text = open(cfg_path).read()
+ except OSError:
+ return False
+ return any(re.search(rf'#define\s+{m}\s+\(?\s*0*[1-9]', text) for m in macros)
+
+
+def class_examples(macros, role: str, repo_root: str) -> set:
+ """Tests (from role's + dual lists) whose example config enables any macro."""
+ pools = {'device': device_tests + dual_tests, 'host': host_test + dual_tests}
+ out = set()
+ for test in pools[role]:
+ cfg = os.path.join(repo_root, 'examples', test, 'src', 'tusb_config.h')
+ if _config_enables(cfg, macros):
+ out.add(test)
+ return out
+
+
+class _Sel:
+ """Accumulates contributions. board->set(tests) plus 'all-board' markers."""
+ def __init__(self):
+ self.full = False
+ self.by_board = {} # name -> set of tests, or 'all'
+ self.roles = set() # roles touched by any contribution
+ self.reasons = []
+
+ def add(self, boards, tests, reason):
+ """tests: 'all' or iterable of test paths."""
+ self.reasons.append(reason)
+ for b in boards:
+ cur = self.by_board.get(b)
+ if tests == 'all' or cur == 'all':
+ self.by_board[b] = 'all'
+ else:
+ self.by_board[b] = (cur or set()) | set(tests)
+
+ def force_full(self, reason):
+ self.full = True
+ self.reasons.append(reason)
+
+
+def _classify_one(path, repo_root, roster_boards, s: _Sel):
+ base = os.path.basename(path)
+ if _NONCODE_RE.match(path):
+ s.reasons.append(f'{path}: non-code, no contribution')
+ return
+ if _FULL_RE.match(path):
+ s.force_full(f'{path}: core/infra -> full matrix')
+ return
+
+ m = re.match(r'src/portable/((?:[^/]+/)?[^/]+)/', path)
+ if m:
+ port = m.group(1)
+ if re.match(r'(dcd_|.*_device)', base):
+ roles = {'device'}
+ elif re.match(r'(hcd_|.*_host)', base):
+ roles = {'host'}
+ else:
+ roles = {'device', 'host'}
+ fams = port_families(port, repo_root)
+ boards = [b['name'] for b in roster_boards
+ if board_family(b['name'], repo_root) in fams and (board_roles(b) & roles)]
+ tests = [t for r in roles for t in ALL_TESTS[r]] + dual_tests
+ s.roles.update(roles)
+ s.add(boards, tests, f'{path}: port {port} -> families {sorted(fams)} -> boards {boards} ({"/".join(sorted(roles))})')
+ return
+
+ m = re.match(r'src/class/([^/]+)/', path)
+ if m:
+ cls = m.group(1)
+ if re.search(r'_device\.[ch]$', base):
+ roles = {'device'}
+ elif re.search(r'_host\.[ch]$', base):
+ roles = {'host'}
+ else:
+ roles = {'device', 'host'}
+ # macro names per role
+ def macros(prefix):
+ if cls == 'net':
+ return [f'CFG_{prefix}_{m2}' for m2 in NET_MACROS]
+ if cls == 'dfu':
+ if base.startswith('dfu_rt'):
+ return [f'CFG_{prefix}_DFU_RUNTIME']
+ if base.startswith('dfu_device') or base.startswith('dfu_host'):
+ return [f'CFG_{prefix}_DFU']
+ return [f'CFG_{prefix}_DFU', f'CFG_{prefix}_DFU_RUNTIME']
+ return [f'CFG_{prefix}_{cls.upper()}']
+ tests = set()
+ if 'device' in roles:
+ tests |= class_examples(macros('TUD'), 'device', repo_root)
+ if 'host' in roles:
+ tests |= class_examples(macros('TUH'), 'host', repo_root)
+ boards = [b['name'] for b in roster_boards if board_roles(b) & roles]
+ s.roles.update(roles)
+ s.add(boards, tests, f'{path}: class {cls} -> {sorted(tests)} ({"/".join(sorted(roles))})')
+ return
+
+ m = re.match(r'src/(device|host)/', path)
+ if m:
+ role = m.group(1)
+ boards = [b['name'] for b in roster_boards if role in board_roles(b)]
+ s.roles.add(role)
+ s.add(boards, ALL_TESTS[role] + dual_tests, f'{path}: core {role} stack -> all {role} tests')
+ return
+
+ m = re.match(r'hw/bsp/([^/]+)/(?:boards/([^/]+)/)?', path)
+ if m:
+ fam, brd = m.group(1), m.group(2)
+ if brd:
+ boards = [b['name'] for b in roster_boards if b['name'] == brd]
+ why = f'{path}: bsp board {brd}'
+ else:
+ boards = [b['name'] for b in roster_boards
+ if board_family(b['name'], repo_root) == fam]
+ why = f'{path}: bsp family {fam}'
+ s.roles.update(('device', 'host'))
+ s.add(boards, 'all', f'{why} -> boards {boards}')
+ return
+
+ m = re.match(r'examples/(device|host|dual)/([^/]+)/', path)
+ if m:
+ test = f'{m.group(1)}/{m.group(2)}'
+ known = any(test in pool for pool in ALL_TESTS.values())
+ if known:
+ boards = [b['name'] for b in roster_boards]
+ role = test_role(test)
+ s.roles.update(('device', 'host') if role == 'dual' else (role,))
+ s.add(boards, [test], f'{path}: example -> {test} on all boards')
+ else:
+ s.reasons.append(f'{path}: example not in HIL lists, no contribution')
+ return
+
+ s.force_full(f'{path}: unclassified -> full matrix')
+
+
+def classify(changed_files, repo_root, rosters):
+ all_boards = []
+ seen = set()
+ for _, boards in rosters:
+ for b in boards:
+ if b['name'] not in seen:
+ seen.add(b['name'])
+ all_boards.append(b)
+
+ s = _Sel()
+ for path in changed_files:
+ _classify_one(path, repo_root, all_boards, s)
+ if s.full:
+ break
+
+ if s.full:
+ return {'full': True, 'boards': {b['name']: 'all' for b in all_boards},
+ 'reasons': s.reasons}
+
+ # role pruning: single-role selections drop the other role's tests and boards
+ by_name = {b['name']: b for b in all_boards}
+ out = {}
+ for name, tests in s.by_board.items():
+ allowed = board_tests(by_name[name])
+ if tests == 'all':
+ kept = list(allowed)
+ else:
+ kept = [t for t in allowed if t in tests]
+ if s.roles and s.roles != {'device', 'host'}:
+ role = next(iter(s.roles))
+ kept = [t for t in kept if test_role(t) in (role, 'dual')]
+ if kept:
+ out[name] = 'all' if set(kept) == set(allowed) else sorted(kept)
+ return {'full': False, 'boards': out, 'reasons': s.reasons}
+```
+
+- [ ] **Step 6: Run tests to verify they pass**
+
+Run: `python3 test/hil/test_hil_select.py`
+Expected: all tests PASS (OK line). Iterate on the engine (not the tests) until green; if a
+test premise contradicts the repo (e.g. a family name), verify against the tree and fix the
+test only with evidence noted in your report.
+
+- [ ] **Step 7: Commit**
+
+```bash
+git add test/hil/hil_examples.py test/hil/hil_select.py test/hil/test_hil_select.py test/hil/hil_test.py test/hil/hil_ci.sh
+git commit -m "hil: add PR-diff selection engine (hil_select) with shared example lists"
+```
+
+---
+
+### Task 2: CLI + args emission
+
+**Files:**
+- Modify: `test/hil/hil_select.py` (add `selection_args`, `main`)
+- Modify: `test/hil/test_hil_select.py` (add CLI/args tests)
+
+**Interfaces:**
+- Consumes: Task 1's `classify` and roster shapes.
+- Produces:
+ - `selection_args(sel: dict, rosters) -> dict` mapping each config path's basename to the
+ `hil_test.py` argument string for that rig: for each selected board ON that roster,
+ `-b <name>`, plus `-bt <name>:<t1>,<t2>` when the board's entry is a list (not 'all').
+ Empty string when no selected board is on that roster. When `sel['full']`, every roster
+ board gets bare `-b`? NO — full means "today's behavior": `selection_args` returns `''`
+ for every config (no filtering args at all).
+ - CLI: `python3 test/hil/hil_select.py [--base REF | --diff-file PATH] CONFIG...` printing
+ the JSON `{'full', 'boards', 'args', 'reasons'}` to stdout, reasons also to stderr
+ (one line each, prefixed `hil_select: `). Non-zero exit only on operational errors
+ (bad ref, unreadable config) — never on an empty selection.
+
+- [ ] **Step 1: Add failing CLI/args tests to `test_hil_select.py`**
+
+```python
+class TestArgsEmission(unittest.TestCase):
+ def test_args_for_scoped_selection(self):
+ s = sel(['src/portable/raspberrypi/rp2040/dcd_rp2040.c'])
+ args = hil_select.selection_args(s, ROSTERS)
+ a = args['tinyusb.json']
+ self.assertIn('-b raspberry_pi_pico', a)
+ self.assertNotIn('stm32f407disco', a)
+ self.assertIn('-bt raspberry_pi_pico:', a) # device-only subset of a device+host board
+
+ def test_args_full_is_empty(self):
+ s = sel(['tools/random_new_script.py'])
+ self.assertEqual(hil_select.selection_args(s, ROSTERS), {'tinyusb.json': ''})
+
+ def test_args_all_board_gets_bare_b(self):
+ s = sel(['hw/bsp/rp2040/boards/raspberry_pi_pico/board.h'])
+ a = hil_select.selection_args(s, ROSTERS)['tinyusb.json']
+ self.assertIn('-b raspberry_pi_pico', a)
+ self.assertNotIn('-bt', a)
+
+ def test_cli_diff_file(self):
+ import subprocess, tempfile, json as j
+ with tempfile.NamedTemporaryFile('w', suffix='.txt', delete=False) as f:
+ f.write('src/class/cdc/cdc_device.c\n')
+ path = f.name
+ r = subprocess.run([sys.executable, os.path.join(REPO, 'test/hil/hil_select.py'),
+ '--diff-file', path, os.path.join(REPO, 'test/hil/tinyusb.json')],
+ capture_output=True, text=True)
+ self.assertEqual(r.returncode, 0, r.stderr)
+ out = j.loads(r.stdout)
+ self.assertFalse(out['full'])
+ self.assertIn('tinyusb.json', out['args'])
+ self.assertTrue(any('cdc_device' in line for line in out['reasons']))
+ os.unlink(path)
+```
+
+- [ ] **Step 2: Run to verify the new tests fail**
+
+Run: `python3 test/hil/test_hil_select.py 2>&1 | tail -3`
+Expected: failures/errors mentioning `selection_args`.
+
+- [ ] **Step 3: Implement `selection_args` and `main`**
+
+Append to `hil_select.py`:
+
+```python
+def selection_args(sel, rosters):
+ args = {}
+ for cfg_path, boards in rosters:
+ key = os.path.basename(cfg_path)
+ if sel['full']:
+ args[key] = ''
+ continue
+ parts = []
+ for b in boards:
+ chosen = sel['boards'].get(b['name'])
+ if chosen is None:
+ continue
+ parts.append(f'-b {b["name"]}')
+ if chosen != 'all':
+ parts.append(f'-bt {b["name"]}:{",".join(chosen)}')
+ args[key] = ' '.join(parts)
+ return args
+
+
+def changed_files_from_git(base, repo_root):
+ mb = subprocess.run(['git', 'merge-base', 'HEAD', base], cwd=repo_root,
+ capture_output=True, text=True, check=True).stdout.strip()
+ diff = subprocess.run(['git', 'diff', '--name-only', f'{mb}..HEAD'], cwd=repo_root,
+ capture_output=True, text=True, check=True).stdout
+ return [l for l in diff.splitlines() if l.strip()]
+
+
+def main():
+ ap = argparse.ArgumentParser(description=__doc__)
+ g = ap.add_mutually_exclusive_group(required=True)
+ g.add_argument('--base', help='git ref to diff against (merge-base..HEAD)')
+ g.add_argument('--diff-file', help='newline-separated changed-file list')
+ ap.add_argument('configs', nargs='+', help='rig roster JSON file(s)')
+ a = ap.parse_args()
+
+ repo_root = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
+ rosters = []
+ for c in a.configs:
+ with open(c) as f:
+ rosters.append((c, json.load(f)['boards']))
+
+ files = (open(a.diff_file).read().splitlines() if a.diff_file
+ else changed_files_from_git(a.base, repo_root))
+ files = [f for f in files if f.strip()]
+
+ s = classify(files, repo_root, rosters)
+ s['args'] = selection_args(s, rosters)
+ for r in s['reasons']:
+ print(f'hil_select: {r}', file=sys.stderr)
+ print(json.dumps(s))
+
+
+if __name__ == '__main__':
+ main()
+```
+
+(The `parts.append f'...'` line above is pseudo-highlighted; write valid Python:
+`parts.append(f'-bt {b["name"]}:{",".join(chosen)}')`.)
+
+- [ ] **Step 4: Run the full suite**
+
+Run: `python3 test/hil/test_hil_select.py && chmod +x test/hil/hil_select.py`
+Expected: OK.
+
+- [ ] **Step 5: Smoke against the real repo state**
+
+Run: `python3 test/hil/hil_select.py --base HEAD test/hil/tinyusb.json test/hil/hfp.json`
+Expected: empty diff ⇒ `{"full": false, "boards": {}, "args": {"tinyusb.json": "", "hfp.json": ""}, ...}` exit 0.
+Then: `printf 'src/portable/wch/dcd_ch32_usbfs.c\n' > /tmp/d.txt && python3 test/hil/hil_select.py --diff-file /tmp/d.txt test/hil/tinyusb.json | python3 -m json.tool | head -20`
+Expected: only WCH-family boards (nanoch32v203, ch32v103r_r1_1v0, ch32v307v_r1_1v0 — whichever reference that port) with device tests.
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add test/hil/hil_select.py test/hil/test_hil_select.py
+git commit -m "hil: hil_select CLI with per-rig hil_test argument emission"
+```
+
+---
+
+### Task 3: hil_ci_set_matrix --select + build.yml wiring
+
+**Files:**
+- Modify: `test/hil/hil_ci_set_matrix.py`
+- Modify: `.github/workflows/build.yml` (set-matrix job; hil-build consumers unchanged; hil-tinyusb + hil-tinyusb-esp steps)
+
+**Interfaces:**
+- Consumes: Task 2's CLI JSON (`full`, `boards`, `args`).
+- Produces:
+ - `hil_ci_set_matrix.py [--select JSON_STRING] CONFIG...`: with `--select` and
+ `full == false`, boards not in `select['boards']` are skipped when building the toolchain
+ buckets; otherwise identical behavior. Buckets stay present (possibly `[]`) so
+ `fromJSON(...)[toolchain]` keeps resolving.
+ - set-matrix outputs: `hil_select_json` (compact selection), `hil_args_tinyusb`,
+ `hil_args_hfp`, `hil_run_tinyusb`, `hil_run_hfp` (string 'true'/'false').
+
+- [ ] **Step 1: Add `--select` to `hil_ci_set_matrix.py`**
+
+In `main()` add:
+
+```python
+ parser.add_argument('--select', help='hil_select.py JSON; scopes boards when full=false')
+```
+
+and after parsing:
+
+```python
+ selected = None
+ sel = json.loads(args.select) if args.select else None
+ if sel and not sel.get('full'):
+ selected = set(sel.get('boards', {}))
+```
+
+then inside the per-board loop, first line:
+
+```python
+ if selected is not None and board['name'] not in selected:
+ continue
+```
+
+- [ ] **Step 2: Verify byte-identical without --select and scoped with it**
+
+Run: `python3 test/hil/hil_ci_set_matrix.py test/hil/tinyusb.json test/hil/hfp.json > /tmp/m1.json && git stash -q && python3 test/hil/hil_ci_set_matrix.py test/hil/tinyusb.json test/hil/hfp.json > /tmp/m0.json && git stash pop -q && diff /tmp/m0.json /tmp/m1.json && echo identical`
+Expected: `identical`.
+Then: `python3 test/hil/hil_ci_set_matrix.py --select '{"full": false, "boards": {"raspberry_pi_pico": "all"}}' test/hil/tinyusb.json test/hil/hfp.json`
+Expected: JSON whose `arm-gcc` list contains only the raspberry_pi_pico entry, `riscv-gcc`/`esp-idf` = [].
+
+- [ ] **Step 3: Wire set-matrix in `.github/workflows/build.yml`**
+
+In the `set-matrix` job: give the checkout full history and add the selection step between
+checkout and matrix generation; make the HIL matrix use it:
+
+```yaml
+ - name: Checkout TinyUSB
+ uses: actions/checkout@v6
+ with:
+ fetch-depth: 0
+
+ - name: HIL selection (PR only)
+ id: hil-select
+ if: github.event_name == 'pull_request'
+ run: |
+ python3 test/hil/test_hil_select.py
+ SELECT_JSON=$(python3 test/hil/hil_select.py --base "origin/${{ github.base_ref }}" test/hil/tinyusb.json test/hil/hfp.json)
+ echo "select=$SELECT_JSON" >> $GITHUB_OUTPUT
+ python3 - "$SELECT_JSON" >> $GITHUB_OUTPUT <<'EOF'
+ import json, sys
+ s = json.loads(sys.argv[1])
+ args = s.get('args', {})
+ for cfg, key in (('tinyusb.json', 'tinyusb'), ('hfp.json', 'hfp')):
+ a = args.get(cfg, '')
+ run = 'true' if (s['full'] or a) else 'false'
+ print(f'args_{key}={a}')
+ print(f'run_{key}={run}')
+ EOF
+```
+
+and in the existing "Generate matrix json" step, change the HIL line to:
+
+```yaml
+ # HIL matrix (merged from tinyusb + hifiphile configs), scoped on PRs
+ SELECT='${{ steps.hil-select.outputs.select }}'
+ HIL_MATRIX_JSON=$(python test/hil/hil_ci_set_matrix.py ${SELECT:+--select "$SELECT"} test/hil/tinyusb.json test/hil/hfp.json)
+```
+
+Add to the job's `outputs:` block:
+
+```yaml
+ hil_args_tinyusb: ${{ steps.hil-select.outputs.args_tinyusb }}
+ hil_args_hfp: ${{ steps.hil-select.outputs.args_hfp }}
+ hil_run_tinyusb: ${{ steps.hil-select.outputs.run_tinyusb }}
+ hil_run_hfp: ${{ steps.hil-select.outputs.run_hfp }}
+```
+
+(On non-PR events the step is skipped: outputs are empty strings — the consumers below treat
+empty `run_*` as 'true' and empty args as no filtering, i.e. today's behavior.)
+
+- [ ] **Step 4: Wire the rig jobs**
+
+In the `hil-tinyusb` job (the matrixed one covering both rigs), find the step that runs
+`hil_test.py --retry 1 ${{ matrix.test_args }} ${{ env.HIL_JSON }} $RERUN_ARGS` (~line 360)
+and change the step's `run:` to select per-rig args and honor the skip flag:
+
+```yaml
+ run: |
+ case "$HIL_JSON" in
+ *tinyusb.json) SEL_ARGS='${{ needs.set-matrix.outputs.hil_args_tinyusb }}'; SEL_RUN='${{ needs.set-matrix.outputs.hil_run_tinyusb }}' ;;
+ *hfp.json) SEL_ARGS='${{ needs.set-matrix.outputs.hil_args_hfp }}'; SEL_RUN='${{ needs.set-matrix.outputs.hil_run_hfp }}' ;;
+ esac
+ if [ "$SEL_RUN" = "false" ]; then echo "HIL skipped by PR selection (no affected boards on this rig)"; exit 0; fi
+ python3 test/hil/hil_test.py --retry 1 ${{ matrix.test_args }} $SEL_ARGS ${{ env.HIL_JSON }} $RERUN_ARGS
+```
+
+Apply the same pattern to the second `hil_test.py` invocation at ~line 423 (`hil-tinyusb-esp`,
+which is tinyusb-rig only: use the `hil_args_tinyusb`/`hil_run_tinyusb` outputs directly, no
+case needed) and to the hfp job's direct `python3 test/hil/hil_test.py hfp.json` call at
+~line 487 (use `hil_args_hfp`/`hil_run_hfp`). Preserve each step's existing surrounding lines
+(report-dir env, RERUN_ARGS logic) — only inject the SEL_ARGS/SEL_RUN mechanics.
+
+- [ ] **Step 5: Validate the YAML and the exact shell locally**
+
+Run: `pre-commit run check-yaml --files .github/workflows/build.yml && python3 -c "import yaml; yaml.safe_load(open('.github/workflows/build.yml')); print('yaml ok')"`
+Expected: `yaml ok` (pyyaml is available; if not, `pip install --user pyyaml` first).
+Also simulate the selection step's python inline script:
+`SELECT_JSON=$(python3 test/hil/hil_select.py --diff-file /tmp/d.txt test/hil/tinyusb.json test/hil/hfp.json) && python3 -c "import json,sys; s=json.loads(sys.argv[1]); print(s['args'])" "$SELECT_JSON"`
+Expected: the args dict prints.
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add test/hil/hil_ci_set_matrix.py .github/workflows/build.yml
+git commit -m "ci: scope HIL build+test matrix by PR diff via hil_select"
+```
+
+---
+
+### Task 4: pre-pr + hil skill docs, final validation
+
+**Files:**
+- Modify: `.claude/skills/pre-pr/SKILL.md` (mapping section delegates to the selector)
+- Modify: `.claude/skills/hil/SKILL.md` (document the selector for manual runs)
+
+**Interfaces:**
+- Consumes: Task 2's CLI.
+
+- [ ] **Step 1: Rewrite pre-pr's "2. Map changes to boards" section**
+
+Replace the section's grep heuristics (keep its numbered-section structure and the roster/cap
+policy) with:
+
+```markdown
+## 2. Map changes to boards
+
+- `python3 test/hil/hil_select.py --base $BASE test/hil/tinyusb.json` → JSON with the affected
+ rig boards (`boards`) and per-file `reasons`. `full: true` means a broad/infra change.
+- Build-board sampling: from the selection's boards (or, when `full`, the representative set
+ `stm32f407disco` + `raspberry_pi_pico`), pick ONE board per family, preferring rig-roster
+ boards; cap at 4 and tell the user which families the cap dropped. The boards list must
+ NEVER end up empty — final fallback is `[stm32f407disco]`.
+- A `full: true` selection or an empty one (docs-only) keeps today's behavior: minimal
+ software-only gate for docs-only, representative set otherwise.
+```
+
+- [ ] **Step 2: Add a short "PR-scoped selection" note to the hil skill**
+
+Append to `.claude/skills/hil/SKILL.md` after the pool-check section:
+
+```markdown
+## PR-scoped selection
+
+`test/hil/hil_select.py` maps a diff to affected boards/tests (used by CI on PRs; fail-open
+to the full matrix). Manual use:
+
+```bash
+ARGS=$(python3 test/hil/hil_select.py --base master test/hil/tinyusb.json | python3 -c "import json,sys; print(json.load(sys.stdin)['args']['tinyusb.json'])")
+python3 test/hil/hil_test.py -B examples $ARGS test/hil/tinyusb.json
+```
+
+Unit suite: `python3 test/hil/test_hil_select.py` (no hardware).
+```
+
+- [ ] **Step 3: Full validation sweep**
+
+Run: `python3 test/hil/test_hil_select.py && python3 -m py_compile test/hil/hil_select.py test/hil/hil_examples.py test/hil/hil_ci_set_matrix.py test/hil/hil_test.py && python3 test/hil/hil_test.py --help >/dev/null && pre-commit run --files $(git diff --name-only claude/hil-pool-check..HEAD) && echo ALL-GREEN`
+Expected: `ALL-GREEN`.
+
+- [ ] **Step 4: Real-diff spot checks (acceptance)**
+
+Run each and eyeball the JSON (record outputs in your report):
+```bash
+for f in 'src/portable/raspberrypi/rp2040/dcd_rp2040.c' 'src/device/usbd.c' 'src/class/cdc/cdc_device.c' 'src/host/usbh.c'; do
+ printf '%s\n' "$f" > /tmp/d.txt
+ echo "=== $f"; python3 test/hil/hil_select.py --diff-file /tmp/d.txt test/hil/tinyusb.json test/hil/hfp.json 2>/dev/null | python3 -m json.tool | sed -n '1,25p'
+done
+```
+Expected: matches the spec's acceptance examples (pico-family only / all-device / CDC examples
+only / host side only, with hfp.json args populated only where hfp boards qualify).
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add .claude/skills/pre-pr/SKILL.md .claude/skills/hil/SKILL.md
+git commit -m "docs: pre-pr and hil skill use hil_select for PR-scoped boards"
+```
diff --git a/docs/superpowers/plans/2026-08-15-ci-hs-reset-edges.md b/docs/superpowers/plans/2026-08-15-ci-hs-reset-edges.md
new file mode 100644
index 000000000..ec0834e55
--- /dev/null
+++ b/docs/superpowers/plans/2026-08-15-ci-hs-reset-edges.md
@@ -0,0 +1,782 @@
+# Bus-Reset Edge Events + Review Fix Wave Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Give the device stack a "bus reset started" event so ci_hs can tell usbd to stand down at the URI interrupt instead of up to 50 ms later, and clear the ten findings agreed from the max review.
+
+**Architecture:** `DCD_EVENT_BUS_RESET` splits into `DCD_EVENT_BUS_RESET_START` / `_END` with a compatibility alias, so every other port stays byte-identical. `dcd_ci_hs.c`'s `bus_reset()` splits along the register/software line — registers at URI (`_START`), software structures at the port-change ending the reset (`_END`) — which eliminates the window where usbd believes it is configured over zeroed queue heads. A single bounded-flush helper absorbs the five flush sites. Seven mechanical fixes follow.
+
+**Tech Stack:** C99, TinyUSB device stack (`src/device/`), ChipIdea HS DCD (`src/portable/chipidea/ci_hs/`), NXP IP3511 DCD (`src/portable/nxp/lpc_ip3511/`), CMake+Ninja and Make builds, J-Link flashing, `test/hil/` HIL harness.
+
+## Global Constraints
+
+- Branch `fix-ci-hs` in worktree `/home/hathach/.herdr/worktrees/tinyusb/fix-ci-hs`. Do NOT push; the user pushes.
+- C99, 2-space indent, no tabs. Match each file's surrounding style (`dcd_lpc_ip3511.c` mixes styles — follow the immediate neighbourhood).
+- Commit messages: imperative mood, no `Co-Authored-By:` or `Claude-Session:` trailers (repo rule: hathach is sole author).
+- The repo pre-commit hook (trailing-whitespace, end-of-file-fixer, codespell, unique-PIDs, ceedling unit tests) must pass. If it rewrites a file, re-stage and retry the commit once.
+- Comments: short, only the non-obvious "why". Cite manuals as `UM10503 25.10.3` / `Errata LPC546xx USB.13` style — never `ES_` prefixes.
+- Never edit anything under `hw/mcu/` or `lib/` (vendor code).
+- Build commands used throughout (each ~30-60 s):
+ `cmake --build examples/cmake-build-<board>` for `mimxrt1064_evk`, `lpcxpresso18s37`, `lpcxpresso11u37`, `lpcxpresso55s28`.
+- Design source of truth: `docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md`.
+
+## File Structure
+
+| File | Responsibility in this plan |
+|---|---|
+| `src/device/dcd.h` | Event enum + compatibility alias + contract comment |
+| `src/device/usbd.c` | Handle both reset edges; log strings; stop breakpointing on DCD refusal |
+| `src/portable/chipidea/ci_hs/dcd_ci_hs.c` | Flush helper; `bus_reset()` split; setup-flush wait; `dcd_set_address`; RESUME guard |
+| `src/portable/nxp/lpc_ip3511/dcd_lpc_ip3511.c` | Torn-setup delivery; USB.13 TODO token |
+| `hw/bsp/lpc55/boards/lpcxpresso55s28/board.cmake` | Delete dead RHPORT block |
+| `hw/bsp/lpc11/boards/lpcxpresso11u37/lpc11u37.ld` | Correct stale comment; relabel ASSERT |
+
+Tasks 1-3 are ordered (each builds on the previous); Tasks 4-6 are independent of each other.
+
+---
+
+### Task 1: Split the bus-reset event into START/END edges
+
+**Files:**
+- Modify: `src/device/dcd.h` (enum at lines 23-34; contract comment above it)
+- Modify: `src/device/usbd.c` (`_usbd_event_str[]` at line 457; the `DCD_EVENT_BUS_RESET` case at line 700)
+
+**Interfaces:**
+- Produces: `DCD_EVENT_BUS_RESET_START` and `DCD_EVENT_BUS_RESET_END` enum members; `#define DCD_EVENT_BUS_RESET DCD_EVENT_BUS_RESET_END`. Task 2 emits `_START` via the existing `dcd_event_bus_signal(uint8_t rhport, dcd_eventid_t eid, bool in_isr)` and `_END` via the existing `dcd_event_bus_reset(uint8_t rhport, tusb_speed_t speed, bool in_isr)`.
+
+- [ ] **Step 1: Replace the enum member in `src/device/dcd.h`**
+
+Replace:
+
+```c
+typedef enum {
+ DCD_EVENT_INVALID = 0, // 0
+ DCD_EVENT_BUS_RESET, // 1
+ DCD_EVENT_UNPLUGGED, // 2
+ DCD_EVENT_SOF, // 3
+ DCD_EVENT_SUSPEND, // 4 TODO LPM Sleep L1 support
+ DCD_EVENT_RESUME, // 5
+ DCD_EVENT_SETUP_RECEIVED, // 6
+ DCD_EVENT_XFER_COMPLETE, // 7
+ USBD_EVENT_FUNC_CALL, // 8 Not an DCD event, just a convenient way to defer ISR function
+ DCD_EVENT_COUNT
+} dcd_eventid_t;
+```
+
+with:
+
+```c
+// Bus reset is reported as two edges. BUS_RESET_START is optional: a controller that
+// cannot tell the edges apart emits only BUS_RESET_END, which stays self-sufficient (it
+// performs the full teardown with or without a preceding START). Emit START when reset
+// signaling is detected - the link is unusable and the speed is not negotiated yet - so
+// the stack stops using endpoints immediately instead of at the end of the reset.
+typedef enum {
+ DCD_EVENT_INVALID = 0, // 0
+ DCD_EVENT_BUS_RESET_START, // 1
+ DCD_EVENT_BUS_RESET_END, // 2 with negotiated speed
+ DCD_EVENT_UNPLUGGED, // 3
+ DCD_EVENT_SOF, // 4
+ DCD_EVENT_SUSPEND, // 5 TODO LPM Sleep L1 support
+ DCD_EVENT_RESUME, // 6
+ DCD_EVENT_SETUP_RECEIVED, // 7
+ DCD_EVENT_XFER_COMPLETE, // 8
+ USBD_EVENT_FUNC_CALL, // 9 Not an DCD event, just a convenient way to defer ISR function
+ DCD_EVENT_COUNT
+} dcd_eventid_t;
+
+#define DCD_EVENT_BUS_RESET DCD_EVENT_BUS_RESET_END // backward compatibility
+```
+
+- [ ] **Step 2: Update the log-string table in `src/device/usbd.c`**
+
+At line 457 the table is indexed by event id and MUST stay in enum order. Replace the
+`"Bus Reset",` entry (line 459) with two entries:
+
+```c
+ "Bus Reset Start",
+ "Bus Reset End",
+```
+
+- [ ] **Step 3: Handle both edges in the usbd task loop**
+
+Replace the case at `src/device/usbd.c:700`:
+
+```c
+ case DCD_EVENT_BUS_RESET:
+ TU_LOG_USBD(": %s Speed\r\n", tu_str_speed[event.bus_reset.speed]);
+ usbd_reset(event.rhport);
+ _usbd_dev.speed = event.bus_reset.speed;
+ break;
+```
+
+with:
+
+```c
+ case DCD_EVENT_BUS_RESET_START:
+ TU_LOG_USBD("\r\n");
+ usbd_reset(event.rhport);
+ break;
+
+ case DCD_EVENT_BUS_RESET_END:
+ TU_LOG_USBD(": %s Speed\r\n", tu_str_speed[event.bus_reset.speed]);
+ // TODO a DCD that reports both edges pays for two teardowns: track a per-rhport
+ // "start seen" flag and skip this reset, keeping it for the single-event DCDs.
+ usbd_reset(event.rhport);
+ _usbd_dev.speed = event.bus_reset.speed;
+ break;
+```
+
+- [ ] **Step 4: Verify legacy ports still build (the alias must carry them)**
+
+Run:
+
+```bash
+cd examples && cmake -B cmake-build-stm32f407disco -DBOARD=stm32f407disco -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel . && cmake --build cmake-build-stm32f407disco
+```
+
+Expected: builds clean. This board's DCD (dwc2) still calls `dcd_event_bus_reset()`, which
+now resolves to `_END` through the unchanged helper — proving the alias works.
+
+- [ ] **Step 5: Verify the unit tests still build and pass**
+
+Run: `cd test/unit-test && ceedling test:all`
+Expected: all tests pass (they reference `DCD_EVENT_BUS_RESET` via the alias).
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add src/device/dcd.h src/device/usbd.c
+git commit -m "usbd: split bus reset into start/end edge events
+
+A DCD that can see reset signaling begin has no way to say so: the only
+event carries the negotiated speed, which does not exist until the reset
+ends. On ChipIdea that leaves the stack believing it is configured for the
+whole reset window (3 ms minimum, tens of ms in practice) while the
+controller has already torn its endpoints down.
+
+Add DCD_EVENT_BUS_RESET_START for the leading edge and rename the existing
+event to DCD_EVENT_BUS_RESET_END, keeping DCD_EVENT_BUS_RESET as an alias
+so every other port and the unit tests are untouched. START is optional and
+END stays self-sufficient, so single-event drivers keep working unchanged."
+```
+
+---
+
+### Task 2: Split ci_hs `bus_reset()` across the two edges, behind one flush helper
+
+**Files:**
+- Modify: `src/portable/chipidea/ci_hs/dcd_ci_hs.c` (`bus_reset()`; `dcd_deinit()`; `dcd_edpt_iso_activate()`; the `INTR_RESET` and `INTR_PORT_CHANGE` branches of `dcd_int_handler()`)
+
+**Interfaces:**
+- Consumes: `DCD_EVENT_BUS_RESET_START` (Task 1), `dcd_event_bus_signal()`, `dcd_event_bus_reset()`.
+- Produces: `static bool flush_endpoints(ci_hs_regs_t *dcd_reg, uint32_t mask)` — writes `ENDPTFLUSH = mask`, spins bounded by `CI_HS_BUSY_SPIN`, returns `true` if the bits cleared. Used by Task 3.
+
+- [ ] **Step 1: Add the flush helper next to `bus_reset()`**
+
+Insert above `bus_reset()`:
+
+```c
+// Flush endpoint buffers and wait for the controller to acknowledge. Callers proceed
+// regardless of the result; the bound only prevents an ISR-context hang on dead hardware.
+static bool flush_endpoints(ci_hs_regs_t *dcd_reg, uint32_t mask) {
+ dcd_reg->ENDPTFLUSH = mask;
+ uint32_t guard = CI_HS_BUSY_SPIN;
+ while (dcd_reg->ENDPTFLUSH & mask) {
+ if (!guard--) {
+ return false;
+ }
+ }
+ return true;
+}
+```
+
+- [ ] **Step 2: Split `bus_reset()` into begin/complete**
+
+Replace the whole `bus_reset()` function with these two. `bus_reset_begin()` keeps only
+register work; `bus_reset_complete()` owns everything that touches `_dcd_data`:
+
+```c
+/// Register-side reset handling, must run inside the reset window (UM10503 25.10.3)
+static void bus_reset_begin(uint8_t rhport) {
+ ci_hs_regs_t *dcd_reg = CI_HS_REG(rhport);
+
+ // The reset value for all endpoint types is the control endpoint. If one endpoint
+ // direction is enabled and the paired endpoint of opposite direction is disabled, then the
+ // endpoint type of the unused direction must be changed from the control type to any other
+ // type (e.g. bulk). Leaving an un-configured endpoint control will cause undefined behavior
+ // for the data PID tracking on the active endpoint.
+ const uint8_t ep_count = ci_ep_count(dcd_reg);
+ for (uint8_t i = 1; i < ep_count; i++) {
+ dcd_reg->ENDPTCTRL[i] = ENDPTCTRL_RESET_MASK;
+ }
+
+ //------------- Clear All Registers -------------//
+ dcd_reg->ENDPTNAK = dcd_reg->ENDPTNAK;
+ dcd_reg->ENDPTNAKEN = 0;
+ dcd_reg->ENDPTSETUPSTAT = dcd_reg->ENDPTSETUPSTAT;
+ dcd_reg->ENDPTCOMPLETE = dcd_reg->ENDPTCOMPLETE;
+
+ uint32_t guard = CI_HS_BUSY_SPIN;
+ while (dcd_reg->ENDPTPRIME && guard--) {}
+ flush_endpoints(dcd_reg, 0xFFFFFFFF);
+}
+
+/// Software-side reset handling, deferred to the port change ending the reset so the queue
+/// heads stay coherent until the stack is told - and so a prime issued by a task that had
+/// not yet seen BUS_RESET_START is flushed here rather than surviving re-enumeration.
+static void bus_reset_complete(uint8_t rhport) {
+ ci_hs_regs_t *dcd_reg = CI_HS_REG(rhport);
+ flush_endpoints(dcd_reg, 0xFFFFFFFF);
+
+ //------------- Queue Head & Queue TD -------------//
+ tu_memclr(&_dcd_data, sizeof(dcd_data_t));
+
+ //------------- Set up Control Endpoints (0 OUT, 1 IN) -------------//
+ _dcd_data.qhd[0][0].zero_length_termination = _dcd_data.qhd[0][1].zero_length_termination = 1;
+ _dcd_data.qhd[0][0].max_packet_size = _dcd_data.qhd[0][1].max_packet_size = CFG_TUD_ENDPOINT0_SIZE;
+ _dcd_data.qhd[0][0].qtd_overlay.next = _dcd_data.qhd[0][1].qtd_overlay.next = QTD_NEXT_INVALID;
+
+ _dcd_data.qhd[0][0].int_on_setup = 1; // OUT only
+
+ dcd_dcache_clean_invalidate(&_dcd_data, sizeof(dcd_data_t));
+}
+```
+
+- [ ] **Step 3: Route the two ISR branches to the new functions**
+
+In `dcd_int_handler()`, the `INTR_RESET` branch becomes:
+
+```c
+ if (int_status & INTR_RESET) {
+ bus_reset_begin(rhport);
+ _port_change_reason[rhport] = PORT_CHANGE_REASON_RESET;
+ dcd_event_bus_signal(rhport, DCD_EVENT_BUS_RESET_START, true);
+ }
+```
+
+and inside the `INTR_PORT_CHANGE` branch, the reset arm (the `else` of the resume test)
+becomes:
+
+```c
+ } else {
+ bus_reset_complete(rhport);
+ // PSPD: 0 full, 1 low, 2 high, 3 undefined (treated as full)
+ const uint32_t pspd = (dcd_reg->PORTSC1 & PORTSC1_PORT_SPEED) >> PORTSC1_PORT_SPEED_POS;
+ const tusb_speed_t speed = (pspd == 1) ? TUSB_SPEED_LOW : (pspd == 2) ? TUSB_SPEED_HIGH : TUSB_SPEED_FULL;
+ dcd_event_bus_reset(rhport, speed, true);
+ }
+```
+
+Delete the now-unused EP0 `ENDPTFLUSH` line that previously sat at the top of that arm —
+`bus_reset_complete()` flushes all endpoints.
+
+- [ ] **Step 4: Route the remaining flush sites through the helper**
+
+In `dcd_deinit()`, replace the flush block with:
+
+```c
+ // flush all endpoints
+ uint32_t guard = CI_HS_BUSY_SPIN;
+ while (dcd_reg->ENDPTPRIME && guard--) {}
+ flush_endpoints(dcd_reg, 0xFFFFFFFF);
+```
+
+In `dcd_edpt_iso_activate()`, replace the flush + spin with:
+
+```c
+ // Flush EP
+ flush_endpoints(dcd_reg, TU_BIT(epnum + (dir ? 16 : 0)));
+```
+
+- [ ] **Step 5: Build both ci_hs board families**
+
+Run:
+
+```bash
+cmake --build examples/cmake-build-mimxrt1064_evk && cmake --build examples/cmake-build-lpcxpresso18s37
+```
+
+Expected: both succeed with no new warnings.
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add src/portable/chipidea/ci_hs/dcd_ci_hs.c
+git commit -m "dcd(ci_hs): report bus reset start at URI, finish at port change
+
+The RM wants the reset cleanup inside the reset window, but the negotiated
+speed only exists once the port reaches its operational state, so the stack
+was told nothing for the whole window - it kept believing it was configured
+while the queue heads had been zeroed under it, and a transfer a class
+driver started in that gap stayed primed across re-enumeration.
+
+Split the work along the register/software line: bus_reset_begin() does the
+register cleanup at URI and signals BUS_RESET_START, bus_reset_complete()
+re-flushes, resets the queue heads and reports BUS_RESET_END with the final
+speed at the port change. Zeroing the queue heads now happens in the same
+breath as telling the stack, and the second flush retires anything primed
+in between.
+
+Fold the five hand-rolled endpoint flushes into one bounded helper while
+the reset path is open."
+```
+
+---
+
+### Task 3: Make the setup-time EP0 flush wait, and stop dropping the SET_ADDRESS status prime
+
+**Files:**
+- Modify: `src/portable/chipidea/ci_hs/dcd_ci_hs.c` (`dcd_set_address()`; the `ENDPTSETUPSTAT` branch inside `dcd_int_handler()`)
+
+**Interfaces:**
+- Consumes: `flush_endpoints()` (Task 2); `qhd_start_xfer()` returning `bool`, already propagated by `dcd_edpt_xfer()`.
+
+- [ ] **Step 1: Wait for the setup-time flush to complete**
+
+In the ISR's setup branch, replace the fire-and-forget flush line
+
+```c
+ dcd_reg->ENDPTFLUSH = TU_BIT(0) | TU_BIT(16);
+```
+
+with
+
+```c
+ // Wait it out: the flush retires a status/handshake phase left primed by the previous
+ // control sequence (UM10503 25.10.8.1.1), and an unfinished flush would otherwise
+ // still be asserted when the task primes the response to this setup and would retire
+ // that instead. A flush waits for any packet already in progress - microseconds at
+ // high speed - and the guard caps wedged hardware.
+ flush_endpoints(dcd_reg, TU_BIT(0) | TU_BIT(16));
+```
+
+- [ ] **Step 2: Honour the status-prime result in `dcd_set_address`**
+
+Replace the body of `dcd_set_address()`:
+
+```c
+void dcd_set_address(uint8_t rhport, uint8_t dev_addr) {
+ // Response with status first before changing device address. A refused prime means a new
+ // setup superseded this transfer; staging an address whose ACK will never arrive would
+ // leave the device answering on it, so only arm the address when the status went out.
+ if (dcd_edpt_xfer(rhport, tu_edpt_addr(0, TUSB_DIR_IN), NULL, 0, false)) {
+ ci_hs_regs_t *dcd_reg = CI_HS_REG(rhport);
+ dcd_reg->DEVICEADDR = (dev_addr << 25) | TU_BIT(24);
+ }
+}
+```
+
+- [ ] **Step 3: Build and commit**
+
+Run: `cmake --build examples/cmake-build-mimxrt1064_evk && cmake --build examples/cmake-build-lpcxpresso18s37`
+Expected: both succeed.
+
+```bash
+git add src/portable/chipidea/ci_hs/dcd_ci_hs.c
+git commit -m "dcd(ci_hs): wait out the setup flush, honour the set-address prime
+
+The flush issued on every new setup was fire-and-forget. A flush waits for
+a packet already in progress, so it could still be asserted when the task
+primed the response to that setup and retire the fresh prime instead -
+leaving EP0 silent until the host gave up.
+
+dcd_set_address() also armed DEVICEADDR unconditionally, but the status
+prime can now be refused when a newer setup supersedes the transfer; the
+address was then staged behind an ACK that never came and the device sat at
+address 0. Only arm it when the status transfer actually started."
+```
+
+---
+
+### Task 4: Emit RESUME only when the port really left suspend
+
+**Files:**
+- Modify: `src/portable/chipidea/ci_hs/dcd_ci_hs.c` (the resume arm of the `INTR_PORT_CHANGE` branch in `dcd_int_handler()`)
+
+**Interfaces:** none consumed or produced.
+
+- [ ] **Step 1: Restore the hardware guard**
+
+In the `INTR_PORT_CHANGE` branch, the resume arm currently reads:
+
+```c
+ if (pci_reason == PORT_CHANGE_REASON_SUSPEND) {
+ dcd_event_bus_signal(rhport, DCD_EVENT_RESUME, true);
+ } else {
+```
+
+Replace that condition with one that also consults live hardware:
+
+```c
+ if (pci_reason == PORT_CHANGE_REASON_SUSPEND) {
+ // Only when the port actually left suspend: a starved snapshot can hold the resume's
+ // port change together with a second suspend, and reporting a resume there would
+ // leave the stack awake on a sleeping bus with no further event to correct it.
+ if (!(dcd_reg->PORTSC1 & PORTSC1_SUSPEND)) {
+ dcd_event_bus_signal(rhport, DCD_EVENT_RESUME, true);
+ }
+ } else {
+```
+
+- [ ] **Step 2: Build and commit**
+
+Run: `cmake --build examples/cmake-build-mimxrt1064_evk && cmake --build examples/cmake-build-lpcxpresso18s37`
+Expected: both succeed.
+
+```bash
+git add src/portable/chipidea/ci_hs/dcd_ci_hs.c
+git commit -m "dcd(ci_hs): only report resume when the port left suspend
+
+A suspend, resume and second suspend collapsed into one interrupt pass
+queued suspend then resume from the recorded cause alone, so the stack
+ended up awake while the bus was still suspended and nothing arrived to
+correct it. Consult PORTSC1 before reporting the resume."
+```
+
+---
+
+### Task 5: ip3511 — never deliver a knowingly-torn setup packet
+
+**Files:**
+- Modify: `src/portable/nxp/lpc_ip3511/dcd_lpc_ip3511.c` (setup branch of `dcd_int_handler()`; the `dcd_edpt_clear_stall()` comment)
+
+**Interfaces:** none consumed or produced.
+
+- [ ] **Step 1: Deliver only when the copy is known good**
+
+Replace:
+
+```c
+ // a SETUP that raced in after the acks (its bit0 consumed above, this copy possibly torn):
+ // its latch is visible again - re-raise the endpoint interrupt so the next pass redelivers
+ // the newer payload
+ if (dcd_reg->DEVCMDSTAT & DEVCMDSTAT_SETUP_RECEIVED_MASK) {
+ dcd_reg->INTSETSTAT = TU_BIT(0);
+ }
+
+ dcd_event_setup_received(rhport, setup_copy, true);
+```
+
+with:
+
+```c
+ // a SETUP that raced in after the acks (its bit0 consumed above) makes this copy suspect:
+ // its latch is visible again, so re-raise the endpoint interrupt and let the next pass
+ // deliver the newer payload rather than passing up bytes that may be torn between the two
+ if (dcd_reg->DEVCMDSTAT & DEVCMDSTAT_SETUP_RECEIVED_MASK) {
+ dcd_reg->INTSETSTAT = TU_BIT(0);
+ } else {
+ dcd_event_setup_received(rhport, setup_copy, true);
+ }
+```
+
+- [ ] **Step 2: Add the TODO token to the USB.13 deferral**
+
+In `dcd_edpt_clear_stall()`, change the caveat's opening line from
+
+```c
+ // Known caveat (Errata LPC546xx USB.13, same semantics in UM11126): with RF/TV preserved at 1, TR
+```
+
+to
+
+```c
+ // TODO implement the Errata LPC546xx USB.13 work-around (same semantics in UM11126): with RF/TV preserved at 1, TR
+```
+
+- [ ] **Step 3: Build and commit**
+
+Run: `cmake --build examples/cmake-build-lpcxpresso11u37 && cmake --build examples/cmake-build-lpcxpresso55s28`
+Expected: both succeed.
+
+```bash
+git add src/portable/nxp/lpc_ip3511/dcd_lpc_ip3511.c
+git commit -m "dcd(ip3511): drop a setup packet the hardware may have overwritten
+
+The handler already notices when a new setup landed while it was copying
+the previous one, and re-raises the endpoint interrupt so the newer payload
+is delivered next pass - but it then passed the suspect copy up anyway.
+Usually harmless, since the redelivery supersedes it, but if that second
+event cannot be queued the torn bytes are processed as a real request.
+Deliver the copy only when no newer setup is pending."
+```
+
+---
+
+### Task 6: BSP cleanups — dead RHPORT block and the stale linker comment
+
+**Files:**
+- Modify: `hw/bsp/lpc55/boards/lpcxpresso55s28/board.cmake`
+- Modify: `hw/bsp/lpc11/boards/lpcxpresso11u37/lpc11u37.ld`
+
+**Interfaces:** none consumed or produced.
+
+- [ ] **Step 1: Delete the redundant RHPORT block**
+
+`hw/bsp/lpc55/family.cmake` already applies the identical guarded defaults (`RHPORT_DEVICE 1`,
+`RHPORT_HOST 0`) after including the board file, so remove these lines from
+`board.cmake` entirely:
+
+```cmake
+# device highspeed, host fullspeed; guarded so a -D override on the cmake command line wins
+if (NOT DEFINED RHPORT_DEVICE)
+ set(RHPORT_DEVICE 1)
+endif ()
+if (NOT DEFINED RHPORT_HOST)
+ set(RHPORT_HOST 0)
+endif ()
+```
+
+Leave `board.mk`'s `RHPORT_DEVICE ?= 1` / `RHPORT_HOST ?= 0` alone — `?=` is the idiomatic
+Make form and matches sibling boards.
+
+- [ ] **Step 2: Prove the defaults and the override still work**
+
+Run:
+
+```bash
+cd examples && rm -rf /tmp/rh-default /tmp/rh-override
+cmake -B /tmp/rh-default -DBOARD=lpcxpresso55s28 -G Ninja . > /tmp/rh-default.log 2>&1
+grep -m1 "RHPORT_DEVICE" /tmp/rh-default.log || cmake -B /tmp/rh-default -DBOARD=lpcxpresso55s28 -G Ninja -LA . | grep -E "^RHPORT_(DEVICE|HOST)"
+cmake -B /tmp/rh-override -DBOARD=lpcxpresso55s28 -DRHPORT_DEVICE=0 -DRHPORT_HOST=1 -G Ninja -LA . | grep -E "^RHPORT_(DEVICE|HOST)"
+```
+
+Expected: the default configure yields device 1 / host 0; the override configure yields
+device 0 / host 1. Then rebuild the real tree: `cmake --build cmake-build-lpcxpresso55s28`.
+
+- [ ] **Step 3: Correct the linker-script comment and relabel the ASSERT**
+
+In `lpc11u37.ld`, replace the comment block above `__user_stack_top` and the ASSERT with:
+
+```text
+ /* Main (MSP/ISR) stack lives at the top of the USB SRAM bank: the 8K main bank is packed so
+ tight that only ~280 B remained above .bss, and ISR frames overflowed into the topmost task
+ stack (cdc_msc_freertos hard fault). Nothing else is placed in this bank in either build
+ system, so the stack owns all 2 KB; the ASSERT is future-proofing in case USB buffers are
+ ever mapped here again. */
+ __user_stack_top = ORIGIN(RamUsb2) + LENGTH(RamUsb2);
+ ASSERT(__user_stack_top - (ADDR(.noinit_RAM2) + SIZEOF(.noinit_RAM2)) >= 0x200,
+ "main stack headroom in RamUsb2 below 512 bytes")
+```
+
+- [ ] **Step 4: Build both build systems for lpc11u37**
+
+Run:
+
+```bash
+cmake --build examples/cmake-build-lpcxpresso11u37
+cd examples/device/cdc_msc_freertos && make -j8 BOARD=lpcxpresso11u37 all && cd ../../..
+```
+
+Expected: both succeed.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add hw/bsp/lpc55/boards/lpcxpresso55s28/board.cmake hw/bsp/lpc11/boards/lpcxpresso11u37/lpc11u37.ld
+git commit -m "bsp: drop duplicated lpc55s28 rhport defaults, fix lpc11u37 comment
+
+hw/bsp/lpc55/family.cmake already applies the same guarded rhport defaults
+after including the board file, so the board-level copy only added a second
+place to keep in sync.
+
+The lpc11u37 linker comment still described USB buffers living in RamUsb2,
+a placement the same branch removed; nothing lands there now, so say so and
+label the headroom assert as future-proofing."
+```
+
+---
+
+### Task 7: Stop halting the target when a DCD legitimately refuses a transfer
+
+**Files:**
+- Modify: `src/device/usbd.c` (`usbd_edpt_xfer()` failure arm)
+
+**Interfaces:** none consumed or produced.
+
+- [ ] **Step 1: Remove the breakpoint from the DCD-refusal path**
+
+Replace the failure arm of `usbd_edpt_xfer()`:
+
+```c
+ } else {
+ // DCD error, mark endpoint as ready to allow next transfer
+ _usbd_dev.ep_status[epnum][dir] &= (uint8_t) ~(TU_EDPT_STATE_BUSY | TU_EDPT_STATE_CLAIMED);
+ TU_LOG_USBD("FAILED\r\n");
+ TU_BREAKPOINT();
+ return false;
+ }
+```
+
+with:
+
+```c
+ } else {
+ // Driver refused the transfer, mark endpoint as ready to allow next transfer. This is a
+ // recoverable condition (e.g. a new setup superseding a control response), not a bug, so
+ // do not break into the debugger - TU_BREAKPOINT() halts the CPU whenever a probe is
+ // attached, which on a test rig is always.
+ _usbd_dev.ep_status[epnum][dir] &= (uint8_t) ~(TU_EDPT_STATE_BUSY | TU_EDPT_STATE_CLAIMED);
+ TU_LOG_USBD("FAILED\r\n");
+ return false;
+ }
+```
+
+- [ ] **Step 2: Confirm no other stack path relies on that breakpoint**
+
+Run: `grep -n "TU_BREAKPOINT" src/device/*.c src/device/*.h`
+Expected: no remaining hits inside `usbd_edpt_xfer`; other occurrences (if any) are in
+unrelated assert macros and stay as they are.
+
+- [ ] **Step 3: Build and run unit tests**
+
+Run:
+
+```bash
+cmake --build examples/cmake-build-mimxrt1064_evk
+cd test/unit-test && ceedling test:all && cd ../..
+```
+
+Expected: build succeeds, all unit tests pass.
+
+- [ ] **Step 4: Commit**
+
+```bash
+git add src/device/usbd.c
+git commit -m "usbd: do not breakpoint when a driver refuses a transfer
+
+TU_BREAKPOINT() is not gated on CFG_TUSB_DEBUG - it halts the CPU whenever
+a debugger is attached, which on a test rig is always. A driver declining a
+transfer is recoverable (a new setup superseding a control response, for
+one) and the endpoint is already released for the retry, so a halted target
+turns a self-healing case into a dead board."
+```
+
+---
+
+### Task 8: Full validation on hardware
+
+**Files:** none modified — this task produces the evidence for the PR description.
+
+**Interfaces:** consumes the firmware built by Tasks 1-7.
+
+- [ ] **Step 1: Software gate**
+
+Run:
+
+```bash
+pre-commit run --all-files
+cd examples
+for b in mimxrt1064_evk lpcxpresso18s37 lpcxpresso11u37 lpcxpresso55s28; do
+ rm -rf cmake-build-$b && cmake -B cmake-build-$b -DBOARD=$b -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel . && cmake --build cmake-build-$b || echo "FAILED $b"
+done
+cd ..
+```
+
+Expected: pre-commit all green; all four boards build every example.
+
+- [ ] **Step 2: Make-build regression checks**
+
+Run:
+
+```bash
+cd examples/host/cdc_msc_hid && make -j8 BOARD=lpcxpresso55s28 all && cd ../../..
+cd examples/device/cdc_msc_throughput && make -j8 BOARD=lpcxpresso11u37 all && cd ../../..
+```
+
+Expected: both link (these two were broken earlier in the branch and are the regression
+canaries for the BSP changes).
+
+- [ ] **Step 3: Flash with verification (mandatory)**
+
+The mimxrt1064_evk has twice accepted a flash that silently did not take, so every load in
+this task uses `verifyfile`. For each board, write a J-Link script of this shape and run it:
+
+```
+r
+h
+loadfile examples/cmake-build-<board>/device/usbtest/usbtest.elf
+verifyfile examples/cmake-build-<board>/device/usbtest/usbtest.elf
+r
+g
+qc
+```
+
+Probes and devices: `mimxrt1064_evk` = `-USB 000725299165 -device MIMXRT1064xxx6A`,
+`lpcxpresso55s28` = `-USB 000727031389 -device LPC55S28`,
+`lpcxpresso11u37` = `-USB 000724441579 -device LPC11U37/401`.
+Invoke as `JLinkExe <probe/device args> -if swd -speed 4000 -autoconnect 1 -NoGui 1 -CommandFile <script>`.
+Expected: `Verify` reports O.K. and the board re-enumerates as `cafe:4010` with its own
+serial before any test runs.
+
+- [ ] **Step 4: HIL batteries and stress**
+
+Hold each board's lock for its own leg (`python3 test/hil/hil_lock.py hold <board> --reason "reset-edge validation"`,
+release after), never run two batteries at once, and abort if CI is active
+(`pgrep -f "hil_test.py [-]-retry"`).
+
+```bash
+# per board: full battery
+timeout 700 python3 test/hil/usbtest.py --serial <serial> --json --keep-binding --timeout 60
+
+# mimxrt1064_evk only: queued-control stress and the unlink storm
+for i in $(seq 1 50); do timeout 200 python3 test/hil/usbtest.py --serial BAE96FB95AFA6DBB8F00005002001200 --tests 9,10 --json --keep-binding --timeout 60 > /dev/null || break; done
+for i in $(seq 1 10); do timeout 300 python3 test/hil/usbtest.py --serial BAE96FB95AFA6DBB8F00005002001200 --tests 11,12,24 --json --keep-binding --timeout 60 > /dev/null || break; done
+```
+
+Serials: 1064 `BAE96FB95AFA6DBB8F00005002001200`, 55s28 `2BF1839A7D51F553A15AB03FD08F70AB`,
+11u37 `17121919`.
+Expected: 30/30 on all three boards, 50/50 and 10/10 loops, and
+`ps -eo stat,comm | awk '$1 ~ /^D/'` empty after each leg.
+
+- [ ] **Step 5: Reset-path evidence with logging**
+
+Build and flash `device/cdc_msc` for `mimxrt1064_evk` with `-DLOG=2 -DLOGGER=rtt`, capture
+RTT during one unplug/replug cycle (`timeout 20s JLinkRTTClient > /tmp/reset.log`), then:
+
+```bash
+grep -cE "Bus Reset Start" /tmp/reset.log
+grep -cE "Bus Reset End" /tmp/reset.log
+grep -c "Resume" /tmp/reset.log
+```
+
+Expected: equal non-zero counts for start and end (one pair per enumeration) and no
+`Resume` lines during a plain plug-in.
+
+- [ ] **Step 6: Suspend/resume pairing**
+
+With the same RTT build attached, suspend the port from the host and resume it:
+
+```bash
+# find the 1064's busport, then:
+echo auto | sudo tee /sys/bus/usb/devices/<busport>/power/control
+sleep 5
+echo on | sudo tee /sys/bus/usb/devices/<busport>/power/control
+```
+
+Expected in the log: one `Suspend` followed by one `Resume`, and no `Bus Reset` of either
+edge from the suspend cycle alone.
+
+- [ ] **Step 7: Record the evidence**
+
+Append the numbers from Steps 1-6 to the PR description draft. No commit.
+
+## Self-Review
+
+**Spec coverage:** §1 event split → Task 1. §2 ci_hs bus_reset split → Task 2. §3 flush
+helper → Task 2 (Steps 1, 4). §4 mechanical: setup-flush wait and `dcd_set_address` → Task 3;
+RESUME guard → Task 4; ip3511 torn setup and USB.13 TODO → Task 5; usbd breakpoint → Task 7;
+BSP pair → Task 6. Verification matrix → Task 8 (legacy-DCD build guard is Task 1 Step 4).
+Deferred items are deliberately absent from every task. No gaps.
+
+**Placeholder scan:** no TBD/TODO-as-placeholder; the two literal `TODO` strings are
+deliverable code comments (Task 1 Step 3, Task 5 Step 2). Every code step carries the exact
+text to write; every run step carries the command and expected result.
+
+**Type consistency:** `flush_endpoints(ci_hs_regs_t *dcd_reg, uint32_t mask) -> bool` is
+defined in Task 2 Step 1 and used with that exact signature in Task 2 Steps 2/4 and Task 3
+Step 1. `DCD_EVENT_BUS_RESET_START` / `_END` are defined in Task 1 and used in Task 2 Step 3
+via `dcd_event_bus_signal()` / `dcd_event_bus_reset()`, whose signatures are quoted in Task 1's
+Interfaces block. `bus_reset_begin()` / `bus_reset_complete()` are defined and called with
+matching names in Task 2.
diff --git a/docs/superpowers/plans/2026-08-16-drop-ep0-prime-verify.md b/docs/superpowers/plans/2026-08-16-drop-ep0-prime-verify.md
new file mode 100644
index 000000000..aa999c9e3
--- /dev/null
+++ b/docs/superpowers/plans/2026-08-16-drop-ep0-prime-verify.md
@@ -0,0 +1,314 @@
+# Drop the EP0 Post-Prime Verify Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Remove the EP0 post-prime verification that was built on a theory the RT106x endpoint-conflict errata has superseded, and prove on hardware that nothing depended on it.
+
+**Architecture:** One deletion in `qhd_start_xfer()`, then a rebase onto current master, then an A/B validation whose "with it" arm is already banked (10x 30/30 batteries plus 40 targeted loops on 2026-08-16). No interfaces change: the pre-prime setup-lockout guard keeps `qhd_start_xfer()` returning `bool`, so `dcd_set_address()`'s gating and usbd's failure path stay exactly as they are.
+
+**Tech Stack:** C99, TinyUSB ChipIdea HS DCD (`src/portable/chipidea/ci_hs/`), CMake+Ninja and Make builds, J-Link (JLinkExe V9.66), `test/hil/usbtest.py` driving the Linux testusb battery.
+
+## Global Constraints
+
+- Branch `fix-ci-hs` in worktree `/home/hathach/.herdr/worktrees/tinyusb/fix-ci-hs`. Do NOT push; the user pushes.
+- C99, 2-space indent. Commit messages imperative, no `Co-Authored-By:` or `Claude-Session:` trailers (repo rule: hathach is sole author).
+- Pre-commit hook (trailing-whitespace, end-of-file-fixer, codespell, unique-PIDs, ceedling) must pass; if it rewrites a file, re-stage and retry the commit once.
+- Never edit anything under `hw/mcu/` or `lib/` (vendor code).
+- Rig etiquette: hold the board lock for hardware work (`python3 test/hil/hil_lock.py hold <board> --reason "..."`, release after); abort if CI is active (`pgrep -f "hil_test.py [-]-retry"`); NEVER use `uhubctl`, `pci-reset` or `pci-rebind`; never touch the actions-runner.
+- JLinkExe on this rig is **V9.66 and has no `verifyfile` command** — use `loadfile` (built-in Program & Verify) plus a mandatory enumeration check.
+- Board facts: `mimxrt1064_evk`, serial `BAE96FB95AFA6DBB8F00005002001200`, J-Link probe `000725299165`, device `MIMXRT1064xxx6A`, expected `cafe:4010`.
+- Design source of truth: `docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md`.
+
+## File Structure
+
+| File | Responsibility in this plan |
+|---|---|
+| `src/portable/chipidea/ci_hs/dcd_ci_hs.c` | The only code change: delete the post-prime block in `qhd_start_xfer()` |
+
+Tasks 2 and 3 change no files; they rebase and validate.
+
+---
+
+### Task 1: Delete the EP0 post-prime verify
+
+**Files:**
+- Modify: `src/portable/chipidea/ci_hs/dcd_ci_hs.c` (the tail of `qhd_start_xfer()`)
+
+**Interfaces:**
+- Produces: `qhd_start_xfer()` keeps its existing signature `static bool qhd_start_xfer(uint8_t rhport, uint8_t epnum, uint8_t dir)` and still returns `false` from the pre-prime setup-lockout guard. No caller changes.
+
+- [ ] **Step 1: Apply the deletion**
+
+In `qhd_start_xfer()`, replace this (everything from the prime write to the closing `return true;`):
+
+```c
+ // start transfer
+ const uint32_t prime_bit = TU_BIT(epnum + (dir ? 16 : 0));
+ dcd_reg->ENDPTPRIME = prime_bit;
+
+ if (epnum == 0) {
+ // RM (RT1050 RM Executing a Transfer / UM10503 25.10.8): after priming EP0 the DCD must
+ // verify the prime completed - ENDPTPRIME bit clear AND the buffer reported ready in
+ // ENDPTSTAT - because the controller silently cancels an EP0 prime when a SETUP arrives
+ // during the prime operation. An undetected drop NAK-parks the endpoint forever: usbd never
+ // re-primes a busy endpoint. A very fast transfer may already have completed and retired the
+ // ENDPTSTAT bit, so ENDPTCOMPLETE also counts as the prime having taken.
+ uint32_t guard = CI_HS_BUSY_SPIN;
+ while (dcd_reg->ENDPTPRIME & prime_bit) {
+ if (!guard--) {
+ dcd_reg->ENDPTFLUSH = prime_bit; // never leave a wedged prime armed over a freed buffer
+ return false;
+ }
+ }
+ // Fail only when the cancel-cause is visibly pending: a completed transfer can have both
+ // status bits already retired by the ISR, and a cancel whose SETUP the ISR consumed is
+ // re-driven by that queued SETUP event anyway.
+ if (!((dcd_reg->ENDPTSTAT | dcd_reg->ENDPTCOMPLETE) & prime_bit) &&
+ (dcd_reg->ENDPTSETUPSTAT & TU_BIT(0))) {
+ return false; // prime cancelled (setup mid-prime): the pending SETUP re-drives EP0
+ }
+ }
+ return true;
+```
+
+with:
+
+```c
+ // start transfer
+ dcd_reg->ENDPTPRIME = TU_BIT(epnum + (dir ? 16 : 0));
+ return true;
+```
+
+Leave the `if (epnum == 0)` setup-lockout block ABOVE the prime write completely untouched —
+that one spins on `ENDPTSETUPSTAT` before priming and is required by UM10503 25.10.8.1.1
+step 4.
+
+- [ ] **Step 2: Confirm nothing else referenced the removed code**
+
+Run:
+
+```bash
+grep -n "ENDPTSTAT\|ENDPTCOMPLETE\|prime_bit" src/portable/chipidea/ci_hs/dcd_ci_hs.c
+```
+
+Expected: no `prime_bit` hits at all; `ENDPTCOMPLETE` hits only in `bus_reset_begin()` and the
+`INTR_USB` branch of `dcd_int_handler()`; `ENDPTSTAT` hits only in `ci_hs_type.h`-style register
+declarations if any appear — none inside `qhd_start_xfer()`.
+
+- [ ] **Step 3: Build both ci_hs board families**
+
+Run:
+
+```bash
+cmake --build examples/cmake-build-mimxrt1064_evk && cmake --build examples/cmake-build-lpcxpresso18s37
+```
+
+Expected: both succeed, no new warnings (in particular no "unused variable" for anything the
+deletion orphaned).
+
+- [ ] **Step 4: Commit**
+
+```bash
+git add src/portable/chipidea/ci_hs/dcd_ci_hs.c
+git commit -m "dcd(ci_hs): drop the EP0 post-prime verify
+
+The verify came from a theory that a setup arriving mid-prime silently
+cancels an EP0 prime, which was how the recurring wedge on the test rig
+looked at the time. The wedge turned out to be Errata i.MX RT1064_A
+ERR050101: with an isochronous IN endpoint active, an IN token to that
+endpoint number on another device sharing the host unprimes one of our OUT
+endpoints, undetectably and with no interrupt. Moving the usbtest iso IN
+endpoint clear of the conflict fixed it - 340 runs where the board used to
+wedge within hours.
+
+The capture that motivated the verify (EP0 status stage armed but unprimed,
+device a control transfer ahead of the host) is explained by that errata
+just as well, because it covers control OUT endpoints and a control status
+stage is one. So the verify has no independent evidence behind it, while it
+does cost two register spins on every EP0 transfer and can misread a
+transfer the interrupt handler already completed as a cancelled prime.
+
+The setup-lockout check before priming stays - that one is in the manual."
+```
+
+---
+
+### Task 2: Rebase onto current master and re-run the software gates
+
+**Files:** none modified by hand.
+
+**Interfaces:** none.
+
+- [ ] **Step 1: Rebase**
+
+Master has advanced (midi2/usbtmc/video changes) since this branch last rebased. Validating a
+tree that is not the one being merged would be a false pass.
+
+```bash
+git fetch origin master
+git rebase origin/master
+```
+
+Expected: clean rebase. If a conflict appears in `src/portable/chipidea/ci_hs/dcd_ci_hs.c` or
+`src/device/usbd.c`, resolve it hunk-by-hunk keeping BOTH sides' intent (never `git checkout
+--theirs/--ours` on a whole file), then `git rebase --continue`.
+
+- [ ] **Step 2: Rebuild everything from scratch**
+
+```bash
+cd examples
+for b in mimxrt1064_evk lpcxpresso18s37 lpcxpresso11u37 lpcxpresso55s28; do
+ rm -rf cmake-build-$b
+ cmake -B cmake-build-$b -DBOARD=$b -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel . && cmake --build cmake-build-$b || echo "FAILED $b"
+done
+cd ..
+```
+
+Expected: all four boards build every example, no "FAILED" line.
+
+- [ ] **Step 3: Make link canaries**
+
+```bash
+cd examples/host/cdc_msc_hid && make -j8 BOARD=lpcxpresso55s28 all && cd ../../..
+cd examples/device/cdc_msc_throughput && make -j8 BOARD=lpcxpresso11u37 all && cd ../../..
+```
+
+Expected: both link. These two were broken earlier in the branch's life and are the regression
+canaries for the BSP changes.
+
+- [ ] **Step 4: Unit tests and pre-commit**
+
+```bash
+cd test/unit-test && ceedling test:all && cd ../..
+pre-commit run --all-files
+```
+
+Expected: all unit tests pass; every pre-commit hook passes.
+
+- [ ] **Step 5: No commit**
+
+This task produces no commit of its own — the rebase rewrites existing commits and the builds
+are throwaway. Record the resulting HEAD hash in the report for Task 3 to reference.
+
+---
+
+### Task 3: Hardware A/B on mimxrt1064_evk
+
+**Files:** none modified — this task produces the evidence.
+
+**Interfaces:** consumes the firmware built in Task 2 at
+`examples/cmake-build-mimxrt1064_evk/device/usbtest/usbtest.elf`.
+
+Only this board is tested: it is the sole ci_hs board on the rig. The lpcxpresso55s28 and
+lpcxpresso11u37 run the ip3511 driver, which this change does not touch.
+
+- [ ] **Step 1: Preconditions**
+
+```bash
+pgrep -f "hil_test.py [-]-retry" && echo "CI ACTIVE - wait" || echo "CI idle"
+ps -eo stat,pid,etimes,comm | awk '$1 ~ /^D/'
+python3 test/hil/hil_lock.py hold mimxrt1064_evk --reason "prime-verify removal A/B"
+```
+
+Expected: CI idle, no pre-existing D-state processes, lock acquired. If CI is active, wait for
+it to drain rather than running concurrently.
+
+- [ ] **Step 2: Flash with verification**
+
+```bash
+cat > /tmp/pv.jlink <<'EOF'
+r
+h
+loadfile examples/cmake-build-mimxrt1064_evk/device/usbtest/usbtest.elf
+r
+g
+qc
+EOF
+JLinkExe -device MIMXRT1064xxx6A -if SWD -speed 4000 -SelectEmuBySN 000725299165 \
+ -autoconnect 1 -nogui 1 -CommandFile /tmp/pv.jlink
+```
+
+Expected: `Program & Verify` reports O.K.
+
+- [ ] **Step 3: Confirm the right image is actually running**
+
+```bash
+sleep 5
+grep -l BAE96FB95AFA6DBB8F00005002001200 /sys/bus/usb/devices/*/serial
+sudo lsusb -v -d cafe:4010 2>/dev/null | grep -A3 "Isochronous" | grep bEndpointAddress
+```
+
+Expected: the board is present, and the iso IN endpoint reads **0x87**. If it reads 0x83 the
+flash did not take (this board has silently no-op'd a flash twice) — reflash and re-check
+before running anything.
+
+- [ ] **Step 4: 5x full battery**
+
+```bash
+for i in $(seq 1 5); do
+ timeout 700 python3 test/hil/usbtest.py --serial BAE96FB95AFA6DBB8F00005002001200 \
+ --json --keep-binding --timeout 60 2>/dev/null | python3 -c "
+import json,sys
+d=json.load(sys.stdin)
+bad=[str(c['num']) for c in d['cases'] if c['status']!='PASS']
+print(f\"run: {d['passed']}/30 speed={d['speed']}\" + (' FAILED:'+','.join(bad) if bad else ''))
+"
+ ps -eo stat,pid,etimes,comm | awk '$1 ~ /^D/ && $4=="testusb"'
+done
+```
+
+Expected: five lines each reading `30/30 speed=480`, and no testusb D-state line between runs.
+
+- [ ] **Step 5: 15x control-focused loop**
+
+These are the paths the removed verify actually protected — queued control, the ch9 subset, and
+both ctrl_out cases. A full battery samples each only once per run.
+
+```bash
+PASS=0
+for i in $(seq 1 15); do
+ timeout 300 python3 test/hil/usbtest.py --serial BAE96FB95AFA6DBB8F00005002001200 \
+ --tests 9,10,14,21 --json --keep-binding --timeout 60 >/dev/null 2>&1 && PASS=$((PASS+1)) || { echo "FAILED at iteration $i"; break; }
+ D=$(ps -eo stat,comm | awk '$1 ~ /^D/ && $2=="testusb"' | wc -l)
+ [ "$D" != "0" ] && { echo "D-STATE at iteration $i"; break; }
+done
+echo "control loops: $PASS/15"
+```
+
+Expected: `control loops: 15/15`, no FAILED or D-STATE line.
+
+- [ ] **Step 6: Release the lock and record**
+
+```bash
+python3 test/hil/hil_lock.py release mimxrt1064_evk
+ps -eo stat,pid,etimes,comm | awk '$1 ~ /^D/'
+```
+
+Expected: lock released, no leftover D-state.
+
+**Acceptance:** 5/5 batteries at 30/30, 15/15 control loops, no `testusb` D-state outliving its
+case runtime.
+
+**Rollback trigger:** any control-case failure (errno 110 or 71 on cases 9, 10, 14, 21) or a
+lingering D-state means the verify was load-bearing after all. In that case: `git revert` the
+Task 1 commit, re-run Steps 4-5 to confirm the failure disappears, and record the result — that
+is a finding worth keeping, not a setback to hide.
+
+---
+
+## Self-Review
+
+**Spec coverage:** the spec's change section → Task 1; "rebase first, then rebuild" → Task 2
+Steps 1-2; software gates → Task 2 Steps 3-4; hardware preconditions, verified flash and the
+0x87 descriptor check → Task 3 Steps 1-3; 5x battery and 15x control loop → Task 3 Steps 4-5;
+acceptance and rollback trigger → Task 3's closing block. The spec's "deliberately kept" list is
+enforced negatively by Task 1 Step 1's instruction to leave the setup-lockout block untouched
+and by Task 1 Step 2's grep. No gaps.
+
+**Placeholder scan:** no TBD/TODO/"handle edge cases"; every step carries its exact command or
+code and its expected result.
+
+**Type consistency:** `qhd_start_xfer(uint8_t rhport, uint8_t epnum, uint8_t dir) -> bool` is
+unchanged by this plan and no caller is touched, so there are no cross-task signatures to
+reconcile. The only removed identifier, `prime_bit`, is local to the deleted block and Task 1
+Step 2 greps to confirm it has no remaining references.
diff --git a/docs/superpowers/specs/2026-07-09-claude-agents-workflows-design.md b/docs/superpowers/specs/2026-07-09-claude-agents-workflows-design.md
index 63788720c..3035723c4 100644
--- a/docs/superpowers/specs/2026-07-09-claude-agents-workflows-design.md
+++ b/docs/superpowers/specs/2026-07-09-claude-agents-workflows-design.md
@@ -29,10 +29,12 @@ Layered: **agents** (who does the work, with baked-in domain knowledge) ×
### Worker agents — `.claude/agents/*.md`
-Tiered models (owner revision 2026-07-09; originally all-opus): `port-dev`
-and `driver-reviewer` on **opus** at **xhigh**; `hil-operator`, `pr-monitor`
-and `static-analyzer` on **sonnet**; `builder` on **haiku** (mechanical,
-log-heavy).
+Tiered models (owner revision 2026-07-09; originally all-opus): `port-dev`,
+`driver-reviewer` and `target-debugger` on **opus** at **xhigh**;
+`hil-operator`, `pr-monitor` and `static-analyzer` on **sonnet**; `builder`
+on **haiku** (mechanical, log-heavy). The registry has no effort field —
+xhigh is requested per `agent()` call by whichever workflow or session spawns
+the agent.
| Agent | Effort | Role |
|---|---|---|
@@ -40,6 +42,7 @@ log-heavy).
| `port-dev` | xhigh | Implement one well-scoped change in one port / file set. Follows repo rules: C99, 2-space indent, snake_case, `TU_ASSERT`, no dynamic allocation, ISR work deferred to task context. Runs `clang-format` (repo `.clang-format`) on touched files before finishing. Cross-checks the MCU datasheet in `$HOME/Documents/calibre-library` when changing dcd/hcd register logic. Verifies with a targeted build of one board using the port. Returns `{item, diffstat, buildOk, notes}`. |
| `driver-reviewer` | xhigh | Review one dcd/hcd directory against dimensions: correctness, ISR safety, register use vs. datasheet AND MCU errata (calibre library; missing erratum workarounds are findings), style. Returns structured findings `{file, line, snippet, why, severity, confidence}` — coverage-first (report everything; filtering happens downstream). |
| `hil-operator` | default | All rig interaction — the actions-runner service is NEVER stopped; per-board flock locks arbitrate with concurrent CI. `hil_test.py` runs rely on its per-board self-locking; manual hardware work (JLink/GDB, usbtest, serial) is wrapped in `test/hil/board_lock.py hold/release`; rig-wide ops (uhubctl, pci-rebind) require `hold --all`; on wedge `usb_recover.sh` + dmesg. Used strictly serially — never two instances concurrently. |
+| `target-debugger` | xhigh | Root-cause one USB misbehavior on one board by instrumenting the device side (TU_LOG/RTT, RAM ring-buffer trace, GDB autopsy, J-Link PC-sampling) with dual-side host+target capture, per `.claude/skills/usb-target-debug/SKILL.md`, plus wire-level capture via the ataradov hardware tap (`.claude/skills/usb-sniffer/SKILL.md`) when the host side can't see or is disputed. Deliberately serial loop under one held board lock (released around `hil_test.py` runs, which self-lock); strictly one instance. Diagnosis standard: evidence shows the mechanism, or a fix flips the ORIGINAL failing case on hardware; stops after two evidence-free cycles with a partial report. Hard rule "fix stays, probe goes, re-verify clean": instrumentation reverted, candidate fix left uncommitted and re-verified on a clean build, pristine firmware reflashed before lock release. Returns `{board, bug, diagnosis, confirmed, ruledOut[], evidence[], fixDiffstat, fixVerified, instrumentationReverted, lockReleased, notes}`. |
| `pr-monitor` | default | Triage one GitHub PR via `gh`: check CI status (`gh pr checks`), read failing run logs and classify each failure infra/flake vs real; re-run infra failures (`gh run rerun --failed`); harvest automated review comments (Codex/Copilot/Claude bots — knows their signals: Codex posts a "Didn't find any major issues" issue comment when clean; Copilot drops out of `requested_reviewers` when done; bot logins differ across APIs); adversarially validate each finding against the actual code. Returns structured triage `{ci: {status, infraRerun[], realFailures[]}, findings: [{source, file, line, claim, verdict, fixHint}]}`. Read/triage/re-run/reply only — never edits code. |
| `static-analyzer` | low | Run PVS-Studio (SAST + MISRA C:2023/C++:2008) for one board: build with exported `compile_commands.json` (via `run_pvs.sh` solo, or a dedicated `cmake-build-pvs` dir when parallel builders run), analyze against `.PVS-Studio/.pvsconfig`, gate on diagnostics in files changed vs a base ref. Returns `{pass, ga1, ga2, changedFindings[], detail}`; `pass=false` only on GA:1 in changed files or tool failure. Read-only. |
@@ -120,8 +123,8 @@ carries the judgment; JS carries the orchestration.
## Model & effort policy
-- Tiered worker models: `port-dev`/`driver-reviewer` **opus** `xhigh`;
- `hil-operator`/`pr-monitor` **sonnet**; `builder` **haiku**.
+- Tiered worker models: `port-dev`/`driver-reviewer`/`target-debugger` **opus**
+ `xhigh`; `hil-operator`/`pr-monitor` **sonnet**; `builder` **haiku**.
- Inline workflow stages: unit/size **haiku**; pvs **sonnet** (low effort);
pr-babysit push/replies **sonnet**.
- Agent frontmatter `model:` is canonical for `agentType` calls; it is read
diff --git a/docs/superpowers/specs/2026-07-23-esp-target-debug-design.md b/docs/superpowers/specs/2026-07-23-esp-target-debug-design.md
new file mode 100644
index 000000000..491466992
--- /dev/null
+++ b/docs/superpowers/specs/2026-07-23-esp-target-debug-design.md
@@ -0,0 +1,102 @@
+# esp-target-debug Skill Design
+
+Backend skill for debugging TinyUSB firmware on Espressif targets (rig:
+`espressif_p4_function_ev`, `espressif_s3_devkitm`) via the chips' **built-in
+USB-Serial-JTAG**, with external JTAG documented as a TODO until the rig has
+an adapter. Companion to `target-debug`, which keeps the architecture-neutral
+methodology (intrusiveness ladder, board locks, dual-side capture, diagnosis
+standards) — this skill is the Espressif toolchain/probe backend, the same
+boundary that makes `usb-kernel-debug` its own skill.
+
+## Goals
+
+- An agent can attach, halt, backtrace, set breakpoints/watchpoints, list
+ FreeRTOS threads, and capture logs on the rig's P4 **while TinyUSB device
+ traffic is live** — every recipe hardware-verified before landing unmarked
+ (the `target-debug` ethos).
+- The S3's USB-SJ/OTG PHY conflict is mapped precisely, not hand-waved:
+ verified working via `board_test` (TinyUSB off — PHY free), verified failure
+ mode with a USB device example, external-JTAG escape hatch documented as
+ TODO.
+
+## Non-goals (deferred)
+
+- External JTAG bring-up (no adapter on the rig) — TODO section with S3 JTAG
+ pin notes (GPIO39-42) and openocd-esp32 adapter support pointers.
+- Xtensa/S3 full parity under live USB traffic (needs external JTAG).
+- ETM-class instruction trace; SystemView tooling beyond an apptrace spike.
+
+## Architecture
+
+New skill `.claude/skills/esp-target-debug/SKILL.md`; two integration edits:
+
+- `target-debug` gains a 2-line pointer under the probe-mapping bullets:
+ Espressif boards use a different toolchain, probe model, and trace story —
+ read `esp-target-debug`.
+- `target-debugger` agent table gains an `esp-target-debug` row (name-only,
+ aligned columns, per the established conventions).
+
+Skill content (order):
+
+1. **Role + defer line** — methodology lives in `target-debug`; this file is
+ the Espressif backend. Built-in USB-SJ now; external JTAG TODO.
+2. **PHY-conflict map** —
+ - S3: USB-SJ and OTG share one PHY (GPIO19/20). TinyUSB claiming the PHY
+ drops JTAG-over-USB mid-session: JTAG works for non-USB examples
+ (`board_test`), dies for USB device examples (verified boundary, exact
+ symptom recorded). External JTAG = the future escape hatch (TODO).
+ - P4: OTG-HS has a dedicated HS PHY; USB-SJ is separate — JTAG and the
+ TinyUSB DUT port coexist (verified). USB-SJ doubles as a live log
+ console during device traffic — the TU_LOG-equivalent channel.
+3. **Toolchain & attach** — `. $HOME/code/esp-idf/export.sh` provides
+ `openocd-esp32` + `riscv32-esp-elf-gdb` (P4) / `xtensa-esp32s3-elf-gdb`
+ (S3). Rig path is raw openocd (HIL firmware isn't an idf project on disk):
+ `openocd -f board/esp32p4-builtin.cfg` with `adapter serial <uid>` (USB-SJ
+ is VID 303A:1001; uid = the `flasher.uid` already in `tinyusb.json`), gdb
+ on :3333. `idf.py openocd` / `idf.py gdb` noted for idf-project work.
+4. **Technique mapping table** (aligned) — ARM technique → Espressif
+ equivalent:
+
+ | target-debug technique | Espressif backend |
+ |---|---|
+ | GDB autopsy, bp/wp | same flow; RISC-V trigger module (P4) / Xtensa 2 bp + 2 wp (S3); budget read verified on P4 |
+ | Vector catch | none — breakpoint the panic handler; decode `mcause`/`mepc`/`mtval` (P4) |
+ | SWO / DWT data trace | none — apptrace over JTAG is the analog (gated spike; lands `(untested)` if it fails) |
+ | RTT / TU_LOG | USB-SJ console — on P4 it coexists with DUT traffic |
+ | FreeRTOS threads | native in openocd-esp32 — `info threads` out of the box |
+ | verifybin | `esptool.py verify_flash` |
+
+5. **Rig discipline deltas** — same `board_lock.py` protocol; flasher is
+ esptool (serial-port-by-uid); reflash pristine before release; one client
+ per USB-SJ device.
+6. **External JTAG — TODO** — S3 JTAG pins, adapter classes openocd-esp32
+ supports, and the efuse caveat (JTAG pin selection), unverified.
+
+## Verification gates (execution order)
+
+All under board locks, serial, evidence in commit messages:
+
+1. **P4 coexistence (headline)**: flash a device example, confirm enumeration
+ + traffic on the DUT port, then attach openocd+gdb over USB-SJ →
+ halt, `bt`, resume — device stays functional after resume.
+2. **P4 budget**: read trigger/watchpoint counts via openocd/gdb; set a
+ hardware watchpoint on a TinyUSB variable, confirm hit.
+3. **P4 threads**: `info threads` lists ESP-IDF tasks (usbd task visible).
+4. **P4 console**: capture USB-SJ console log output during device traffic.
+5. **P4 apptrace spike (gated)**: bounded attempt; verified recipe or
+ `(untested)` tag.
+6. **S3 boundary**: `board_test` flashed → attach works (halt+bt); then a USB
+ device example → record the exact JTAG failure symptom when the PHY
+ switches. No further S3 work (external JTAG TODO).
+
+## Constraints
+
+- Worktree `claude/improve-debug-skill-agent`; commit per gate; pre-commit
+ before each; no Co-Authored-By trailers.
+- Formatting conventions already established: aligned table columns,
+ skill-name-only cross references, bullets over run-on paragraphs.
+- Espressif builds need `export.sh` first (CLAUDE.md); P4/S3 examples build
+ via idf.py — reuse existing HIL-built firmware where possible instead of
+ rebuilding.
+- Hardware-verify-before-landing: unverified content ships tagged
+ `(untested)` or not at all.
diff --git a/docs/superpowers/specs/2026-07-28-hil-test-refactor-design.md b/docs/superpowers/specs/2026-07-28-hil-test-refactor-design.md
new file mode 100644
index 000000000..3cd202d95
--- /dev/null
+++ b/docs/superpowers/specs/2026-07-28-hil-test-refactor-design.md
@@ -0,0 +1,141 @@
+# hil_test.py refactor: test core + infra helpers
+
+**Date:** 2026-07-28
+**Branch:** `claude/hil-test-split` (based on `claude/hil-pool-check`, which adds `pool_check.py`)
+
+## Motivation
+
+`test/hil/hil_test.py` is 2370 lines mixing five concerns: board-lock protocol, per-controller
+scheduling permits, flash/reset backends, the actual per-example tests, and orchestration/report/CLI.
+The lock protocol additionally exists in three copies (`hil_test.py`, `board_lock.py`,
+`.claude/skills/hil/pool_check.py`), which has already produced drift (pool_check's copy lacks
+hil_test's fail-open and error guards). Splitting the infrastructure out makes `hil_test.py`
+test-focused and gives external tools (pool_check) one canonical import for locks, permits, and
+flashing.
+
+## Goal / non-goals
+
+**Goal:** behavior-preserving code motion. `hil_test.py`'s CLI, arguments, output, report format,
+and runtime behavior stay byte-identical. One deliberate user-visible change: the operator lock CLI
+moves from `board_lock.py` to `hil_lock.py` (same subcommands, same behavior); `board_lock.py` is
+deleted.
+
+**Non-goals (explicit follow-ups, not this change):**
+- The 15 pool_check findings from the 2026-07-28 code review (exception isolation, park-on-failure,
+ espressif coverage, probe-recovery criterion, etc.).
+- pool_check adopting `flash_permit` controller budgeting (enabled by this split).
+- Any change to lock semantics, permit widths, flash behavior, or test logic.
+
+## Resulting layout (`test/hil/`)
+
+| File | ~Lines | Role |
+|---|---|---|
+| `hil_test.py` | 1600 | tests + orchestration + report + CLI (unchanged interface) |
+| `hil_lock.py` (new) | 420 | board-lock protocol + controller permits + operator CLI |
+| `hil_flash.py` (new) | 250 | `run_cmd` + flash/reset backends + `find_firmware` |
+| `board_lock.py` | deleted | superseded by `hil_lock.py` |
+
+Import graph: `hil_test` → {`hil_lock`, `hil_flash`}; the helpers import nothing local (no cycles).
+`pool_check.py` imports all three.
+
+## hil_lock.py
+
+Docstring states the scope: board locks + controller flash/battery permits; the CLI manages board
+locks only (permits are in-process semaphores with no CLI meaning).
+
+**Flock core** (protocol defined once; moved from `board_lock.py`/`hil_test.py`):
+- `BOARD_LOCK_DIR = '/tmp/tinyusb-hil-locks'`, `lock_path(board)`
+- `CI_REASON = 'hil_test.py'` — the release-protected holder tag (release refuses to kill it)
+- `flock_nb(board) -> fh` — `os.open(O_RDWR|O_CREAT, 0o666)` **without O_TRUNC** (a losing racer
+ must not wipe the winner's record), `fdopen('r+')`, `LOCK_EX|LOCK_NB`; raises `OSError` when held
+- `write_record(fh, reason)` — truncate+seek+`json.dump({pid, reason, since})`+flush
+- `clear_record(fh)` — truncate(0), swallow OSError (records stay truthful on release)
+- `read_record(board) -> dict | None` — today's `board_lock.read_info`
+- `acquire_board_lock(board, reason=CI_REASON) -> fh | None` — today's `hil_test.acquire_board_lock`
+ with a `reason` parameter: `HIL_NO_BOARD_LOCK=1` bypass, fail-open with warning on lock-dir
+ OSError, `RuntimeError` carrying holder info on conflict
+
+**Controller permits** (moved verbatim from `hil_test.py`):
+- `FLASH_PARALLEL`, `USBTEST_PARALLEL`, `CONTROLLER_SLOTS` (env-overridable as today)
+- `controller_of(uid)`, `controller_slot(pci)`, `controller_permit`, `flash_permit(uid)`,
+ `usbtest_permit(uid)`
+- Per-worker globals (`usbtest_sems`, `flash_sems`, `controller_map`, `controller_meta`,
+ `controller_hints`) set by a new `init_scheduling(sems, fsems, cmap, cmeta, hints)` hook that
+ `hil_test.init_worker` calls from the Pool initializer. `controller_permit`'s PROFILE logging
+ calls back through a module-level `log = print`-style hook that `hil_test` points at `log_line`
+ during `init_scheduling` (keeps helpers free of hil_test imports). The `PROFILE` env flag
+ (`HIL_PROFILE=1`) is read independently in `hil_lock` at import, same derivation as today.
+
+**Operator CLI** (moved verbatim from `board_lock.py`): `hold`/`release`/`status` subcommands with
+the daemon-holder machinery (double-fork, setsid, stdio detach, success pipe, SIGTERM bow-out),
+release policy (probe the flock; protect `CI_REASON` holders; SIGTERM other recorded pids),
+`is_locked` pid-liveness, `--all`/`--config` roster handling. The hold/release/status internals
+switch to the flock-core helpers above; observable behavior unchanged.
+
+## hil_flash.py
+
+Moved verbatim from `hil_test.py`:
+- `CMD_TIMEOUT` (env-overridable), `run_cmd(cmd, cwd, timeout)`, `cmd_stdout_text(out)`
+- `OPENCOD_ADI_PATH`, `TINYUSB_ROOT`
+- All backends: `flash_jlink`/`reset_jlink`, `flash_stlink`/`reset_stlink`,
+ `flash_stflash`/`reset_stflash`, `flash_openocd`/`reset_openocd`,
+ `flash_openocd_wch`/`reset_openocd_wch`, `flash_openocd_adi`/`reset_openocd_adi`,
+ `flash_wlink_rs`/`reset_wlink_rs`, `flash_esptool`/`reset_esptool`,
+ `flash_uniflash`/`reset_uniflash`, `flash_lm4flash`/`reset_lm4flash`
+- `find_firmware(variant, example)`
+- `get_serial_dev(id, vendor_str, product_str, ifnum)` — moves here (not hil_test) because
+ `flash_esptool` calls it; keeping it test-side would create a helper→hil_test import cycle.
+ Tests call `hil_flash.get_serial_dev`.
+- Module globals `build_dir = 'cmake-build'` and `verbose = False`, set by callers exactly as the
+ `hil_test` globals are today (`hil_test.main` sets them from argparse; pool_check sets them
+ directly). `run_cmd`'s verbose echo reads `hil_flash.verbose`.
+
+Dispatch in callers stays string-based: `getattr(hil_flash, f'flash_{flasher["name"].lower()}')`.
+
+## hil_test.py (what remains)
+
+Config TypedDicts (`Board`, `FlasherCfg`, …), device-node lookup except `get_serial_dev`
+(`get_disk_dev`, `get_hid_dev`, `get_alsa_capture_dev`, `open_serial_dev`, `serial_write_all`,
+`read_disk_file`, `open_mtp_dev`, `get_printer_dev`/`open_printer_dev`), enum-timeout globals +
+`wait_until`,
+`log_line`/print-lock, `compact_output`, all `test_*` functions, test lists, `test_example`,
+`build_board`, `test_board`, report rendering/accumulation, `main`. Call sites use explicit
+module-qualified names (`hil_lock.flash_permit(...)`, `hil_flash.run_cmd(...)`) so provenance is
+greppable; no `from … import *`-style mirroring.
+
+`init_worker` keeps its signature (Pool initargs unchanged) and forwards the scheduling state to
+`hil_lock.init_scheduling(...)`.
+
+## Consumer updates (same commit)
+
+- **`.claude/skills/hil/pool_check.py`** — drop its private `lock_board`/`unlock_board` in favor of
+ `hil_lock.flock_nb` + `write_record(fh, 'pool_check')` (+ `clear_record` on release; deliberately NOT `acquire_board_lock`, whose HIL_NO_BOARD_LOCK bypass and fail-open behavior pool_check must not inherit); import
+ flashers/`find_firmware`/`get_serial_dev`/`cmd_stdout_text`/`TINYUSB_ROOT`/`build_dir` from
+ `hil_flash`; `BOARD_LOCK_DIR` references move to `hil_lock`. pool_check then imports **only**
+ `hil_lock` + `hil_flash` (no `hil_test`), so its `pymtp` stub shim is deleted — that shim existed
+ solely because importing `hil_test` pulls in libmtp.
+- **`test/hil/hil_ci.sh`** — the scp list is currently `hil_test.py`, `pymtp.py`, `$CONFIG`; add
+ `hil_lock.py` and `hil_flash.py` (hil_test cannot even import without them). `board_lock.py` was
+ never in the list.
+- **Docs rename `board_lock.py` → `hil_lock.py`** (live docs only): `.claude/skills/hil/SKILL.md`,
+ `.claude/agents/hil-operator.md`, `.claude/agents/target-debugger.md`,
+ `.claude/skills/etm-trace/SKILL.md`, `.claude/skills/usb-kernel-recover/SKILL.md`,
+ `.claude/skills/target-debug/SKILL.md`. Historical `docs/superpowers/{plans,specs}` stay as
+ records.
+- **CI workflow** — untouched (invokes `hil_test.py` CLI only).
+
+## Verification
+
+1. `python3 -m py_compile` on all three modules + pool_check.
+2. `hil_lock.py hold/status/release` interplay: hold, conflicting hold, status listing, release,
+ protection of a `CI_REASON` record, stale-record cleanup.
+3. `pool_check.py --scan-only`, then a single flash board (e.g. `-b stm32f407disco`).
+4. Full `hil_test.py -b stm32f407disco -B examples tinyusb.json` on the rig; compare the report
+ row and log shape against a pre-refactor run.
+5. `pre-commit run` on all touched files.
+
+## Sequencing
+
+Lands on top of `claude/hil-pool-check`. After merge, fix the pool_check review findings as a
+separate change on the new module boundaries, and update agent-memory references to
+`board_lock.py`.
diff --git a/docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md b/docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md
new file mode 100644
index 000000000..898b3c8ab
--- /dev/null
+++ b/docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md
@@ -0,0 +1,179 @@
+# PR-scoped HIL selection: helper/hil_select.py
+
+**Date:** 2026-07-29
+**Branch:** `claude/hil-select` (based on `claude/hil-pool-check`, which carries the
+hil_lock/hil_flash split and the current rig rosters)
+
+## Motivation
+
+Every PR currently builds and runs the full HIL matrix (both rigs, every roster board, every
+test). Most PRs touch one port or one class: a `dcd_rp2040` change cannot affect an STM32 board,
+a `cdc_device.c` change cannot affect an MSC-only example, and a device-stack change cannot
+affect host tests. Scoping HIL to the affected boards/tests cuts CI wall time and rig wear
+without losing relevant coverage.
+
+## Goal / non-goals
+
+**Goal:** a shared selector that maps a PR diff to (boards, per-board test lists), wired into
+CI's `set-matrix` on `pull_request` events (pruning both `hil-build` and the rig jobs) and
+callable locally (pre-pr, manual runs). Scoping may only shrink coverage when the mapping is
+confident; every uncertainty widens to the full matrix.
+
+**Non-goals:**
+- Variant-level selection (all variants of a selected board run).
+- Scoping the non-HIL build jobs (cmake/CircleCI one-per-family builds are independent build
+ coverage and stay untouched).
+- Scoping push/master/scheduled runs (always full).
+- Changing hil_test.py behavior (the selector only *composes* existing `-b`/`-bt` args).
+
+## Component: `test/hil/helper/hil_select.py`
+
+Stdlib-only, importable and CLI. Lives beside the harness so `hil_ci.sh` copies are unaffected
+(it runs on the GitHub runner / dev PC, not on the rig). It must NOT import `hil_test.py`
+(which drags pyserial/pymtp onto the bare GitHub runner): the three test lists
+(`device_tests`, `dual_tests`, `host_test`) move verbatim into the stdlib-only
+`test/hil/helper/hil_util.py` that both `hil_test.py` and `hil_select.py` import (behavior
+preserving; `hil_ci.sh` copies the whole `helper/` directory).
+
+```
+python3 test/hil/helper/hil_select.py --base <ref> [--diff-file <path>] CONFIG.json [CONFIG.json...]
+```
+
+- `--base REF`: changed files = `git diff --name-only $(git merge-base HEAD REF)..HEAD`
+ (mirrors pre-pr). `--diff-file`: newline-separated file list instead of git (unit tests, CI
+ reuse of a precomputed diff).
+- Output (stdout, JSON):
+
+```json
+{
+ "full": false,
+ "boards": {"raspberry_pi_pico": "all", "stm32f407disco": ["device/cdc_msc", "device/cdc_dual_ports"]},
+ "args": {"tinyusb.json": "-b raspberry_pi_pico -b stm32f407disco -bt stm32f407disco:device/cdc_msc,device/cdc_dual_ports",
+ "hfp.json": ""},
+ "reasons": ["src/portable/raspberrypi/rp2040/dcd_rp2040.c: port rp2040 -> family rp2040 -> boards [raspberry_pi_pico, ...] (device role)"]
+}
+```
+
+- `full: true` ⇒ `boards`/`args` cover the entire rosters (identical to today's behavior).
+- `args` maps each input config file to the hil_test.py argument string for that rig: `-b` per
+ selected board on that roster, plus `-bt BOARD:t1,t2` for boards with a restricted test list
+ ("all" boards get bare `-b`). An empty string means: nothing on this rig is affected — the
+ rig job is skipped for this PR.
+- Per-file reasoning lines (`file → rule → contribution`) go in `reasons` and to stderr, so the
+ CI log answers "why did/didn't HIL run X" without archaeology.
+
+## Classification rules
+
+Each changed file yields a contribution; the selection is the union. Any file matching no rule
+sets `full: true` (fail-open). Rules, first match wins:
+
+1. **Non-code:** `docs/**`, `.claude/**` (except the workflows below via rule 8), `*.md`,
+ `*.rst`, `LICENSE*` → contributes nothing.
+2. **Port:** `src/portable/<vendor>/<ip>/**` (or single-level `src/portable/<name>/**`).
+ Role from basename: `dcd_*`/`*_device*` → device; `hcd_*`/`*_host*` → host; anything else
+ (shared port files, e.g. `dwc2/dwc2_common.c`) → both. Families = directories of
+ `hw/bsp/*/family.cmake|family.mk` whose text references `<vendor>/<ip>` (pre-pr's grep),
+ boards = those families' entries on the input rosters. Tests = all tests of that role
+ (device_tests / host_test from hil_test.py's lists; dual_tests count as both roles).
+3. **Class:** `src/class/<c>/*_device.*` → all device-capable roster boards; tests = the
+ device/dual examples in hil_test.py's lists whose `examples/<role>/<ex>/src/tusb_config.h`
+ defines `CFG_TUD_<C>` with a nonzero value (derived at runtime; `<C>` = upper-cased class
+ dir, with the map `musb→n/a`-style exceptions NOT needed — class dirs and config macros
+ share names: cdc, msc, hid, midi, audio, video, vendor, usbtmc, mtp, printer. Two
+ exceptions: in class dir `dfu`, `dfu_rt_device.*` maps to CFG_TUD_DFU_RUNTIME and
+ `dfu_device.*` to CFG_TUD_DFU; class dir `net` maps to CFG_TUD_ECM_RNDIS|CFG_TUD_NCM.) `*_host.*` analogously via `CFG_TUH_<C>`. Shared class files (e.g. `cdc.h`) →
+ both roles' matching examples. A class with zero matching examples contributes nothing
+ (known path, does not force full).
+4. **Core role:** `src/device/**` → all device-capable boards, all device tests (+dual);
+ `src/host/**` → all host-capable boards, all host tests (+dual).
+5. **Core common:** `src/common/**`, `src/osal/**`, `src/tusb.c`, `src/tusb.h`,
+ `src/tusb_option.h` → full.
+6. **BSP:** `hw/bsp/<family>/**` → that family's roster boards, all their tests;
+ `hw/bsp/<family>/boards/<board>/**` narrows to that board if it is on a roster, and
+ contributes nothing when it is not (an off-rig board cannot be HIL-tested; known path,
+ does not force full).
+ Family-agnostic BSP files (`hw/bsp/board_api.h`, `hw/bsp/board.c`, ansi_escape.h) → full.
+7. **Example:** `examples/<role>/<ex>/**` → all roster boards, tests = that example if present
+ in hil_test.py's lists, else contributes nothing. `examples/build_system/**`, top-level
+ `examples/CMakeLists.txt` → full. `examples/device/board_test/**` → full: it is the park
+ firmware hil_test.py flashes on every board (variant boundary + teardown), not a test.
+8. **Harness/infra:** `test/hil/**`, `.github/workflows/build*.yml`,
+ `.github/actions/**`, `tools/build.py`, `tools/get_deps.py`, `tools/cmake/**`,
+ `hw/mcu/**`, `lib/**` → full.
+9. **Everything else** (`test/unit-test/**`, `tools/**` not above, unknown paths) → full.
+ (Unit-test-only changes could safely skip HIL, but per the fail-open stance anything not
+ explicitly classified widens; narrowing rule 9 is a later refinement.)
+
+**Role pruning:** after the union, if only device-role contributions exist, host-only boards
+drop out and host tests are stripped from mixed boards (vice versa for host-only changes).
+Dual tests survive either role. Board capability (device/host) comes from the roster entry's
+`tests` flags/only-list, same logic hil_test.py uses.
+
+**No-rig-coverage case:** a cleanly classified change whose boards intersect a roster to the
+empty set yields an empty `args` string for that rig and a stderr line saying so — the rig job
+is skipped, not widened (running unrelated boards would test nothing relevant).
+
+**Roster source:** `config['boards']` only (boards-skip stays parked).
+
+## CI wiring (`.github/workflows/build.yml`)
+
+- `set-matrix` (PR events only): after generating today's matrices, run
+ `helper/hil_select.py --base origin/${{ github.base_ref }} test/hil/tinyusb.json test/hil/hfp.json`
+ (checkout with enough history to reach the merge base: `fetch-depth: 0` on this one job, or
+ an explicit `git fetch origin $BASE_REF`). New job outputs: `hil_select_full`,
+ `hil_args_tinyusb`, `hil_args_hfp`, plus the selected-board list consumed by the matrix
+ generator. Non-PR events: skip the selector, outputs default to full/empty-args-means-all.
+- `hil_ci_set_matrix.py` gains `--select '<json>'`: when given and `full` is false, it emits
+ build entries only for selected boards (per config). Untouched otherwise.
+- `hil-tinyusb` job (one matrixed job covering both rigs, selected by `matrix.hil_json`): a
+ step picks the rig's selector args in shell (`case "$HIL_JSON" in ...`) from the set-matrix
+ outputs and either appends them to the `hil_test.py` invocation or exits the step early with
+ a "HIL skipped by selection" log line when that rig has nothing to run (`run` flag output
+ false). The separate `hil-tinyusb-esp` job (esptool split) gets the same treatment with the
+ tinyusb args. Non-PR events: outputs default to run=true with empty args (today's behavior).
+- The `--flasher`/`--exclude-flasher` split in the existing matrix `test_args` composes fine
+ with `-b` (hil_test.py applies both filters).
+
+## Local use
+
+- pre-pr's "Map changes to boards" step delegates to
+ `python3 test/hil/helper/hil_select.py --base $BASE test/hil/tinyusb.json` and derives its
+ one-board-per-family sample from the selector's board set (its capping/sampling policy is
+ unchanged — the selector provides the affected set, pre-pr samples it).
+- Manual: `python3 test/hil/hil_test.py -B examples $(python3 test/hil/helper/hil_select.py --base master test/hil/tinyusb.json | jq -r '.args["tinyusb.json"]') test/hil/tinyusb.json`
+ — documented in the hil skill.
+
+## Testing
+
+`test/hil/test/test_hil_select.py` — stdlib `unittest`, no hardware, injected diffs via
+`--diff-file`/API. Cases (the acceptance examples):
+1. `src/portable/raspberrypi/rp2040/dcd_rp2040.c` → only rp2040-family roster boards, device
+ tests only, host-only boards absent, `full` false.
+2. `src/device/usbd.c` → every device-capable board on both rosters, all device tests + dual,
+ no host-only board, no host tests.
+3. `src/class/cdc/cdc_device.c` → only examples with CFG_TUD_CDC enabled (must include
+ device/cdc_msc and device/cdc_dual_ports; must exclude device/msc_dual_lun and all
+ host tests).
+4. `src/class/msc/msc_host.c` → host-capable boards only, host examples with CFG_TUH_MSC.
+5. `tools/random_new_script.py` → `full: true`.
+6. `docs/foo.rst` alone → contributes nothing ⇒ empty selection, `full` false, all `args`
+ empty (CI additionally has check-paths gating; the selector's answer is still honest).
+7. `hw/bsp/rp2040/family.cmake` → rp2040-family boards, all their tests.
+8. Mixed device+host diff → no pruning (both roles present).
+The suite runs in `set-matrix` before the selector is used, and locally via
+`python3 test/hil/test/test_hil_select.py`.
+
+## Safety properties
+
+- Fail-open: unknown/infra paths ⇒ full matrix; selector crash in CI ⇒ job fails visibly
+ (never silently skips HIL).
+- Only `pull_request` events are scoped.
+- The selection JSON + per-file reasons are printed in the job log for audit.
+- hil_test.py errors on `-b` names not in the config — the selector only emits roster names,
+ and the unit suite locks that invariant.
+
+## Sequencing
+
+Lands on `claude/hil-select` on top of the pool-check/split stack. Follow-ups it does not
+include: narrowing rule 9 for unit-test-only changes; variant-level selection; pre-pr skill
+text update ships in the same change (its mapping section shrinks to a selector call).
diff --git a/docs/superpowers/specs/2026-07-30-hil-usbtest-fleet-wedge-design.md b/docs/superpowers/specs/2026-07-30-hil-usbtest-fleet-wedge-design.md
new file mode 100644
index 000000000..3ed0c1519
--- /dev/null
+++ b/docs/superpowers/specs/2026-07-30-hil-usbtest-fleet-wedge-design.md
@@ -0,0 +1,236 @@
+# HIL fleet-wedge containment
+
+Date: 2026-07-30
+Status: implemented, then superseded in part — addendum last checked 2026-08-12
+against the shipped code; where they disagree the CODE and the usb-kernel-recover
+skill win, never this document.
+
+- **Pool guard.** A single constant, not the flat 4200s below and not a derivation:
+ `POOL_TIMEOUT = pos_int_env('HIL_POOL_TIMEOUT', 3600)`. A per-controller model briefly
+ lived here and was removed -- it under-modelled the flash phase and could INVERT
+ (adding a usbtest board lowered the guard, because the derived value fell below the
+ baseline it was meant to raise). The guard's only job is to stop a wedged pool short
+ of the job ceiling so the report still gets written; predicting a healthy run's
+ duration is a different problem. `pos_int_env` warns only on a non-integer or a value
+ <= 0: there is NO upper clamp and no warning above any threshold, so a pin larger than
+ a job ceiling silently restores the inversion this work removed.
+- **Job ceilings.** 90/90/120 min (build.yml), not 60/60/90 and not the 85/115 below.
+ They must clear the 3600s guard plus the pre-pool checkout/artifact merge and the
+ post-guard sweep and report upload. No job pins `HIL_POOL_TIMEOUT`.
+- **Battery budgets.** `USBTEST_BATTERY_BUDGET` 260s, `USBTEST_RECOVERY_BUDGET` 250s.
+ The 200s-with-a-197s-floor derivation recorded here was never shipped; the floor
+ assertion was removed with it.
+- **HUNG recovery.** Reflash of the DUT through its roster flasher
+ (`usbtest.py --recover-board/--recover-fw`), not the root-cycle-first recovery in
+ section 1d — replaced after the 2026-08-11 ppps measurement (uhubctl never cuts
+ VBUS; root-cycle is probe-only). Since 2026-08-12 the reflash is SKIPPED
+ when `hil_flash.convoy_safe(board['flasher'])` is false (usbtest.py:675): the flasher
+ would enumerate by opening usbfs nodes, block on the same convoy, and become a second
+ stray rather than clear the first. A holder that owns the device lock inside a driver
+ ioctl is terminal either way -- a reflash only produces a disconnect, and
+ `usb_disconnect()` needs that same lock -- and that state needs a reboot.
+
+Step 0 done — the host was rebooted 2026-07-30 14:11 and the rig
+came back clean. The device that triggered this incident was removed from the rig, so
+only the containment work remains relevant.
+Rig: `ci.lan` (Proxmox guest on `pve.lan`)
+
+## Problem
+
+On 2026-07-29/30 every board in the `ci.lan` usbtest fleet failed, `openocd` processes
+landed in uninterruptible sleep, and no subsequent HIL run could start. Two GitHub
+Actions runs were stranded: `30484641269` sat `in_progress` for over eight hours
+(past GitHub's own 360-minute default), and `30485082274` sat `queued` behind it from
+2026-07-29 19:35 UTC onward. Both report directories were written empty.
+
+A reboot of the `ci` guest at 10:48 did not clear the condition: the same kernel state
+re-formed at 10:52:23.
+
+## Root cause
+
+Five layers, each independently observable.
+
+### 1. A permanently wedged hub worker holds a root-hub device lock
+
+A device that repeatedly re-asserts connect while failing to enumerate keeps
+`hub_event()` busy, and `hub_event()` holds `usb_lock_device(hdev)` on its hub for its
+whole run (hub.c:5896/5989). The `usb_hub_wq` worker sits in `hub_port_reset`, so that
+hub's `device_lock` is effectively never released:
+
+```
+kworker/14:6+usb_hub_wq (state D, 400+ s)
+ msleep+0x2b
+ hub_port_reset+0x1a4 [usbcore]
+ hub_event+0x727 [usbcore]
+```
+
+`usb usbN-portM: Cannot enable. Maybe the USB cable is bad?` is logged every four seconds
+for as long as it lasts.
+
+Verified against hub.c v6.12.96 rather than inferred: the kernel does **not** retry
+without bound, and root and downstream ports are bounded identically —
+`hub_port_reset()` tries `PORT_RESET_TRIES` then logs that message (hub.c:3149),
+`hub_port_connect()` wraps it in `PORT_INIT_TRIES` = 4 and disables the port on give-up
+(hub.c:5455/5619). A count in the thousands is therefore that many separate connect
+events, not one runaway loop, and it indicts the device rather than the port.
+
+### 2. A parked board storms the second controller
+
+`ra6m5_ek` (`test/hil/tinyusb.json`, uid `8419032D32363657364EF4622D294B4E`, at
+`13-3.3`) runs dfu firmware (`cafe:400b`) and re-enumerates every 1-2 seconds
+continuously, wrapping the entire bus-13 devnum space (`...120 -> 127 -> 4 -> 6 -> 10`).
+This is standing `hub_event` and Address-Device pressure on controller `03:00.0`,
+concurrent with parallel usbtest batteries on the same silicon.
+
+The board is already listed in `boards-skip`, which is precisely why it storms:
+`boards-skip` stops testing a board but never parks it, so it keeps running whatever
+firmware it last received. Park-flash only runs as teardown of a board that actually
+executed tests.
+
+### 3. The kernel `usbtest` control-queue case waits without a timeout
+
+`test_ctrl_queue` blocks on an untimed `wait_for_completion()` while `usbdev_ioctl`
+holds the DUT's `device_lock`:
+
+```
+wait_for_completion+0x8a <- no _timeout variant
+test_ctrl_queue+0x4ab [usbtest]
+usbtest_do_ioctl+0x501 [usbtest]
+usbdev_ioctl+0x6b8 [usbcore]
+```
+
+`--timeout 60` in `test/hil/usbtest.py` is a subprocess timeout only. `SIGKILL` is not
+delivered to a task in uninterruptible sleep. `usbtest.py` already recognises this and
+reports `HUNG`, then calls `usb_recover.sh root-cycle`.
+
+### 4. openocd inherits the convoy and the whole fleet dies
+
+Once a device lock is stuck, `port_event()` takes a child device's lock to warm-reset
+it and blocks while still holding its hub's lock. Any later
+`open("/dev/bus/usb/BBB/DDD")` against such a device blocks uninterruptibly:
+
+```
+usbdev_open+0xdc [usbcore] -> __mutex_lock
+chrdev_open -> do_sys_openat2 -> __x64_sys_openat
+```
+
+That is the state of the three `openocd` processes at 04:16:51 (pids 207921, 207987,
+208034) — the flasher, unkillable. Because one controller carries two buses, a single
+convoy takes out every board on both, which is why the failure presents as the entire
+fleet.
+
+The existing `HUNG` recovery cannot help here. A root-port VBUS cycle frees a
+*device-lock* holder; it cannot free a lock held by a stuck *hub worker*, and on this
+rig the cycle lands on the controller that is already wedged.
+
+### 5. Nothing bounds the damage, so one bad run becomes a CI outage
+
+- `hil-tinyusb` and `hil-tinyusb-esp` in `.github/workflows/build.yml` carry no
+ `timeout-minutes`. Only `hil-hfp-iar` does.
+- `ci.lan` runs a single runner service, so there is one job slot.
+- `test/hil/hil_test.py` bounds the pool with `POOL_TIMEOUT` (4200 s), and that guard
+ fires correctly — but the recovery path does not survive a D-state worker:
+
+```python
+with Pool(processes=os.cpu_count() or 1, initializer=init_worker, initargs=initargs) as pool:
+ async_ret = pool.map_async(test_board, config_boards)
+ try:
+ mret = async_ret.get(timeout=POOL_TIMEOUT)
+ except MpTimeoutError:
+ pool.terminate()
+ pool.join() # blocks forever: a D-state worker never reaps
+ raise RuntimeError(f'HIL worker pool timed out after {POOL_TIMEOUT}s')
+```
+
+`multiprocessing` joins workers unbounded, so both `pool.terminate()` and
+`pool.join()` hang, as does the `with Pool(...)` exit on the success path. Normal
+`hil-tinyusb (tinyusb.json)` runs take 10-20 minutes; one recent run took 71.3
+minutes, which is the 70-minute guard firing and succeeding. The eight-hour run is the
+pathological case.
+
+## Design
+
+### Step 0 — recovery (manual prerequisite)
+
+Power-cycle the PVE **host**, not the `ci` guest. A guest reboot is not sufficient;
+hubs latch up across the PCIe reset, which the 10:48 reboot demonstrated. Nothing
+below can be verified until the rig is clean.
+
+### Section 1 — CI containment
+
+**1a. Two layered timers.** An inner guard inside `hil_test.py` (`POOL_TIMEOUT`, 70 min)
+that fails gracefully -- it writes a report naming the timeout and the dispatched boards,
+shuts the pool down and exits -- and an outer `timeout-minutes` per rig job (85 for the
+hil-tinyusb jobs; 115 for hil-hfp-iar, which also builds four boards with IAR in the same
+job) as the backstop for when even exiting cannot free the runner. The ceiling must stay
+ABOVE the inner guard, or GitHub kills the job before the report is written.
+
+> **Corrected after measurement.** An earlier revision cut the guard to 30 min on the
+> reading that real runs take 9-17 min and everything longer was the old guard firing.
+> That was wrong. `hil_lock.py` records 22.2/14.3/12.5/10.8 min at usbtest width 1/2/3/4,
+> and raising the per-battery budget to 380s made hung boards cost more again. The 30 min
+> guard then fired on 5 of the last 8 HIL job executions across both rigs, and because
+> `map_async` is all-or-nothing each of those runs published a banner instead of any
+> per-board result. Restored to 4200s, the value whose original rationale -- usbtest
+> batteries are serialized fleet-wide, lengthening the tail -- was correct.
+
+**1b. Bound the pool shutdown.** Add a helper to `test/hil/hil_test.py`:
+
+```python
+def _shutdown_pool(pool, grace=30):
+ """terminate() a Pool without ever blocking forever: multiprocessing joins its
+ workers unbounded, and a worker in uninterruptible sleep (wedged usbfs) never
+ reaps -- which would hold the runner's only job slot indefinitely."""
+ t = threading.Thread(target=pool.terminate, daemon=True)
+ t.start()
+ t.join(grace)
+ return not t.is_alive()
+```
+
+On the `MpTimeoutError` path: write the report first, recording the boards that never
+reported so the run stops producing an empty report directory; then `_shutdown_pool`;
+then `os._exit(1)` if it did not return. The hard exit is the point — it is the only
+way past a kernel-side unkillable child. Use the same helper for the `with Pool(...)`
+exit path.
+
+**1c. Pre-flight rig health check.** `check_rig_health()` runs before the build and
+**never aborts**. It probes `/proc` unprivileged (dmesg is restricted on the rig) for a
+wedged `usb_hub_wq` worker, and reports a `/proc` too restricted to trust as its own
+distinct cause rather than as a diagnosed fault.
+
+It is deliberately non-fatal: the rig is unattended and every remedy for a real wedge is
+manual, so aborting would not fix anything -- it would discard the per-board results the
+run can still collect and leave CI red until a human noticed. It emits a GitHub
+`::error::` annotation and continues. The automatic containment is 1a and 1b, which bound
+a stuck run and explain it without anyone touching the rig.
+
+**1d. Order the recovery correctly.** In `test/hil/usbtest.py`, attempt
+`usb_recover.sh root-cycle` FIRST on a `HUNG` case, and only check for a wedged hub worker
+*afterwards*.
+
+> **Corrected during implementation.** This section originally said to check for a wedged
+> worker *before* the cycle and skip it on a hit. That is backwards. Our own stuck
+> `testusb` holds the DUT's device lock, so any port event drives a hub worker into
+> `usb_lock_device()` on it -- uninterruptible, so it reads `D` in ~100% of samples and the
+> confirmation window makes the wrong verdict *more* confident, not less. Cutting VBUS is
+> precisely what completes the in-flight URB, returns the ioctl and frees that worker, so
+> gating on that signature would suppress the recovery in the exact ordering it exists for.
+> A worker still wedged after the cycle is the genuinely unrecoverable case, and that is
+> what the code now reports.
+
+## Verification
+
+- Unit-test `shutdown_pool` and the `hil_health` detectors against a synthetic `/proc`.
+ A real wedge cannot be manufactured on demand, so they are tested against fabricated
+ inputs rather than live hardware.
+- Confirm the detectors flag a genuinely wedged rig, and return clean on a healthy one.
+- One clean full-fleet `hil_test.py` run to prove `check_rig_health` does not
+ false-abort.
+
+## Out of scope
+
+- **`ra6m5_ek` park and its dfu reset loop.** Dropped by decision. Consequence: the
+ layer-2 devnum storm remains as standing pressure on controller `03:00.0`. Unplugging
+ the board or flashing `board_test` by hand resolves it without any code change.
+- **An unattended PVE watchdog** that detects the wedge and power-cycles the host.
+ Declined: more moving parts, and it can cut a running CI job.
diff --git a/docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md b/docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md
new file mode 100644
index 000000000..e01831d34
--- /dev/null
+++ b/docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md
@@ -0,0 +1,162 @@
+# Bus-reset edge events + review fix wave — design
+
+Date: 2026-08-15
+Branch: `fix-ci-hs` (unpushed, 6 commits over master `53fef2833`)
+
+## Problem
+
+A max-effort review of the branch produced 15 findings. Four are regressions the branch
+itself introduced; the rest are pre-existing or cross-cutting. The load-bearing one:
+
+`dcd_ci_hs.c` now runs the RM-prescribed reset cleanup at the URI (reset-start) interrupt
+but does not tell usbd until the Port Change Detect that ends the reset. For the whole
+reset window — a minimum of 3 ms, typically 10–50 ms — usbd still believes the device is
+configured while the DCD's queue heads have been zeroed. A class driver writing in that
+window (`tud_hid_n_report()`, `tud_cdc_write_flush()`) primes a disabled endpoint over a
+zeroed dQH, *after* the cleanup's flush, so the stale prime survives re-enumeration over a
+buffer usbd has already released. On a 600 MHz M7 that window is enormous. Master had no
+gap: cleanup and event were adjacent statements.
+
+The stack has no way to express "reset started" — `DCD_EVENT_BUS_RESET` carries the
+negotiated speed, which does not exist until the reset ends. That missing vocabulary is
+the actual defect; the driver-level workarounds considered (deferring the memclr, guarding
+primes with a private flag) only shrink the window.
+
+## Design
+
+### 1. Stack: split the bus-reset event into two edges
+
+`src/device/dcd.h`:
+
+```c
+DCD_EVENT_BUS_RESET_START, // reset signaling detected; bus unusable, speed unknown
+DCD_EVENT_BUS_RESET_END, // reset complete; .bus_reset.speed is final
+...
+#define DCD_EVENT_BUS_RESET DCD_EVENT_BUS_RESET_END // backward compatibility
+```
+
+No new helper: `dcd_event_bus_reset(rhport, speed, in_isr)` keeps its name and emits
+`_END`, so every other port is bit-identical to today; `_START` uses the existing
+payload-free `dcd_event_bus_signal()`. The alias keeps unit-test/fuzz references
+compiling.
+
+**Contract (documented in `dcd.h`):** `_START` is optional. A DCD that cannot distinguish
+the two edges emits only `_END`, which stays self-sufficient — it performs the full
+teardown with or without a preceding `_START`.
+
+`src/device/usbd.c`:
+- `case DCD_EVENT_BUS_RESET_START:` → `usbd_reset(rhport)` only; speed untouched.
+- `case DCD_EVENT_BUS_RESET_END:` → unchanged (`usbd_reset()` + latch speed).
+- `_usbd_event_str[]` gains both names.
+- `TODO:` note that a DCD signalling both edges should not pay for two teardowns — track
+ a per-rhport "start seen" flag and skip the redundant `usbd_reset()` in `_END`, keeping
+ the unconditional teardown for the legacy single-event path.
+
+Cost, accepted deliberately: one extra queued event and one extra `usbd_reset()` per
+enumeration on ci_hs only, bounded at one per reset against a default
+`CFG_TUD_TASK_QUEUE_SZ` of 16 (queue pressure is the failure PR #3817 fixed, hence the
+explicit note).
+
+### 2. ci_hs: split `bus_reset()` along the register/software line
+
+- **`bus_reset_begin()` — at URI, inside the reset window (UM10503 25.10.3):** ENDPTCTRL
+ type-reset loop, `ENDPTNAK`/`ENDPTNAKEN`, `ENDPTSETUPSTAT` and `ENDPTCOMPLETE`
+ write-back clears, bounded `ENDPTPRIME` drain, `ENDPTFLUSH` all. Emit `_START`.
+ Registers only — nothing in `_dcd_data` is touched, so no software structure is pulled
+ out from under a task mid-`dcd_edpt_xfer`.
+- **`bus_reset_complete()` — at the PCI ending the reset:** re-flush, `tu_memclr(&_dcd_data)`,
+ EP0 queue-head re-init, dcache clean. Emit `_END` with the final PSPD speed.
+
+Two properties fall out: the re-flush kills any prime armed during the window without a
+new state flag, and the memclr now happens at the same instant usbd is told, so the
+"configured over zeroed queue heads" mismatch is eliminated rather than shrunk. Residual
+exposure (a task priming exactly as the ISR memclrs) equals master's.
+
+The reason-dispatch (`pci_reason`, suspend/URI ordering) is unchanged; only the reset
+case's body moves.
+
+### 3. ci_hs: one bounded-flush helper
+
+Extract `flush_endpoints(dcd_reg, mask)` — writes `ENDPTFLUSH = mask`, spins bounded by
+`CI_HS_BUSY_SPIN` until those bits clear, returns `true` if they cleared — and route all
+five flush sites through it (`bus_reset_begin`, `bus_reset_complete`, `dcd_deinit`,
+`dcd_edpt_iso_activate`, the setup-time EP0 flush). The unified part is the mechanism
+(one bound, one spin idiom, one return convention); callers keep their existing reactions,
+all of which currently proceed regardless, and that stays true here — no caller gains new
+error handling in this wave. Without this, §2 adds a fifth site to a file that already
+carried four hand-rolled variants.
+
+### 4. Mechanical fixes
+
+`dcd_ci_hs.c`
+- Setup-time EP0 flush waits for completion (via §3's helper) before the SETUP event is
+ queued, so the flush can no longer still be asserted when the task primes the response —
+ which also dissolves its interaction with the post-prime verify. This adds a bounded
+ spin in ISR context; the RM notes a flush waits out any packet already in progress, so
+ the wait is one packet time (microseconds at HS) and the existing `CI_HS_BUSY_SPIN`
+ bound caps the pathological case, consistent with the file's other flush sites.
+- `dcd_set_address()` writes `DEVICEADDR` only if the status-ZLP prime took. A refused
+ prime means a newer SETUP superseded the transfer; staging an address whose ACK will
+ never arrive is wrong.
+- Emit `DCD_EVENT_RESUME` only when `!(PORTSC1 & PORTSC1_SUSPEND)` (restores master's
+ hardware guard, lost in the rework).
+
+`dcd_lpc_ip3511.c`
+- Deliver the setup copy only when known-good:
+ `if (latch still set) { INTSETSTAT = TU_BIT(0); } else { dcd_event_setup_received(...); }`.
+- `TODO:` token on the USB.13 deferral so backlog sweeps surface it.
+
+`usbd.c`
+- The DCD-refusal path in `usbd_edpt_xfer` stops routing through the breakpoint-carrying
+ assert: a DCD declining a prime is documented and self-healing, not a programming error,
+ and `TU_BREAKPOINT()` is not gated on `CFG_TUSB_DEBUG` — with a probe attached (always,
+ on the rig) it halts the target. Log and return false instead.
+
+BSP
+- Delete the seven-line RHPORT block in `lpcxpresso55s28/board.cmake` (byte-identical to
+ `family.cmake`'s own guards; `board.mk`'s `?=` stays as the idiomatic Make form).
+- `lpc11u37.ld`: correct the stale comment (nothing lands in RamUsb2 in either build
+ system now — the stack owns the whole bank) and keep the ASSERT, re-labelled as
+ future-proofing.
+
+## Findings improved for free (documented, no code)
+
+A reset that starts and never completes — cable pulled mid-reset — now delivers `_START`
+and tears usbd down, where before usbd stayed configured on a dead bus. This softens both
+the adjudicated UNPLUGGED-removal finding and the deferred aborted-reset item: a stray
+later PCI delivering `_END` becomes harmless (usbd already torn down, just latches a
+speed) instead of deconfiguring a live device. True detach detection still requires OTGSC
+B-session-valid VBUS sensing — board-dependent, still a follow-up.
+
+## Explicitly deferred
+
+- Prime verification generalized to all endpoints and all causes (RM 25.10.8.2); the
+ EP0/SETUP-gated form stays, its flush interaction fixed by §4.
+- usbd discards `usbd_control_xfer_cb`/`tud_control_xfer` returns — cross-DCD behavior
+ change needing its own regression pass, despite `usbd.c` being open here.
+- Timed-out flush still proceeds to the memclr (now confined to one helper).
+- LPC55S2x USB.3 FORCE_FS workaround; iso-IN 1023 enforcement; 8-byte OUT-spill
+ enforcement; USB.13 INTONNAK workaround.
+- Gating `TU_BREAKPOINT()` on `CFG_TUSB_DEBUG` stack-wide.
+- Unguarded `set()` RHPORT knobs in ~14 sibling `board.cmake` files.
+
+## Verification
+
+1. `pre-commit run --all-files`; builds for mimxrt1064_evk, lpcxpresso18s37,
+ lpcxpresso11u37, lpcxpresso55s28, plus Make link checks for the two previously-broken
+ targets (`host/cdc_msc_hid` on 55s28, `device/cdc_msc_throughput` on 11u37).
+2. Cross-DCD build guard: one non-ci_hs, non-ip3511 board (e.g. `stm32f407disco`) to prove
+ the `DCD_EVENT_BUS_RESET` alias keeps legacy ports compiling untouched.
+3. HIL on byte-verified flash (`verifyfile` on every J-Link load — the 1064's silent
+ flash no-op has struck twice): usbtest 30/30 on mimxrt1064_evk, lpcxpresso55s28,
+ lpcxpresso11u37; 50× case-9/10 loops on the 1064; 10× case-11/12/24 unlink loops.
+4. Reset-path specific: confirm HS enumeration (480) and, with `LOG=2`, that a single
+ enumeration shows exactly one `_START`/`_END` pair and no spurious RESUME.
+5. Suspend/resume exercise on the 1064 (host-side autosuspend on the port) confirming
+ `SUSPEND`/`RESUME` pairing and no reset misclassification.
+
+## Success criteria
+
+All four regressions closed, no new findings in a scoped re-review of the wave diff, every
+listed HIL result green on verified flash, and legacy DCDs provably untouched (alias build
+check + unchanged `_END` semantics).
diff --git a/docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md b/docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md
new file mode 100644
index 000000000..cc1840972
--- /dev/null
+++ b/docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md
@@ -0,0 +1,90 @@
+# Drop the EP0 post-prime verify — design
+
+Date: 2026-08-16
+Branch: `fix-ci-hs` (unpushed, 19 commits over merge-base `53fef2833`)
+
+## Context
+
+The branch grew while chasing a wedge on `mimxrt1064_evk`: the board would stop answering a
+host transfer, the URB would never complete, `testusb` would block uninterruptibly and the
+whole rig would follow it down. Eight occurrences over four days, across the Linux usbtest
+battery's queued control and bulk tests.
+
+The cause turned out to be silicon: **Errata i.MX RT1064_A / RT1060_A ERR050101**. While an
+isochronous IN endpoint is active, an IN token addressed to that same endpoint number on
+another device sharing the host silently unprimes one of this device's OUT endpoints —
+control, bulk, interrupt or isochronous. NXP states it cannot be detected by software and
+raises no interrupt. Moving the usbtest example's iso IN endpoint from 3 to 7 (commit
+`42870b15b`) cleared it: 340 consecutive wedge-free runs, where the board previously
+re-wedged within hours.
+
+Before that was known, an earlier theory — a SETUP arriving mid-prime silently cancelling an
+EP0 prime — produced a post-prime verification block in `qhd_start_xfer()`. That theory's
+supporting capture (EP0's status ZLP armed but unprimed, the device a control transfer ahead
+of the host) is explained by ERR050101 just as well, because the errata explicitly covers
+*control* OUT endpoints and a control status stage **is** an OUT endpoint. The generalized
+version of that verify was already reverted (`565bb0d99`) as both regression-prone and aimed
+at a failure the vendor documents as undetectable in software. This spec removes what
+remains of it.
+
+## Change
+
+Delete the post-prime block in `qhd_start_xfer()` (`src/portable/chipidea/ci_hs/dcd_ci_hs.c`):
+the bounded `ENDPTPRIME` drain, the `ENDPTFLUSH`-on-timeout, and the
+`ENDPTSTAT | ENDPTCOMPLETE` / `ENDPTSETUPSTAT` verdict. The tail becomes:
+
+```c
+ // start transfer
+ dcd_reg->ENDPTPRIME = TU_BIT(epnum + (dir ? 16 : 0));
+ return true;
+```
+
+This removes two register spins and four volatile reads from every EP0 transfer, and with
+them the false-fail path a reviewer flagged: a transfer the interrupt handler has already
+completed reads identically to a cancelled prime.
+
+## Deliberately kept
+
+- **The pre-prime setup-lockout guard** directly above it — UM10503 25.10.8.1.1 step 4
+ verbatim ("Before priming for status/handshake phases ensure that ENDPTSETUPSTAT is '0'"),
+ and older than the wedge theory. It also keeps `qhd_start_xfer()` returning `bool`, so
+ `dcd_set_address()`'s gating and the usbd breakpoint removal stay meaningful — no cascade.
+- **The setup-time EP0 flush and its completion wait** — the flush is the 25.10.8.1.1 step-3
+ remark; the wait exists because an unfinished flush can retire a freshly primed response,
+ an interaction independent of the verify.
+- **The `BUS_RESET_START`/`END` split** and the rest of the review-driven hardening.
+- Everything hardware-proven: the rf_tv fix, the lpc11u37 stack move, the lpc55s28
+ onboarding, the lpc55 Make OHCI link, and the ERR050101 endpoint move itself.
+
+The commit message records the corrected attribution of the handoff capture, so the next
+reader does not re-derive the superseded theory from the same evidence.
+
+## Validation
+
+The "with it" arm is already banked from 2026-08-16: 10x 30/30 batteries plus 15x TEST 27,
+15x tests 9/10 and 10x tests 11/12/24, all clean. This is the second half of an A/B.
+
+1. **Rebase onto current master first** (master has moved: midi2/usbtmc/video), then rebuild —
+ otherwise the validated tree is not the tree that merges.
+2. **Software gates:** `pre-commit run --all-files`; full example builds for
+ mimxrt1064_evk, lpcxpresso18s37, lpcxpresso11u37, lpcxpresso55s28; the two Make link
+ canaries (`host/cdc_msc_hid` on lpcxpresso55s28, `device/cdc_msc_throughput` on
+ lpcxpresso11u37); `ceedling test:all`.
+3. **Hardware — mimxrt1064_evk only.** It is the only ci_hs board on the rig; the other two
+ run ip3511, which this change does not touch. Preconditions: CI idle
+ (`pgrep -f "hil_test.py [-]-retry"`), board lock held for the whole run. Flash with
+ `loadfile` (its built-in Program & Verify — JLinkExe V9.66 has no `verifyfile`), then
+ confirm re-enumeration as `cafe:4010` with serial `BAE96FB95AFA6DBB8F00005002001200`, and
+ confirm `lsusb -v` still reports the iso IN endpoint as **0x87** so a stale image cannot
+ masquerade as a pass.
+4. **Runs:** 5x the full 30-case battery, then 15x `--tests 9,10,14,21` (queued control, ch9
+ subset, both ctrl_out cases) — the control paths the verify actually protected, which a
+ plain battery samples only once per run. Print a `testusb` D-state scan after every
+ iteration.
+
+**Acceptance:** 5/5 batteries at 30/30, 15/15 loops, and no `testusb` D-state outliving its
+case runtime.
+
+**Rollback trigger:** any control-case failure (errno 110 or 71 on cases 9, 10, 14, 21) or a
+lingering D-state means the verify was load-bearing after all — restore it and record that
+result in the commit message. A negative result is a finding, not a setback.