summaryrefslogtreecommitdiff
path: root/docs/superpowers
diff options
context:
space:
mode:
authorHiFiPHile <[email protected]>2026-08-25 09:55:24 +0200
committerHiFiPHile <[email protected]>2026-09-02 10:50:02 +0200
commita0d3de76869ce728c7c1d9713085bc970da39671 (patch)
treebd24f3efb60daeda2e0aaf4df08da97504aafbf3 /docs/superpowers
parent278c0531a07e4fe8112ef91d0d2dbba8ef0a40b3 (diff)
parent5c0e31cdabaf37f14e1f5e988a020abfc1000495 (diff)
Merge branch 'master' into agent/fix-dwc2-host-fifo-allocation
Diffstat (limited to 'docs/superpowers')
-rw-r--r--docs/superpowers/followup/pr3803-flasher-recover.md280
-rw-r--r--docs/superpowers/followup/pr3803-hil-blindness-reporting.md185
-rw-r--r--docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md118
-rw-r--r--docs/superpowers/followup/pr3803-pci-rebind-stranding.md157
-rw-r--r--docs/superpowers/followup/pr3803-usbtest-recovery-reserve.md175
-rw-r--r--docs/superpowers/plans/2026-08-15-ci-hs-reset-edges.md782
-rw-r--r--docs/superpowers/plans/2026-08-16-drop-ep0-prime-verify.md314
-rw-r--r--docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md22
-rw-r--r--docs/superpowers/specs/2026-07-30-hil-usbtest-fleet-wedge-design.md236
-rw-r--r--docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md162
-rw-r--r--docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md90
11 files changed, 2510 insertions, 11 deletions
diff --git a/docs/superpowers/followup/pr3803-flasher-recover.md b/docs/superpowers/followup/pr3803-flasher-recover.md
new file mode 100644
index 000000000..e9fff7480
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-flasher-recover.md
@@ -0,0 +1,280 @@
+# `flasher_recover` Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Give the 15 HIL boards whose flasher cannot reach its probe past a poisoned usbfs
+node a second, convoy-safe flasher used only for recovery.
+
+**Architecture:** An optional roster key `flasher_recover` beside `flasher`.
+`hil_flash.recover_flasher(board)` picks it when present; `hil_test` substitutes it into the
+`--recover-board` JSON so `usbtest.py` never learns a second entry exists. Delivery over
+openocd's jlink driver is convoy-safe by construction, but the flash command form must
+differ from the one `flash_openocd` uses, so the recovery gets its own flasher name.
+
+**Tech Stack:** Python 3.13 stdlib, openocd 0.12.0+dev (build 0ce743125 on ci.lan),
+libjaylink, J-Link probes.
+
+## Global Constraints
+
+- Roster JSON: `test/hil/tinyusb.json`. `flasher_recover` is OPTIONAL; absent means today's
+ behaviour (`recover_flasher` returns the primary).
+- Never change the shape of `board['flasher']` — it is read as a dict in `hil_flash`,
+ `hil_test`, `usbtest`, `hil_pool_check`, `hil_select` and the roster lint, and is shipped
+ as JSON to a subprocess.
+- Flasher dispatch is by name: `getattr(hil_flash, f'flash_{name}')` / `reset_{name}`.
+- `RECOVER_FLASH_TIMEOUT = 90`, `RECOVER_RESET_TIMEOUT = 30` (`usbtest.py`). Any board whose
+ flash cannot finish inside 90 s is not a candidate.
+- Tests run offline: `cd test/hil && python3 test/test_hil_select.py`.
+
+## What is already established
+
+**Landed on PR #3803 and inert without roster entries:** `hil_flash.recover_flasher()`,
+`convoy_safe()` accepting openocd-over-jlink, `hil_test` substituting the recovery flasher
+into `--recover-board`, and `test_hil_select.FlasherRecoverEntry` (4 tests).
+
+**Verified in source:**
+- openocd's jlink driver ignores `adapter usb vid_pid` — `jlink.c` never reads
+ `adapter_usb_get_vids/pids`; selection is `adapter serial` / USB address / usb location.
+ Do NOT lint a jlink recovery entry for `vid_pid`.
+- It is convoy-safe anyway: libjaylink `discovery_usb.c` returns early unless
+ `idVendor == 0x1366` and the PID is in its table, and only THEN calls `libusb_open`. A
+ wedged `cafe:4010` DUT is never opened.
+- CMSIS-DAP stays pin-gated: `cmsis_dap_usb_bulk.c:107` skips before `libusb_open`, and
+ `id_filter` is only `vids[0] || pids[0]`.
+
+**Measured on ci.lan 2026-08-17**, base args
+`-f interface/jlink.cfg -c "transport select swd" -c "adapter speed 4000" -f target/<cfg>`:
+
+| Board | target cfg | flash | reset |
+|--------------------------|--------------|-------|-------|
+| stm32f407disco | stm32f4x | OK | OK |
+| stm32f072disco | stm32f0x | OK | OK |
+| stm32f723disco | stm32f7x | OK | OK |
+| stm32l476disco | stm32l4x | OK | OK |
+| feather_nrf52840_express | nrf52 | OK | OK |
+| metro_m4_express | atsame5x | OK | OK |
+| frdm_k64f | k60 | OK | OK |
+
+`frdm_k64f` is host-only (`tests.device == false`) — verify its reset over UART
+(`/dev/serial/by-id/usb-SEGGER_J-Link_000621000000-if00`), never by USB disconnect.
+
+**Excluded, with reasons:** `lpcxpresso11u37` — 118 s for 24 KB at 1 MHz with a verify
+mismatch, versus 0.277 s via JLinkExe; cannot fit `RECOVER_FLASH_TIMEOUT`.
+`mimxrt1064_evk`, `ra4m1_ek`, `nrf54lm20dk` — no target config exists in this openocd
+build, so they cannot be covered at all. **The board that wedges most (mimxrt1064_evk) is
+therefore still uncovered by this work.**
+
+**The blocker this plan solves:** `flash_openocd` issues `program <fw> verify reset exit`,
+which fails over the jlink transport on BOTH families tried (`stm32f4x`, `stm32f0x`) with
+`Examination failed` → `auto_probe failed`, with or without a preceding `init; reset halt`.
+Every successful flash above used the explicit sequence in Task 1.
+
+**Why this is a separate PR:** it adds a roster capability and a new flasher backend, which
+is a different scope from containing a wedge; and it needs bench time on seven boards.
+
+## File Structure
+
+- `test/hil/hil_flash.py` — add `flash_openocd_seq` / `reset_openocd_seq`; extend
+ `convoy_safe` to accept the new name. This is the only file that learns the command form.
+- `test/hil/tinyusb.json` — seven `flasher_recover` entries.
+- `test/hil/test/test_hil_select.py` — extend `FlasherRecoverEntry`; add a roster lint.
+
+---
+
+### Task 1: `openocd_seq` flasher backend
+
+**Files:**
+- Modify: `test/hil/hil_flash.py` (beside `flash_openocd`, ~line 100)
+- Test: `test/hil/test/test_hil_select.py`
+
+**Interfaces:**
+- Consumes: `_openocd_cmd_base(flasher)`, `hil_util.run_cmd`.
+- Produces: `flash_openocd_seq(board, firmware, timeout=None)`,
+ `reset_openocd_seq(board, timeout=None)`, both returning
+ `subprocess.CompletedProcess`; `convoy_safe()` returns True for
+ `{'name': 'openocd_seq', 'args': '...interface/jlink.cfg...'}`.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_openocd_seq_is_convoy_safe_over_jlink(self):
+ self.assertTrue(hil_flash.convoy_safe(
+ {'name': 'openocd_seq', 'args': '-f interface/jlink.cfg -f target/stm32f4x.cfg'}))
+
+ def test_openocd_seq_uses_explicit_flash_commands_not_program(self):
+ """`program` fails over the jlink transport: Examination failed -> auto_probe
+ failed, measured on stm32f4x and stm32f0x."""
+ seen = {}
+ real = hil_util.run_cmd
+ hil_util.run_cmd = lambda cmd, **k: seen.setdefault('cmd', cmd) or real('true')
+ try:
+ hil_flash.flash_openocd_seq(
+ {'flasher': {'name': 'openocd_seq', 'uid': 'X', 'args': '-f interface/jlink.cfg'}},
+ '/tmp/fw.elf', timeout=5)
+ finally:
+ hil_util.run_cmd = real
+ self.assertIn('flash write_image erase /tmp/fw.elf', seen['cmd'])
+ self.assertIn('verify_image /tmp/fw.elf', seen['cmd'])
+ self.assertNotIn('program ', seen['cmd'])
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_select.py FlasherRecoverEntry -v`
+Expected: FAIL — `module 'hil_flash' has no attribute 'flash_openocd_seq'`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def flash_openocd_seq(board, firmware, timeout=None):
+ # Explicit commands, NOT `program`: over the jlink transport `program` fails at the
+ # flash bank probe ("Examination failed" -> "auto_probe failed"), measured on
+ # stm32f4x and stm32f0x, with or without a preceding reset halt. This sequence
+ # succeeded on all seven candidate boards.
+ flasher = board['flasher']
+ verify = f' -c "verify_image {firmware}"' if flasher.get('verify', True) else ''
+ return hil_util.run_cmd(
+ f'{_openocd_cmd_base(flasher)} -c "init" -c "reset halt" '
+ f'-c "flash write_image erase {firmware}"{verify} -c "reset run" -c "shutdown"',
+ timeout=timeout)
+
+
+def reset_openocd_seq(board, timeout=None):
+ flasher = board['flasher']
+ return hil_util.run_cmd(
+ f'{_openocd_cmd_base(flasher)} -c "init" -c "reset run" -c "shutdown"',
+ timeout=timeout)
+```
+
+In `convoy_safe`, replace `if name != 'openocd':` with:
+
+```python
+ if name not in ('openocd', 'openocd_seq'):
+ return False
+```
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_select.py FlasherRecoverEntry -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/hil_flash.py test/hil/test/test_hil_select.py
+git commit -m "hil: add openocd_seq flasher for convoy-safe recovery delivery"
+```
+
+---
+
+### Task 2: Roster entries for the seven validated boards
+
+**Files:**
+- Modify: `test/hil/tinyusb.json`
+- Test: `test/hil/test/test_hil_select.py`
+
+**Interfaces:**
+- Consumes: `flash_openocd_seq` / `reset_openocd_seq` from Task 1.
+- Produces: seven boards for which `hil_flash.convoy_safe(hil_flash.recover_flasher(b))`
+ is True.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_roster_recover_entries_are_convoy_safe_and_named_openocd_seq(self):
+ import json, pathlib
+ roster = json.loads((pathlib.Path(__file__).parent.parent / 'tinyusb.json').read_text())
+ recover = [b for b in roster['boards'] if 'flasher_recover' in b]
+ self.assertGreaterEqual(len(recover), 7)
+ for b in recover:
+ f = b['flasher_recover']
+ self.assertEqual(f['name'], 'openocd_seq', b['name'])
+ self.assertIn('interface/jlink.cfg', f['args'], b['name'])
+ self.assertIn('adapter speed', f['args'], b['name']) # required; see below
+ self.assertTrue(hil_flash.convoy_safe(f), b['name'])
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_select.py FlasherRecoverEntry -v`
+Expected: FAIL — `0 >= 7`
+
+- [ ] **Step 3: Add the entries**
+
+`adapter speed` is REQUIRED: without it examination fails outright on the jlink driver.
+Add to each board below, using the SAME `uid` as its primary jlink entry:
+
+```json
+"flasher_recover": {
+ "name": "openocd_seq",
+ "uid": "<same probe serial as flasher.uid>",
+ "args": "-f interface/jlink.cfg -c \"transport select swd\" -c \"adapter speed 4000\" -f target/<cfg>.cfg"
+}
+```
+
+| Board | `uid` | `<cfg>` |
+|--------------------------|----------------|-----------|
+| stm32f407disco | 000773661813 | stm32f4x |
+| stm32f072disco | 779541626 | stm32f0x |
+| stm32f723disco | 000776606156 | stm32f7x |
+| stm32l476disco | 777632258 | stm32l4x |
+| feather_nrf52840_express | 681295394 | nrf52 |
+| metro_m4_express | 123456 | atsame5x |
+| frdm_k64f | 000621000000 | k60 |
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_select.py -v`
+Expected: PASS, and no other selector test regresses.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/tinyusb.json test/hil/test/test_hil_select.py
+git commit -m "hil: give seven J-Link boards a convoy-safe recovery flasher"
+```
+
+---
+
+### Task 3: Bench validation on the rig
+
+**Files:** none — this task produces evidence, not code.
+
+- [ ] **Step 1: Confirm the rig is idle and take the locks**
+
+```bash
+ssh [email protected] 'if pgrep -f "[h]il_test.py" >/dev/null; then echo BUSY; exit 1; fi'
+ssh [email protected] 'cd ~/actions-runner/_work/tinyusb/tinyusb && \
+ nohup timeout 900 python3 test/hil/helper/hil_lock.py hold <boards...> --reason "flasher_recover validation" &'
+```
+
+Guard with `if`, never `cmd && echo || echo` — that form only gates the echo and will take
+locks during a live CI run.
+
+- [ ] **Step 2: For each board, flash then reset through the recovery entry**
+
+```bash
+python3 test/hil/hil_test.py -b <board> test/hil/tinyusb.json # normal path still works
+```
+
+Then force the recovery path by running usbtest with the recovery flags and a firmware that
+hangs a case, or drive `hil_flash.flash_openocd_seq` / `reset_openocd_seq` directly.
+
+- [ ] **Step 3: Verify**
+
+Device boards: `sudo dmesg` shows `USB disconnect` then a fresh enumeration.
+`frdm_k64f`: UART shows the boot banner (see above).
+Every flash must finish well inside `RECOVER_FLASH_TIMEOUT` (90 s).
+
+- [ ] **Step 4: Release locks and record the results in the PR body**
+
+---
+
+## Out of scope, and why
+
+- **`mimxrt1064_evk`** needs an i.MX RT target config that this openocd build does not
+ have. Sourcing or writing one is its own investigation; until then the board with the
+ most wedges has no automated recovery.
+- **Changing `flash_openocd`** to the explicit form would cover these boards without a new
+ name, but `program` is what nine pinned CMSIS-DAP boards use in CI daily and no CMSIS-DAP
+ image could be built in the originating worktree (no pico-sdk) to re-validate it.
diff --git a/docs/superpowers/followup/pr3803-hil-blindness-reporting.md b/docs/superpowers/followup/pr3803-hil-blindness-reporting.md
new file mode 100644
index 000000000..69ff939b0
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-hil-blindness-reporting.md
@@ -0,0 +1,185 @@
+# Blindness Reporting Gaps Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Make a HIL worker's sysfs blindness reach the report in the two cases where it
+currently does not — an untested producer, and a board that raises.
+
+**Architecture:** A worker returns `hil_util.sysfs_blind()` as the last field of its result
+tuple; `_blind_note()` turns that into a report banner. Two holes: nothing tests the
+producer, and a board that raises returns no tuple at all, so its blindness is lost.
+
+**Tech Stack:** Python 3.13 stdlib, multiprocessing Pool with `maxtasksperchild=1`.
+
+## Global Constraints
+
+- A blind worker answers `SYSFS_UNKNOWN` for every attribute, so its "device not found"
+ means "could not tell". The report must say so or a red cell reads as a broken board.
+- `maxtasksperchild=1`: one worker per board, so the flag is per-board and must not be
+ smeared across boards.
+- Tests: `cd test/hil && python3 test/test_hil_bounded.py`.
+
+## What is already established
+
+- `hil_test.test_board` returns `(..., hil_util.sysfs_blind(), stray)`; `_blind_note(mret)`
+ renders the banner; wired into all three report paths.
+- **The producer is provably untested**: replacing `hil_util.sysfs_blind()` with `False` in
+ the return leaves all tests green. Nothing drives `test_board` — it needs a board dict, a
+ real flock, a flasher and `test_example` per test.
+- Blindness fired for real on ci.lan: four workers went blind in one run, and cells failed
+ *because* of it (`Printer device not found ... (this worker is blind)`).
+
+**Why this is a separate PR:** closing it means making `test_board` testable, which is a
+refactor of the harness's orchestration layer — a different scope from the containment
+work, and the reason the gap was accepted rather than papered over.
+
+## File Structure
+
+- `test/hil/hil_test.py` — extract the result-tuple assembly from `test_board` so it can be
+ built and asserted without running a board; carry blindness out of the raise path.
+- `test/hil/test/test_hil_bounded.py` — tests for both.
+
+---
+
+### Task 1: Make the result tuple assembly testable
+
+**Files:**
+- Modify: `test/hil/hil_test.py` (`test_board`, the `return (name, err_count, ...)` at the
+ end of the try block)
+- Test: `test/hil/test/test_hil_bounded.py`
+
+**Interfaces:**
+- Produces: `_board_result(name, err_count, failed_tests, rows, t_total, board_wide_fail)`
+ returning the 7-tuple `(name, err_count, failed, rows, t_total, blind, stray)`, reading
+ `hil_util.sysfs_blind()` and `hil_health.kill_own_children()` itself.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+class BoardResultCarriesBlindness(unittest.TestCase):
+ def test_a_blind_worker_reports_it(self):
+ from helper import hil_util, hil_health
+ self.addCleanup(setattr, hil_util, 'sysfs_blind', hil_util.sysfs_blind)
+ self.addCleanup(setattr, hil_health, 'kill_own_children', hil_health.kill_own_children)
+ hil_util.sysfs_blind = lambda: True
+ hil_health.kill_own_children = lambda: 0
+ row = hil_test._board_result('b', 0, [], [], 1.0, False)
+ self.assertTrue(row[5], 'blindness did not reach the result tuple')
+ self.assertIn('b', hil_test._blind_note([row]))
+
+ def test_a_sighted_worker_does_not(self):
+ from helper import hil_util, hil_health
+ self.addCleanup(setattr, hil_util, 'sysfs_blind', hil_util.sysfs_blind)
+ self.addCleanup(setattr, hil_health, 'kill_own_children', hil_health.kill_own_children)
+ hil_util.sysfs_blind = lambda: False
+ hil_health.kill_own_children = lambda: 0
+ row = hil_test._board_result('b', 0, [], [], 1.0, False)
+ self.assertFalse(row[5])
+ self.assertEqual(hil_test._blind_note([row]), '')
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py BoardResultCarriesBlindness -v`
+Expected: FAIL — `module 'hil_test' has no attribute '_board_result'`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def _board_result(name, err_count, failed_tests, rows, t_total, board_wide_fail):
+ """Assemble a worker's result tuple. Separate from test_board so the two fields only
+ the WORKER can answer -- its process-global blindness latch and what it could not kill
+ -- are testable without running a board."""
+ stray = hil_health.kill_own_children()
+ return (name, err_count, [] if board_wide_fail else sorted(set(failed_tests)),
+ rows, t_total, hil_util.sysfs_blind(), stray)
+```
+
+Replace the tail of `test_board` with:
+
+```python
+ return _board_result(name, err_count, failed_tests, rows, t_total, board_wide_fail)
+```
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py -v`
+Expected: PASS, and the existing `BlindWorkerReachesTheReport` tests still pass.
+
+- [ ] **Step 5: Verify the mutation is now caught**
+
+Replace `hil_util.sysfs_blind()` with `False` inside `_board_result` and re-run; the suite
+MUST fail. Restore it.
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add test/hil/hil_test.py test/hil/test/test_hil_bounded.py
+git commit -m "test/hil: make the worker result tuple testable, covering blindness"
+```
+
+---
+
+### Task 2: Carry blindness out of the worker-raise path
+
+**Files:**
+- Modify: `test/hil/hil_test.py` (`test_board`'s except/finally, and `main`'s worker-raise
+ handler that builds synthetic rows)
+- Test: `test/hil/test/test_hil_bounded.py`
+
+**Interfaces:**
+- Consumes: `_board_result` from Task 1.
+- Produces: a board that raises still contributes a row whose blindness field is accurate.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_a_board_that_raises_still_reports_blindness(self):
+ """The result tuple is returned inside a try whose finally only releases the lock,
+ so a board that dies by exception contributed nothing -- and its blindness, the
+ thing that most explains its failure, was lost with it."""
+ from helper import hil_util
+ self.addCleanup(setattr, hil_util, 'sysfs_blind', hil_util.sysfs_blind)
+ hil_util.sysfs_blind = lambda: True
+ row = hil_test._board_result_on_error('b', RuntimeError('boom'))
+ self.assertTrue(row[5])
+ self.assertIn('b', hil_test._blind_note([row]))
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py BoardResultCarriesBlindness -v`
+Expected: FAIL — no `_board_result_on_error`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def _board_result_on_error(name, exc):
+ """A row for a board that died by exception. err_count 1, no per-test detail, but the
+ blindness and stray fields are still accurate -- they explain the failure more often
+ than the exception text does."""
+ rows = [(name, {BOUNDARY_CELL: f'{REPORT_CELL["fail"]} {type(exc).__name__}'}, None)]
+ return _board_result(name, 1, [], rows, 0.0, True)
+```
+
+Wrap the body of `test_board` so the exception path returns it instead of propagating.
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/hil_test.py test/hil/test/test_hil_bounded.py
+git commit -m "test/hil: keep a raising board's blindness in the report"
+```
+
+---
+
+## Caution
+
+`test_board`'s `finally` releases the board flock. Any restructuring MUST keep that
+release on every path, including the new error path — a leaked flock locks the board until
+the host reboots.
diff --git a/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md b/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md
new file mode 100644
index 000000000..fe377f741
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md
@@ -0,0 +1,118 @@
+# IAR HIL Leg Re-run Spec Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Let the `hil-hfp-iar` CI leg re-run only its failed boards, as the other two HIL
+legs already do.
+
+**Architecture:** `hil_test.py` writes a `<config>.failed` spec into `HIL_REPORT_DIR`; a
+workflow step reads it on the next attempt and passes the boards back as arguments. The IAR
+leg passes `--retry 1` like the others but sets no `HIL_REPORT_DIR` and has no read-back
+step, so its spec is written into the workspace and never read.
+
+**Tech Stack:** GitHub Actions YAML, self-hosted runner.
+
+## Global Constraints
+
+- `.github/workflows/build.yml`. The two working legs are `hil-tinyusb` (matrix) — see its
+ `Set HIL report dir (per run+job; persists across run attempts)` and `Get re-run spec from
+ previous attempt` steps — and they are the pattern to copy.
+- The report dir must be keyed by run id AND job so a matrix leg does not collide with
+ another, and must survive across run attempts (that is the whole point).
+- The IAR leg is the only HIL job that BUILDS inline; its `Build` step is bounded at
+ `timeout-minutes: 30` under a 120-minute job ceiling. Do not disturb that.
+
+## What is already established
+
+- Verified by reading the workflow: `hil-hfp-iar` has neither `HIL_REPORT_DIR` nor a
+ `Get re-run spec` step, while passing `--retry 1`.
+- Consequence: a GitHub re-run of that job re-tests its whole matrix. **This is not a
+ regression** — that leg never had the mechanism — and the unread spec costs only a file.
+- The report artifact upload for that leg is named `hil-report-hfp-iar`.
+
+**Why this is a separate PR:** it is CI plumbing with no code change, it needs a real
+re-run on the self-hosted runner to prove, and it duplicates ~15 lines of workflow that
+would be better factored — a decision worth making on its own.
+
+## File Structure
+
+- `.github/workflows/build.yml` — the `hil-hfp-iar` job only.
+
+---
+
+### Task 1: Give the IAR leg a persistent report dir and a re-run spec
+
+**Files:**
+- Modify: `.github/workflows/build.yml` (job `hil-hfp-iar`)
+
+**Interfaces:**
+- Consumes: `hil_test.py`'s existing `--report-dir` / `.failed` behaviour — no code change.
+- Produces: `env.HIL_REPORT_DIR` for the job, and `$RERUN_ARGS` for the test step.
+
+- [ ] **Step 1: Copy the two steps from `hil-tinyusb`, before the Build step**
+
+```yaml
+ - name: Set HIL report dir (per run+job; persists across run attempts)
+ run: |
+ BASE=$HOME/hil-reports
+ echo "HIL_REPORT_DIR=$BASE/${GITHUB_RUN_ID}-hfp-iar" >> "$GITHUB_ENV"
+
+ - name: Get re-run spec from previous attempt
+ run: |
+ SPEC="$HIL_REPORT_DIR/hfp.json.failed"
+ if [ -f "$SPEC" ]; then
+ echo "RERUN_ARGS=$(cat "$SPEC")" >> "$GITHUB_ENV"
+ echo "re-running only: $(cat "$SPEC")"
+ fi
+```
+
+Match the exact spec filename `hil_test.py` writes for this leg's config — read
+`_write_failed_spec` and the `failed_fname` construction rather than assuming.
+
+- [ ] **Step 2: Pass the spec to the test step**
+
+```yaml
+ python3 test/hil/hil_test.py --retry 1 $SEL_ARGS hfp.json $RERUN_ARGS
+```
+
+`--retry 1` stays FIRST so argparse's last-wins keeps any explicit override working.
+
+- [ ] **Step 3: Point the artifact upload at the report dir**
+
+```yaml
+ path: ${{ env.HIL_REPORT_DIR }}/hil_report.md
+```
+
+- [ ] **Step 4: Validate the YAML**
+
+Run: `python3 -c "import yaml,sys; d=yaml.safe_load(open('.github/workflows/build.yml')); j=d['jobs']['hil-hfp-iar']; print(j['timeout-minutes'], [s.get('name') for s in j['steps']])"`
+Expected: the ceiling is still 120, the Build step still carries `timeout-minutes: 30`, and
+the two new steps appear before Build.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add .github/workflows/build.yml
+git commit -m "ci: let the IAR HIL leg re-run only its failed boards"
+```
+
+---
+
+### Task 2: Prove it on a real re-run
+
+**Files:** none — evidence only.
+
+- [ ] **Step 1:** Push and let `hil-hfp-iar` run to a failure (or force one).
+- [ ] **Step 2:** Confirm `$HIL_REPORT_DIR/hfp.json.failed` exists on the runner after the
+ job.
+- [ ] **Step 3:** Use GitHub's "Re-run failed jobs" and confirm the log line
+ `re-running only: ...` and that only those boards are tested.
+- [ ] **Step 4:** Record the run URL in the PR body.
+
+---
+
+## Consider first
+
+Three jobs would then carry the same ~15 lines. Factoring them into a composite action, or
+computing the report dir inside `hil_test.py` from `GITHUB_RUN_ID`, may be the better
+change — decide that before copying the block a third time.
diff --git a/docs/superpowers/followup/pr3803-pci-rebind-stranding.md b/docs/superpowers/followup/pr3803-pci-rebind-stranding.md
new file mode 100644
index 000000000..de1f7163b
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-pci-rebind-stranding.md
@@ -0,0 +1,157 @@
+# `pci-rebind` Stranding Investigation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Settle when a PCI unbind/rebind of an xHCI controller strands it driverless, so
+the `usb-kernel-recover` skill can state a rule instead of a hypothesis.
+
+**Architecture:** No product code. This is a controlled reproduction against the rig's
+kernel, ending in a documentation change and — if the boundary turns out to be
+detectable — a guard in `usb_recover.sh`.
+
+**Tech Stack:** Linux 6.12.96 (ci.lan), Renesas uPD720201 xHCI, `usb_recover.sh`.
+
+## Global Constraints
+
+- ci.lan is a live CI rig. Take every affected board's lock first
+ (`hil_lock.py hold --all --reason ...`) and confirm no `hil_test.py` is running, with an
+ `if`, not an `&&` chain.
+- A stranded controller takes every fixture on it offline; recovery is
+ `usb_recover.sh pci-bind <addr>` or, failing that, a PVE **host** power cycle — an
+ operator action. Do not start this without being able to reach the host.
+- The rig has two Renesas controllers plus an AMD one; pick the controller with the fewest
+ fixtures for the experiment.
+
+## What is already established
+
+**The skill claimed, unconditionally, that `pci-rebind`'s re-bind hangs on the D-state URB
+and leaves the controller with no driver.** That claim was generalised from ONE observation
+and was used to delete `pci-rebind` and `pci-bind` from `usb_recover.sh` entirely.
+
+**It was refuted in the field on 2026-08-17.** After `hub-cycle 17-2.7` failed to clear a
+wedge, `pci-rebind 0000:05:00.0` recovered the controller in about one second:
+
+```
+02:34:41 remove, state 4 / USB bus 18 deregistered
+02:34:41 remove, state 1 / USB bus 17 deregistered
+02:34:42 xHCI Host Controller / new USB bus registered, assigned bus number 1
+02:34:42 new USB bus registered, assigned bus number 2
+```
+
+Both actions were restored, with the guidance scoped to failure mode: **dead controller →
+use it; device-lock convoy → do not**. Buses renumbered 17/18 → 1/2, which is why rig-wide
+operations need every board's lock.
+
+**What is NOT known:** why the earlier attempt stranded and this one did not. The leading
+hypothesis is that it turns on whether a live D-state URB exists **on that controller** at
+the moment of the re-bind — but in the 02:34 incident the wedged board (17-2.7) was on that
+very controller, which weakens it. An alternative is that `hub-cycle` had already cleared
+the holder, leaving only a dead controller.
+
+**Why this is a separate PR:** it is an experiment that risks taking the rig offline, and
+its output is a documentation change plus possibly a guard — a different scope from any
+code change.
+
+## File Structure
+
+- `.claude/skills/usb-kernel-recover/SKILL.md` — replace the hypothesis in section 3b and
+ the Common-mistakes entry with whatever the experiment establishes.
+- `.claude/skills/usb-kernel-recover/scripts/usb_recover.sh` — only if the boundary is
+ detectable from userspace.
+
+---
+
+### Task 1: Reproduce a controller-scoped D-state wedge
+
+**Files:** none.
+
+- [ ] **Step 1: Establish the safety net**
+
+```bash
+ssh [email protected] 'if pgrep -f "[h]il_test.py" >/dev/null; then echo BUSY; exit 1; fi'
+# hold ALL boards on the target controller
+```
+
+Confirm host access to pve.lan before continuing.
+
+- [ ] **Step 2: Create a wedge deliberately**
+
+Run `usbtest.py` against a board known to hang (`mimxrt1064_evk` has wedged eight times,
+TEST 9/10/24/27), or drive `testusb` directly until a case does not return.
+
+- [ ] **Step 3: Confirm the holder and its controller**
+
+```bash
+ps -eo pid,stat,etimes,wchan:22,args | awk '$2 ~ /D/'
+sudo cat /proc/<pid>/stack # usbdev_ioctl + [usbtest] = the owner
+readlink -f /sys/bus/usb/devices/usb<N> # bus -> PCI addr
+```
+
+Record whether the holder is on the SAME controller you will rebind.
+
+---
+
+### Task 2: Rebind and record the outcome
+
+**Files:** none.
+
+- [ ] **Step 1: Rebind, with a bounded observer**
+
+```bash
+timeout 120 sudo usb_recover.sh pci-rebind <addr>; echo "rc=$?"
+```
+
+- [ ] **Step 2: Record which of the three outcomes occurred**
+
+1. Re-bind completes, controller recovers (as on 2026-08-17).
+2. Re-bind hangs; `/sys/bus/pci/devices/<addr>/driver` is gone → **stranded**.
+3. Re-bind completes but the wedge persists.
+
+Capture `sudo journalctl -k --since ...` around the attempt either way.
+
+- [ ] **Step 3: If stranded, recover**
+
+```bash
+sudo usb_recover.sh pci-bind <addr>
+```
+
+If that hangs too, the only remaining step is a PVE host power cycle — an operator action.
+
+- [ ] **Step 4: Repeat at least three times**
+
+One observation is what produced the wrong rule in the first place. Vary whether a D-state
+holder is live on that controller at rebind time; that is the hypothesis under test.
+
+---
+
+### Task 3: Write down what was learned
+
+**Files:**
+- Modify: `.claude/skills/usb-kernel-recover/SKILL.md`
+
+- [ ] **Step 1: Replace section 3b's scoping with the measured rule**
+
+State the condition under which stranding occurs, with the journal lines. If the experiment
+does NOT reproduce stranding, say that too, with the attempt count — "not reproduced in N
+attempts" is a better record than an unexplained warning.
+
+- [ ] **Step 2: If the boundary is detectable, guard the script**
+
+For example, refuse `pci-rebind` when a D-state holder exists on that controller, since the
+holder is enumerable from `/proc` and the controller from `readlink`. Only add this if the
+experiment shows it predicts the outcome.
+
+- [ ] **Step 3: Commit**
+
+```bash
+git add .claude/skills/usb-kernel-recover/
+git commit -m "skills: replace the pci-rebind stranding hypothesis with measurement"
+```
+
+---
+
+## Abort criteria
+
+Stop and hand back to the operator if: a rebind strands the controller and `pci-bind` does
+not recover it; `uhubctl` starts hanging (the convoy has spread to the hub locks); or a CI
+run starts while the rig is in a broken state.
diff --git a/docs/superpowers/followup/pr3803-usbtest-recovery-reserve.md b/docs/superpowers/followup/pr3803-usbtest-recovery-reserve.md
new file mode 100644
index 000000000..eb8959520
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-usbtest-recovery-reserve.md
@@ -0,0 +1,175 @@
+# usbtest Recovery Reserve Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Make the post-hang recovery reserve a derived, asserted property instead of an
+accident of four independently-set constants.
+
+**Architecture:** `hil_test` passes `--budget` and `--outer-timeout` to `usbtest.py`, which
+decides at runtime whether a recovery still fits. Today the reserve survives only because
+the four numbers happen to line up; nothing ties them together or fails when they stop.
+
+**Tech Stack:** Python 3.13 stdlib.
+
+## Global Constraints
+
+- `usbtest.py`: `RECOVER_FLASH_TIMEOUT = 90`, `RECOVER_RESET_TIMEOUT = 30`.
+- `hil_test.py`: `USBTEST_BATTERY_BUDGET = 260`, `USBTEST_RECOVERY_BUDGET = 250`,
+ `USBTEST_OVERSHOOT = 120`; `outer = BATTERY_BUDGET + (RECOVERY_BUDGET if recovery else
+ OVERSHOOT)`, used for both the child's `--outer-timeout` and the parent's `run_cmd` bound.
+- All five are env-overridable via `hil_util.pos_int_env`, so a rig can change them.
+- Tests: `cd test/hil && python3 test/test_hil_health.py` and `test_hil_bounded.py`.
+
+## What is already established
+
+The reserve holds at the shipped values, checked by hand:
+
+- The battery checks its budget BEFORE dispatching a case, so it can overshoot by one
+ case — worst case `260 + 60 + 5 = 325 s`.
+- Recovery is gated on `_time_left() >= RECOVER_RESET_TIMEOUT`, where
+ `_time_left() = outer_timeout - elapsed - 35`; with `outer = 510` that allows recovery
+ until `elapsed = 445 s`, and the reflash until `385 s`.
+- So ~60 s of margin survives, and recovery does fire.
+
+**The defect is structural, not arithmetic:** lower `--outer-timeout`, raise `--timeout`, or
+raise `USBTEST_BATTERY_BUDGET` via the env and the reserve silently disappears. The failure
+mode is a skipped reflash that leaves the D-state holder for the next job — the exact thing
+the containment exists to prevent — with no error anywhere.
+
+**Why this is a separate PR:** it changes the timing contract between `hil_test` and
+`usbtest.py`, which affects every board's run duration, so it wants its own review and a
+full rig run.
+
+## File Structure
+
+- `test/hil/usbtest.py` — a `reserve_ok()` predicate plus a startup assertion.
+- `test/hil/hil_test.py` — derive the battery budget from the outer bound rather than
+ setting both independently.
+- `test/hil/test/test_hil_health.py` — tests.
+
+---
+
+### Task 1: Assert the reserve at startup
+
+**Files:**
+- Modify: `test/hil/usbtest.py` (constants block, and `main()` after argparse)
+- Test: `test/hil/test/test_hil_health.py`
+
+**Interfaces:**
+- Produces: `usbtest.reserve_ok(budget, outer, case_timeout)` returning bool.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+class RecoveryReserveIsChecked(unittest.TestCase):
+ """The battery may overshoot its budget by ONE already-started case, so the outer bound
+ must leave room for that overshoot AND a bounded recovery afterwards."""
+
+ def setUp(self):
+ import usbtest
+ self.u = usbtest
+
+ def test_the_shipped_numbers_leave_room(self):
+ self.assertTrue(self.u.reserve_ok(budget=260, outer=510, case_timeout=60))
+
+ def test_a_tighter_outer_bound_is_rejected(self):
+ self.assertFalse(self.u.reserve_ok(budget=260, outer=380, case_timeout=60))
+
+ def test_a_longer_case_timeout_is_rejected(self):
+ self.assertFalse(self.u.reserve_ok(budget=260, outer=510, case_timeout=200))
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_hil_health.py RecoveryReserveIsChecked -v`
+Expected: FAIL — `module 'usbtest' has no attribute 'reserve_ok'`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def reserve_ok(budget: int, outer: int, case_timeout: int) -> bool:
+ """Does `outer` leave room for the battery's worst case AND a bounded recovery?
+
+ The budget is checked BEFORE dispatch, so the battery can run to
+ `budget + case_timeout + 5` (the +5 is run_case's reap). _time_left() subtracts a
+ further 35 s of fixed tail. A reflash needs RECOVER_FLASH_TIMEOUT beyond that.
+ """
+ worst_case_end = budget + case_timeout + 5
+ return outer - worst_case_end - 35 >= RECOVER_FLASH_TIMEOUT
+```
+
+In `main()`, after parsing args:
+
+```python
+ if args.budget and args.outer_timeout and not reserve_ok(
+ args.budget, args.outer_timeout, args.timeout):
+ print(f'warning: --outer-timeout {args.outer_timeout} leaves no room for a bounded '
+ f'recovery after a --budget {args.budget} battery with --timeout '
+ f'{args.timeout} cases; a HUNG board will be left wedged', file=sys.stderr)
+```
+
+Warn, do not exit: a caller that deliberately runs without recovery is legitimate.
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_hil_health.py RecoveryReserveIsChecked -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/usbtest.py test/hil/test/test_hil_health.py
+git commit -m "usbtest: check the recovery reserve instead of assuming it"
+```
+
+---
+
+### Task 2: Derive the outer bound from one place
+
+**Files:**
+- Modify: `test/hil/hil_test.py` (constants block ~line 227, and `test_device_usbtest`)
+- Test: `test/hil/test/test_hil_bounded.py`
+
+**Interfaces:**
+- Consumes: `usbtest.reserve_ok` semantics (duplicate the arithmetic, do not import
+ usbtest — `hil_test` must not import it).
+- Produces: an assertion at module import that the shipped constants satisfy the reserve.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_the_shipped_constants_satisfy_the_reserve(self):
+ """Whatever the env overrides, the pair hil_test computes must leave recovery room:
+ outer - (budget + case_timeout + 5) - 35 >= 90."""
+ outer = hil_test.USBTEST_BATTERY_BUDGET + hil_test.USBTEST_RECOVERY_BUDGET
+ self.assertGreaterEqual(outer - (hil_test.USBTEST_BATTERY_BUDGET + 60 + 5) - 35, 90)
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Temporarily set `HIL_USBTEST_RECOVERY_BUDGET=100` and run; expect FAIL. Unset.
+
+- [ ] **Step 3: Add the guard**
+
+```python
+# The recovery reserve is a PROPERTY of these two, not a coincidence: the battery may
+# overshoot its budget by one already-started case (checked before dispatch), and a bounded
+# reflash needs 90 s after a 35 s fixed tail. Env overrides make this checkable at import
+# rather than discoverable when a wedge is left unrecovered.
+if USBTEST_RECOVERY_BUDGET - 60 - 5 - 35 < 90:
+ print(f'warning: HIL_USBTEST_RECOVERY_BUDGET={USBTEST_RECOVERY_BUDGET} leaves no room '
+ f'for a bounded reflash after a one-case overshoot; HUNG boards will stay wedged',
+ file=sys.stderr)
+```
+
+- [ ] **Step 4: Run tests to verify they pass**
+
+Run: `cd test/hil && python3 test/test_hil_bounded.py -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/hil_test.py test/hil/test/test_hil_bounded.py
+git commit -m "hil: warn when the timeout constants leave no recovery reserve"
+```
diff --git a/docs/superpowers/plans/2026-08-15-ci-hs-reset-edges.md b/docs/superpowers/plans/2026-08-15-ci-hs-reset-edges.md
new file mode 100644
index 000000000..ec0834e55
--- /dev/null
+++ b/docs/superpowers/plans/2026-08-15-ci-hs-reset-edges.md
@@ -0,0 +1,782 @@
+# Bus-Reset Edge Events + Review Fix Wave Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Give the device stack a "bus reset started" event so ci_hs can tell usbd to stand down at the URI interrupt instead of up to 50 ms later, and clear the ten findings agreed from the max review.
+
+**Architecture:** `DCD_EVENT_BUS_RESET` splits into `DCD_EVENT_BUS_RESET_START` / `_END` with a compatibility alias, so every other port stays byte-identical. `dcd_ci_hs.c`'s `bus_reset()` splits along the register/software line — registers at URI (`_START`), software structures at the port-change ending the reset (`_END`) — which eliminates the window where usbd believes it is configured over zeroed queue heads. A single bounded-flush helper absorbs the five flush sites. Seven mechanical fixes follow.
+
+**Tech Stack:** C99, TinyUSB device stack (`src/device/`), ChipIdea HS DCD (`src/portable/chipidea/ci_hs/`), NXP IP3511 DCD (`src/portable/nxp/lpc_ip3511/`), CMake+Ninja and Make builds, J-Link flashing, `test/hil/` HIL harness.
+
+## Global Constraints
+
+- Branch `fix-ci-hs` in worktree `/home/hathach/.herdr/worktrees/tinyusb/fix-ci-hs`. Do NOT push; the user pushes.
+- C99, 2-space indent, no tabs. Match each file's surrounding style (`dcd_lpc_ip3511.c` mixes styles — follow the immediate neighbourhood).
+- Commit messages: imperative mood, no `Co-Authored-By:` or `Claude-Session:` trailers (repo rule: hathach is sole author).
+- The repo pre-commit hook (trailing-whitespace, end-of-file-fixer, codespell, unique-PIDs, ceedling unit tests) must pass. If it rewrites a file, re-stage and retry the commit once.
+- Comments: short, only the non-obvious "why". Cite manuals as `UM10503 25.10.3` / `Errata LPC546xx USB.13` style — never `ES_` prefixes.
+- Never edit anything under `hw/mcu/` or `lib/` (vendor code).
+- Build commands used throughout (each ~30-60 s):
+ `cmake --build examples/cmake-build-<board>` for `mimxrt1064_evk`, `lpcxpresso18s37`, `lpcxpresso11u37`, `lpcxpresso55s28`.
+- Design source of truth: `docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md`.
+
+## File Structure
+
+| File | Responsibility in this plan |
+|---|---|
+| `src/device/dcd.h` | Event enum + compatibility alias + contract comment |
+| `src/device/usbd.c` | Handle both reset edges; log strings; stop breakpointing on DCD refusal |
+| `src/portable/chipidea/ci_hs/dcd_ci_hs.c` | Flush helper; `bus_reset()` split; setup-flush wait; `dcd_set_address`; RESUME guard |
+| `src/portable/nxp/lpc_ip3511/dcd_lpc_ip3511.c` | Torn-setup delivery; USB.13 TODO token |
+| `hw/bsp/lpc55/boards/lpcxpresso55s28/board.cmake` | Delete dead RHPORT block |
+| `hw/bsp/lpc11/boards/lpcxpresso11u37/lpc11u37.ld` | Correct stale comment; relabel ASSERT |
+
+Tasks 1-3 are ordered (each builds on the previous); Tasks 4-6 are independent of each other.
+
+---
+
+### Task 1: Split the bus-reset event into START/END edges
+
+**Files:**
+- Modify: `src/device/dcd.h` (enum at lines 23-34; contract comment above it)
+- Modify: `src/device/usbd.c` (`_usbd_event_str[]` at line 457; the `DCD_EVENT_BUS_RESET` case at line 700)
+
+**Interfaces:**
+- Produces: `DCD_EVENT_BUS_RESET_START` and `DCD_EVENT_BUS_RESET_END` enum members; `#define DCD_EVENT_BUS_RESET DCD_EVENT_BUS_RESET_END`. Task 2 emits `_START` via the existing `dcd_event_bus_signal(uint8_t rhport, dcd_eventid_t eid, bool in_isr)` and `_END` via the existing `dcd_event_bus_reset(uint8_t rhport, tusb_speed_t speed, bool in_isr)`.
+
+- [ ] **Step 1: Replace the enum member in `src/device/dcd.h`**
+
+Replace:
+
+```c
+typedef enum {
+ DCD_EVENT_INVALID = 0, // 0
+ DCD_EVENT_BUS_RESET, // 1
+ DCD_EVENT_UNPLUGGED, // 2
+ DCD_EVENT_SOF, // 3
+ DCD_EVENT_SUSPEND, // 4 TODO LPM Sleep L1 support
+ DCD_EVENT_RESUME, // 5
+ DCD_EVENT_SETUP_RECEIVED, // 6
+ DCD_EVENT_XFER_COMPLETE, // 7
+ USBD_EVENT_FUNC_CALL, // 8 Not an DCD event, just a convenient way to defer ISR function
+ DCD_EVENT_COUNT
+} dcd_eventid_t;
+```
+
+with:
+
+```c
+// Bus reset is reported as two edges. BUS_RESET_START is optional: a controller that
+// cannot tell the edges apart emits only BUS_RESET_END, which stays self-sufficient (it
+// performs the full teardown with or without a preceding START). Emit START when reset
+// signaling is detected - the link is unusable and the speed is not negotiated yet - so
+// the stack stops using endpoints immediately instead of at the end of the reset.
+typedef enum {
+ DCD_EVENT_INVALID = 0, // 0
+ DCD_EVENT_BUS_RESET_START, // 1
+ DCD_EVENT_BUS_RESET_END, // 2 with negotiated speed
+ DCD_EVENT_UNPLUGGED, // 3
+ DCD_EVENT_SOF, // 4
+ DCD_EVENT_SUSPEND, // 5 TODO LPM Sleep L1 support
+ DCD_EVENT_RESUME, // 6
+ DCD_EVENT_SETUP_RECEIVED, // 7
+ DCD_EVENT_XFER_COMPLETE, // 8
+ USBD_EVENT_FUNC_CALL, // 9 Not an DCD event, just a convenient way to defer ISR function
+ DCD_EVENT_COUNT
+} dcd_eventid_t;
+
+#define DCD_EVENT_BUS_RESET DCD_EVENT_BUS_RESET_END // backward compatibility
+```
+
+- [ ] **Step 2: Update the log-string table in `src/device/usbd.c`**
+
+At line 457 the table is indexed by event id and MUST stay in enum order. Replace the
+`"Bus Reset",` entry (line 459) with two entries:
+
+```c
+ "Bus Reset Start",
+ "Bus Reset End",
+```
+
+- [ ] **Step 3: Handle both edges in the usbd task loop**
+
+Replace the case at `src/device/usbd.c:700`:
+
+```c
+ case DCD_EVENT_BUS_RESET:
+ TU_LOG_USBD(": %s Speed\r\n", tu_str_speed[event.bus_reset.speed]);
+ usbd_reset(event.rhport);
+ _usbd_dev.speed = event.bus_reset.speed;
+ break;
+```
+
+with:
+
+```c
+ case DCD_EVENT_BUS_RESET_START:
+ TU_LOG_USBD("\r\n");
+ usbd_reset(event.rhport);
+ break;
+
+ case DCD_EVENT_BUS_RESET_END:
+ TU_LOG_USBD(": %s Speed\r\n", tu_str_speed[event.bus_reset.speed]);
+ // TODO a DCD that reports both edges pays for two teardowns: track a per-rhport
+ // "start seen" flag and skip this reset, keeping it for the single-event DCDs.
+ usbd_reset(event.rhport);
+ _usbd_dev.speed = event.bus_reset.speed;
+ break;
+```
+
+- [ ] **Step 4: Verify legacy ports still build (the alias must carry them)**
+
+Run:
+
+```bash
+cd examples && cmake -B cmake-build-stm32f407disco -DBOARD=stm32f407disco -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel . && cmake --build cmake-build-stm32f407disco
+```
+
+Expected: builds clean. This board's DCD (dwc2) still calls `dcd_event_bus_reset()`, which
+now resolves to `_END` through the unchanged helper — proving the alias works.
+
+- [ ] **Step 5: Verify the unit tests still build and pass**
+
+Run: `cd test/unit-test && ceedling test:all`
+Expected: all tests pass (they reference `DCD_EVENT_BUS_RESET` via the alias).
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add src/device/dcd.h src/device/usbd.c
+git commit -m "usbd: split bus reset into start/end edge events
+
+A DCD that can see reset signaling begin has no way to say so: the only
+event carries the negotiated speed, which does not exist until the reset
+ends. On ChipIdea that leaves the stack believing it is configured for the
+whole reset window (3 ms minimum, tens of ms in practice) while the
+controller has already torn its endpoints down.
+
+Add DCD_EVENT_BUS_RESET_START for the leading edge and rename the existing
+event to DCD_EVENT_BUS_RESET_END, keeping DCD_EVENT_BUS_RESET as an alias
+so every other port and the unit tests are untouched. START is optional and
+END stays self-sufficient, so single-event drivers keep working unchanged."
+```
+
+---
+
+### Task 2: Split ci_hs `bus_reset()` across the two edges, behind one flush helper
+
+**Files:**
+- Modify: `src/portable/chipidea/ci_hs/dcd_ci_hs.c` (`bus_reset()`; `dcd_deinit()`; `dcd_edpt_iso_activate()`; the `INTR_RESET` and `INTR_PORT_CHANGE` branches of `dcd_int_handler()`)
+
+**Interfaces:**
+- Consumes: `DCD_EVENT_BUS_RESET_START` (Task 1), `dcd_event_bus_signal()`, `dcd_event_bus_reset()`.
+- Produces: `static bool flush_endpoints(ci_hs_regs_t *dcd_reg, uint32_t mask)` — writes `ENDPTFLUSH = mask`, spins bounded by `CI_HS_BUSY_SPIN`, returns `true` if the bits cleared. Used by Task 3.
+
+- [ ] **Step 1: Add the flush helper next to `bus_reset()`**
+
+Insert above `bus_reset()`:
+
+```c
+// Flush endpoint buffers and wait for the controller to acknowledge. Callers proceed
+// regardless of the result; the bound only prevents an ISR-context hang on dead hardware.
+static bool flush_endpoints(ci_hs_regs_t *dcd_reg, uint32_t mask) {
+ dcd_reg->ENDPTFLUSH = mask;
+ uint32_t guard = CI_HS_BUSY_SPIN;
+ while (dcd_reg->ENDPTFLUSH & mask) {
+ if (!guard--) {
+ return false;
+ }
+ }
+ return true;
+}
+```
+
+- [ ] **Step 2: Split `bus_reset()` into begin/complete**
+
+Replace the whole `bus_reset()` function with these two. `bus_reset_begin()` keeps only
+register work; `bus_reset_complete()` owns everything that touches `_dcd_data`:
+
+```c
+/// Register-side reset handling, must run inside the reset window (UM10503 25.10.3)
+static void bus_reset_begin(uint8_t rhport) {
+ ci_hs_regs_t *dcd_reg = CI_HS_REG(rhport);
+
+ // The reset value for all endpoint types is the control endpoint. If one endpoint
+ // direction is enabled and the paired endpoint of opposite direction is disabled, then the
+ // endpoint type of the unused direction must be changed from the control type to any other
+ // type (e.g. bulk). Leaving an un-configured endpoint control will cause undefined behavior
+ // for the data PID tracking on the active endpoint.
+ const uint8_t ep_count = ci_ep_count(dcd_reg);
+ for (uint8_t i = 1; i < ep_count; i++) {
+ dcd_reg->ENDPTCTRL[i] = ENDPTCTRL_RESET_MASK;
+ }
+
+ //------------- Clear All Registers -------------//
+ dcd_reg->ENDPTNAK = dcd_reg->ENDPTNAK;
+ dcd_reg->ENDPTNAKEN = 0;
+ dcd_reg->ENDPTSETUPSTAT = dcd_reg->ENDPTSETUPSTAT;
+ dcd_reg->ENDPTCOMPLETE = dcd_reg->ENDPTCOMPLETE;
+
+ uint32_t guard = CI_HS_BUSY_SPIN;
+ while (dcd_reg->ENDPTPRIME && guard--) {}
+ flush_endpoints(dcd_reg, 0xFFFFFFFF);
+}
+
+/// Software-side reset handling, deferred to the port change ending the reset so the queue
+/// heads stay coherent until the stack is told - and so a prime issued by a task that had
+/// not yet seen BUS_RESET_START is flushed here rather than surviving re-enumeration.
+static void bus_reset_complete(uint8_t rhport) {
+ ci_hs_regs_t *dcd_reg = CI_HS_REG(rhport);
+ flush_endpoints(dcd_reg, 0xFFFFFFFF);
+
+ //------------- Queue Head & Queue TD -------------//
+ tu_memclr(&_dcd_data, sizeof(dcd_data_t));
+
+ //------------- Set up Control Endpoints (0 OUT, 1 IN) -------------//
+ _dcd_data.qhd[0][0].zero_length_termination = _dcd_data.qhd[0][1].zero_length_termination = 1;
+ _dcd_data.qhd[0][0].max_packet_size = _dcd_data.qhd[0][1].max_packet_size = CFG_TUD_ENDPOINT0_SIZE;
+ _dcd_data.qhd[0][0].qtd_overlay.next = _dcd_data.qhd[0][1].qtd_overlay.next = QTD_NEXT_INVALID;
+
+ _dcd_data.qhd[0][0].int_on_setup = 1; // OUT only
+
+ dcd_dcache_clean_invalidate(&_dcd_data, sizeof(dcd_data_t));
+}
+```
+
+- [ ] **Step 3: Route the two ISR branches to the new functions**
+
+In `dcd_int_handler()`, the `INTR_RESET` branch becomes:
+
+```c
+ if (int_status & INTR_RESET) {
+ bus_reset_begin(rhport);
+ _port_change_reason[rhport] = PORT_CHANGE_REASON_RESET;
+ dcd_event_bus_signal(rhport, DCD_EVENT_BUS_RESET_START, true);
+ }
+```
+
+and inside the `INTR_PORT_CHANGE` branch, the reset arm (the `else` of the resume test)
+becomes:
+
+```c
+ } else {
+ bus_reset_complete(rhport);
+ // PSPD: 0 full, 1 low, 2 high, 3 undefined (treated as full)
+ const uint32_t pspd = (dcd_reg->PORTSC1 & PORTSC1_PORT_SPEED) >> PORTSC1_PORT_SPEED_POS;
+ const tusb_speed_t speed = (pspd == 1) ? TUSB_SPEED_LOW : (pspd == 2) ? TUSB_SPEED_HIGH : TUSB_SPEED_FULL;
+ dcd_event_bus_reset(rhport, speed, true);
+ }
+```
+
+Delete the now-unused EP0 `ENDPTFLUSH` line that previously sat at the top of that arm —
+`bus_reset_complete()` flushes all endpoints.
+
+- [ ] **Step 4: Route the remaining flush sites through the helper**
+
+In `dcd_deinit()`, replace the flush block with:
+
+```c
+ // flush all endpoints
+ uint32_t guard = CI_HS_BUSY_SPIN;
+ while (dcd_reg->ENDPTPRIME && guard--) {}
+ flush_endpoints(dcd_reg, 0xFFFFFFFF);
+```
+
+In `dcd_edpt_iso_activate()`, replace the flush + spin with:
+
+```c
+ // Flush EP
+ flush_endpoints(dcd_reg, TU_BIT(epnum + (dir ? 16 : 0)));
+```
+
+- [ ] **Step 5: Build both ci_hs board families**
+
+Run:
+
+```bash
+cmake --build examples/cmake-build-mimxrt1064_evk && cmake --build examples/cmake-build-lpcxpresso18s37
+```
+
+Expected: both succeed with no new warnings.
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add src/portable/chipidea/ci_hs/dcd_ci_hs.c
+git commit -m "dcd(ci_hs): report bus reset start at URI, finish at port change
+
+The RM wants the reset cleanup inside the reset window, but the negotiated
+speed only exists once the port reaches its operational state, so the stack
+was told nothing for the whole window - it kept believing it was configured
+while the queue heads had been zeroed under it, and a transfer a class
+driver started in that gap stayed primed across re-enumeration.
+
+Split the work along the register/software line: bus_reset_begin() does the
+register cleanup at URI and signals BUS_RESET_START, bus_reset_complete()
+re-flushes, resets the queue heads and reports BUS_RESET_END with the final
+speed at the port change. Zeroing the queue heads now happens in the same
+breath as telling the stack, and the second flush retires anything primed
+in between.
+
+Fold the five hand-rolled endpoint flushes into one bounded helper while
+the reset path is open."
+```
+
+---
+
+### Task 3: Make the setup-time EP0 flush wait, and stop dropping the SET_ADDRESS status prime
+
+**Files:**
+- Modify: `src/portable/chipidea/ci_hs/dcd_ci_hs.c` (`dcd_set_address()`; the `ENDPTSETUPSTAT` branch inside `dcd_int_handler()`)
+
+**Interfaces:**
+- Consumes: `flush_endpoints()` (Task 2); `qhd_start_xfer()` returning `bool`, already propagated by `dcd_edpt_xfer()`.
+
+- [ ] **Step 1: Wait for the setup-time flush to complete**
+
+In the ISR's setup branch, replace the fire-and-forget flush line
+
+```c
+ dcd_reg->ENDPTFLUSH = TU_BIT(0) | TU_BIT(16);
+```
+
+with
+
+```c
+ // Wait it out: the flush retires a status/handshake phase left primed by the previous
+ // control sequence (UM10503 25.10.8.1.1), and an unfinished flush would otherwise
+ // still be asserted when the task primes the response to this setup and would retire
+ // that instead. A flush waits for any packet already in progress - microseconds at
+ // high speed - and the guard caps wedged hardware.
+ flush_endpoints(dcd_reg, TU_BIT(0) | TU_BIT(16));
+```
+
+- [ ] **Step 2: Honour the status-prime result in `dcd_set_address`**
+
+Replace the body of `dcd_set_address()`:
+
+```c
+void dcd_set_address(uint8_t rhport, uint8_t dev_addr) {
+ // Response with status first before changing device address. A refused prime means a new
+ // setup superseded this transfer; staging an address whose ACK will never arrive would
+ // leave the device answering on it, so only arm the address when the status went out.
+ if (dcd_edpt_xfer(rhport, tu_edpt_addr(0, TUSB_DIR_IN), NULL, 0, false)) {
+ ci_hs_regs_t *dcd_reg = CI_HS_REG(rhport);
+ dcd_reg->DEVICEADDR = (dev_addr << 25) | TU_BIT(24);
+ }
+}
+```
+
+- [ ] **Step 3: Build and commit**
+
+Run: `cmake --build examples/cmake-build-mimxrt1064_evk && cmake --build examples/cmake-build-lpcxpresso18s37`
+Expected: both succeed.
+
+```bash
+git add src/portable/chipidea/ci_hs/dcd_ci_hs.c
+git commit -m "dcd(ci_hs): wait out the setup flush, honour the set-address prime
+
+The flush issued on every new setup was fire-and-forget. A flush waits for
+a packet already in progress, so it could still be asserted when the task
+primed the response to that setup and retire the fresh prime instead -
+leaving EP0 silent until the host gave up.
+
+dcd_set_address() also armed DEVICEADDR unconditionally, but the status
+prime can now be refused when a newer setup supersedes the transfer; the
+address was then staged behind an ACK that never came and the device sat at
+address 0. Only arm it when the status transfer actually started."
+```
+
+---
+
+### Task 4: Emit RESUME only when the port really left suspend
+
+**Files:**
+- Modify: `src/portable/chipidea/ci_hs/dcd_ci_hs.c` (the resume arm of the `INTR_PORT_CHANGE` branch in `dcd_int_handler()`)
+
+**Interfaces:** none consumed or produced.
+
+- [ ] **Step 1: Restore the hardware guard**
+
+In the `INTR_PORT_CHANGE` branch, the resume arm currently reads:
+
+```c
+ if (pci_reason == PORT_CHANGE_REASON_SUSPEND) {
+ dcd_event_bus_signal(rhport, DCD_EVENT_RESUME, true);
+ } else {
+```
+
+Replace that condition with one that also consults live hardware:
+
+```c
+ if (pci_reason == PORT_CHANGE_REASON_SUSPEND) {
+ // Only when the port actually left suspend: a starved snapshot can hold the resume's
+ // port change together with a second suspend, and reporting a resume there would
+ // leave the stack awake on a sleeping bus with no further event to correct it.
+ if (!(dcd_reg->PORTSC1 & PORTSC1_SUSPEND)) {
+ dcd_event_bus_signal(rhport, DCD_EVENT_RESUME, true);
+ }
+ } else {
+```
+
+- [ ] **Step 2: Build and commit**
+
+Run: `cmake --build examples/cmake-build-mimxrt1064_evk && cmake --build examples/cmake-build-lpcxpresso18s37`
+Expected: both succeed.
+
+```bash
+git add src/portable/chipidea/ci_hs/dcd_ci_hs.c
+git commit -m "dcd(ci_hs): only report resume when the port left suspend
+
+A suspend, resume and second suspend collapsed into one interrupt pass
+queued suspend then resume from the recorded cause alone, so the stack
+ended up awake while the bus was still suspended and nothing arrived to
+correct it. Consult PORTSC1 before reporting the resume."
+```
+
+---
+
+### Task 5: ip3511 — never deliver a knowingly-torn setup packet
+
+**Files:**
+- Modify: `src/portable/nxp/lpc_ip3511/dcd_lpc_ip3511.c` (setup branch of `dcd_int_handler()`; the `dcd_edpt_clear_stall()` comment)
+
+**Interfaces:** none consumed or produced.
+
+- [ ] **Step 1: Deliver only when the copy is known good**
+
+Replace:
+
+```c
+ // a SETUP that raced in after the acks (its bit0 consumed above, this copy possibly torn):
+ // its latch is visible again - re-raise the endpoint interrupt so the next pass redelivers
+ // the newer payload
+ if (dcd_reg->DEVCMDSTAT & DEVCMDSTAT_SETUP_RECEIVED_MASK) {
+ dcd_reg->INTSETSTAT = TU_BIT(0);
+ }
+
+ dcd_event_setup_received(rhport, setup_copy, true);
+```
+
+with:
+
+```c
+ // a SETUP that raced in after the acks (its bit0 consumed above) makes this copy suspect:
+ // its latch is visible again, so re-raise the endpoint interrupt and let the next pass
+ // deliver the newer payload rather than passing up bytes that may be torn between the two
+ if (dcd_reg->DEVCMDSTAT & DEVCMDSTAT_SETUP_RECEIVED_MASK) {
+ dcd_reg->INTSETSTAT = TU_BIT(0);
+ } else {
+ dcd_event_setup_received(rhport, setup_copy, true);
+ }
+```
+
+- [ ] **Step 2: Add the TODO token to the USB.13 deferral**
+
+In `dcd_edpt_clear_stall()`, change the caveat's opening line from
+
+```c
+ // Known caveat (Errata LPC546xx USB.13, same semantics in UM11126): with RF/TV preserved at 1, TR
+```
+
+to
+
+```c
+ // TODO implement the Errata LPC546xx USB.13 work-around (same semantics in UM11126): with RF/TV preserved at 1, TR
+```
+
+- [ ] **Step 3: Build and commit**
+
+Run: `cmake --build examples/cmake-build-lpcxpresso11u37 && cmake --build examples/cmake-build-lpcxpresso55s28`
+Expected: both succeed.
+
+```bash
+git add src/portable/nxp/lpc_ip3511/dcd_lpc_ip3511.c
+git commit -m "dcd(ip3511): drop a setup packet the hardware may have overwritten
+
+The handler already notices when a new setup landed while it was copying
+the previous one, and re-raises the endpoint interrupt so the newer payload
+is delivered next pass - but it then passed the suspect copy up anyway.
+Usually harmless, since the redelivery supersedes it, but if that second
+event cannot be queued the torn bytes are processed as a real request.
+Deliver the copy only when no newer setup is pending."
+```
+
+---
+
+### Task 6: BSP cleanups — dead RHPORT block and the stale linker comment
+
+**Files:**
+- Modify: `hw/bsp/lpc55/boards/lpcxpresso55s28/board.cmake`
+- Modify: `hw/bsp/lpc11/boards/lpcxpresso11u37/lpc11u37.ld`
+
+**Interfaces:** none consumed or produced.
+
+- [ ] **Step 1: Delete the redundant RHPORT block**
+
+`hw/bsp/lpc55/family.cmake` already applies the identical guarded defaults (`RHPORT_DEVICE 1`,
+`RHPORT_HOST 0`) after including the board file, so remove these lines from
+`board.cmake` entirely:
+
+```cmake
+# device highspeed, host fullspeed; guarded so a -D override on the cmake command line wins
+if (NOT DEFINED RHPORT_DEVICE)
+ set(RHPORT_DEVICE 1)
+endif ()
+if (NOT DEFINED RHPORT_HOST)
+ set(RHPORT_HOST 0)
+endif ()
+```
+
+Leave `board.mk`'s `RHPORT_DEVICE ?= 1` / `RHPORT_HOST ?= 0` alone — `?=` is the idiomatic
+Make form and matches sibling boards.
+
+- [ ] **Step 2: Prove the defaults and the override still work**
+
+Run:
+
+```bash
+cd examples && rm -rf /tmp/rh-default /tmp/rh-override
+cmake -B /tmp/rh-default -DBOARD=lpcxpresso55s28 -G Ninja . > /tmp/rh-default.log 2>&1
+grep -m1 "RHPORT_DEVICE" /tmp/rh-default.log || cmake -B /tmp/rh-default -DBOARD=lpcxpresso55s28 -G Ninja -LA . | grep -E "^RHPORT_(DEVICE|HOST)"
+cmake -B /tmp/rh-override -DBOARD=lpcxpresso55s28 -DRHPORT_DEVICE=0 -DRHPORT_HOST=1 -G Ninja -LA . | grep -E "^RHPORT_(DEVICE|HOST)"
+```
+
+Expected: the default configure yields device 1 / host 0; the override configure yields
+device 0 / host 1. Then rebuild the real tree: `cmake --build cmake-build-lpcxpresso55s28`.
+
+- [ ] **Step 3: Correct the linker-script comment and relabel the ASSERT**
+
+In `lpc11u37.ld`, replace the comment block above `__user_stack_top` and the ASSERT with:
+
+```text
+ /* Main (MSP/ISR) stack lives at the top of the USB SRAM bank: the 8K main bank is packed so
+ tight that only ~280 B remained above .bss, and ISR frames overflowed into the topmost task
+ stack (cdc_msc_freertos hard fault). Nothing else is placed in this bank in either build
+ system, so the stack owns all 2 KB; the ASSERT is future-proofing in case USB buffers are
+ ever mapped here again. */
+ __user_stack_top = ORIGIN(RamUsb2) + LENGTH(RamUsb2);
+ ASSERT(__user_stack_top - (ADDR(.noinit_RAM2) + SIZEOF(.noinit_RAM2)) >= 0x200,
+ "main stack headroom in RamUsb2 below 512 bytes")
+```
+
+- [ ] **Step 4: Build both build systems for lpc11u37**
+
+Run:
+
+```bash
+cmake --build examples/cmake-build-lpcxpresso11u37
+cd examples/device/cdc_msc_freertos && make -j8 BOARD=lpcxpresso11u37 all && cd ../../..
+```
+
+Expected: both succeed.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add hw/bsp/lpc55/boards/lpcxpresso55s28/board.cmake hw/bsp/lpc11/boards/lpcxpresso11u37/lpc11u37.ld
+git commit -m "bsp: drop duplicated lpc55s28 rhport defaults, fix lpc11u37 comment
+
+hw/bsp/lpc55/family.cmake already applies the same guarded rhport defaults
+after including the board file, so the board-level copy only added a second
+place to keep in sync.
+
+The lpc11u37 linker comment still described USB buffers living in RamUsb2,
+a placement the same branch removed; nothing lands there now, so say so and
+label the headroom assert as future-proofing."
+```
+
+---
+
+### Task 7: Stop halting the target when a DCD legitimately refuses a transfer
+
+**Files:**
+- Modify: `src/device/usbd.c` (`usbd_edpt_xfer()` failure arm)
+
+**Interfaces:** none consumed or produced.
+
+- [ ] **Step 1: Remove the breakpoint from the DCD-refusal path**
+
+Replace the failure arm of `usbd_edpt_xfer()`:
+
+```c
+ } else {
+ // DCD error, mark endpoint as ready to allow next transfer
+ _usbd_dev.ep_status[epnum][dir] &= (uint8_t) ~(TU_EDPT_STATE_BUSY | TU_EDPT_STATE_CLAIMED);
+ TU_LOG_USBD("FAILED\r\n");
+ TU_BREAKPOINT();
+ return false;
+ }
+```
+
+with:
+
+```c
+ } else {
+ // Driver refused the transfer, mark endpoint as ready to allow next transfer. This is a
+ // recoverable condition (e.g. a new setup superseding a control response), not a bug, so
+ // do not break into the debugger - TU_BREAKPOINT() halts the CPU whenever a probe is
+ // attached, which on a test rig is always.
+ _usbd_dev.ep_status[epnum][dir] &= (uint8_t) ~(TU_EDPT_STATE_BUSY | TU_EDPT_STATE_CLAIMED);
+ TU_LOG_USBD("FAILED\r\n");
+ return false;
+ }
+```
+
+- [ ] **Step 2: Confirm no other stack path relies on that breakpoint**
+
+Run: `grep -n "TU_BREAKPOINT" src/device/*.c src/device/*.h`
+Expected: no remaining hits inside `usbd_edpt_xfer`; other occurrences (if any) are in
+unrelated assert macros and stay as they are.
+
+- [ ] **Step 3: Build and run unit tests**
+
+Run:
+
+```bash
+cmake --build examples/cmake-build-mimxrt1064_evk
+cd test/unit-test && ceedling test:all && cd ../..
+```
+
+Expected: build succeeds, all unit tests pass.
+
+- [ ] **Step 4: Commit**
+
+```bash
+git add src/device/usbd.c
+git commit -m "usbd: do not breakpoint when a driver refuses a transfer
+
+TU_BREAKPOINT() is not gated on CFG_TUSB_DEBUG - it halts the CPU whenever
+a debugger is attached, which on a test rig is always. A driver declining a
+transfer is recoverable (a new setup superseding a control response, for
+one) and the endpoint is already released for the retry, so a halted target
+turns a self-healing case into a dead board."
+```
+
+---
+
+### Task 8: Full validation on hardware
+
+**Files:** none modified — this task produces the evidence for the PR description.
+
+**Interfaces:** consumes the firmware built by Tasks 1-7.
+
+- [ ] **Step 1: Software gate**
+
+Run:
+
+```bash
+pre-commit run --all-files
+cd examples
+for b in mimxrt1064_evk lpcxpresso18s37 lpcxpresso11u37 lpcxpresso55s28; do
+ rm -rf cmake-build-$b && cmake -B cmake-build-$b -DBOARD=$b -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel . && cmake --build cmake-build-$b || echo "FAILED $b"
+done
+cd ..
+```
+
+Expected: pre-commit all green; all four boards build every example.
+
+- [ ] **Step 2: Make-build regression checks**
+
+Run:
+
+```bash
+cd examples/host/cdc_msc_hid && make -j8 BOARD=lpcxpresso55s28 all && cd ../../..
+cd examples/device/cdc_msc_throughput && make -j8 BOARD=lpcxpresso11u37 all && cd ../../..
+```
+
+Expected: both link (these two were broken earlier in the branch and are the regression
+canaries for the BSP changes).
+
+- [ ] **Step 3: Flash with verification (mandatory)**
+
+The mimxrt1064_evk has twice accepted a flash that silently did not take, so every load in
+this task uses `verifyfile`. For each board, write a J-Link script of this shape and run it:
+
+```
+r
+h
+loadfile examples/cmake-build-<board>/device/usbtest/usbtest.elf
+verifyfile examples/cmake-build-<board>/device/usbtest/usbtest.elf
+r
+g
+qc
+```
+
+Probes and devices: `mimxrt1064_evk` = `-USB 000725299165 -device MIMXRT1064xxx6A`,
+`lpcxpresso55s28` = `-USB 000727031389 -device LPC55S28`,
+`lpcxpresso11u37` = `-USB 000724441579 -device LPC11U37/401`.
+Invoke as `JLinkExe <probe/device args> -if swd -speed 4000 -autoconnect 1 -NoGui 1 -CommandFile <script>`.
+Expected: `Verify` reports O.K. and the board re-enumerates as `cafe:4010` with its own
+serial before any test runs.
+
+- [ ] **Step 4: HIL batteries and stress**
+
+Hold each board's lock for its own leg (`python3 test/hil/hil_lock.py hold <board> --reason "reset-edge validation"`,
+release after), never run two batteries at once, and abort if CI is active
+(`pgrep -f "hil_test.py [-]-retry"`).
+
+```bash
+# per board: full battery
+timeout 700 python3 test/hil/usbtest.py --serial <serial> --json --keep-binding --timeout 60
+
+# mimxrt1064_evk only: queued-control stress and the unlink storm
+for i in $(seq 1 50); do timeout 200 python3 test/hil/usbtest.py --serial BAE96FB95AFA6DBB8F00005002001200 --tests 9,10 --json --keep-binding --timeout 60 > /dev/null || break; done
+for i in $(seq 1 10); do timeout 300 python3 test/hil/usbtest.py --serial BAE96FB95AFA6DBB8F00005002001200 --tests 11,12,24 --json --keep-binding --timeout 60 > /dev/null || break; done
+```
+
+Serials: 1064 `BAE96FB95AFA6DBB8F00005002001200`, 55s28 `2BF1839A7D51F553A15AB03FD08F70AB`,
+11u37 `17121919`.
+Expected: 30/30 on all three boards, 50/50 and 10/10 loops, and
+`ps -eo stat,comm | awk '$1 ~ /^D/'` empty after each leg.
+
+- [ ] **Step 5: Reset-path evidence with logging**
+
+Build and flash `device/cdc_msc` for `mimxrt1064_evk` with `-DLOG=2 -DLOGGER=rtt`, capture
+RTT during one unplug/replug cycle (`timeout 20s JLinkRTTClient > /tmp/reset.log`), then:
+
+```bash
+grep -cE "Bus Reset Start" /tmp/reset.log
+grep -cE "Bus Reset End" /tmp/reset.log
+grep -c "Resume" /tmp/reset.log
+```
+
+Expected: equal non-zero counts for start and end (one pair per enumeration) and no
+`Resume` lines during a plain plug-in.
+
+- [ ] **Step 6: Suspend/resume pairing**
+
+With the same RTT build attached, suspend the port from the host and resume it:
+
+```bash
+# find the 1064's busport, then:
+echo auto | sudo tee /sys/bus/usb/devices/<busport>/power/control
+sleep 5
+echo on | sudo tee /sys/bus/usb/devices/<busport>/power/control
+```
+
+Expected in the log: one `Suspend` followed by one `Resume`, and no `Bus Reset` of either
+edge from the suspend cycle alone.
+
+- [ ] **Step 7: Record the evidence**
+
+Append the numbers from Steps 1-6 to the PR description draft. No commit.
+
+## Self-Review
+
+**Spec coverage:** §1 event split → Task 1. §2 ci_hs bus_reset split → Task 2. §3 flush
+helper → Task 2 (Steps 1, 4). §4 mechanical: setup-flush wait and `dcd_set_address` → Task 3;
+RESUME guard → Task 4; ip3511 torn setup and USB.13 TODO → Task 5; usbd breakpoint → Task 7;
+BSP pair → Task 6. Verification matrix → Task 8 (legacy-DCD build guard is Task 1 Step 4).
+Deferred items are deliberately absent from every task. No gaps.
+
+**Placeholder scan:** no TBD/TODO-as-placeholder; the two literal `TODO` strings are
+deliverable code comments (Task 1 Step 3, Task 5 Step 2). Every code step carries the exact
+text to write; every run step carries the command and expected result.
+
+**Type consistency:** `flush_endpoints(ci_hs_regs_t *dcd_reg, uint32_t mask) -> bool` is
+defined in Task 2 Step 1 and used with that exact signature in Task 2 Steps 2/4 and Task 3
+Step 1. `DCD_EVENT_BUS_RESET_START` / `_END` are defined in Task 1 and used in Task 2 Step 3
+via `dcd_event_bus_signal()` / `dcd_event_bus_reset()`, whose signatures are quoted in Task 1's
+Interfaces block. `bus_reset_begin()` / `bus_reset_complete()` are defined and called with
+matching names in Task 2.
diff --git a/docs/superpowers/plans/2026-08-16-drop-ep0-prime-verify.md b/docs/superpowers/plans/2026-08-16-drop-ep0-prime-verify.md
new file mode 100644
index 000000000..aa999c9e3
--- /dev/null
+++ b/docs/superpowers/plans/2026-08-16-drop-ep0-prime-verify.md
@@ -0,0 +1,314 @@
+# Drop the EP0 Post-Prime Verify Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Remove the EP0 post-prime verification that was built on a theory the RT106x endpoint-conflict errata has superseded, and prove on hardware that nothing depended on it.
+
+**Architecture:** One deletion in `qhd_start_xfer()`, then a rebase onto current master, then an A/B validation whose "with it" arm is already banked (10x 30/30 batteries plus 40 targeted loops on 2026-08-16). No interfaces change: the pre-prime setup-lockout guard keeps `qhd_start_xfer()` returning `bool`, so `dcd_set_address()`'s gating and usbd's failure path stay exactly as they are.
+
+**Tech Stack:** C99, TinyUSB ChipIdea HS DCD (`src/portable/chipidea/ci_hs/`), CMake+Ninja and Make builds, J-Link (JLinkExe V9.66), `test/hil/usbtest.py` driving the Linux testusb battery.
+
+## Global Constraints
+
+- Branch `fix-ci-hs` in worktree `/home/hathach/.herdr/worktrees/tinyusb/fix-ci-hs`. Do NOT push; the user pushes.
+- C99, 2-space indent. Commit messages imperative, no `Co-Authored-By:` or `Claude-Session:` trailers (repo rule: hathach is sole author).
+- Pre-commit hook (trailing-whitespace, end-of-file-fixer, codespell, unique-PIDs, ceedling) must pass; if it rewrites a file, re-stage and retry the commit once.
+- Never edit anything under `hw/mcu/` or `lib/` (vendor code).
+- Rig etiquette: hold the board lock for hardware work (`python3 test/hil/hil_lock.py hold <board> --reason "..."`, release after); abort if CI is active (`pgrep -f "hil_test.py [-]-retry"`); NEVER use `uhubctl`, `pci-reset` or `pci-rebind`; never touch the actions-runner.
+- JLinkExe on this rig is **V9.66 and has no `verifyfile` command** — use `loadfile` (built-in Program & Verify) plus a mandatory enumeration check.
+- Board facts: `mimxrt1064_evk`, serial `BAE96FB95AFA6DBB8F00005002001200`, J-Link probe `000725299165`, device `MIMXRT1064xxx6A`, expected `cafe:4010`.
+- Design source of truth: `docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md`.
+
+## File Structure
+
+| File | Responsibility in this plan |
+|---|---|
+| `src/portable/chipidea/ci_hs/dcd_ci_hs.c` | The only code change: delete the post-prime block in `qhd_start_xfer()` |
+
+Tasks 2 and 3 change no files; they rebase and validate.
+
+---
+
+### Task 1: Delete the EP0 post-prime verify
+
+**Files:**
+- Modify: `src/portable/chipidea/ci_hs/dcd_ci_hs.c` (the tail of `qhd_start_xfer()`)
+
+**Interfaces:**
+- Produces: `qhd_start_xfer()` keeps its existing signature `static bool qhd_start_xfer(uint8_t rhport, uint8_t epnum, uint8_t dir)` and still returns `false` from the pre-prime setup-lockout guard. No caller changes.
+
+- [ ] **Step 1: Apply the deletion**
+
+In `qhd_start_xfer()`, replace this (everything from the prime write to the closing `return true;`):
+
+```c
+ // start transfer
+ const uint32_t prime_bit = TU_BIT(epnum + (dir ? 16 : 0));
+ dcd_reg->ENDPTPRIME = prime_bit;
+
+ if (epnum == 0) {
+ // RM (RT1050 RM Executing a Transfer / UM10503 25.10.8): after priming EP0 the DCD must
+ // verify the prime completed - ENDPTPRIME bit clear AND the buffer reported ready in
+ // ENDPTSTAT - because the controller silently cancels an EP0 prime when a SETUP arrives
+ // during the prime operation. An undetected drop NAK-parks the endpoint forever: usbd never
+ // re-primes a busy endpoint. A very fast transfer may already have completed and retired the
+ // ENDPTSTAT bit, so ENDPTCOMPLETE also counts as the prime having taken.
+ uint32_t guard = CI_HS_BUSY_SPIN;
+ while (dcd_reg->ENDPTPRIME & prime_bit) {
+ if (!guard--) {
+ dcd_reg->ENDPTFLUSH = prime_bit; // never leave a wedged prime armed over a freed buffer
+ return false;
+ }
+ }
+ // Fail only when the cancel-cause is visibly pending: a completed transfer can have both
+ // status bits already retired by the ISR, and a cancel whose SETUP the ISR consumed is
+ // re-driven by that queued SETUP event anyway.
+ if (!((dcd_reg->ENDPTSTAT | dcd_reg->ENDPTCOMPLETE) & prime_bit) &&
+ (dcd_reg->ENDPTSETUPSTAT & TU_BIT(0))) {
+ return false; // prime cancelled (setup mid-prime): the pending SETUP re-drives EP0
+ }
+ }
+ return true;
+```
+
+with:
+
+```c
+ // start transfer
+ dcd_reg->ENDPTPRIME = TU_BIT(epnum + (dir ? 16 : 0));
+ return true;
+```
+
+Leave the `if (epnum == 0)` setup-lockout block ABOVE the prime write completely untouched —
+that one spins on `ENDPTSETUPSTAT` before priming and is required by UM10503 25.10.8.1.1
+step 4.
+
+- [ ] **Step 2: Confirm nothing else referenced the removed code**
+
+Run:
+
+```bash
+grep -n "ENDPTSTAT\|ENDPTCOMPLETE\|prime_bit" src/portable/chipidea/ci_hs/dcd_ci_hs.c
+```
+
+Expected: no `prime_bit` hits at all; `ENDPTCOMPLETE` hits only in `bus_reset_begin()` and the
+`INTR_USB` branch of `dcd_int_handler()`; `ENDPTSTAT` hits only in `ci_hs_type.h`-style register
+declarations if any appear — none inside `qhd_start_xfer()`.
+
+- [ ] **Step 3: Build both ci_hs board families**
+
+Run:
+
+```bash
+cmake --build examples/cmake-build-mimxrt1064_evk && cmake --build examples/cmake-build-lpcxpresso18s37
+```
+
+Expected: both succeed, no new warnings (in particular no "unused variable" for anything the
+deletion orphaned).
+
+- [ ] **Step 4: Commit**
+
+```bash
+git add src/portable/chipidea/ci_hs/dcd_ci_hs.c
+git commit -m "dcd(ci_hs): drop the EP0 post-prime verify
+
+The verify came from a theory that a setup arriving mid-prime silently
+cancels an EP0 prime, which was how the recurring wedge on the test rig
+looked at the time. The wedge turned out to be Errata i.MX RT1064_A
+ERR050101: with an isochronous IN endpoint active, an IN token to that
+endpoint number on another device sharing the host unprimes one of our OUT
+endpoints, undetectably and with no interrupt. Moving the usbtest iso IN
+endpoint clear of the conflict fixed it - 340 runs where the board used to
+wedge within hours.
+
+The capture that motivated the verify (EP0 status stage armed but unprimed,
+device a control transfer ahead of the host) is explained by that errata
+just as well, because it covers control OUT endpoints and a control status
+stage is one. So the verify has no independent evidence behind it, while it
+does cost two register spins on every EP0 transfer and can misread a
+transfer the interrupt handler already completed as a cancelled prime.
+
+The setup-lockout check before priming stays - that one is in the manual."
+```
+
+---
+
+### Task 2: Rebase onto current master and re-run the software gates
+
+**Files:** none modified by hand.
+
+**Interfaces:** none.
+
+- [ ] **Step 1: Rebase**
+
+Master has advanced (midi2/usbtmc/video changes) since this branch last rebased. Validating a
+tree that is not the one being merged would be a false pass.
+
+```bash
+git fetch origin master
+git rebase origin/master
+```
+
+Expected: clean rebase. If a conflict appears in `src/portable/chipidea/ci_hs/dcd_ci_hs.c` or
+`src/device/usbd.c`, resolve it hunk-by-hunk keeping BOTH sides' intent (never `git checkout
+--theirs/--ours` on a whole file), then `git rebase --continue`.
+
+- [ ] **Step 2: Rebuild everything from scratch**
+
+```bash
+cd examples
+for b in mimxrt1064_evk lpcxpresso18s37 lpcxpresso11u37 lpcxpresso55s28; do
+ rm -rf cmake-build-$b
+ cmake -B cmake-build-$b -DBOARD=$b -G Ninja -DCMAKE_BUILD_TYPE=MinSizeRel . && cmake --build cmake-build-$b || echo "FAILED $b"
+done
+cd ..
+```
+
+Expected: all four boards build every example, no "FAILED" line.
+
+- [ ] **Step 3: Make link canaries**
+
+```bash
+cd examples/host/cdc_msc_hid && make -j8 BOARD=lpcxpresso55s28 all && cd ../../..
+cd examples/device/cdc_msc_throughput && make -j8 BOARD=lpcxpresso11u37 all && cd ../../..
+```
+
+Expected: both link. These two were broken earlier in the branch's life and are the regression
+canaries for the BSP changes.
+
+- [ ] **Step 4: Unit tests and pre-commit**
+
+```bash
+cd test/unit-test && ceedling test:all && cd ../..
+pre-commit run --all-files
+```
+
+Expected: all unit tests pass; every pre-commit hook passes.
+
+- [ ] **Step 5: No commit**
+
+This task produces no commit of its own — the rebase rewrites existing commits and the builds
+are throwaway. Record the resulting HEAD hash in the report for Task 3 to reference.
+
+---
+
+### Task 3: Hardware A/B on mimxrt1064_evk
+
+**Files:** none modified — this task produces the evidence.
+
+**Interfaces:** consumes the firmware built in Task 2 at
+`examples/cmake-build-mimxrt1064_evk/device/usbtest/usbtest.elf`.
+
+Only this board is tested: it is the sole ci_hs board on the rig. The lpcxpresso55s28 and
+lpcxpresso11u37 run the ip3511 driver, which this change does not touch.
+
+- [ ] **Step 1: Preconditions**
+
+```bash
+pgrep -f "hil_test.py [-]-retry" && echo "CI ACTIVE - wait" || echo "CI idle"
+ps -eo stat,pid,etimes,comm | awk '$1 ~ /^D/'
+python3 test/hil/hil_lock.py hold mimxrt1064_evk --reason "prime-verify removal A/B"
+```
+
+Expected: CI idle, no pre-existing D-state processes, lock acquired. If CI is active, wait for
+it to drain rather than running concurrently.
+
+- [ ] **Step 2: Flash with verification**
+
+```bash
+cat > /tmp/pv.jlink <<'EOF'
+r
+h
+loadfile examples/cmake-build-mimxrt1064_evk/device/usbtest/usbtest.elf
+r
+g
+qc
+EOF
+JLinkExe -device MIMXRT1064xxx6A -if SWD -speed 4000 -SelectEmuBySN 000725299165 \
+ -autoconnect 1 -nogui 1 -CommandFile /tmp/pv.jlink
+```
+
+Expected: `Program & Verify` reports O.K.
+
+- [ ] **Step 3: Confirm the right image is actually running**
+
+```bash
+sleep 5
+grep -l BAE96FB95AFA6DBB8F00005002001200 /sys/bus/usb/devices/*/serial
+sudo lsusb -v -d cafe:4010 2>/dev/null | grep -A3 "Isochronous" | grep bEndpointAddress
+```
+
+Expected: the board is present, and the iso IN endpoint reads **0x87**. If it reads 0x83 the
+flash did not take (this board has silently no-op'd a flash twice) — reflash and re-check
+before running anything.
+
+- [ ] **Step 4: 5x full battery**
+
+```bash
+for i in $(seq 1 5); do
+ timeout 700 python3 test/hil/usbtest.py --serial BAE96FB95AFA6DBB8F00005002001200 \
+ --json --keep-binding --timeout 60 2>/dev/null | python3 -c "
+import json,sys
+d=json.load(sys.stdin)
+bad=[str(c['num']) for c in d['cases'] if c['status']!='PASS']
+print(f\"run: {d['passed']}/30 speed={d['speed']}\" + (' FAILED:'+','.join(bad) if bad else ''))
+"
+ ps -eo stat,pid,etimes,comm | awk '$1 ~ /^D/ && $4=="testusb"'
+done
+```
+
+Expected: five lines each reading `30/30 speed=480`, and no testusb D-state line between runs.
+
+- [ ] **Step 5: 15x control-focused loop**
+
+These are the paths the removed verify actually protected — queued control, the ch9 subset, and
+both ctrl_out cases. A full battery samples each only once per run.
+
+```bash
+PASS=0
+for i in $(seq 1 15); do
+ timeout 300 python3 test/hil/usbtest.py --serial BAE96FB95AFA6DBB8F00005002001200 \
+ --tests 9,10,14,21 --json --keep-binding --timeout 60 >/dev/null 2>&1 && PASS=$((PASS+1)) || { echo "FAILED at iteration $i"; break; }
+ D=$(ps -eo stat,comm | awk '$1 ~ /^D/ && $2=="testusb"' | wc -l)
+ [ "$D" != "0" ] && { echo "D-STATE at iteration $i"; break; }
+done
+echo "control loops: $PASS/15"
+```
+
+Expected: `control loops: 15/15`, no FAILED or D-STATE line.
+
+- [ ] **Step 6: Release the lock and record**
+
+```bash
+python3 test/hil/hil_lock.py release mimxrt1064_evk
+ps -eo stat,pid,etimes,comm | awk '$1 ~ /^D/'
+```
+
+Expected: lock released, no leftover D-state.
+
+**Acceptance:** 5/5 batteries at 30/30, 15/15 control loops, no `testusb` D-state outliving its
+case runtime.
+
+**Rollback trigger:** any control-case failure (errno 110 or 71 on cases 9, 10, 14, 21) or a
+lingering D-state means the verify was load-bearing after all. In that case: `git revert` the
+Task 1 commit, re-run Steps 4-5 to confirm the failure disappears, and record the result — that
+is a finding worth keeping, not a setback to hide.
+
+---
+
+## Self-Review
+
+**Spec coverage:** the spec's change section → Task 1; "rebase first, then rebuild" → Task 2
+Steps 1-2; software gates → Task 2 Steps 3-4; hardware preconditions, verified flash and the
+0x87 descriptor check → Task 3 Steps 1-3; 5x battery and 15x control loop → Task 3 Steps 4-5;
+acceptance and rollback trigger → Task 3's closing block. The spec's "deliberately kept" list is
+enforced negatively by Task 1 Step 1's instruction to leave the setup-lockout block untouched
+and by Task 1 Step 2's grep. No gaps.
+
+**Placeholder scan:** no TBD/TODO/"handle edge cases"; every step carries its exact command or
+code and its expected result.
+
+**Type consistency:** `qhd_start_xfer(uint8_t rhport, uint8_t epnum, uint8_t dir) -> bool` is
+unchanged by this plan and no caller is touched, so there are no cross-task signatures to
+reconcile. The only removed identifier, `prime_bit`, is local to the deleted block and Task 1
+Step 2 greps to confirm it has no remaining references.
diff --git a/docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md b/docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md
index 8158758bc..898b3c8ab 100644
--- a/docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md
+++ b/docs/superpowers/specs/2026-07-29-hil-pr-scoped-selection-design.md
@@ -1,4 +1,4 @@
-# PR-scoped HIL selection: hil_select.py
+# PR-scoped HIL selection: helper/hil_select.py
**Date:** 2026-07-29
**Branch:** `claude/hil-select` (based on `claude/hil-pool-check`, which carries the
@@ -26,17 +26,17 @@ confident; every uncertainty widens to the full matrix.
- Scoping push/master/scheduled runs (always full).
- Changing hil_test.py behavior (the selector only *composes* existing `-b`/`-bt` args).
-## Component: `test/hil/hil_select.py`
+## Component: `test/hil/helper/hil_select.py`
Stdlib-only, importable and CLI. Lives beside the harness so `hil_ci.sh` copies are unaffected
(it runs on the GitHub runner / dev PC, not on the rig). It must NOT import `hil_test.py`
(which drags pyserial/pymtp onto the bare GitHub runner): the three test lists
-(`device_tests`, `dual_tests`, `host_test`) move verbatim into a tiny stdlib-only
-`test/hil/hil_examples.py` that both `hil_test.py` and `hil_select.py` import (behavior
-preserving; `hil_ci.sh` scp list gains the new file).
+(`device_tests`, `dual_tests`, `host_test`) move verbatim into the stdlib-only
+`test/hil/helper/hil_util.py` that both `hil_test.py` and `hil_select.py` import (behavior
+preserving; `hil_ci.sh` copies the whole `helper/` directory).
```
-python3 test/hil/hil_select.py --base <ref> [--diff-file <path>] CONFIG.json [CONFIG.json...]
+python3 test/hil/helper/hil_select.py --base <ref> [--diff-file <path>] CONFIG.json [CONFIG.json...]
```
- `--base REF`: changed files = `git diff --name-only $(git merge-base HEAD REF)..HEAD`
@@ -118,7 +118,7 @@ is skipped, not widened (running unrelated boards would test nothing relevant).
## CI wiring (`.github/workflows/build.yml`)
- `set-matrix` (PR events only): after generating today's matrices, run
- `hil_select.py --base origin/${{ github.base_ref }} test/hil/tinyusb.json test/hil/hfp.json`
+ `helper/hil_select.py --base origin/${{ github.base_ref }} test/hil/tinyusb.json test/hil/hfp.json`
(checkout with enough history to reach the merge base: `fetch-depth: 0` on this one job, or
an explicit `git fetch origin $BASE_REF`). New job outputs: `hil_select_full`,
`hil_args_tinyusb`, `hil_args_hfp`, plus the selected-board list consumed by the matrix
@@ -137,15 +137,15 @@ is skipped, not widened (running unrelated boards would test nothing relevant).
## Local use
- pre-pr's "Map changes to boards" step delegates to
- `python3 test/hil/hil_select.py --base $BASE test/hil/tinyusb.json` and derives its
+ `python3 test/hil/helper/hil_select.py --base $BASE test/hil/tinyusb.json` and derives its
one-board-per-family sample from the selector's board set (its capping/sampling policy is
unchanged — the selector provides the affected set, pre-pr samples it).
-- Manual: `python3 test/hil/hil_test.py -B examples $(python3 test/hil/hil_select.py --base master test/hil/tinyusb.json | jq -r '.args["tinyusb.json"]') test/hil/tinyusb.json`
+- Manual: `python3 test/hil/hil_test.py -B examples $(python3 test/hil/helper/hil_select.py --base master test/hil/tinyusb.json | jq -r '.args["tinyusb.json"]') test/hil/tinyusb.json`
— documented in the hil skill.
## Testing
-`test/hil/test_hil_select.py` — stdlib `unittest`, no hardware, injected diffs via
+`test/hil/test/test_hil_select.py` — stdlib `unittest`, no hardware, injected diffs via
`--diff-file`/API. Cases (the acceptance examples):
1. `src/portable/raspberrypi/rp2040/dcd_rp2040.c` → only rp2040-family roster boards, device
tests only, host-only boards absent, `full` false.
@@ -161,7 +161,7 @@ is skipped, not widened (running unrelated boards would test nothing relevant).
7. `hw/bsp/rp2040/family.cmake` → rp2040-family boards, all their tests.
8. Mixed device+host diff → no pruning (both roles present).
The suite runs in `set-matrix` before the selector is used, and locally via
-`python3 test/hil/test_hil_select.py`.
+`python3 test/hil/test/test_hil_select.py`.
## Safety properties
diff --git a/docs/superpowers/specs/2026-07-30-hil-usbtest-fleet-wedge-design.md b/docs/superpowers/specs/2026-07-30-hil-usbtest-fleet-wedge-design.md
new file mode 100644
index 000000000..3ed0c1519
--- /dev/null
+++ b/docs/superpowers/specs/2026-07-30-hil-usbtest-fleet-wedge-design.md
@@ -0,0 +1,236 @@
+# HIL fleet-wedge containment
+
+Date: 2026-07-30
+Status: implemented, then superseded in part — addendum last checked 2026-08-12
+against the shipped code; where they disagree the CODE and the usb-kernel-recover
+skill win, never this document.
+
+- **Pool guard.** A single constant, not the flat 4200s below and not a derivation:
+ `POOL_TIMEOUT = pos_int_env('HIL_POOL_TIMEOUT', 3600)`. A per-controller model briefly
+ lived here and was removed -- it under-modelled the flash phase and could INVERT
+ (adding a usbtest board lowered the guard, because the derived value fell below the
+ baseline it was meant to raise). The guard's only job is to stop a wedged pool short
+ of the job ceiling so the report still gets written; predicting a healthy run's
+ duration is a different problem. `pos_int_env` warns only on a non-integer or a value
+ <= 0: there is NO upper clamp and no warning above any threshold, so a pin larger than
+ a job ceiling silently restores the inversion this work removed.
+- **Job ceilings.** 90/90/120 min (build.yml), not 60/60/90 and not the 85/115 below.
+ They must clear the 3600s guard plus the pre-pool checkout/artifact merge and the
+ post-guard sweep and report upload. No job pins `HIL_POOL_TIMEOUT`.
+- **Battery budgets.** `USBTEST_BATTERY_BUDGET` 260s, `USBTEST_RECOVERY_BUDGET` 250s.
+ The 200s-with-a-197s-floor derivation recorded here was never shipped; the floor
+ assertion was removed with it.
+- **HUNG recovery.** Reflash of the DUT through its roster flasher
+ (`usbtest.py --recover-board/--recover-fw`), not the root-cycle-first recovery in
+ section 1d — replaced after the 2026-08-11 ppps measurement (uhubctl never cuts
+ VBUS; root-cycle is probe-only). Since 2026-08-12 the reflash is SKIPPED
+ when `hil_flash.convoy_safe(board['flasher'])` is false (usbtest.py:675): the flasher
+ would enumerate by opening usbfs nodes, block on the same convoy, and become a second
+ stray rather than clear the first. A holder that owns the device lock inside a driver
+ ioctl is terminal either way -- a reflash only produces a disconnect, and
+ `usb_disconnect()` needs that same lock -- and that state needs a reboot.
+
+Step 0 done — the host was rebooted 2026-07-30 14:11 and the rig
+came back clean. The device that triggered this incident was removed from the rig, so
+only the containment work remains relevant.
+Rig: `ci.lan` (Proxmox guest on `pve.lan`)
+
+## Problem
+
+On 2026-07-29/30 every board in the `ci.lan` usbtest fleet failed, `openocd` processes
+landed in uninterruptible sleep, and no subsequent HIL run could start. Two GitHub
+Actions runs were stranded: `30484641269` sat `in_progress` for over eight hours
+(past GitHub's own 360-minute default), and `30485082274` sat `queued` behind it from
+2026-07-29 19:35 UTC onward. Both report directories were written empty.
+
+A reboot of the `ci` guest at 10:48 did not clear the condition: the same kernel state
+re-formed at 10:52:23.
+
+## Root cause
+
+Five layers, each independently observable.
+
+### 1. A permanently wedged hub worker holds a root-hub device lock
+
+A device that repeatedly re-asserts connect while failing to enumerate keeps
+`hub_event()` busy, and `hub_event()` holds `usb_lock_device(hdev)` on its hub for its
+whole run (hub.c:5896/5989). The `usb_hub_wq` worker sits in `hub_port_reset`, so that
+hub's `device_lock` is effectively never released:
+
+```
+kworker/14:6+usb_hub_wq (state D, 400+ s)
+ msleep+0x2b
+ hub_port_reset+0x1a4 [usbcore]
+ hub_event+0x727 [usbcore]
+```
+
+`usb usbN-portM: Cannot enable. Maybe the USB cable is bad?` is logged every four seconds
+for as long as it lasts.
+
+Verified against hub.c v6.12.96 rather than inferred: the kernel does **not** retry
+without bound, and root and downstream ports are bounded identically —
+`hub_port_reset()` tries `PORT_RESET_TRIES` then logs that message (hub.c:3149),
+`hub_port_connect()` wraps it in `PORT_INIT_TRIES` = 4 and disables the port on give-up
+(hub.c:5455/5619). A count in the thousands is therefore that many separate connect
+events, not one runaway loop, and it indicts the device rather than the port.
+
+### 2. A parked board storms the second controller
+
+`ra6m5_ek` (`test/hil/tinyusb.json`, uid `8419032D32363657364EF4622D294B4E`, at
+`13-3.3`) runs dfu firmware (`cafe:400b`) and re-enumerates every 1-2 seconds
+continuously, wrapping the entire bus-13 devnum space (`...120 -> 127 -> 4 -> 6 -> 10`).
+This is standing `hub_event` and Address-Device pressure on controller `03:00.0`,
+concurrent with parallel usbtest batteries on the same silicon.
+
+The board is already listed in `boards-skip`, which is precisely why it storms:
+`boards-skip` stops testing a board but never parks it, so it keeps running whatever
+firmware it last received. Park-flash only runs as teardown of a board that actually
+executed tests.
+
+### 3. The kernel `usbtest` control-queue case waits without a timeout
+
+`test_ctrl_queue` blocks on an untimed `wait_for_completion()` while `usbdev_ioctl`
+holds the DUT's `device_lock`:
+
+```
+wait_for_completion+0x8a <- no _timeout variant
+test_ctrl_queue+0x4ab [usbtest]
+usbtest_do_ioctl+0x501 [usbtest]
+usbdev_ioctl+0x6b8 [usbcore]
+```
+
+`--timeout 60` in `test/hil/usbtest.py` is a subprocess timeout only. `SIGKILL` is not
+delivered to a task in uninterruptible sleep. `usbtest.py` already recognises this and
+reports `HUNG`, then calls `usb_recover.sh root-cycle`.
+
+### 4. openocd inherits the convoy and the whole fleet dies
+
+Once a device lock is stuck, `port_event()` takes a child device's lock to warm-reset
+it and blocks while still holding its hub's lock. Any later
+`open("/dev/bus/usb/BBB/DDD")` against such a device blocks uninterruptibly:
+
+```
+usbdev_open+0xdc [usbcore] -> __mutex_lock
+chrdev_open -> do_sys_openat2 -> __x64_sys_openat
+```
+
+That is the state of the three `openocd` processes at 04:16:51 (pids 207921, 207987,
+208034) — the flasher, unkillable. Because one controller carries two buses, a single
+convoy takes out every board on both, which is why the failure presents as the entire
+fleet.
+
+The existing `HUNG` recovery cannot help here. A root-port VBUS cycle frees a
+*device-lock* holder; it cannot free a lock held by a stuck *hub worker*, and on this
+rig the cycle lands on the controller that is already wedged.
+
+### 5. Nothing bounds the damage, so one bad run becomes a CI outage
+
+- `hil-tinyusb` and `hil-tinyusb-esp` in `.github/workflows/build.yml` carry no
+ `timeout-minutes`. Only `hil-hfp-iar` does.
+- `ci.lan` runs a single runner service, so there is one job slot.
+- `test/hil/hil_test.py` bounds the pool with `POOL_TIMEOUT` (4200 s), and that guard
+ fires correctly — but the recovery path does not survive a D-state worker:
+
+```python
+with Pool(processes=os.cpu_count() or 1, initializer=init_worker, initargs=initargs) as pool:
+ async_ret = pool.map_async(test_board, config_boards)
+ try:
+ mret = async_ret.get(timeout=POOL_TIMEOUT)
+ except MpTimeoutError:
+ pool.terminate()
+ pool.join() # blocks forever: a D-state worker never reaps
+ raise RuntimeError(f'HIL worker pool timed out after {POOL_TIMEOUT}s')
+```
+
+`multiprocessing` joins workers unbounded, so both `pool.terminate()` and
+`pool.join()` hang, as does the `with Pool(...)` exit on the success path. Normal
+`hil-tinyusb (tinyusb.json)` runs take 10-20 minutes; one recent run took 71.3
+minutes, which is the 70-minute guard firing and succeeding. The eight-hour run is the
+pathological case.
+
+## Design
+
+### Step 0 — recovery (manual prerequisite)
+
+Power-cycle the PVE **host**, not the `ci` guest. A guest reboot is not sufficient;
+hubs latch up across the PCIe reset, which the 10:48 reboot demonstrated. Nothing
+below can be verified until the rig is clean.
+
+### Section 1 — CI containment
+
+**1a. Two layered timers.** An inner guard inside `hil_test.py` (`POOL_TIMEOUT`, 70 min)
+that fails gracefully -- it writes a report naming the timeout and the dispatched boards,
+shuts the pool down and exits -- and an outer `timeout-minutes` per rig job (85 for the
+hil-tinyusb jobs; 115 for hil-hfp-iar, which also builds four boards with IAR in the same
+job) as the backstop for when even exiting cannot free the runner. The ceiling must stay
+ABOVE the inner guard, or GitHub kills the job before the report is written.
+
+> **Corrected after measurement.** An earlier revision cut the guard to 30 min on the
+> reading that real runs take 9-17 min and everything longer was the old guard firing.
+> That was wrong. `hil_lock.py` records 22.2/14.3/12.5/10.8 min at usbtest width 1/2/3/4,
+> and raising the per-battery budget to 380s made hung boards cost more again. The 30 min
+> guard then fired on 5 of the last 8 HIL job executions across both rigs, and because
+> `map_async` is all-or-nothing each of those runs published a banner instead of any
+> per-board result. Restored to 4200s, the value whose original rationale -- usbtest
+> batteries are serialized fleet-wide, lengthening the tail -- was correct.
+
+**1b. Bound the pool shutdown.** Add a helper to `test/hil/hil_test.py`:
+
+```python
+def _shutdown_pool(pool, grace=30):
+ """terminate() a Pool without ever blocking forever: multiprocessing joins its
+ workers unbounded, and a worker in uninterruptible sleep (wedged usbfs) never
+ reaps -- which would hold the runner's only job slot indefinitely."""
+ t = threading.Thread(target=pool.terminate, daemon=True)
+ t.start()
+ t.join(grace)
+ return not t.is_alive()
+```
+
+On the `MpTimeoutError` path: write the report first, recording the boards that never
+reported so the run stops producing an empty report directory; then `_shutdown_pool`;
+then `os._exit(1)` if it did not return. The hard exit is the point — it is the only
+way past a kernel-side unkillable child. Use the same helper for the `with Pool(...)`
+exit path.
+
+**1c. Pre-flight rig health check.** `check_rig_health()` runs before the build and
+**never aborts**. It probes `/proc` unprivileged (dmesg is restricted on the rig) for a
+wedged `usb_hub_wq` worker, and reports a `/proc` too restricted to trust as its own
+distinct cause rather than as a diagnosed fault.
+
+It is deliberately non-fatal: the rig is unattended and every remedy for a real wedge is
+manual, so aborting would not fix anything -- it would discard the per-board results the
+run can still collect and leave CI red until a human noticed. It emits a GitHub
+`::error::` annotation and continues. The automatic containment is 1a and 1b, which bound
+a stuck run and explain it without anyone touching the rig.
+
+**1d. Order the recovery correctly.** In `test/hil/usbtest.py`, attempt
+`usb_recover.sh root-cycle` FIRST on a `HUNG` case, and only check for a wedged hub worker
+*afterwards*.
+
+> **Corrected during implementation.** This section originally said to check for a wedged
+> worker *before* the cycle and skip it on a hit. That is backwards. Our own stuck
+> `testusb` holds the DUT's device lock, so any port event drives a hub worker into
+> `usb_lock_device()` on it -- uninterruptible, so it reads `D` in ~100% of samples and the
+> confirmation window makes the wrong verdict *more* confident, not less. Cutting VBUS is
+> precisely what completes the in-flight URB, returns the ioctl and frees that worker, so
+> gating on that signature would suppress the recovery in the exact ordering it exists for.
+> A worker still wedged after the cycle is the genuinely unrecoverable case, and that is
+> what the code now reports.
+
+## Verification
+
+- Unit-test `shutdown_pool` and the `hil_health` detectors against a synthetic `/proc`.
+ A real wedge cannot be manufactured on demand, so they are tested against fabricated
+ inputs rather than live hardware.
+- Confirm the detectors flag a genuinely wedged rig, and return clean on a healthy one.
+- One clean full-fleet `hil_test.py` run to prove `check_rig_health` does not
+ false-abort.
+
+## Out of scope
+
+- **`ra6m5_ek` park and its dfu reset loop.** Dropped by decision. Consequence: the
+ layer-2 devnum storm remains as standing pressure on controller `03:00.0`. Unplugging
+ the board or flashing `board_test` by hand resolves it without any code change.
+- **An unattended PVE watchdog** that detects the wedge and power-cycles the host.
+ Declined: more moving parts, and it can cut a running CI job.
diff --git a/docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md b/docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md
new file mode 100644
index 000000000..e01831d34
--- /dev/null
+++ b/docs/superpowers/specs/2026-08-15-ci-hs-reset-edges-design.md
@@ -0,0 +1,162 @@
+# Bus-reset edge events + review fix wave — design
+
+Date: 2026-08-15
+Branch: `fix-ci-hs` (unpushed, 6 commits over master `53fef2833`)
+
+## Problem
+
+A max-effort review of the branch produced 15 findings. Four are regressions the branch
+itself introduced; the rest are pre-existing or cross-cutting. The load-bearing one:
+
+`dcd_ci_hs.c` now runs the RM-prescribed reset cleanup at the URI (reset-start) interrupt
+but does not tell usbd until the Port Change Detect that ends the reset. For the whole
+reset window — a minimum of 3 ms, typically 10–50 ms — usbd still believes the device is
+configured while the DCD's queue heads have been zeroed. A class driver writing in that
+window (`tud_hid_n_report()`, `tud_cdc_write_flush()`) primes a disabled endpoint over a
+zeroed dQH, *after* the cleanup's flush, so the stale prime survives re-enumeration over a
+buffer usbd has already released. On a 600 MHz M7 that window is enormous. Master had no
+gap: cleanup and event were adjacent statements.
+
+The stack has no way to express "reset started" — `DCD_EVENT_BUS_RESET` carries the
+negotiated speed, which does not exist until the reset ends. That missing vocabulary is
+the actual defect; the driver-level workarounds considered (deferring the memclr, guarding
+primes with a private flag) only shrink the window.
+
+## Design
+
+### 1. Stack: split the bus-reset event into two edges
+
+`src/device/dcd.h`:
+
+```c
+DCD_EVENT_BUS_RESET_START, // reset signaling detected; bus unusable, speed unknown
+DCD_EVENT_BUS_RESET_END, // reset complete; .bus_reset.speed is final
+...
+#define DCD_EVENT_BUS_RESET DCD_EVENT_BUS_RESET_END // backward compatibility
+```
+
+No new helper: `dcd_event_bus_reset(rhport, speed, in_isr)` keeps its name and emits
+`_END`, so every other port is bit-identical to today; `_START` uses the existing
+payload-free `dcd_event_bus_signal()`. The alias keeps unit-test/fuzz references
+compiling.
+
+**Contract (documented in `dcd.h`):** `_START` is optional. A DCD that cannot distinguish
+the two edges emits only `_END`, which stays self-sufficient — it performs the full
+teardown with or without a preceding `_START`.
+
+`src/device/usbd.c`:
+- `case DCD_EVENT_BUS_RESET_START:` → `usbd_reset(rhport)` only; speed untouched.
+- `case DCD_EVENT_BUS_RESET_END:` → unchanged (`usbd_reset()` + latch speed).
+- `_usbd_event_str[]` gains both names.
+- `TODO:` note that a DCD signalling both edges should not pay for two teardowns — track
+ a per-rhport "start seen" flag and skip the redundant `usbd_reset()` in `_END`, keeping
+ the unconditional teardown for the legacy single-event path.
+
+Cost, accepted deliberately: one extra queued event and one extra `usbd_reset()` per
+enumeration on ci_hs only, bounded at one per reset against a default
+`CFG_TUD_TASK_QUEUE_SZ` of 16 (queue pressure is the failure PR #3817 fixed, hence the
+explicit note).
+
+### 2. ci_hs: split `bus_reset()` along the register/software line
+
+- **`bus_reset_begin()` — at URI, inside the reset window (UM10503 25.10.3):** ENDPTCTRL
+ type-reset loop, `ENDPTNAK`/`ENDPTNAKEN`, `ENDPTSETUPSTAT` and `ENDPTCOMPLETE`
+ write-back clears, bounded `ENDPTPRIME` drain, `ENDPTFLUSH` all. Emit `_START`.
+ Registers only — nothing in `_dcd_data` is touched, so no software structure is pulled
+ out from under a task mid-`dcd_edpt_xfer`.
+- **`bus_reset_complete()` — at the PCI ending the reset:** re-flush, `tu_memclr(&_dcd_data)`,
+ EP0 queue-head re-init, dcache clean. Emit `_END` with the final PSPD speed.
+
+Two properties fall out: the re-flush kills any prime armed during the window without a
+new state flag, and the memclr now happens at the same instant usbd is told, so the
+"configured over zeroed queue heads" mismatch is eliminated rather than shrunk. Residual
+exposure (a task priming exactly as the ISR memclrs) equals master's.
+
+The reason-dispatch (`pci_reason`, suspend/URI ordering) is unchanged; only the reset
+case's body moves.
+
+### 3. ci_hs: one bounded-flush helper
+
+Extract `flush_endpoints(dcd_reg, mask)` — writes `ENDPTFLUSH = mask`, spins bounded by
+`CI_HS_BUSY_SPIN` until those bits clear, returns `true` if they cleared — and route all
+five flush sites through it (`bus_reset_begin`, `bus_reset_complete`, `dcd_deinit`,
+`dcd_edpt_iso_activate`, the setup-time EP0 flush). The unified part is the mechanism
+(one bound, one spin idiom, one return convention); callers keep their existing reactions,
+all of which currently proceed regardless, and that stays true here — no caller gains new
+error handling in this wave. Without this, §2 adds a fifth site to a file that already
+carried four hand-rolled variants.
+
+### 4. Mechanical fixes
+
+`dcd_ci_hs.c`
+- Setup-time EP0 flush waits for completion (via §3's helper) before the SETUP event is
+ queued, so the flush can no longer still be asserted when the task primes the response —
+ which also dissolves its interaction with the post-prime verify. This adds a bounded
+ spin in ISR context; the RM notes a flush waits out any packet already in progress, so
+ the wait is one packet time (microseconds at HS) and the existing `CI_HS_BUSY_SPIN`
+ bound caps the pathological case, consistent with the file's other flush sites.
+- `dcd_set_address()` writes `DEVICEADDR` only if the status-ZLP prime took. A refused
+ prime means a newer SETUP superseded the transfer; staging an address whose ACK will
+ never arrive is wrong.
+- Emit `DCD_EVENT_RESUME` only when `!(PORTSC1 & PORTSC1_SUSPEND)` (restores master's
+ hardware guard, lost in the rework).
+
+`dcd_lpc_ip3511.c`
+- Deliver the setup copy only when known-good:
+ `if (latch still set) { INTSETSTAT = TU_BIT(0); } else { dcd_event_setup_received(...); }`.
+- `TODO:` token on the USB.13 deferral so backlog sweeps surface it.
+
+`usbd.c`
+- The DCD-refusal path in `usbd_edpt_xfer` stops routing through the breakpoint-carrying
+ assert: a DCD declining a prime is documented and self-healing, not a programming error,
+ and `TU_BREAKPOINT()` is not gated on `CFG_TUSB_DEBUG` — with a probe attached (always,
+ on the rig) it halts the target. Log and return false instead.
+
+BSP
+- Delete the seven-line RHPORT block in `lpcxpresso55s28/board.cmake` (byte-identical to
+ `family.cmake`'s own guards; `board.mk`'s `?=` stays as the idiomatic Make form).
+- `lpc11u37.ld`: correct the stale comment (nothing lands in RamUsb2 in either build
+ system now — the stack owns the whole bank) and keep the ASSERT, re-labelled as
+ future-proofing.
+
+## Findings improved for free (documented, no code)
+
+A reset that starts and never completes — cable pulled mid-reset — now delivers `_START`
+and tears usbd down, where before usbd stayed configured on a dead bus. This softens both
+the adjudicated UNPLUGGED-removal finding and the deferred aborted-reset item: a stray
+later PCI delivering `_END` becomes harmless (usbd already torn down, just latches a
+speed) instead of deconfiguring a live device. True detach detection still requires OTGSC
+B-session-valid VBUS sensing — board-dependent, still a follow-up.
+
+## Explicitly deferred
+
+- Prime verification generalized to all endpoints and all causes (RM 25.10.8.2); the
+ EP0/SETUP-gated form stays, its flush interaction fixed by §4.
+- usbd discards `usbd_control_xfer_cb`/`tud_control_xfer` returns — cross-DCD behavior
+ change needing its own regression pass, despite `usbd.c` being open here.
+- Timed-out flush still proceeds to the memclr (now confined to one helper).
+- LPC55S2x USB.3 FORCE_FS workaround; iso-IN 1023 enforcement; 8-byte OUT-spill
+ enforcement; USB.13 INTONNAK workaround.
+- Gating `TU_BREAKPOINT()` on `CFG_TUSB_DEBUG` stack-wide.
+- Unguarded `set()` RHPORT knobs in ~14 sibling `board.cmake` files.
+
+## Verification
+
+1. `pre-commit run --all-files`; builds for mimxrt1064_evk, lpcxpresso18s37,
+ lpcxpresso11u37, lpcxpresso55s28, plus Make link checks for the two previously-broken
+ targets (`host/cdc_msc_hid` on 55s28, `device/cdc_msc_throughput` on 11u37).
+2. Cross-DCD build guard: one non-ci_hs, non-ip3511 board (e.g. `stm32f407disco`) to prove
+ the `DCD_EVENT_BUS_RESET` alias keeps legacy ports compiling untouched.
+3. HIL on byte-verified flash (`verifyfile` on every J-Link load — the 1064's silent
+ flash no-op has struck twice): usbtest 30/30 on mimxrt1064_evk, lpcxpresso55s28,
+ lpcxpresso11u37; 50× case-9/10 loops on the 1064; 10× case-11/12/24 unlink loops.
+4. Reset-path specific: confirm HS enumeration (480) and, with `LOG=2`, that a single
+ enumeration shows exactly one `_START`/`_END` pair and no spurious RESUME.
+5. Suspend/resume exercise on the 1064 (host-side autosuspend on the port) confirming
+ `SUSPEND`/`RESUME` pairing and no reset misclassification.
+
+## Success criteria
+
+All four regressions closed, no new findings in a scoped re-review of the wave diff, every
+listed HIL result green on verified flash, and legacy DCDs provably untouched (alias build
+check + unchanged `_END` semantics).
diff --git a/docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md b/docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md
new file mode 100644
index 000000000..cc1840972
--- /dev/null
+++ b/docs/superpowers/specs/2026-08-16-drop-ep0-prime-verify-design.md
@@ -0,0 +1,90 @@
+# Drop the EP0 post-prime verify — design
+
+Date: 2026-08-16
+Branch: `fix-ci-hs` (unpushed, 19 commits over merge-base `53fef2833`)
+
+## Context
+
+The branch grew while chasing a wedge on `mimxrt1064_evk`: the board would stop answering a
+host transfer, the URB would never complete, `testusb` would block uninterruptibly and the
+whole rig would follow it down. Eight occurrences over four days, across the Linux usbtest
+battery's queued control and bulk tests.
+
+The cause turned out to be silicon: **Errata i.MX RT1064_A / RT1060_A ERR050101**. While an
+isochronous IN endpoint is active, an IN token addressed to that same endpoint number on
+another device sharing the host silently unprimes one of this device's OUT endpoints —
+control, bulk, interrupt or isochronous. NXP states it cannot be detected by software and
+raises no interrupt. Moving the usbtest example's iso IN endpoint from 3 to 7 (commit
+`42870b15b`) cleared it: 340 consecutive wedge-free runs, where the board previously
+re-wedged within hours.
+
+Before that was known, an earlier theory — a SETUP arriving mid-prime silently cancelling an
+EP0 prime — produced a post-prime verification block in `qhd_start_xfer()`. That theory's
+supporting capture (EP0's status ZLP armed but unprimed, the device a control transfer ahead
+of the host) is explained by ERR050101 just as well, because the errata explicitly covers
+*control* OUT endpoints and a control status stage **is** an OUT endpoint. The generalized
+version of that verify was already reverted (`565bb0d99`) as both regression-prone and aimed
+at a failure the vendor documents as undetectable in software. This spec removes what
+remains of it.
+
+## Change
+
+Delete the post-prime block in `qhd_start_xfer()` (`src/portable/chipidea/ci_hs/dcd_ci_hs.c`):
+the bounded `ENDPTPRIME` drain, the `ENDPTFLUSH`-on-timeout, and the
+`ENDPTSTAT | ENDPTCOMPLETE` / `ENDPTSETUPSTAT` verdict. The tail becomes:
+
+```c
+ // start transfer
+ dcd_reg->ENDPTPRIME = TU_BIT(epnum + (dir ? 16 : 0));
+ return true;
+```
+
+This removes two register spins and four volatile reads from every EP0 transfer, and with
+them the false-fail path a reviewer flagged: a transfer the interrupt handler has already
+completed reads identically to a cancelled prime.
+
+## Deliberately kept
+
+- **The pre-prime setup-lockout guard** directly above it — UM10503 25.10.8.1.1 step 4
+ verbatim ("Before priming for status/handshake phases ensure that ENDPTSETUPSTAT is '0'"),
+ and older than the wedge theory. It also keeps `qhd_start_xfer()` returning `bool`, so
+ `dcd_set_address()`'s gating and the usbd breakpoint removal stay meaningful — no cascade.
+- **The setup-time EP0 flush and its completion wait** — the flush is the 25.10.8.1.1 step-3
+ remark; the wait exists because an unfinished flush can retire a freshly primed response,
+ an interaction independent of the verify.
+- **The `BUS_RESET_START`/`END` split** and the rest of the review-driven hardening.
+- Everything hardware-proven: the rf_tv fix, the lpc11u37 stack move, the lpc55s28
+ onboarding, the lpc55 Make OHCI link, and the ERR050101 endpoint move itself.
+
+The commit message records the corrected attribution of the handoff capture, so the next
+reader does not re-derive the superseded theory from the same evidence.
+
+## Validation
+
+The "with it" arm is already banked from 2026-08-16: 10x 30/30 batteries plus 15x TEST 27,
+15x tests 9/10 and 10x tests 11/12/24, all clean. This is the second half of an A/B.
+
+1. **Rebase onto current master first** (master has moved: midi2/usbtmc/video), then rebuild —
+ otherwise the validated tree is not the tree that merges.
+2. **Software gates:** `pre-commit run --all-files`; full example builds for
+ mimxrt1064_evk, lpcxpresso18s37, lpcxpresso11u37, lpcxpresso55s28; the two Make link
+ canaries (`host/cdc_msc_hid` on lpcxpresso55s28, `device/cdc_msc_throughput` on
+ lpcxpresso11u37); `ceedling test:all`.
+3. **Hardware — mimxrt1064_evk only.** It is the only ci_hs board on the rig; the other two
+ run ip3511, which this change does not touch. Preconditions: CI idle
+ (`pgrep -f "hil_test.py [-]-retry"`), board lock held for the whole run. Flash with
+ `loadfile` (its built-in Program & Verify — JLinkExe V9.66 has no `verifyfile`), then
+ confirm re-enumeration as `cafe:4010` with serial `BAE96FB95AFA6DBB8F00005002001200`, and
+ confirm `lsusb -v` still reports the iso IN endpoint as **0x87** so a stale image cannot
+ masquerade as a pass.
+4. **Runs:** 5x the full 30-case battery, then 15x `--tests 9,10,14,21` (queued control, ch9
+ subset, both ctrl_out cases) — the control paths the verify actually protected, which a
+ plain battery samples only once per run. Print a `testusb` D-state scan after every
+ iteration.
+
+**Acceptance:** 5/5 batteries at 30/30, 15/15 loops, and no `testusb` D-state outliving its
+case runtime.
+
+**Rollback trigger:** any control-case failure (errno 110 or 71 on cases 9, 10, 14, 21) or a
+lingering D-state means the verify was load-bearing after all — restore it and record that
+result in the commit message. A negative result is a finding, not a setback.