summaryrefslogtreecommitdiff
path: root/docs/superpowers/followup
diff options
context:
space:
mode:
Diffstat (limited to 'docs/superpowers/followup')
-rw-r--r--docs/superpowers/followup/pr3803-flasher-recover.md280
-rw-r--r--docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md118
-rw-r--r--docs/superpowers/followup/pr3803-pci-rebind-stranding.md157
-rw-r--r--docs/superpowers/followup/pr3840-mret-board-result.md99
-rw-r--r--docs/superpowers/followup/pr3840-skill-md-no-boards-drift.md38
-rw-r--r--docs/superpowers/followup/pr3840-write-report-atomicity.md30
-rw-r--r--docs/superpowers/followup/pr3851-msc-host-tur-retry.md149
-rw-r--r--docs/superpowers/followup/pr3853-board-putchar-logger.md57
-rw-r--r--docs/superpowers/followup/pr3853-rtt-harness-adoption.md62
9 files changed, 990 insertions, 0 deletions
diff --git a/docs/superpowers/followup/pr3803-flasher-recover.md b/docs/superpowers/followup/pr3803-flasher-recover.md
new file mode 100644
index 000000000..1f71c990f
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-flasher-recover.md
@@ -0,0 +1,280 @@
+# `flasher_recover` Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Give the 15 HIL boards whose flasher cannot reach its probe past a poisoned usbfs
+node a second, convoy-safe flasher used only for recovery.
+
+**Architecture:** An optional roster key `flasher_recover` beside `flasher`.
+`hil_flash.recover_flasher(board)` picks it when present; `hil_test` substitutes it into the
+`--recover-board` JSON so `usbtest.py` never learns a second entry exists. Delivery over
+openocd's jlink driver is convoy-safe by construction, but the flash command form must
+differ from the one `flash_openocd` uses, so the recovery gets its own flasher name.
+
+**Tech Stack:** Python 3.13 stdlib, openocd 0.12.0+dev (build 0ce743125 on ci.lan),
+libjaylink, J-Link probes.
+
+## Global Constraints
+
+- Roster JSON: `test/hil/tinyusb.json`. `flasher_recover` is OPTIONAL; absent means today's
+ behaviour (`recover_flasher` returns the primary).
+- Never change the shape of `board['flasher']` — it is read as a dict in `hil_flash`,
+ `hil_test`, `usbtest`, `hil_pool_check`, `ci_select` and the roster lint, and is shipped
+ as JSON to a subprocess.
+- Flasher dispatch is by name: `getattr(hil_flash, f'flash_{name}')` / `reset_{name}`.
+- `RECOVER_FLASH_TIMEOUT = 90`, `RECOVER_RESET_TIMEOUT = 30` (`usbtest.py`). Any board whose
+ flash cannot finish inside 90 s is not a candidate.
+- Tests run offline: `cd test/hil && python3 test/test_ci_select.py`.
+
+## What is already established
+
+**Landed on PR #3803 and inert without roster entries:** `hil_flash.recover_flasher()`,
+`convoy_safe()` accepting openocd-over-jlink, `hil_test` substituting the recovery flasher
+into `--recover-board`, and `test_ci_select.FlasherRecoverEntry` (4 tests).
+
+**Verified in source:**
+- openocd's jlink driver ignores `adapter usb vid_pid` — `jlink.c` never reads
+ `adapter_usb_get_vids/pids`; selection is `adapter serial` / USB address / usb location.
+ Do NOT lint a jlink recovery entry for `vid_pid`.
+- It is convoy-safe anyway: libjaylink `discovery_usb.c` returns early unless
+ `idVendor == 0x1366` and the PID is in its table, and only THEN calls `libusb_open`. A
+ wedged `cafe:4010` DUT is never opened.
+- CMSIS-DAP stays pin-gated: `cmsis_dap_usb_bulk.c:107` skips before `libusb_open`, and
+ `id_filter` is only `vids[0] || pids[0]`.
+
+**Measured on ci.lan 2026-08-17**, base args
+`-f interface/jlink.cfg -c "transport select swd" -c "adapter speed 4000" -f target/<cfg>`:
+
+| Board | target cfg | flash | reset |
+|--------------------------|--------------|-------|-------|
+| stm32f407disco | stm32f4x | OK | OK |
+| stm32f072disco | stm32f0x | OK | OK |
+| stm32f723disco | stm32f7x | OK | OK |
+| stm32l476disco | stm32l4x | OK | OK |
+| feather_nrf52840_express | nrf52 | OK | OK |
+| metro_m4_express | atsame5x | OK | OK |
+| frdm_k64f | k60 | OK | OK |
+
+`frdm_k64f` is host-only (`tests.device == false`) — verify its reset over UART
+(`/dev/serial/by-id/usb-SEGGER_J-Link_000621000000-if00`), never by USB disconnect.
+
+**Excluded, with reasons:** `lpcxpresso11u37` — 118 s for 24 KB at 1 MHz with a verify
+mismatch, versus 0.277 s via JLinkExe; cannot fit `RECOVER_FLASH_TIMEOUT`.
+`mimxrt1064_evk`, `ra4m1_ek`, `nrf54lm20dk` — no target config exists in this openocd
+build, so they cannot be covered at all. **The board that wedges most (mimxrt1064_evk) is
+therefore still uncovered by this work.**
+
+**The blocker this plan solves:** `flash_openocd` issues `program <fw> verify reset exit`,
+which fails over the jlink transport on BOTH families tried (`stm32f4x`, `stm32f0x`) with
+`Examination failed` → `auto_probe failed`, with or without a preceding `init; reset halt`.
+Every successful flash above used the explicit sequence in Task 1.
+
+**Why this is a separate PR:** it adds a roster capability and a new flasher backend, which
+is a different scope from containing a wedge; and it needs bench time on seven boards.
+
+## File Structure
+
+- `test/hil/hil_flash.py` — add `flash_openocd_seq` / `reset_openocd_seq`; extend
+ `convoy_safe` to accept the new name. This is the only file that learns the command form.
+- `test/hil/tinyusb.json` — seven `flasher_recover` entries.
+- `test/hil/test/test_ci_select.py` — extend `FlasherRecoverEntry`; add a roster lint.
+
+---
+
+### Task 1: `openocd_seq` flasher backend
+
+**Files:**
+- Modify: `test/hil/hil_flash.py` (beside `flash_openocd`, ~line 100)
+- Test: `test/hil/test/test_ci_select.py`
+
+**Interfaces:**
+- Consumes: `_openocd_cmd_base(flasher)`, `hil_util.run_cmd`.
+- Produces: `flash_openocd_seq(board, firmware, timeout=None)`,
+ `reset_openocd_seq(board, timeout=None)`, both returning
+ `subprocess.CompletedProcess`; `convoy_safe()` returns True for
+ `{'name': 'openocd_seq', 'args': '...interface/jlink.cfg...'}`.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_openocd_seq_is_convoy_safe_over_jlink(self):
+ self.assertTrue(hil_flash.convoy_safe(
+ {'name': 'openocd_seq', 'args': '-f interface/jlink.cfg -f target/stm32f4x.cfg'}))
+
+ def test_openocd_seq_uses_explicit_flash_commands_not_program(self):
+ """`program` fails over the jlink transport: Examination failed -> auto_probe
+ failed, measured on stm32f4x and stm32f0x."""
+ seen = {}
+ real = hil_util.run_cmd
+ hil_util.run_cmd = lambda cmd, **k: seen.setdefault('cmd', cmd) or real('true')
+ try:
+ hil_flash.flash_openocd_seq(
+ {'flasher': {'name': 'openocd_seq', 'uid': 'X', 'args': '-f interface/jlink.cfg'}},
+ '/tmp/fw.elf', timeout=5)
+ finally:
+ hil_util.run_cmd = real
+ self.assertIn('flash write_image erase /tmp/fw.elf', seen['cmd'])
+ self.assertIn('verify_image /tmp/fw.elf', seen['cmd'])
+ self.assertNotIn('program ', seen['cmd'])
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_ci_select.py FlasherRecoverEntry -v`
+Expected: FAIL — `module 'hil_flash' has no attribute 'flash_openocd_seq'`
+
+- [ ] **Step 3: Write minimal implementation**
+
+```python
+def flash_openocd_seq(board, firmware, timeout=None):
+ # Explicit commands, NOT `program`: over the jlink transport `program` fails at the
+ # flash bank probe ("Examination failed" -> "auto_probe failed"), measured on
+ # stm32f4x and stm32f0x, with or without a preceding reset halt. This sequence
+ # succeeded on all seven candidate boards.
+ flasher = board['flasher']
+ verify = f' -c "verify_image {firmware}"' if flasher.get('verify', True) else ''
+ return hil_util.run_cmd(
+ f'{_openocd_cmd_base(flasher)} -c "init" -c "reset halt" '
+ f'-c "flash write_image erase {firmware}"{verify} -c "reset run" -c "shutdown"',
+ timeout=timeout)
+
+
+def reset_openocd_seq(board, timeout=None):
+ flasher = board['flasher']
+ return hil_util.run_cmd(
+ f'{_openocd_cmd_base(flasher)} -c "init" -c "reset run" -c "shutdown"',
+ timeout=timeout)
+```
+
+In `convoy_safe`, replace `if name != 'openocd':` with:
+
+```python
+ if name not in ('openocd', 'openocd_seq'):
+ return False
+```
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_ci_select.py FlasherRecoverEntry -v`
+Expected: PASS
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/hil_flash.py test/hil/test/test_ci_select.py
+git commit -m "hil: add openocd_seq flasher for convoy-safe recovery delivery"
+```
+
+---
+
+### Task 2: Roster entries for the seven validated boards
+
+**Files:**
+- Modify: `test/hil/tinyusb.json`
+- Test: `test/hil/test/test_ci_select.py`
+
+**Interfaces:**
+- Consumes: `flash_openocd_seq` / `reset_openocd_seq` from Task 1.
+- Produces: seven boards for which `hil_flash.convoy_safe(hil_flash.recover_flasher(b))`
+ is True.
+
+- [ ] **Step 1: Write the failing test**
+
+```python
+ def test_roster_recover_entries_are_convoy_safe_and_named_openocd_seq(self):
+ import json, pathlib
+ roster = json.loads((pathlib.Path(__file__).parent.parent / 'tinyusb.json').read_text())
+ recover = [b for b in roster['boards'] if 'flasher_recover' in b]
+ self.assertGreaterEqual(len(recover), 7)
+ for b in recover:
+ f = b['flasher_recover']
+ self.assertEqual(f['name'], 'openocd_seq', b['name'])
+ self.assertIn('interface/jlink.cfg', f['args'], b['name'])
+ self.assertIn('adapter speed', f['args'], b['name']) # required; see below
+ self.assertTrue(hil_flash.convoy_safe(f), b['name'])
+```
+
+- [ ] **Step 2: Run test to verify it fails**
+
+Run: `cd test/hil && python3 test/test_ci_select.py FlasherRecoverEntry -v`
+Expected: FAIL — `0 >= 7`
+
+- [ ] **Step 3: Add the entries**
+
+`adapter speed` is REQUIRED: without it examination fails outright on the jlink driver.
+Add to each board below, using the SAME `uid` as its primary jlink entry:
+
+```json
+"flasher_recover": {
+ "name": "openocd_seq",
+ "uid": "<same probe serial as flasher.uid>",
+ "args": "-f interface/jlink.cfg -c \"transport select swd\" -c \"adapter speed 4000\" -f target/<cfg>.cfg"
+}
+```
+
+| Board | `uid` | `<cfg>` |
+|--------------------------|----------------|-----------|
+| stm32f407disco | 000773661813 | stm32f4x |
+| stm32f072disco | 779541626 | stm32f0x |
+| stm32f723disco | 000776606156 | stm32f7x |
+| stm32l476disco | 777632258 | stm32l4x |
+| feather_nrf52840_express | 681295394 | nrf52 |
+| metro_m4_express | 123456 | atsame5x |
+| frdm_k64f | 000621000000 | k60 |
+
+- [ ] **Step 4: Run test to verify it passes**
+
+Run: `cd test/hil && python3 test/test_ci_select.py -v`
+Expected: PASS, and no other selector test regresses.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add test/hil/tinyusb.json test/hil/test/test_ci_select.py
+git commit -m "hil: give seven J-Link boards a convoy-safe recovery flasher"
+```
+
+---
+
+### Task 3: Bench validation on the rig
+
+**Files:** none — this task produces evidence, not code.
+
+- [ ] **Step 1: Confirm the rig is idle and take the locks**
+
+```bash
+ssh [email protected] 'if pgrep -f "[h]il_test.py" >/dev/null; then echo BUSY; exit 1; fi'
+ssh [email protected] 'cd ~/actions-runner/_work/tinyusb/tinyusb && \
+ nohup timeout 900 python3 test/hil/helper/hil_lock.py hold <boards...> --reason "flasher_recover validation" &'
+```
+
+Guard with `if`, never `cmd && echo || echo` — that form only gates the echo and will take
+locks during a live CI run.
+
+- [ ] **Step 2: For each board, flash then reset through the recovery entry**
+
+```bash
+python3 test/hil/hil_test.py -b <board> test/hil/tinyusb.json # normal path still works
+```
+
+Then force the recovery path by running usbtest with the recovery flags and a firmware that
+hangs a case, or drive `hil_flash.flash_openocd_seq` / `reset_openocd_seq` directly.
+
+- [ ] **Step 3: Verify**
+
+Device boards: `sudo dmesg` shows `USB disconnect` then a fresh enumeration.
+`frdm_k64f`: UART shows the boot banner (see above).
+Every flash must finish well inside `RECOVER_FLASH_TIMEOUT` (90 s).
+
+- [ ] **Step 4: Release locks and record the results in the PR body**
+
+---
+
+## Out of scope, and why
+
+- **`mimxrt1064_evk`** needs an i.MX RT target config that this openocd build does not
+ have. Sourcing or writing one is its own investigation; until then the board with the
+ most wedges has no automated recovery.
+- **Changing `flash_openocd`** to the explicit form would cover these boards without a new
+ name, but `program` is what nine pinned CMSIS-DAP boards use in CI daily and no CMSIS-DAP
+ image could be built in the originating worktree (no pico-sdk) to re-validate it.
diff --git a/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md b/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md
new file mode 100644
index 000000000..fe377f741
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-hil-iar-rerun-spec.md
@@ -0,0 +1,118 @@
+# IAR HIL Leg Re-run Spec Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Let the `hil-hfp-iar` CI leg re-run only its failed boards, as the other two HIL
+legs already do.
+
+**Architecture:** `hil_test.py` writes a `<config>.failed` spec into `HIL_REPORT_DIR`; a
+workflow step reads it on the next attempt and passes the boards back as arguments. The IAR
+leg passes `--retry 1` like the others but sets no `HIL_REPORT_DIR` and has no read-back
+step, so its spec is written into the workspace and never read.
+
+**Tech Stack:** GitHub Actions YAML, self-hosted runner.
+
+## Global Constraints
+
+- `.github/workflows/build.yml`. The two working legs are `hil-tinyusb` (matrix) — see its
+ `Set HIL report dir (per run+job; persists across run attempts)` and `Get re-run spec from
+ previous attempt` steps — and they are the pattern to copy.
+- The report dir must be keyed by run id AND job so a matrix leg does not collide with
+ another, and must survive across run attempts (that is the whole point).
+- The IAR leg is the only HIL job that BUILDS inline; its `Build` step is bounded at
+ `timeout-minutes: 30` under a 120-minute job ceiling. Do not disturb that.
+
+## What is already established
+
+- Verified by reading the workflow: `hil-hfp-iar` has neither `HIL_REPORT_DIR` nor a
+ `Get re-run spec` step, while passing `--retry 1`.
+- Consequence: a GitHub re-run of that job re-tests its whole matrix. **This is not a
+ regression** — that leg never had the mechanism — and the unread spec costs only a file.
+- The report artifact upload for that leg is named `hil-report-hfp-iar`.
+
+**Why this is a separate PR:** it is CI plumbing with no code change, it needs a real
+re-run on the self-hosted runner to prove, and it duplicates ~15 lines of workflow that
+would be better factored — a decision worth making on its own.
+
+## File Structure
+
+- `.github/workflows/build.yml` — the `hil-hfp-iar` job only.
+
+---
+
+### Task 1: Give the IAR leg a persistent report dir and a re-run spec
+
+**Files:**
+- Modify: `.github/workflows/build.yml` (job `hil-hfp-iar`)
+
+**Interfaces:**
+- Consumes: `hil_test.py`'s existing `--report-dir` / `.failed` behaviour — no code change.
+- Produces: `env.HIL_REPORT_DIR` for the job, and `$RERUN_ARGS` for the test step.
+
+- [ ] **Step 1: Copy the two steps from `hil-tinyusb`, before the Build step**
+
+```yaml
+ - name: Set HIL report dir (per run+job; persists across run attempts)
+ run: |
+ BASE=$HOME/hil-reports
+ echo "HIL_REPORT_DIR=$BASE/${GITHUB_RUN_ID}-hfp-iar" >> "$GITHUB_ENV"
+
+ - name: Get re-run spec from previous attempt
+ run: |
+ SPEC="$HIL_REPORT_DIR/hfp.json.failed"
+ if [ -f "$SPEC" ]; then
+ echo "RERUN_ARGS=$(cat "$SPEC")" >> "$GITHUB_ENV"
+ echo "re-running only: $(cat "$SPEC")"
+ fi
+```
+
+Match the exact spec filename `hil_test.py` writes for this leg's config — read
+`_write_failed_spec` and the `failed_fname` construction rather than assuming.
+
+- [ ] **Step 2: Pass the spec to the test step**
+
+```yaml
+ python3 test/hil/hil_test.py --retry 1 $SEL_ARGS hfp.json $RERUN_ARGS
+```
+
+`--retry 1` stays FIRST so argparse's last-wins keeps any explicit override working.
+
+- [ ] **Step 3: Point the artifact upload at the report dir**
+
+```yaml
+ path: ${{ env.HIL_REPORT_DIR }}/hil_report.md
+```
+
+- [ ] **Step 4: Validate the YAML**
+
+Run: `python3 -c "import yaml,sys; d=yaml.safe_load(open('.github/workflows/build.yml')); j=d['jobs']['hil-hfp-iar']; print(j['timeout-minutes'], [s.get('name') for s in j['steps']])"`
+Expected: the ceiling is still 120, the Build step still carries `timeout-minutes: 30`, and
+the two new steps appear before Build.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add .github/workflows/build.yml
+git commit -m "ci: let the IAR HIL leg re-run only its failed boards"
+```
+
+---
+
+### Task 2: Prove it on a real re-run
+
+**Files:** none — evidence only.
+
+- [ ] **Step 1:** Push and let `hil-hfp-iar` run to a failure (or force one).
+- [ ] **Step 2:** Confirm `$HIL_REPORT_DIR/hfp.json.failed` exists on the runner after the
+ job.
+- [ ] **Step 3:** Use GitHub's "Re-run failed jobs" and confirm the log line
+ `re-running only: ...` and that only those boards are tested.
+- [ ] **Step 4:** Record the run URL in the PR body.
+
+---
+
+## Consider first
+
+Three jobs would then carry the same ~15 lines. Factoring them into a composite action, or
+computing the report dir inside `hil_test.py` from `GITHUB_RUN_ID`, may be the better
+change — decide that before copying the block a third time.
diff --git a/docs/superpowers/followup/pr3803-pci-rebind-stranding.md b/docs/superpowers/followup/pr3803-pci-rebind-stranding.md
new file mode 100644
index 000000000..de1f7163b
--- /dev/null
+++ b/docs/superpowers/followup/pr3803-pci-rebind-stranding.md
@@ -0,0 +1,157 @@
+# `pci-rebind` Stranding Investigation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Settle when a PCI unbind/rebind of an xHCI controller strands it driverless, so
+the `usb-kernel-recover` skill can state a rule instead of a hypothesis.
+
+**Architecture:** No product code. This is a controlled reproduction against the rig's
+kernel, ending in a documentation change and — if the boundary turns out to be
+detectable — a guard in `usb_recover.sh`.
+
+**Tech Stack:** Linux 6.12.96 (ci.lan), Renesas uPD720201 xHCI, `usb_recover.sh`.
+
+## Global Constraints
+
+- ci.lan is a live CI rig. Take every affected board's lock first
+ (`hil_lock.py hold --all --reason ...`) and confirm no `hil_test.py` is running, with an
+ `if`, not an `&&` chain.
+- A stranded controller takes every fixture on it offline; recovery is
+ `usb_recover.sh pci-bind <addr>` or, failing that, a PVE **host** power cycle — an
+ operator action. Do not start this without being able to reach the host.
+- The rig has two Renesas controllers plus an AMD one; pick the controller with the fewest
+ fixtures for the experiment.
+
+## What is already established
+
+**The skill claimed, unconditionally, that `pci-rebind`'s re-bind hangs on the D-state URB
+and leaves the controller with no driver.** That claim was generalised from ONE observation
+and was used to delete `pci-rebind` and `pci-bind` from `usb_recover.sh` entirely.
+
+**It was refuted in the field on 2026-08-17.** After `hub-cycle 17-2.7` failed to clear a
+wedge, `pci-rebind 0000:05:00.0` recovered the controller in about one second:
+
+```
+02:34:41 remove, state 4 / USB bus 18 deregistered
+02:34:41 remove, state 1 / USB bus 17 deregistered
+02:34:42 xHCI Host Controller / new USB bus registered, assigned bus number 1
+02:34:42 new USB bus registered, assigned bus number 2
+```
+
+Both actions were restored, with the guidance scoped to failure mode: **dead controller →
+use it; device-lock convoy → do not**. Buses renumbered 17/18 → 1/2, which is why rig-wide
+operations need every board's lock.
+
+**What is NOT known:** why the earlier attempt stranded and this one did not. The leading
+hypothesis is that it turns on whether a live D-state URB exists **on that controller** at
+the moment of the re-bind — but in the 02:34 incident the wedged board (17-2.7) was on that
+very controller, which weakens it. An alternative is that `hub-cycle` had already cleared
+the holder, leaving only a dead controller.
+
+**Why this is a separate PR:** it is an experiment that risks taking the rig offline, and
+its output is a documentation change plus possibly a guard — a different scope from any
+code change.
+
+## File Structure
+
+- `.claude/skills/usb-kernel-recover/SKILL.md` — replace the hypothesis in section 3b and
+ the Common-mistakes entry with whatever the experiment establishes.
+- `.claude/skills/usb-kernel-recover/scripts/usb_recover.sh` — only if the boundary is
+ detectable from userspace.
+
+---
+
+### Task 1: Reproduce a controller-scoped D-state wedge
+
+**Files:** none.
+
+- [ ] **Step 1: Establish the safety net**
+
+```bash
+ssh [email protected] 'if pgrep -f "[h]il_test.py" >/dev/null; then echo BUSY; exit 1; fi'
+# hold ALL boards on the target controller
+```
+
+Confirm host access to pve.lan before continuing.
+
+- [ ] **Step 2: Create a wedge deliberately**
+
+Run `usbtest.py` against a board known to hang (`mimxrt1064_evk` has wedged eight times,
+TEST 9/10/24/27), or drive `testusb` directly until a case does not return.
+
+- [ ] **Step 3: Confirm the holder and its controller**
+
+```bash
+ps -eo pid,stat,etimes,wchan:22,args | awk '$2 ~ /D/'
+sudo cat /proc/<pid>/stack # usbdev_ioctl + [usbtest] = the owner
+readlink -f /sys/bus/usb/devices/usb<N> # bus -> PCI addr
+```
+
+Record whether the holder is on the SAME controller you will rebind.
+
+---
+
+### Task 2: Rebind and record the outcome
+
+**Files:** none.
+
+- [ ] **Step 1: Rebind, with a bounded observer**
+
+```bash
+timeout 120 sudo usb_recover.sh pci-rebind <addr>; echo "rc=$?"
+```
+
+- [ ] **Step 2: Record which of the three outcomes occurred**
+
+1. Re-bind completes, controller recovers (as on 2026-08-17).
+2. Re-bind hangs; `/sys/bus/pci/devices/<addr>/driver` is gone → **stranded**.
+3. Re-bind completes but the wedge persists.
+
+Capture `sudo journalctl -k --since ...` around the attempt either way.
+
+- [ ] **Step 3: If stranded, recover**
+
+```bash
+sudo usb_recover.sh pci-bind <addr>
+```
+
+If that hangs too, the only remaining step is a PVE host power cycle — an operator action.
+
+- [ ] **Step 4: Repeat at least three times**
+
+One observation is what produced the wrong rule in the first place. Vary whether a D-state
+holder is live on that controller at rebind time; that is the hypothesis under test.
+
+---
+
+### Task 3: Write down what was learned
+
+**Files:**
+- Modify: `.claude/skills/usb-kernel-recover/SKILL.md`
+
+- [ ] **Step 1: Replace section 3b's scoping with the measured rule**
+
+State the condition under which stranding occurs, with the journal lines. If the experiment
+does NOT reproduce stranding, say that too, with the attempt count — "not reproduced in N
+attempts" is a better record than an unexplained warning.
+
+- [ ] **Step 2: If the boundary is detectable, guard the script**
+
+For example, refuse `pci-rebind` when a D-state holder exists on that controller, since the
+holder is enumerable from `/proc` and the controller from `readlink`. Only add this if the
+experiment shows it predicts the outcome.
+
+- [ ] **Step 3: Commit**
+
+```bash
+git add .claude/skills/usb-kernel-recover/
+git commit -m "skills: replace the pci-rebind stranding hypothesis with measurement"
+```
+
+---
+
+## Abort criteria
+
+Stop and hand back to the operator if: a rebind strands the controller and `pci-bind` does
+not recover it; `uhubctl` starts hanging (the convoy has spread to the hub locks); or a CI
+run starts while the rig is in a broken state.
diff --git a/docs/superpowers/followup/pr3840-mret-board-result.md b/docs/superpowers/followup/pr3840-mret-board-result.md
new file mode 100644
index 000000000..77b76b605
--- /dev/null
+++ b/docs/superpowers/followup/pr3840-mret-board-result.md
@@ -0,0 +1,99 @@
+# Give the HIL worker result a name
+
+**Origin:** split out of PR #3840 (making `hil_report.md` a rendering of `hil_report.json`).
+Delete this file when its own PR lands.
+
+> **SUPERSEDED IN PART (2026-08-26).** Written against a 7-field tuple whose index 5 was
+> `blind`. The sysfs blindness subsystem is gone: `test_board` now returns **6** fields with
+> `stray` at index 5, and its board-locked early return is 5 wide. The problem described
+> below is unchanged and still worth fixing — three producers, three widths, and
+> `len(r) > 5 and r[5]` reads a WRONG SLOT rather than raising. But drop the `blind` field
+> from the proposed NamedTuple and re-derive every index from `hil_test.test_board` before
+> executing, or `_stray_note` starts reading a duration as a stray count.
+> `StrayNoteSurvivesTheTupleWidth` pins the current shape.
+
+## What is established
+
+`test_board()` returns a bare tuple that three producers build and fourteen call sites read
+positionally. It has grown 5 → 6 → 7 fields, and the code already works around its own
+shape:
+
+```python
+hil_test.py:1992 dirty = [(r[0], r[6]) for r in mret if len(r) > 6 and r[6]]
+hil_test.py:2014 blind = [r[0] for r in mret if len(r) > 5 and r[5]]
+hil_test.py:2386 for name, _, _, _, dur, *_ in mret:
+hil_report.py:306 for name, _, _, rows, *_ in mret:
+```
+
+Two facts make this worth closing rather than tolerating:
+
+- **The declared type is already wrong.** `hil_test.py:1711` says
+ `tuple[str, int, list[str], list, float]` — five fields — while the main return at `:1872`
+ yields seven (`+ sysfs_blind(), stray`).
+- **A wrong slot is a wrong verdict, not a crash.** Field 5 is `blind`, which decides whether
+ a board's red cells are reported as broken hardware or as "could not tell". Inserting a
+ field mid-tuple makes `r[5]` read the wrong slot and keep running.
+
+It has bitten once already: `test_hil_bounded.py`'s
+`test_both_row_widths_survive_the_report_writers` exists because the blindness flag widened
+the tuple to 6 while the pool-timeout path still synthesised 5-field rows, and *"a
+fixed-width unpack in either one raises INSIDE the containment path, which is where a raise
+costs every board's results."* That is why the unpacks end in `*_`.
+
+## What remains
+
+A `NamedTuple` with defaults. Verified to pickle across the pool boundary and to stay
+fully tuple-compatible — existing `r[0]`, `e[1]`, `for name, _, _, rows, *_` and `len(r)`
+all keep working, so it lands without touching the fourteen consumers:
+
+```python
+class BoardResult(NamedTuple):
+ """What one worker returns. Field ORDER is load-bearing: it is unpacked positionally
+ in a dozen places, and the pool-timeout path synthesises one by hand."""
+ name: str
+ err_count: int
+ failed_tests: list[str]
+ rows: list | None # None from the pool-timeout synthesis, never []
+ duration: float
+ blind: bool = False # defaults, so a synthesised result is full-width
+ stray: int = 0
+```
+
+Then a second, smaller step removes the coupling itself: `accumulate_report` takes
+`[(name, rows)]` pairs instead of `mret`, and `hil_test` does the extraction because it owns
+the shape. One line at each end; the subtle merge logic — stale lock clearing,
+`BOUNDARY_CELL`, `duration=None` preservation — is untouched.
+
+## Sizing
+
+| | Sites |
+|---|---|
+| Producers to convert | 4 (`hil_test.py:1724`, `:1872`, `:2283`, `:2327`) |
+| Arity guards deleted | 2 (`:1992`, `:2014`) |
+| Wrong annotation fixed | 1 (`:1711`) |
+| `hil_report`'s coupled line | 1 (`:306`) |
+| Positional consumers (optional migration) | 14 |
+| **Test fixtures building tuples by hand** | **34** |
+
+Production code is roughly ten changed lines. **The work is dominated by the test
+fixtures**, which is also the risk.
+
+## Do this first, or the refactor is unverifiable
+
+`test_hil_report.py` (27 sites) and `test_hil_bounded.py` (7) construct plain tuples by
+hand — `('boardA', 0, [], [], 1.0, True)`. A producer that forgot to switch to
+`BoardResult`, or a pickling regression, **passes the entire 310-test suite** and surfaces
+only on the rig. Convert the fixtures to build `BoardResult` as task 1, before touching any
+producer. This ordering is not optional.
+
+Second trap: `rows` is `None` on the pool-timeout path (`hil_test.py:2283`), never `[]`, and
+`accumulate_report` guards with `if rows and ...`. A well-meaning `rows: list = []` default
+silently changes that path. Pin it with a test before the conversion.
+
+## Why it was split out
+
+PR #3840 touches the report document. This touches `test_board`'s return and the containment
+paths, where a raise costs every board's results rather than one board's — a different blast
+radius, needing its own review and its own rig run. #3840 is twice-reviewed and dogfooded
+ten times on hardware; folding this in would reset that surface for a latent-trap cleanup
+that is not causing bugs today.
diff --git a/docs/superpowers/followup/pr3840-skill-md-no-boards-drift.md b/docs/superpowers/followup/pr3840-skill-md-no-boards-drift.md
new file mode 100644
index 000000000..a039a8c12
--- /dev/null
+++ b/docs/superpowers/followup/pr3840-skill-md-no-boards-drift.md
@@ -0,0 +1,38 @@
+# `SKILL.md` contradicts the code on no-boards tables
+
+**Origin:** split out of PR #3840, surfaced by its second review round. Delete this file
+when its own PR lands.
+
+`.claude/skills/hil/SKILL.md:150-151` tells the reading agent:
+
+> `**HIL run selected no boards.**` — the filters intersected to nothing, so there is **no
+> table at all**. Report that (and the filter shown), never `"pass": true`.
+
+That was true when the no-boards exit wrote a bare notice. It no longer is. An
+`--accumulate` no-boards run keeps the accumulated rows — deliberately, because wiping them
+destroyed real results — so the artifact now reads:
+
+```
+**HIL run selected no boards.** filters emptied
+
+**✅ 1 passed · ❌ 0 failed · ⚪ 0 skipped · blank not run**
+
+| Board | t | duration |
+...
+```
+
+The behaviour is correct; the documentation is wrong, and wrong in the direction that
+matters. An agent is told to expect no table, sees one, and has no rule for whether those
+rows are reportable. **They are not this run's** — they are a previous attempt's, carried
+forward.
+
+**What remains:** update that bullet to describe both cases — a fresh run has no table, an
+`--accumulate` run shows the previous attempt's rows under the notice and they must not be
+reported as this run's. Add a test asserting the fresh case renders no matrix, so the two
+halves cannot drift again.
+
+## Why it was split out
+
+PR #3840 fixed the findings that changed a verdict. This is a documentation drift: the
+behaviour is correct and the doc describing it is not, so it is better reviewed on its own
+than appended to a branch already carrying a module consolidation.
diff --git a/docs/superpowers/followup/pr3840-write-report-atomicity.md b/docs/superpowers/followup/pr3840-write-report-atomicity.md
new file mode 100644
index 000000000..2094207bb
--- /dev/null
+++ b/docs/superpowers/followup/pr3840-write-report-atomicity.md
@@ -0,0 +1,30 @@
+# `write_report` commits the two artifacts non-atomically
+
+**Origin:** split out of PR #3840, surfaced by its second review round. Delete this file
+when its own PR lands.
+
+```python
+md = render_report(doc) + '\n'
+report_dir.mkdir(parents=True, exist_ok=True)
+(report_dir / REPORT_JSON).write_text(json.dumps(doc, indent=2) + '\n')
+(report_dir / REPORT_MD).write_text(md, encoding='utf-8')
+```
+
+Rendering before writing closed the *render-failure* case: a raise can no longer commit a
+sidecar the markdown contradicts. It does not close the *interrupted-between-writes* case. A
+kill between those two lines leaves the pair disagreeing — and this runs on the containment
+path, on the way to `os._exit`, on a rig whose jobs get cancelled by the GitHub job ceiling.
+
+**What remains:** write both to temp files, then `os.replace` both. The window shrinks from
+two full writes to two renames, and neither file is ever observed half-written. `os.replace`
+is atomic per file on POSIX; the pair is still not transactional, which is acceptable and
+should be said in the docstring rather than implied away.
+
+Worth pairing with a test that kills between the writes — or, more practically, one that
+asserts no partial file is ever visible by checking the temp-then-rename shape directly.
+
+## Why it was split out
+
+A durability edge, not a wrong verdict. PR #3840 closed the render-failure half of this
+(nothing is written until the markdown renders); the interrupted-between-writes half needs
+a temp-then-rename and is better reviewed on its own.
diff --git a/docs/superpowers/followup/pr3851-msc-host-tur-retry.md b/docs/superpowers/followup/pr3851-msc-host-tur-retry.md
new file mode 100644
index 000000000..462b90fdf
--- /dev/null
+++ b/docs/superpowers/followup/pr3851-msc-host-tur-retry.md
@@ -0,0 +1,149 @@
+# MSC host: bound the Test Unit Ready retry loop and act on sense data
+
+> Split out of PR #3851 (`etmtrace-rp2350`, rp2350 ETM trace + stock clocks):
+> a host-stack MSC bug with no relation to that branch's scope.
+
+**Goal:** stop `msch_open`'s enumeration retry from spinning forever when a
+device answers Test Unit Ready with CHECK CONDITION, and use the sense data the
+driver already fetches to decide whether to keep waiting, give up, or report.
+
+---
+
+## What is already established
+
+### The loop is unbounded, and the source says so
+
+`src/class/msc/msc_host.c:445-472` is a two-function cycle with no counter:
+
+```c
+static bool config_test_unit_ready_complete(...) {
+ if (csw->status == 0) {
+ ... tuh_msc_read_capacity(...); // ready -> proceed to mount
+ } else {
+ // Note: During enumeration, some device fails Test Unit Ready and require a few retries
+ // with Request Sense to start working !!
+ // TODO limit number of retries <-- :459, pre-existing
+ TU_LOG_DRV("SCSI Request Sense\r\n");
+ TU_ASSERT(tuh_msc_request_sense(dev_addr, cbw->lun, enum_buf,
+ config_request_sense_complete, 0));
+ }
+ return true;
+}
+
+static bool config_request_sense_complete(...) {
+ TU_ASSERT(csw->status == 0);
+ TU_ASSERT(tuh_msc_test_unit_ready(dev_addr, cbw->lun,
+ config_test_unit_ready_complete, 0)); // :472
+ return true;
+}
+```
+
+Two defects, independent of each other:
+
+1. **No bound.** TUR fail -> Request Sense -> TUR -> ... forever. `tuh_msc_mount_cb()`
+ is never called and the application is never told anything; the device sits
+ enumerated-but-unmounted indefinitely.
+2. **Sense data is fetched and discarded.** `config_request_sense_complete`
+ checks only the CSW status. `enum_buf` holds a `scsi_sense_fixed_resp_t`
+ whose `sense_key` / ASC / ASCQ distinguish "Not Ready — becoming ready"
+ (retry is correct) from "Not Ready — medium not present" (a card reader with
+ no card; retrying can never succeed) from a hard error. The driver cannot
+ currently tell these apart because it never looks.
+
+### Measured on hardware (2026-08-25)
+
+Rig: `raspberry_pi_pico` (RP2040) + Pico-PIO-USB host on GP20/21, probe
+`E6614103E719612F`, console over the probe's CDC. Build:
+`-DCFG_TUH_RPI_PIO_USB=1 -DLOG=2`.
+
+- `examples/host/msc_file_explorer` never mounts. Debug log over ~25 s:
+ **1× `SCSI Test Unit Ready`, 350× `SCSI Request Sense`**, zero
+ `SCSI Read Capacity`, zero mount callbacks. `dd` reports
+ `no MSC device mounted`.
+- **The transfers themselves all succeed** — every CBW/CSW pair logs `OK`
+ (`Queue EP 02 with 31 bytes ... OK`, `Queue EP 81 with 13 bytes ... OK`), so
+ this is a SCSI-state-machine problem, not a bulk-transfer or PIO-USB timing
+ problem.
+- Reproduced with **two different drives** (`24a9:1802` "STORAGE DEVICE" and the
+ drive swapped in after it), so it is not one device's quirk.
+- Control transfers on the same target are fine: `examples/host/device_info`
+ reads full descriptors from the same drive on the same board
+ (`bcdUSB 0210`, `bMaxPacketSize0 64`, i.e. full-speed).
+- **The very same drive mounts and sustains I/O on RP2350**
+ (`pico2_etm_trace` carrier): `msc_file_explorer` + `dd` returns
+ `dd: 524288 bytes in 8448 ms = 62 KB/s`. Confirmed by the maintainer at the
+ bench, so the device is healthy and the "not ready" answer is provoked by
+ something specific to the RP2040 setup.
+- **Bumping Pico-PIO-USB does not fix it.** Retested with upstream HEAD
+ `5a37a66` (10 commits ahead of the pinned `675543b`, including
+ `512d3a2` "Place calc_usb_crc16 in RAM like calc_usb_crc5 and the CRC
+ tables", which looked like a promising RP2040 timing fix, and `cbf055d`
+ transaction-length clamp) via `-DPICO_PIO_USB_PATH=<clone>`: identical
+ failure, no mount.
+- Clock is **not** a factor: identical failure at 120 MHz, 133 MHz and
+ 156 MHz on RP2040 (and on RP2350 all of 120/125/126/138/150/156/162/174/186/240 MHz
+ behave identically).
+
+### What is NOT established
+
+- The actual sense key/ASC/ASCQ the failing drives return — the driver never
+ logs it. **Task 1 below exists to capture it**, and its answer decides whether
+ a bounded retry is sufficient or a "medium not present" path is also needed.
+- **Why the RP2040 setup provokes the not-ready state.** Leading suspect is
+ VBUS quality rather than firmware: the RP2350 carrier feeds J5 through a
+ proper load switch, while the RP2040 rig is a bare Pico whose GP22 "VBUS
+ enable" drives nothing (no load switch on a bare Pico), so the drive is fed
+ directly off the VBUS pin through hookup wire. A bus-powered drive that
+ cannot spin up answers exactly this "not ready" forever. Measure VBUS at the
+ device under load, or retest with a powered hub / self-powered device,
+ BEFORE attributing the stall to the host stack.
+- The actual sense key (Task 1) — still the gate for any policy change.
+
+---
+
+## What remains
+
+### Task 1: Log the sense response (diagnostic, ship-able on its own)
+
+**Files:** `src/class/msc/msc_host.c` (`config_request_sense_complete`, ~:467)
+
+Add a `TU_LOG_DRV` of `sense_key`, `add_sense_code`, `add_sense_qualifier` from
+the fixed-format response in `usbh_get_enum_buf()`. `scsi_sense_fixed_resp_t` is
+already declared in `src/class/msc/msc.h`.
+
+Verify on the rig above: rebuild `msc_file_explorer` with `-DLOG=2`, flash, read
+the probe CDC, and record the triple. Expected candidates:
+`0x02/0x04/0x01` (becoming ready) or `0x02/0x3A/0x00` (medium not present).
+
+### Task 2: Bound the retry
+
+**Files:** `src/class/msc/msc_host.c`, `msch_interface_t` (add a retry counter),
+`src/class/msc/msc_host.h` (a `CFG_TUH_MSC_TUR_RETRY_COUNT`-style knob with a
+sane default; follow the existing `CFG_TUH_MSC_*` naming in
+`src/tusb_option.h`).
+
+On exhaustion, stop the cycle and surface the failure rather than silently
+looping — the application currently has no way to learn the device is stuck.
+
+### Task 3: Decide behaviour per sense key
+
+Gated on Task 1's measurement. At minimum: keep retrying on "becoming ready",
+stop immediately on "medium not present". Do not invent policy for sense keys
+that were not observed.
+
+### Task 4: Regression coverage
+
+`test/unit-test/` has no MSC host suite today; adding one means mocking
+`tuh_msc_*` completions. Confirm with the maintainer whether a unit test or a
+HIL case on a known not-ready device (an empty card reader is the cheap
+reproducer) is the wanted evidence before building either.
+
+---
+
+## Why it was split out
+
+Found while sweeping PIO-USB clocks on the `etmtrace-rp2350` branch, which
+touches only rp2040/rp2350 clock pinning and ETM trace config. This bug is in
+the class-driver layer, affects every MCU running the MSC host, and predates
+that branch (the `// TODO limit number of retries` is already in master). It
+deserves its own PR and its own hardware evidence.
diff --git a/docs/superpowers/followup/pr3853-board-putchar-logger.md b/docs/superpowers/followup/pr3853-board-putchar-logger.md
new file mode 100644
index 000000000..46a4417bd
--- /dev/null
+++ b/docs/superpowers/followup/pr3853-board-putchar-logger.md
@@ -0,0 +1,57 @@
+# `board_putchar` is not LOGGER-aware
+
+**Origin:** surfaced while validating the RTT console in PR #3853 (the `rtt` skill
+promotion), which is harness-only scope. This is a src-level fix to `hw/bsp/board.c`
+that touches every board/logger combination, so it needs its own build sweep rather
+than a drive-by. Delete this file when its own PR lands.
+
+## Established (with evidence)
+
+`hw/bsp/board.c` retargets stdio through `sys_write`/`sys_read`, which are compiled
+per logger: `SEGGER_RTT_Write`/`SEGGER_RTT_Read` under `LOGGER_RTT`, ITM under
+`LOGGER_SWO`, `board_uart_write`/`board_uart_read` by default. The two board-level
+character helpers do not agree:
+
+```c
+168: int board_getchar(void) {
+169: char c;
+170: return (sys_read(0, &c, 1) > 0) ? (int) c : (-1);
+171: }
+172:
+173: int board_putchar(int c) {
+174: if (board_uart_write((const char *)&c, 1) > 0) {
+```
+
+`board_getchar` follows the logger; `board_putchar` always goes to the UART. So with
+`LOGGER=rtt` console input arrives over RTT while the echo goes out the UART.
+
+Measured on ea4088_quickstart (`LOGGER=rtt`, `board_uart_write` is a `-1` stub on
+lpc40): the `board_test` echo vanishes entirely while a `printf` echo — same console,
+same keystroke — comes back byte-for-byte. `LOGGER=swo` has the same asymmetry by
+construction (ITM out of `sys_write`, UART out of `board_putchar`), unverified on
+hardware.
+
+## What remains
+
+Candidate fix: route `board_putchar` through `sys_write(0, ...)` for symmetry with
+`board_getchar`. Two things to settle while doing it:
+
+- `board_putchar` currently passes `&c` of an `int` to a `const char*` — it writes
+ the low byte only on little-endian. Narrow to a `char` local as part of the change.
+- The default (UART) path must keep its current return contract: `board_uart_write`
+ returns negative when the UART is a stub, and the default `sys_write` breaks out of
+ its retry loop on that, returning a short count — so `board_putchar` still has to
+ map "wrote nothing" to `-1`.
+
+## Validation
+
+Build sweep across loggers and families — at minimum one UART board, one
+`LOGGER=rtt` board and one `LOGGER=swo` board — plus a hardware check that the
+`board_test` echo comes back on an RTT board (ea4088_quickstart reproduces the bug
+today) and that a plain UART board's echo is unchanged.
+
+## Why it was split out
+
+PR #3853 promotes a debug-tooling skill and touches `test/hil/*.py` and
+`tools/rtt.py`. A `hw/bsp/board.c` change lands in every example on every board and
+belongs in a review that carries the build evidence for it.
diff --git a/docs/superpowers/followup/pr3853-rtt-harness-adoption.md b/docs/superpowers/followup/pr3853-rtt-harness-adoption.md
new file mode 100644
index 000000000..8f3eae16b
--- /dev/null
+++ b/docs/superpowers/followup/pr3853-rtt-harness-adoption.md
@@ -0,0 +1,62 @@
+# Follow-up: finish RTT-console adoption in the HIL harness
+
+Split out of the `rtt` skill-promotion PR #3853. That PR deliberately ships the skill + CLI and leaves the harness's remaining
+VCOM assumptions in place — converting them is separate test-infra scope that
+deserves its own review and HIL runs. Scope here is `test/hil/*.py` only; the
+src-level `board_putchar` asymmetry this work surfaced has its own handoff
+(`pr3853-board-putchar-logger.md`).
+
+## Established (with evidence)
+
+- `hil_util.JlinkRtt` + `open_board_console()` work end-to-end:
+ ea4088_quickstart runs its host suite over RTT (16 passed / 0 failed / 3
+ skipped, the 'hil: read the host console over RTT when the probe has no VCOM' commit), and the `rtt` skill's boards.md carries the
+ validated matrix.
+- `test_host_device_info` honors `"logger": "rtt"` (hil_test.py, `test_host_device_info`; the eof fail-fast assert sits in its read loop):
+ in RTT mode it resets via the flasher BEFORE opening the console (which
+ then owns the probe; Commander delivers the buffered boot burst) and its
+ read loop fails fast on `JlinkRtt.eof` instead of blaming the board.
+
+## Remaining gaps
+
+1. **`test_host_cdc_msc_hid` and `test_host_msc_file_explorer` (hil_test.py) still call `hil_util.get_serial_dev(flasher["uid"], ...)`
+ directly** — on a `logger: rtt` board with `is_cdc`/`is_msc` fixtures they
+ would fail with the same "No serial device found" the console work fixed
+ for device_info (an interim load-time gate in `hil_test.py` now rejects
+ that combination up front; delete the gate when this lands). Fix: route
+ both through `open_board_console(board)` — but design the conversion
+ reset-aware rather than hand-copying device_info's dual branch: hoist a
+ `reset=` parameter into `open_board_console` that does the per-console
+ ordering itself (RTT: reset via flasher BEFORE opening — the console owns
+ the probe; VCOM: reset after open to catch the banner), and REMOVE the
+ existing post-open `# reset device to catch mount messages` blocks in both
+ tests (grep the marker — line numbers churn) — kept as-is on an RTT board they reset
+ while the console holds the probe. `JlinkRtt` carries input for their
+ menus and implements the `reset_input_buffer()` those tests call.
+2. **`hil_pool_check.check_host_serial` carries its own inline RTT branch**
+ (reset → `JlinkRtt` → poll through `hil_util.strip_banner`) — RTT boards
+ ARE health-checkable today, but the console-opening logic now lives in
+ two places (`open_board_console` in hil_test.py and this branch), each
+ with its own reset-ordering. Fix: hoist `open_board_console()` into
+ `hil_util.py` with the `reset=` parameter from item 1 and collapse
+ pool_check's branch onto it; keep the `do_reset` flush semantics for the
+ VCOM path intact.
+3. **OpenOCD console backend in the harness**: the skill's CLI
+ (`tools/rtt.py --backend openocd`, class
+ `OpenocdRtt` in the same module) is built, deduplicated behind a shared
+ base class next to `JlinkRtt` in `tools/rtt.py`, re-exported by
+ `hil_util`, and hardware-validated (all 20 rig boards through the CLI on
+ both backends, incl. the 8 native-probe ones). What remains is only the
+ `open_board_console` plumbing: choosing `OpenocdRtt` for a
+ `"logger": "rtt"` board with an openocd/stlink flasher needs the per-test
+ flashed-ELF path (for the control-block address) and, for stlink
+ flashers, an openocd target-cfg mapping the roster doesn't carry — until
+ then the config-load gate keeps rejecting non-jlink rtt boards.
+
+## Validation for this follow-up
+
+Run the ea4088 local host suite (a board with a `is_cdc`+`is_msc` capable
+device attached to J3, or the rig's frdm_k64f/mimxrt1064 with a temporary
+`logger: rtt` entry) so cdc_msc_hid and msc_file_explorer actually execute
+over RTT; then a `hil_pool_check.py` pass on a no-VCOM board. Delete this doc
+when the follow-up PR lands.