| Age | Commit message (Collapse) | Author |
|
frdm_k64f as a USB host with a CH9102 CDC (TX-RX loopback) and a Lexar MSC
drive behind a hub; flasher = onboard OpenSDA J-Link. host/cdc_msc_hid passes
(CDC mount+echo, MSC mount + disk-size check). device_info remains a known
device_info/usbh limitation (its synchronous descriptor dump starves a 2nd
device's enumeration) and is not ci_fs-specific.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ExGPLP5eU43LR7o6yYLpNi
|
|
Full-fleet profiling (HIL_PROFILE=1 instrumentation, included) showed each
uPD720201 controller's serialized usbtest battery chain dominates wall time,
and a board whose marginal device port bounces during concurrent batteries
can wedge or kill the controller ("xHCI host not responding to stop endpoint
command"). Every such death traced to mimxrt1015's port (its old "kills the
uPD720201" reputation) - it is removed from the config until recabled;
mimxrt1064's enum-retry stalls were a loose device cable (re-seated).
nrf54lm20dk moves to boards-skip until its failing J-Link probe is replugged.
With the hardware fixed both cards run width-4 batteries plus full flash
churn clean, so scheduling stays simple: two symmetric knobs, flashes and
batteries budgeted per controller.
- schedule_boards(): dispatch boards round-robin across host controllers from
a persisted hint cache (~/.cache/tinyusb-hil/ctrl_cache.json), learned and
merge-on-write refreshed each run (concurrent HIL jobs keep each other's
entries). Only the cached PCI address is consumed - dispatch order and
first-flash budgeting, never battery serialization (batteries resolve live
or fail closed to an all-slot permit).
- HIL_FLASH_PARALLEL (8) and HIL_USBTEST_PARALLEL (4) are budgeted per
controller via lock slots assigned on first sight.
- re-runs: a failed run writes <report dir>/<config>.failed with the exact
re-run spec (--accumulate -b <failed board> -bt <board>:<its failed
tests>) instead of the inverted --skip-board list of everything that
passed; --skip-board is gone, --flasher/--exclude-flasher scope a config
across CI jobs by flasher type (no board names hardcoded in workflows),
and -a/--accumulate merges a re-run into the existing report. The spec is
stamped with GITHUB_RUN_ID and cleared on fresh runs, so a retry can never
consume a spec left behind by a different run's dead or skipped attempt.
- CI: esp-idf firmware builds move out of hil-build into hil-build-esp, and
the esptool-flashed boards run in their own hil-tinyusb-esp job, so the
main hil-tinyusb run starts as soon as the fast toolchains finish instead
of waiting on the slow esp-idf build (an esp toolchain flake previously
skipped the whole rig run). Artifacts are namespaced per toolchain so the
esp job downloads only esp-idf binaries.
- HIL_PROFILE=1: timestamped log lines, per-flash durations, permit-wait
logging, uid->controller map dump for analysis.
- hil_report: per-variant test duration as a dedicated trailing column,
recorded only by full runs.
Validated on the ci rig (fixed seeds 20260716/777, full fleet at 8/4):
738s/780s walls with only known-flake failures and no controller deaths,
vs 1134-1211s serialized-battery baseline.
|
|
and mimxrt1015
edpt_schedule_packets() advanced xfer->buffer past each armed EP0
chunk, so the OUT-complete handler invalidated the cache at the
ADVANCED pointer: one line past the received data. The CPU then read
stale cached bytes instead of the DMA'd packet, and the misplaced
invalidate discarded a dirty line of whatever variable follows the
buffer - random neighbor corruption on every control-OUT data stage.
Found by usbtest ctrl_out (cases 14/21) on espressif_p4_function_ev
with DMA enabled, the first DWC2 target combining buffer DMA with a
data cache: usbd control state wedged after the first control write
(every later request stalled), and one build layout panicked in the
usbd memcpy with a wild pointer.
Rework the EP0 chunk bookkeeping so xfer->buffer always points at the
un-consumed position: the arm no longer advances it; instead the EP0
re-arm paths advance past each completed (full) chunk, invalidating it
first on the OUT side. The final OUT completion invalidates exactly
the received bytes of its last chunk, taken from DOEPDMA ("incremented
on every AHB transaction", databook 7.1.83 - the same semantics the
SETUP path relies on) before dma_setup_prepare() re-targets it. EP0
chunking state (ep0_pending) is now also dropped on bus reset and on
a new SETUP, so a stale latched completion can no longer re-arm EP0
DMA from dead state. No behavior change for targets without dcache.
While root-causing, the FIFO layout was cross-checked against the
DWC2 databook/programming guide v4.20a: the existing GDFIFOCFG
programming (EPInfoBaseAddr = otg_dfifo_depth - 2*ep_count, one SPRAM
word per endpoint direction for buffer DMA) is conformant and needs
no change; the P4 HS instance's reset GDFIFOCFG (0x03800400) merely
reflects a scatter/gather-sized EP_LOC_CNT of 128 that buffer DMA
does not need.
With the fix in place, enable the usbtest battery on the espressif
fleet: tools/build.py allowlists device/usbtest (a plain IDF component
like board_test/video_capture) and both espressif boards' only-lists
gain device/usbtest. Also re-enable device/usbtest on mimxrt1015_evk:
its skip predated the dcd_ci_hs stale-ACTIVE-overlay fix (already on
this branch), which cured the battery that previously killed the
uPD720201 host controller twice (2026-07-11 ROM fw, 2026-07-13 case 27
on fw 2.0.2.6); rig-validated 30/30 three consecutive runs.
Validated on rig (all 30/30): espressif_p4_function_ev(-DMA) (was
22/30 under DMA), espressif_s3_devkitm(-DMA), stm32f723disco(-DMA),
mimxrt1015_evk; p4/s3 slave-mode unaffected (DMA-only code path);
compile-checked stm32h743nucleo +TUD DMA, stm32f407disco,
stm32l476disco (device ports currently on the dead hub).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_017TQZrFfU3K4Y198aLsUpBC
|
|
Pool/config:
- record real uids (ra8m1_ek), enable usbtest for espressif s3/p4, then
park ra6m5_ek and ra8m1_ek in boards-skip (ra6m5's usbtest/MSC traffic
can kill the uPD720201 host on its ROM firmware; ra8m1 USBHS bring-up
pending); max32666/nrf54lm20 stay enabled - their MosChip flakiness
never wedges
- re-enable device/usbtest on HS boards (mimxrt1064, ch32v307) now that
uPD720201 firmware 2.0.2.6 fixes the command-ring death; mimxrt1015
stays skipped - its HS battery killed the controller on both ROM and
2.0.2.6 firmware (board-specific); match the moved host-test bundles
(f723 <-> rt1064); skip never-passing tests on the new
nrf5340dk/nrf54lm20dk boards and the detached pico host bundle, each
documented with a comment
Host-controller quirk gating in usbtest.py (auto-skip, self-heals on a
healthy xHCI):
- MosChip MCS9990 EHCI: case 25 (int-OUT never scheduled, FRINDEX bug)
and case 11 (unlinked reads complete short/EREMOTEIO)
- Renesas uPD720201 xHCI: firmware-gated. The card must run firmware
>= 2.0.2.6 (RAM-uploaded - it reverts to ROM on every power cycle):
on older firmware the command ring dies under unlink stress (a
Configure Endpoint command stops completing; the hub worker deadlocks
holding the device lock; only a host power cycle recovers; three
boards reproduced it). usbtest.py reads the FW version register (PCI
config 0x6c) and refuses to run at all on older firmware - hil_test
surfaces that as a failed test with the reason. On current firmware
the full 30-case battery runs (validated FS+HS: metro_m4, f723,
f723-DMA all 30/30).
Scheduling (hil_test.py):
- Shuffle each (board, variant)'s test order with a seeded RNG
(HIL_SHUFFLE_SEED to replay) so usbtest batteries and flash churn spread
across the timeline instead of convoying on one controller.
- Per-controller usbtest + flash semaphores: HIL_USBTEST_PARALLEL
(default 4) concurrent usbtest batteries and HIL_FLASH_PARALLEL
(default 8) concurrent flashes per host controller. Profiled on
uPD720201 firmware 2.0.2.6 across 8/1..12/8: wall time falls
22.2/14.3/12.5/10.8 min at usbtest width 1/2/3/4 and plateaus there;
zero controller errors everywhere; first battery case failures
(leaf-hub bandwidth stretch) appear at 12/8, and flash width 12 only
amplifies flasher-hub contention flakes - so 8/4 is the optimum. A
separate battery-window flash throttle was profiled and dropped.
- Give every example a unique hardcoded USB PID (0x4001-0x4022, usbtest
keeps 0x4010) instead of the PID_MAP interface bitmap: different
examples now always re-enumerate back-to-back, even on boards whose
CPU reset does not drop D+ (WCH CH58x), so the EXAMPLE_PID table and
same-PID adjacency reordering in hil_test.py are gone; only the
variant-boundary same-example repeat needs a swap.
- Report matrix: stable columns with the metric-bearing tests pinned
first (usbtest, cdc_msc_throughput, msc_file_explorer[_freertos]),
the rest alphabetical.
Fail fast:
- enum wait budget 8 s on the first attempt, 4 s on retries; dfu waits
are deadline-based so dfu-util's own runtime counts against the
budget.
A device-absent failure now costs ~3-5x a passing test (20-30 s)
instead of 10-30x (47-150 s).
- CI runs hil_test with --retry 1 and no in-run second pass: a broken
fixture fails the job fast instead of holding the self-hosted runner
for hours and blocking other PRs' HIL jobs. hil_test still writes the
.skip sidecar, so a manual re-run attempt only retests what failed.
Review fixes (multi-agent adversarial review of this commit):
- tinyusb_win_usbser.inf: the PID rework moved five CDC examples onto
even PIDs the INF's odd-only DeviceList never matched (legacy-Windows
usbser binding) - appended 0x4006/4008/400a/4020/4022 to both lists.
- usbtest example: USBTEST_TIER is now overridable and the descriptors
and pumps are tier-conditional, so a board whose DCD cannot serve a
tier lowers it instead of skipping the whole example - RA2A1 (RUSB2
with no isochronous pipe) builds at tier 3 via its BOARD_ define; the
host battery follows the tier advertised in bcdDevice. Tier-4 output
verified byte-identical after the refactor.
- dynamic_configuration's second config derived USB_PID + 11 = 0x4018,
colliding with net_lwip_webserver - now USB_PID + 0x0100, outside the
per-example space. tools/check_example_pids.py (pre-commit hook)
enforces PID uniqueness incl. derived and literal idProduct values.
- usbtest.py firmware gate: matched by device ID (uPD720201/720202,
both use the 0x6c FW register), and an unreadable version (setpci
missing/denied) now refuses with its own message instead of
masquerading as "firmware 0x00000000"; noted the gate is necessary
but not sufficient (board-specific kills stay per-board skips).
- hil_test: deadline waits use time.monotonic(); multiprocessing
context pinned to fork (raw semaphores in Pool initargs); flash and
usbtest permits unified into one fail-closed, exception-safe
ctrl_permit (unknown controller takes every slot and logs a warning
instead of silently borrowing slot 0); an all-skipped battery
reports as skip, not "0/0" failure; slow-body polls (mtp, printer,
disk read) go through a shared deadline-based wait_until so their
bodies count against the enum budget; throughput's FS detection
compares serials case-insensitively like every other walk; a missing
MSC read-speed line now fails the host msc_file_explorer test
instead of passing with an empty metric.
Hardening:
- fail fast (15 s) when a driver-registry sysfs write blocks: a wedged
device otherwise turns every subsequent battery into an unkillable
D-state writer and silently hangs the whole run
- usb-recover skill: a VM reboot is not a reliable cure (MosChip hubs
latch up across the PCIe reset); full host power cycle is
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01HeF2gZ1M7GWkz6Av4BpKPg
|
|
Binds the kernel usbtest driver (gadget-zero profile), runs the tier-based
battery, auto-recovers kernel-side hangs, and skips cases the host
controller cannot run (MosChip MCS9990 EHCI int-OUT).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01HeF2gZ1M7GWkz6Av4BpKPg
|
|
The board was added to the CI HIL pool with device/cdc_msc_throughput
skipped, so the test had never run on it. Remove the skip to include it
in the device test set.
Verified passing on the ci rig (USBFS, full-speed):
CDC read 639 / write 544 kBps, MSC read 845 / write 575 kBps.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
|
|
|
Fold in the CH58x BSP review fixes:
- family.mk: drop stray trailing backslashes on the last LDFLAGS/SRC_C entries
(harmless -- GNU Make ends the list at the blank line -- but misleading).
- debug_uart.c: uart_write() spun on a full ring buffer with nothing to drain it
(only uart_sync() advances tx_consume), so a burst larger than the buffer
deadlocked. Drain the FIFO while waiting, like uart_sync() does.
- wch-riscv.cfg: move the OpenOCD work area from 0x80000000 (unmapped) to the
0x20000000 SRAM, sized to 32 KB, matching ch32v20x/wch-riscv.cfg.
- family.c: implement board_get_unique_id() from the factory MAC. CH58x is a BLE
part, so a unique 6-byte MAC lives in FlashROM at ROM_CFG_MAC_ADDR; GetMACAddress()
reads it via FLASH_EEPROM_CMD (in libISP583.a), so no extra source file is needed.
The read buffer is TU_ATTR_ALIGNED(4) and 8 bytes, per the SDK's documented
4-byte-aligned, word-granular buffer contract (CH58x_flash.c).
- test/hil/tinyusb.json: key ch582m_evt off this board's actual MAC (D443627B5450)
instead of the fixed placeholder, like every other board.
Verified on ci.lan HIL: ch582m_evt enumerates with serial D443627B5450 and all
device examples pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
|
Add the device-only CH582M-EVT (WCH USBFS via the shared dcd_ch32_usbfs.c),
riscv-gcc, flashed by openocd_wch probe 7FD88F0604B5, to tinyusb.json.
Also reorder device_tests to keep examples sharing a VID:PID non-adjacent:
cdc_msc and cdc_msc_throughput both use cafe:4003, and on boards whose
CPU-reset does not drop D+ (e.g. WCH CH58x via openocd) back-to-back same-PID
firmware leaves the host on the previous example's cached descriptors, so the
new example's CDC never enumerates and the test fails. Moving dfu (cafe:4000)
between them changes the PID and forces the host to re-enumerate.
Remote HIL on ci.lan: all device examples pass, including cdc_msc_throughput
(no skip needed).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
|
Enabling the audio test fleet-wide surfaced failures on esp32-p4/s3 and
metro_m4_express: the UAC mic enumerates but arecord fails the iso IN read
with EIO, while 18 other boards pass strict=1.000.
esp32: root cause is the FreeRTOS tick rate. ESP-IDF defaults
CONFIG_FREERTOS_HZ to 100, so the audio task wakes only every 10 ms and
can't service the 1 ms UAC iso frames -> underrun -> arecord EIO. (The same
dwc2 driver passes on STM32, whose FreeRTOSConfig is 1000 Hz.) Set
CONFIG_FREERTOS_HZ=1000 in the example sdkconfig.defaults; the example
defaults are honored in the generated sdkconfig alongside the BSP's, so this
takes effect.
metro_m4_express (samd51): not tick-rate -- its FreeRTOSConfig is already
1000 Hz like the passing boards -- so it's a separate iso-IN issue, skipped
for now.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
|
Now that CH32V103 USB device works, add the board to the active HIL pool.
It is a WCH RISC-V USBFS part, so it builds under the riscv-gcc bucket;
single config (USBFS only, no fsdev variant).
cdc_msc_throughput is skipped for this board: its device->host CDC bulk-IN
read hard-fails here (a known, pre-existing dcd_ch32_usbfs throughput
limitation, not specific to CH32V103). All other device tests pass on
ci.lan (verified green, 0 failures).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
|
* hil: enable nanoch32v203 in CI with fsdev + usbfs variants
nanoch32v203 was parked in boards-skip; move it into the active pool now
that the board is wired to the ci.lan rig. Cover both USB device IPs as
build variants:
- nanoch32v203-fsdev: RHPORT_DEVICE=0 (USBD / stm32 FSDev IP)
- nanoch32v203-usbfs: RHPORT_DEVICE=1 (WCH USBFS IP)
|
|
* test/hil: replace build.flags_on with named variant schema
Boards declare build variants as `variant: [{name, flags}]` instead of
`build.flags_on`. The variant `name` is the build dir (cmake-build-<name>) and
the HIL report row; `flags` is the raw CFLAGS string (-D...=1) injected via
CFLAGS_CLI. No `variant` => a single build named after the board.
- build.py: --build-name <name> (dir) + --cflag=<token> (raw CFLAGS, repeatable,
=form survives the matrix's shell word-splitting); drop -f1/CFLAGS wrapping.
- hil_ci_set_matrix.py: emit one build arg per variant.
- hil_test.py: iterate variants; report row + build dir = variant name.
- hil_ci.sh: copy all cmake-build-<board>* dirs for -b runs.
- get_deps.py: accept (ignore) --build-name/--cflag from matrix args.
- tinyusb.json: migrate all 6 flags_on boards to variant.
* board_test: park CI build with busy spin instead of wfe
|
|
Add stm32u083nucleo to the active boards in tinyusb.json, flashed via the
stlink flasher (onboard ST-Link + STM32CubeProgrammer); ci's openocd build
has no STM32U0 flash driver. STM32_Programmer_CLI lives in ~/bin on ci,
which the remote `bash -s` shell in hil_ci.sh did not have on PATH, so add
$HOME/bin to its PATH export (matching the GHA runner .path). Verified
remote: 13/13 device tests pass on ci.lan.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
|
* add ek_tm4c123gxl to the hil pool, flashing with lm4flash
|
|
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
|
|
|
|
|
Signed-off-by: HiFiPhile <[email protected]>
|
|
|
|
hil remove hub from pico2 since it is not stable
|
|
|
|
|
|
|
|
across all families
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
`FSDEV_BUS_32BIT` with `CFG_TUSB_FSDEV_32BIT`, adjust data/address stride macros, and refactor register/PMU access for consistency across platforms.
|
|
|
|
|
|
|
|
|
|
|
|
rusb use custom write
enable all hil test for ra4m1
|
|
add build.args option for hil json
add MAX3421_HOST=1 for metro m4 express
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|