PPU testbench status (issue #390)
Status snapshot; details in waves.md, Readme.md
and the wiki pages wiki/soc/ppu1.md / wiki/soc/ppu2.md.
Tests
| Test | Status | Covers |
|---|---|---|
tb_ppu_regs |
✅ ALL PASS | register write + CPU read-back (LCDC/SCY/SCX/BGP/LY), reset release, counters, mode rhythm |
tb_ppu_bg_scanline |
✅ ALL PASS | BG fetch pipeline, pixel stream on LD0/LD1, 456-tick line |
tb_ppu_scroll |
✅ ALL PASS | PPU2 V+SCY / H+SCX scroll adders on the VRAM bus |
tb_ppu_window |
✅ ALL PASS | WIN layer, LCDC.6 window tile map $9C00 |
tb_ppu_bg_win_matrix |
✅ ALL PASS | BG/WIN combination matrix (C1..C11): map selects bit3/bit6 x enables x WY/WX incl. mid-line switch |
tb_ppu_scene |
✅ ALL PASS | synthetic LCD+BG+WIN+OBJ scene (mode rhythm + BG pixel stream) |
tb_ppu_frame |
✅ ALL PASS (slow, ~5 min) | VBlank at LY>=144, V wrap 153->0, ppu_int_vbl, LYC==LY interrupt |
tb_ppu_win_mid |
✅ (aux run for waves) | mid-line BG->WIN switch wave source |
tb_ppu_sprites |
🟡 dev | mode-2 OAM scan runs; sprite store/compare not claimed (see blockers) |
tb_ppu_oam_cpu |
🟡 dev | CPU OAM write lands word-addressed; strobe timing needs SoC model |
tb_ppu_oam_read |
🟡 dev | CPU OAM read return path needs SoC cycle timing |
tb_ppu_dma |
🟡 dev | VRAM->OAM DMA needs the MMIO DMA controller/arbiter model |
tb_ppu_ring_init0 |
🟡 dev (round 22) | no-reset FF init-0 probe: boot-state report + per-line obj_prio_ck / OAM-clock / Y-test counters |
Run everything: run_all.sh (or run_all.bat on Windows); tb_ppu_frame is
slow and included.
Blockers (all documented in wiki/soc/ppu1.md & ppu2.md)
-
Sprite store/compare/pixel path - the per-slot in-use dffr (
PPU2 g611-g628) never clock (obj_prio_ckfrom PPU1 never pulses), so no slot is claimed,sp_bp_cysnever fires, LD stays BG-only. -
The
oamux groups (scanw518vs port-B adderw475) overlap in time (contentionx); dynamic two-phase bus timing is not reproducible statically. -
The mode-2 store window (PPU2
w852 = oam_rd_ck & w209 & w816withw816= AND of the Y-test adder results incl. FF40_D2) never opens with undefined OAM data, so the banks are never enabled (they are NOT held in reset by the flags - corrected round 21). Everything funnels into the Y-test/OAM data path.
Suspected netlist issues - reported to the author, NOT patched here
-
PPU2 g938-g943(ppu2.v:2089-2094): scan-address register, no async reset (nr1 = w149 = const1). -
PPU1 g325/g326(ppu1.v:1510-1511):w530window dividers, no async reset (nr1 = w47 = const1). -
PPU1 g652(ppu1.v:1837): duplicate-looking LCDC.D7 latch (qw149read only byg305). -
PPU2 g419/g421(ppu2.v:1570/1572):oamux groups can be enabled at the same time (scan vs port-B adder) - possible phase/enable issue. -
STAT read-back (round 15/16): at LY==LYC the LYC interrupt fires but
$FF41reads0xC2/0xC3(bit2=0, bit7=1); netlist chaing859/g906/ g280/w546/w736documented in wiki/soc/ppu1.md.
No-reset FF "init-0" probe (round 22) - not the blocker
Per the checklist, the no-reset triggers were pinned to 0 instead of x
in the test (tb_ppu_ring_init0, dev). Result:
-
Power-on: dmglib dffr-family cells declare
initial val = 1'b0, so at boot all no-reset FFs read 0/1 (neverx); the PPU1 sprite ringg286/g287-g289/g325/g326is deterministic and resets (w816dips) at everyh_restart. Only FFs clocked from undefined data keepx: PPU1g882-g889and PPU2 scan-address bitsg942/g943(mode-2 scan). -
Forced-0 run (clones
dmg_*_nr0intemp/nr0:xclocked into a no-reset FF stores 0): scan register fully defined (010000) each line,g882-g889= 0. Result unchanged -obj_prio_ckstill 0 edges/line, PPU1(w228&w229&w241)never 1, PPU2 Y-test AND6 (w816) never high, store windoww8520 edges.
Conclusion: the no-reset-FF x is NOT the cause of the silent
obj_prio_ck (the ring is reset each line but never opens its pulse
window; the PPU2 Y-test never passes with the current oa/port-B phase
timing). The no-reset FFs remain a real silicon concern (undetermined
power-up, no garbage recovery) and stay reported to the author.
Weak / discharge-only oa bus (round 23) - also not the blocker
Per the checklist the oa bus was made "weak": the six oa-chain
inverse-hold nodes (w497/w146/w500/w554/w641/w49, each driven by four
notif0 mux groups) were re-modelled as precharged nodes - pullup keepers +
open-drain drivers (dmg_notif0_od; variant netlists + probes in
temp/weak/, gitignored). Result: the bus and the scan address become fully
deterministic (baseline Y-test group is x ~85% of mode 2); the scan
schedule length is unchanged (39 steps) but the row sequence shifts
(baseline words {2,4,...,78}, weak {7,15,...,79}) - both start above word 0.
Even with a visible sprite in every OAM entry, lines LY 1..15: Y-test
AND6 w816 never high, port B reads zero in all sampled phases, store
window w852 never opens, obj_prio_ck 0 edges/line - in both variants.
Conclusion: oa contention is not the blocker either; the port-B read data
never reaches the Y-test as a passing compare (read phase not sampled
statically, or Y-test term polarity/bounds inverted - see the w852/w817
polarity open question in wiki/soc/ppu2.md). Needs schematic-level two-phase
timing / OAM macro byte mapping (author or msinger ground truth).
OAM A/B inverse-hold model fix + weak buses (round 26)
n_oama/n_oamb are inverse-hold buses (idle = precharge HIGH = data 0);
PPU2's scan capture stores the pad level directly (dmg_latch g733-g748).
Committed oam_ram.v reworked to precharge keepers + discharge-only pads
(pad low for stored 1, hi-Z during writes): x on the ports drops from
320/6000 samples (old strong-~data drive) to 0/6000. All 6 fast tests and
tb_ppu_frame still ALL PASS. Test-only variants in temp/oamweak/ (PPU2
OAM-port drivers open-drain, + weak oa, + swapped port mapping) - in every
combination with a visible sprite in all 40 OAM entries, lines LY 1..15:
Y-test AND6 never high, store window w852 never opens, obj_prio_ck
0 edges/line. Read dumps show defined pad levels while the scan walks words
{5,7,15,...} - the Y bytes (even words under the model layout) are never
presented; the scan word stream / byte<->word<->port mapping remains the
open item.
Weak oa accepted as the default bus model (round 27)
Per the round-23 experiment the oa-chain nodes are now simulated as
discharge-only + keepers by default: gen_weakbus.py produces
ppu2_weakbus.v (24 oa-chain notif0 -> dmg_notif0_od in
bus_weak_cells.v, pullups on w497/w146/w500/w554/w641/w49); every PPU
test compile uses it (bus model only - ppu2.v untouched). oa x:
mode 2 0/1280, idle 0/5168 (residual only in mode 3); scan addresses
defined, words {7,15,...,79}/mode 2. Regression suite 7/7 ALL PASS
(incl. tb_ppu_frame); sprite wave regenerated. Sprite claim still
inert (see blockers above).
Y-test root cause (round 28) - comparator is fine, addressing is not
Decisive probe (temp/oamweak/tb_ytest_all10*): with every OAM byte = 0x10
the Y-test B operand (latched port-B level, dmg_latchnq_comp g204-g250,
en w120) = 0x10, AND6 w816 = 1 and the store window w852 opens 39x
per line on LY 1..7 and correctly fails on LY=8 - the comparator, the
AND6 polarity (pass = 1) and the store logic work. Real content failed
only because port B never presented a Y byte during the mode-2 scan:
default netlist -> oa low bits x (g419/g421 enable overlap) -> no read
(B operand 0x00); weakbus -> contested low bits resolve to odd words
{7,15,...,79} -> port B carries tile bytes (0x01). The remaining blocker
is the oa scan word<->byte<->port mapping / row schedule (even words
2k hold entry-k Y bytes under the model layout), to be pinned against the
schematic or the author's scan-address/oa-phase review. obj_prio_ck
(PPU1) remains flat even when the store fires - the PPU1 side of the
handshake is the next item once addressing is fixed.
Scan-only oa bus - claims fire (round 29)
Variant ppu2_scanonly.v (temp/oamweak; weakbus + the 18 non-scan
oa-chain drivers disabled, only the mode-2 scan group w518 drives):
scan address stable and EVEN (words {2,4,...,78}); with realistic content
(Y=16 in entries 1..39) the Y-test w816 = 1 on LY 1..7 (0 on LY=8) and
the store window w852 opens ~38x per line - slots are claimed. Entry 0
(words 0/1) is not visited by this sequence. obj_prio_ck (PPU1) still
flat. Conclusion: the oa addressing corruption in the static sim comes
from the overlapping oa-chain mux groups (scan w518 vs port-B w475/
CPU w403/store w444, g419/g421); way forward: author makes the groups
phase-exclusive, or the sprite test continues on the scan-only bus model
toward the PPU1 handshake.
Path b taken - sprite pixels reach LD0/LD1 (round 30)
gen_weakbus.py now also generates ppu2_m2only.v (non-scan oa-chain
drivers gated off while ppu_mode2=1; mode-3 store re-read groups intact).
Dev test tb_ppu_sprite_e2e (sprite entry 1: Y=16, X=16, tile 1, BG zero):
obj_prio_ck ~10-11 pulses/line, sp_bp_cys + sprite_x_match pulse,
one in-use flag set; LD shows sprite colour-01 pixels at LX ~7..14 on the
visible rows. Netlist untouched; still dev - remaining: entry-0 words not
visited by the scan sequence, LX off-by-one, author g419/g421 phase review.
What is needed to finish the sprite test
-
Author review/fix of the suspected items above (or confirmation that the no-reset FFs / oa phases are intentional - then the sim must be adjusted instead).
-
After that: re-run
tb_ppu_sprites- expected flow: mode-2 Y-test -> store claim (10 slots) -> mode-3sprite_x_match->obj_prio_ckpulses -> sprite VRAM fetch (sp_bp_cys) -> OBJ pixels on LD0/LD1. -
Promote
tb_ppu_spritesto a regression test with asserts, add its wave images towaves.mdand update the coverage matrix inReadme.md.