A synthesizable SystemVerilog reimplementation of the Texas Instruments TMS34010 Graphics System Processor, initially targeting Intel/Altera Cyclone V. This is FPGA RTL, not a software emulator.
The production-revision TMS34010 GPU scope frozen by Task 0161 is complete
through Task 0174. Every defined graphics, display, interrupt/resume, and
associated original-pin behavior in that scope has primary-source ownership
and self-checking evidence; the complete regression, pinned MAME differential,
preserved TI workloads, and clean Quartus implementation flow all pass at the
signed-off revision. The final matrix and exact boundary are in
docs/gpu_completion_signoff.md. Tasks 0162–0168 closed array semantics,
direction, windowing, coherent checkpoints, PBX resume, and interruptible
LINE. Task 0169 closes the complete production
graphics-instruction/register matrix and adds a 1,186-case generated oracle:
1,176 PPOP/PSIZE/backend cells plus 10 PMASK-read cells. It includes
destination-aligned COLOR0/COLOR1 dithering, physical-word-aligned PMASK,
PMASK-before-transparency ordering, and processed PIXT memory-to-memory
forms. The authoritative matrix is
docs/graphics_conformance.md. Task 0170 closes the complete production
display/video matrix and adds a 577-case generated scheduler oracle. It
corrects DPYADR to decrement its raw SRFADR for both ORG values, retains the
ORG=0 complement at the local-bus pins, defines DPYTAP as physical
column/tap bits, exhaustively crosses sync direction/ENV/interlace/cadence,
and fixes the package BLANK pin to active-low polarity. The authoritative
display ledger is docs/display_conformance.md. Task 0171 recloses
DPYCTL.SRT from every graphics engine through field sequencing,
fixed-priority arbitration, CDC, LRDY/HOLD behavior, and every original-pin
MTR/RTM phase. A full-pin program checks the exact ordered
RTM/MTR/MTR/RTM trace; the all-engine system program checks per-engine cycle
counts; and all 1,186 graphics matrix cases assert exact SRT tagging. The
authoritative integration ledger is docs/srt_conformance.md. Task 0172 adds
a license-separated adapter for exact MAME commit
70725158b4e9d2e1230c0515faec754f9cee86a2 and a deterministic 137-case
graphics corpus. The same bootstrap programs, complete A/B/ST state,
graphics I/O configuration, memory images, and framebuffer windows run on
MAME and RTL. Checked-in MAME and accepted RTL results keep the ordinary
regression offline, while scripts/mame_graphics_diff.sh rebuilds and reruns
the live reference. Every secondary-reference difference is machine-classified
with a minimized reproducer; none is unexplained. The setup and divergence
ledger is docs/mame_graphics_reference.md. Task 0173 adds a
preserved-software integration gate. Five TI-authored
programs—the ROM tutorial, Sample Function Library display-list demo, two
Graphics/Math tests, and 1987 GSP Paint—are hash-verified and loaded from the
pinned original disks, then run to deterministic milestones on RTL and exact
pinned MAME. The ROM tutorial also rebuilds with TI's original DOS
assembler/linker to an identical address/word load image. This found and
fixed complemented SUBI/CMPI extensions, asymmetric MMTM/MMFM masks, the two
remaining field memory-to-memory MOVE forms, and subfield/two-register
processor I/O access. Exact engine result locks and a closed nine-group
ledger leave zero unexplained differences. No extracted executable or arcade
ROM is committed. Provenance and run instructions are in
docs/ti_workloads.md. Task 0124
reconciled the official instruction summary and all remaining system
integration work into docs/completion_audit.md; Task 0125 closed the
logical-status and ANDI/ANDNI semantic findings, and Tasks 0126–0127 landed
the two missing memory-to-memory MOVE forms. Task 0128 implemented the EMU
handshake and completed every row in the official instruction summary. Task
0129 corrected the MOVK/ADDK/SUBK encoded-zero constant to the specified
value 32. Task 0130 corrected all shift encodings, SLA overflow, and
per-instruction status masks. Task 0131 corrected MOVI C preservation for
both immediate widths. Task 0132 resolved the REV value and EXGPC alignment
rules directly against their individual instruction pages. Task 0133 resolved
FILL XY's final linear DADDR writeback and W=0 status preservation. Task
0134 corrected SUBXY's coordinate comparisons to signed 16-bit
semantics. Task 0135 completed the primary-page instruction status audit,
resolved A0009, corrected divide/modulo and odd-result multiply edge cases,
and completed W=3/PIXT graphics-window status behavior. Task 0136 resolved
the remaining field-alignment assumption and added exact
specification-derived sequencing from 1–32-bit architectural fields onto
aligned 16-bit physical words. Task 0137 implemented the source-specific
INTPEND/INTENB contract, synchronized both active-low external interrupt
pins, and exposed synchronous host/display request sidebands. Task 0138
corrected REFCNT to the guide's continuous 14-bit decrementing counter,
integrated it with CONTROL.RR/RM and the I/O register file, and exposed its
refresh request, row, and mode at the core boundary. Task 0139 integrated the
internal noninterlaced video counters/timing registers, corrected DIP to the
start of horizontal blanking, and exported timing outputs. Task 0140
corrected the inherited sync/blank interval endpoints for the guide's
one-VCLK equality-to-output delay. Task 0141 made DPYADR live and added a
held screen-refresh request/acknowledge client with frame reload, line
cadence, and completion-time address updates. Task 0142 completed direct
host-side HSTCTL ownership, HINT, HCS-selected reset halt, and
instruction-boundary HLT/NMI behavior. Task 0143 added the synchronous
HSTADR/HSTDATA indirect-memory engine, including LBL byte ordering,
INCR/INCW address sequencing, prefetch buffering, and held local-word
requests. Task 0144 integrated that engine with the processor-visible I/O
registers and replaced the temporary HSTCTL-only core boundary with one
synchronous four-register host port plus an exposed held local-word client.
Task 0145 added the specification-priority local-bus arbiter, including
held-owner completion, pulsed DRAM-refresh retention, CPU partial-word RMW
reservation, inter-word preemption, and the external-HOLD restart exception.
Task 0146 connected every core memory client through the field sequencer and
arbiter behind one synthesizable functional-system/controller boundary.
Task 0147 added the standalone original-pin local-bus phase engine with
LCLK1/LCLK2, exact address/status multiplexing, LRDY extension, all landed
word/screen/DRAM/I/O cycle families, and eight post-reset RAS-only cycles.
Task 0148 connected that engine to the core-clock memory fabric through a
two-phase coherent command/response bridge, propagated opcode IAQ and
screen-refresh ORG, and added an integrated pin-system wrapper.
Task 0149 routes processor accesses to on-chip registers around field
sequencing into the specified two-clock physical I/O read/write cycles and
commits writes exactly once on returned completion.
Task 0150 routes host-indirect accesses through those same physical cycles,
adds an independent shared-register read view, and commits host-side writes
only after their returned completion.
Task 0151 connects the active-low physical HOLD input to the existing
fixed-priority arbiter through synchronized level handshakes, emits the
Q3/Q4-only active-low HOLDA component, and phases LAD/control output-enable
release and reacquisition at the specified Q2/Q3 boundaries.
Task 0152 synchronizes RUN/EMU, transfers each architectural EMU event into
the 8× domain with a held handshake, and combines exact Q1/Q2 EMUA with
Q3/Q4 HLDA on the original shared output pin.
Task 0153 replaces the integrated wrapper's synchronous host boundary with
the original active-low HCS/HREAD/HWRITE/HLDS/HUDS controls, HFS selection,
byte-lane HD direction, immediate HRDY waits, coherent bundled capture, and
prior-indirect busy backpressure.
Task 0154 closes the remaining ordinary I/O-register reserved behavior:
CONTROL/DPYCTL/DPYTAP masks apply identically to processor and host-indirect
writes, and the four reserved register locations ignore writes and read zero.
Task 0155 introduces the independent VCLK domain: HCOUNT/VCOUNT, timing
compares, DPYADR, and automatic screen scheduling now live there, while
atomic configuration/command/status mailboxes, lossless DIP delivery, and a
bundled held screen-request bridge isolate every core/VCLK crossing.
Task 0156 implements internally generated interlaced video: the odd field
starts at HTOTAL/2, performs the specified VESYNC midline VCOUNT advance,
returns to the even field at the full-line boundary, and applies signed
DUDATE/2 to the DPYSTRT reload preceding that even field.
Task 0157 implements external video synchronization: independently
synchronized active-low HSYNC/VSYNC inputs receive the specified 2.5-VCLK
recognition delay, HTOTAL/VTOTAL provide missing-sync fallbacks, HSD selects
horizontal input or output operation, external interlace uses the documented
horizontal recognition window, and explicit sync output enables propagate to
the pin-system boundary.
Task 0158 consumes DPYCTL.SRT and converts only graphics pixel reads/writes
into the specified explicit VRAM memory-to-register/register-to-memory local
cycles. It adds the exact TR/QE/W/address-status phases, retains ordinary
instruction/data/I/O/host traffic, and avoids unnecessary destination reads
for direct replace operations. This also corrects the completion boundary:
the TMS34010 has no pixel-data output pins; attached VRAM supplies pixels
through its serial port.
Task 0159 adds the Cyclone V realization boundary: wrapped core/bus and
phase-separated video PLLs, independent continuous VCLK, per-domain
active-high reset release, actual HD/LAD/control/sync tri-states, active-low
sync inversion, and a DE10-Nano top-level port surface. The reusable
hierarchy now carries separate core, bus, and video resets so every clock
domain releases synchronously.
Task 0160 adds the complete Quartus Prime Lite 17 project, assigns all 63
top-level ports to DE10-Nano clocks/KEY0/JP1/JP7 pins, constrains every clock
and I/O boundary, and makes map, fit, assembly, TimeQuest, pin, resource, and
CDC report acceptance one deterministic command. Multicorner timing closes
with +0.747 ns worst setup and +0.128 ns worst hold slack; all 27 required
two-stage synchronizers are reported. A final-ratio integration test proves
minimum-interval DRAM refresh completes in at most 11 of the available 32
core clocks when HOLD is inactive and LRDY is bounded.
Task 0161 defines the new functional GPU gate and records all currently known
remaining work: empty-array and terminal-context corrections, directional
PIXBLT, W=1/W=3 array semantics, resumable FILL/PIXBLT/LINE execution,
exhaustive graphics/display conformance, SRT/pin reintegration, MAME
differential testing, TI software workloads, and final regression/Quartus
sign-off. Exact instruction-cycle parity, the optional instruction cache,
first-silicon compatibility, external VRAM serial pixels, later-family
devices, and board analog validation are explicit non-goals.
Task 0162 closes the first GPU finding. Every FILL and PIXBLT encoding now
treats either zero DYDX dimension as an empty, memory-quiescent operation
without implied-register or status changes. Full-color and binary PIXBLT
completion writes SADDR/DADDR as the first pixels of the hypothetical next
rows, while FILL retains its distinct final-row next-X DADDR. Focused coverage
exercises every form, multiple PSIZEs, signed/non-unit pitches, and injected
physical-memory stalls.
Task 0163 implements the complete production-revision PBH/PBV contract.
Full-color PIXBLTs traverse in all four horizontal/vertical directions;
L,L consumes software-selected corners, while L,XY, XY,L, and XY,XY
automatically adjust both source and destination from their default
top-left addresses. Binary-source forms remain direction-independent as
specified. Exact terminal context is decoupled from traversal state, and
stalled-memory tests include safe forward/reverse overlapping copies.
Task 0164 completes W=1 array hit detection as the specified common-rectangle
operation. FILL XY and every XY-destination PIXBLT suppress all pixel traffic,
set V/WVP from intersection, and, on a hit, return the exact intersection
dimensions in DYDX plus its lowest-address or PBH/PBV-selected starting corner
in DADDR. Miss results remain architecturally indeterminate but are
deterministically preserved; zero dimensions retain the Task 0162 no-op
contract.
Task 0165 implements true W=3 array preclipping. FILL XY computes its
effective destination intersection before issuing traffic; B,XY/L,XY/XY,XY
apply identical left/top pixel offsets to source and destination and then
traverse only the effective rectangle in the selected direction. Clipped
pixels generate no source read, destination read, MTR, RTM, or write, and
full exclusion is memory-quiescent. Exact original-array completion context,
V-only status, PBH/PBV, PPOP/PMASK, arbitrary waits, and SRT tagging are
locked by a request-by-request reference model.
Task 0166 makes every FILL and PIXBLT form checkpointable at the specified
destination-word and row boundaries. A quiet seven-cycle sequence publishes
the next source/destination pixels, traversal cursor, effective dimensions,
and saved result/geometry context through B0/B2/B10-B14; its final cycle is
the sole safe interrupt-recognition point. Exact request traces and every
captured restart suffix are reference-checked under direction, W=3,
PPOP/PMASK, SRT, partial-word RMW, and memory stalls.
Task 0167 connects those boundaries to maskable and NMIM=0 nonmaskable entry.
The core stacks the rewound array opcode PC and an ST copy with PBX set,
installs the ordinary clean handler ST, then lets RETI restore PBX and refetch
the opcode. Four quiet B-file read states reconstruct all private FILL/PIXBLT
state from the documented handler-preserved context; successful completion
clears PBX. Every legal checkpoint in all eight forms is interrupted under
three-cycle memory stalls, including repeated entries, with exact
no-duplicate/no-skip traffic and ordinary/LINE PBX-negative controls. Task
0168 completes LINE interruption after every nonfinal pixel. Three quiet
writeback states publish the next d/DADDR/COUNT in B0/B2/B10, after which a
pending interrupt stacks the rewound opcode PC with PBX clear. RETI therefore
uses the ordinary LINE setup to continue from the architectural image.
All signed octants, horizontal/vertical/diagonal lines, repeated DI/NMI,
W=3, RMW, SRT, exact traffic, and stalls are reference-checked; W=1/W=2
violations retain their complete-abort path without a restart checkpoint.
The repository currently contains:
- a multicycle 32-bit core with bit-addressed instruction and data access;
- A/B register files, shared stack pointer, status register, ALU, shifter, multiplier, and multicycle divider;
- the instruction families tracked in
docs/instruction_coverage.md, including field-aware MOVE, stack/trap/interrupt operations, and conditional control; - PIXT, FILL, PIXBLT, DRAV, and LINE graphics datapaths with pixel processing, plane masking, transparency, and all window modes;
- on-chip I/O-register storage plus every maskable pending-source boundary and maskable/nonmaskable entry path, including defined reserved-field and reserved-location behavior;
- architectural reset and illegal-opcode vector entry, including HCS-selected host-present reset halt;
- a synchronous direct-host HSTCTL boundary with complementary host/processor field ownership, active-low HINT, and instruction-boundary HLT;
- an integrated synchronous four-register host engine with shared processor/host HSTADR/HSTDATA state, byte-order triggers, pre-read/post-write incrementing, stalled-request stability, and an exposed local-word client;
- synchronized physical RUN/EMU sampling, exact Q1/Q2 EMUA pulse/halt indication, and resume;
- a synthesizable field-to-word sequencer covering §4.1 alignment cases A–G, partial-word RMW locking/restart, and arbitrary word-side stalls;
- a synthesizable fixed-priority HOLD/screen/DRAM/host/CPU local-cycle arbiter with held grants and a captured DRAM-refresh event;
- an integrated functional-system wrapper that converges CPU/graphics, screen, DRAM-refresh, and host-indirect traffic on one abstract controller;
- a standalone 8×-clock original-pin local-bus engine covering ordinary word, screen-transfer, program-controlled MTR/RTM, RAS-only, CAS-before-RAS, and I/O cycles, including LRDY waits and reset initialization;
- an integrated core-clock-to-8× pin-system wrapper using a lossless MCP command/response CDC, including returned read data, IAQ, and screen ORG;
- processor and host-indirect on-chip I/O access through dedicated RAS/LAL-only physical cycles, including internal read data and completion-qualified register writes;
- active-low physical HOLD sampling and synchronized grant return, with early Q3/Q4 HOLDA indication and explicit Q2/Q3 LAD/control output-enable release/resume sequencing;
- the original shared HLDA/EMUA output, with lossless EMU-event CDC and phase-exclusive EMUA/HLDA selection under simultaneous halt and HOLD;
- the asynchronous original-pin host bus, with legal-access qualification, HCS-triggered HSTCTL delay, coherent register request/response capture, indirect-busy waits, latched read data, and per-byte HD output enables;
- integrated dedicated-VCLK internal/external noninterlaced/interlaced video timing with live HCOUNT/VCOUNT, synchronized active-low sync inputs, DPYCTL.DXV/HSD direction control and output enables, DPYCTL.ENV blanking, field-aware DIP, live raw-decrement DPYADR, half-DUDATE field starts, and held screen-refresh scheduling; coherent MCP configuration/command/status, event, and completed-screen-transaction crossings; plus integrated REFCNT/refresh-request generation;
- DPYCTL.SRT classification for every graphics engine, with explicit program-controlled VRAM MTR/RTM pin cycles and unaffected nonpixel traffic;
- a Cyclone V top-level adapter with vendor-isolated PLLs, three reset conditioners, active-low video-clock/sync/blank phase mapping, and IOE-ready bidirectional host/local/video pads;
- a Quartus Prime Lite 17 project with a complete SDC, deterministic implementation/report validator, 63 fitted 3.3-V LVTTL pins, and archived reproducible sign-off evidence;
- 159 self-checking SystemVerilog testbenches, including non-integer-clock video CDC, cycle-by-cycle internal interlace and external-sync coverage, the generated 577-case display scheduler matrix, end-to-end SRT graphics-cycle coverage, direct FPGA pad/reset checks, and an exhaustive 65,536-opcode static status-policy sweep.
The Task 0160 processor/FPGA baseline and its reproducible Cyclone V implementation flow remain complete. Tasks 0162–0174 close every identified production functional, surviving-software, and final implementation gate, so the Task 0161 scoped GPU-complete claim is restored. A board-level system also needs external VRAM/DRAM or an equivalent memory/video subsystem, level translation, and signal-integrity validation to consume the landed transfer cycles and emit pixels; those surrounding-device responsibilities are not part of this processor.
git submodule update --init --recursive
scripts/lint.sh
scripts/regress.sh
scripts/sim.sh tb_smoke
scripts/sim.sh tb_pixt_win
scripts/sim.sh tb_mame_graphics_replay
scripts/sim.sh tb_ti_workload_replay
# Explicit first-time network/build step for the optional live reference:
scripts/mame_graphics_diff.sh --setup
scripts/ti_workloads.sh --setup --build-rom
QUARTUS_SH=/path/to/quartus-17.0/quartus/bin/quartus_sh \
scripts/synth_quartus.shThe scripts prefer Questa/ModelSim and fall back to Verilator. Testbenches must
print TEST_RESULT: PASS; the simulator exit code alone is not treated as a
pass. scripts/regress.sh discovers and runs all testbenches; set
REGRESS_JOBS to parallelize the Verilator flow. scripts/synth_quartus.sh
requires Quartus Prime Lite 17.0.2 and rebuilds map, fit, assembly, and
TimeQuest results before validating the exact warnings, pin fit, complete
constraints, timing, synchronizers, and resource envelopes.
Before changing RTL, read AGENTS.md, tasks.md,
architecture.md, and the relevant specification in the
pinned third_party/TMS34010_Info submodule.
rtl/— synthesizable package, core, memory sequencing, I/O, video, and refresh RTL.sim/models/— nonsynthesizable behavioral memory model.sim/tb/— focused self-checking testbenches.fpga/— Quartus 17 project, SDC, DE10-Nano pinout, report script, and reproducible implementation evidence.docs/— architecture, assumptions, coverage, memory/timing notes, and the authoritative Cyclone V HDL coding-guideline bundle.docs/completion_audit.md— primary-spec reconciliation and ordered exit gates for project completion.docs/status_audit.md— complete individual-instruction N/C/Z/V policy, undefined-bit handling, and regression evidence.docs/graphics_conformance.md— every production graphics instruction and defined graphics-register field, its primary source, RTL owner, side effects, undefined boundaries, and named regression evidence.docs/srt_conformance.md— graphics-only SRT classification, zero-cycle conditions, continuation semantics, arbitration/CDC behavior, original-pin MTR/RTM phases, and processor/external-VRAM scope.docs/mame_graphics_reference.md— exact MAME provenance/license boundary, deterministic differential corpus, opt-in setup, and the closed secondary-reference divergence ledger.docs/ti_workloads.md— preserved TI media hashes, original-tool ROM rebuild, workload entry/memory maps, deterministic milestones, result locks, closed divergence classes, and optional legal arcade-ROM readiness.scripts/— lint, simulation, and Quartus entry points.tasks.md/changelog.md— task-level design and implementation history.third_party/TMS34010_Info/— pinned primary/reference documentation.