Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

302 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

TMS34010_sv

A synthesizable SystemVerilog reimplementation of the Texas Instruments TMS34010 Graphics System Processor, initially targeting Intel/Altera Cyclone V. This is FPGA RTL, not a software emulator.

Current status

The production-revision TMS34010 GPU scope frozen by Task 0161 is complete through Task 0174. Every defined graphics, display, interrupt/resume, and associated original-pin behavior in that scope has primary-source ownership and self-checking evidence; the complete regression, pinned MAME differential, preserved TI workloads, and clean Quartus implementation flow all pass at the signed-off revision. The final matrix and exact boundary are in docs/gpu_completion_signoff.md. Tasks 0162–0168 closed array semantics, direction, windowing, coherent checkpoints, PBX resume, and interruptible LINE. Task 0169 closes the complete production graphics-instruction/register matrix and adds a 1,186-case generated oracle: 1,176 PPOP/PSIZE/backend cells plus 10 PMASK-read cells. It includes destination-aligned COLOR0/COLOR1 dithering, physical-word-aligned PMASK, PMASK-before-transparency ordering, and processed PIXT memory-to-memory forms. The authoritative matrix is docs/graphics_conformance.md. Task 0170 closes the complete production display/video matrix and adds a 577-case generated scheduler oracle. It corrects DPYADR to decrement its raw SRFADR for both ORG values, retains the ORG=0 complement at the local-bus pins, defines DPYTAP as physical column/tap bits, exhaustively crosses sync direction/ENV/interlace/cadence, and fixes the package BLANK pin to active-low polarity. The authoritative display ledger is docs/display_conformance.md. Task 0171 recloses DPYCTL.SRT from every graphics engine through field sequencing, fixed-priority arbitration, CDC, LRDY/HOLD behavior, and every original-pin MTR/RTM phase. A full-pin program checks the exact ordered RTM/MTR/MTR/RTM trace; the all-engine system program checks per-engine cycle counts; and all 1,186 graphics matrix cases assert exact SRT tagging. The authoritative integration ledger is docs/srt_conformance.md. Task 0172 adds a license-separated adapter for exact MAME commit 70725158b4e9d2e1230c0515faec754f9cee86a2 and a deterministic 137-case graphics corpus. The same bootstrap programs, complete A/B/ST state, graphics I/O configuration, memory images, and framebuffer windows run on MAME and RTL. Checked-in MAME and accepted RTL results keep the ordinary regression offline, while scripts/mame_graphics_diff.sh rebuilds and reruns the live reference. Every secondary-reference difference is machine-classified with a minimized reproducer; none is unexplained. The setup and divergence ledger is docs/mame_graphics_reference.md. Task 0173 adds a preserved-software integration gate. Five TI-authored programs—the ROM tutorial, Sample Function Library display-list demo, two Graphics/Math tests, and 1987 GSP Paint—are hash-verified and loaded from the pinned original disks, then run to deterministic milestones on RTL and exact pinned MAME. The ROM tutorial also rebuilds with TI's original DOS assembler/linker to an identical address/word load image. This found and fixed complemented SUBI/CMPI extensions, asymmetric MMTM/MMFM masks, the two remaining field memory-to-memory MOVE forms, and subfield/two-register processor I/O access. Exact engine result locks and a closed nine-group ledger leave zero unexplained differences. No extracted executable or arcade ROM is committed. Provenance and run instructions are in docs/ti_workloads.md. Task 0124 reconciled the official instruction summary and all remaining system integration work into docs/completion_audit.md; Task 0125 closed the logical-status and ANDI/ANDNI semantic findings, and Tasks 0126–0127 landed the two missing memory-to-memory MOVE forms. Task 0128 implemented the EMU handshake and completed every row in the official instruction summary. Task 0129 corrected the MOVK/ADDK/SUBK encoded-zero constant to the specified value 32. Task 0130 corrected all shift encodings, SLA overflow, and per-instruction status masks. Task 0131 corrected MOVI C preservation for both immediate widths. Task 0132 resolved the REV value and EXGPC alignment rules directly against their individual instruction pages. Task 0133 resolved FILL XY's final linear DADDR writeback and W=0 status preservation. Task 0134 corrected SUBXY's coordinate comparisons to signed 16-bit semantics. Task 0135 completed the primary-page instruction status audit, resolved A0009, corrected divide/modulo and odd-result multiply edge cases, and completed W=3/PIXT graphics-window status behavior. Task 0136 resolved the remaining field-alignment assumption and added exact specification-derived sequencing from 1–32-bit architectural fields onto aligned 16-bit physical words. Task 0137 implemented the source-specific INTPEND/INTENB contract, synchronized both active-low external interrupt pins, and exposed synchronous host/display request sidebands. Task 0138 corrected REFCNT to the guide's continuous 14-bit decrementing counter, integrated it with CONTROL.RR/RM and the I/O register file, and exposed its refresh request, row, and mode at the core boundary. Task 0139 integrated the internal noninterlaced video counters/timing registers, corrected DIP to the start of horizontal blanking, and exported timing outputs. Task 0140 corrected the inherited sync/blank interval endpoints for the guide's one-VCLK equality-to-output delay. Task 0141 made DPYADR live and added a held screen-refresh request/acknowledge client with frame reload, line cadence, and completion-time address updates. Task 0142 completed direct host-side HSTCTL ownership, HINT, HCS-selected reset halt, and instruction-boundary HLT/NMI behavior. Task 0143 added the synchronous HSTADR/HSTDATA indirect-memory engine, including LBL byte ordering, INCR/INCW address sequencing, prefetch buffering, and held local-word requests. Task 0144 integrated that engine with the processor-visible I/O registers and replaced the temporary HSTCTL-only core boundary with one synchronous four-register host port plus an exposed held local-word client. Task 0145 added the specification-priority local-bus arbiter, including held-owner completion, pulsed DRAM-refresh retention, CPU partial-word RMW reservation, inter-word preemption, and the external-HOLD restart exception. Task 0146 connected every core memory client through the field sequencer and arbiter behind one synthesizable functional-system/controller boundary. Task 0147 added the standalone original-pin local-bus phase engine with LCLK1/LCLK2, exact address/status multiplexing, LRDY extension, all landed word/screen/DRAM/I/O cycle families, and eight post-reset RAS-only cycles. Task 0148 connected that engine to the core-clock memory fabric through a two-phase coherent command/response bridge, propagated opcode IAQ and screen-refresh ORG, and added an integrated pin-system wrapper. Task 0149 routes processor accesses to on-chip registers around field sequencing into the specified two-clock physical I/O read/write cycles and commits writes exactly once on returned completion. Task 0150 routes host-indirect accesses through those same physical cycles, adds an independent shared-register read view, and commits host-side writes only after their returned completion. Task 0151 connects the active-low physical HOLD input to the existing fixed-priority arbiter through synchronized level handshakes, emits the Q3/Q4-only active-low HOLDA component, and phases LAD/control output-enable release and reacquisition at the specified Q2/Q3 boundaries. Task 0152 synchronizes RUN/EMU, transfers each architectural EMU event into the 8× domain with a held handshake, and combines exact Q1/Q2 EMUA with Q3/Q4 HLDA on the original shared output pin. Task 0153 replaces the integrated wrapper's synchronous host boundary with the original active-low HCS/HREAD/HWRITE/HLDS/HUDS controls, HFS selection, byte-lane HD direction, immediate HRDY waits, coherent bundled capture, and prior-indirect busy backpressure. Task 0154 closes the remaining ordinary I/O-register reserved behavior: CONTROL/DPYCTL/DPYTAP masks apply identically to processor and host-indirect writes, and the four reserved register locations ignore writes and read zero. Task 0155 introduces the independent VCLK domain: HCOUNT/VCOUNT, timing compares, DPYADR, and automatic screen scheduling now live there, while atomic configuration/command/status mailboxes, lossless DIP delivery, and a bundled held screen-request bridge isolate every core/VCLK crossing. Task 0156 implements internally generated interlaced video: the odd field starts at HTOTAL/2, performs the specified VESYNC midline VCOUNT advance, returns to the even field at the full-line boundary, and applies signed DUDATE/2 to the DPYSTRT reload preceding that even field. Task 0157 implements external video synchronization: independently synchronized active-low HSYNC/VSYNC inputs receive the specified 2.5-VCLK recognition delay, HTOTAL/VTOTAL provide missing-sync fallbacks, HSD selects horizontal input or output operation, external interlace uses the documented horizontal recognition window, and explicit sync output enables propagate to the pin-system boundary. Task 0158 consumes DPYCTL.SRT and converts only graphics pixel reads/writes into the specified explicit VRAM memory-to-register/register-to-memory local cycles. It adds the exact TR/QE/W/address-status phases, retains ordinary instruction/data/I/O/host traffic, and avoids unnecessary destination reads for direct replace operations. This also corrects the completion boundary: the TMS34010 has no pixel-data output pins; attached VRAM supplies pixels through its serial port. Task 0159 adds the Cyclone V realization boundary: wrapped core/bus and phase-separated video PLLs, independent continuous VCLK, per-domain active-high reset release, actual HD/LAD/control/sync tri-states, active-low sync inversion, and a DE10-Nano top-level port surface. The reusable hierarchy now carries separate core, bus, and video resets so every clock domain releases synchronously. Task 0160 adds the complete Quartus Prime Lite 17 project, assigns all 63 top-level ports to DE10-Nano clocks/KEY0/JP1/JP7 pins, constrains every clock and I/O boundary, and makes map, fit, assembly, TimeQuest, pin, resource, and CDC report acceptance one deterministic command. Multicorner timing closes with +0.747 ns worst setup and +0.128 ns worst hold slack; all 27 required two-stage synchronizers are reported. A final-ratio integration test proves minimum-interval DRAM refresh completes in at most 11 of the available 32 core clocks when HOLD is inactive and LRDY is bounded. Task 0161 defines the new functional GPU gate and records all currently known remaining work: empty-array and terminal-context corrections, directional PIXBLT, W=1/W=3 array semantics, resumable FILL/PIXBLT/LINE execution, exhaustive graphics/display conformance, SRT/pin reintegration, MAME differential testing, TI software workloads, and final regression/Quartus sign-off. Exact instruction-cycle parity, the optional instruction cache, first-silicon compatibility, external VRAM serial pixels, later-family devices, and board analog validation are explicit non-goals. Task 0162 closes the first GPU finding. Every FILL and PIXBLT encoding now treats either zero DYDX dimension as an empty, memory-quiescent operation without implied-register or status changes. Full-color and binary PIXBLT completion writes SADDR/DADDR as the first pixels of the hypothetical next rows, while FILL retains its distinct final-row next-X DADDR. Focused coverage exercises every form, multiple PSIZEs, signed/non-unit pitches, and injected physical-memory stalls. Task 0163 implements the complete production-revision PBH/PBV contract. Full-color PIXBLTs traverse in all four horizontal/vertical directions; L,L consumes software-selected corners, while L,XY, XY,L, and XY,XY automatically adjust both source and destination from their default top-left addresses. Binary-source forms remain direction-independent as specified. Exact terminal context is decoupled from traversal state, and stalled-memory tests include safe forward/reverse overlapping copies. Task 0164 completes W=1 array hit detection as the specified common-rectangle operation. FILL XY and every XY-destination PIXBLT suppress all pixel traffic, set V/WVP from intersection, and, on a hit, return the exact intersection dimensions in DYDX plus its lowest-address or PBH/PBV-selected starting corner in DADDR. Miss results remain architecturally indeterminate but are deterministically preserved; zero dimensions retain the Task 0162 no-op contract. Task 0165 implements true W=3 array preclipping. FILL XY computes its effective destination intersection before issuing traffic; B,XY/L,XY/XY,XY apply identical left/top pixel offsets to source and destination and then traverse only the effective rectangle in the selected direction. Clipped pixels generate no source read, destination read, MTR, RTM, or write, and full exclusion is memory-quiescent. Exact original-array completion context, V-only status, PBH/PBV, PPOP/PMASK, arbitrary waits, and SRT tagging are locked by a request-by-request reference model. Task 0166 makes every FILL and PIXBLT form checkpointable at the specified destination-word and row boundaries. A quiet seven-cycle sequence publishes the next source/destination pixels, traversal cursor, effective dimensions, and saved result/geometry context through B0/B2/B10-B14; its final cycle is the sole safe interrupt-recognition point. Exact request traces and every captured restart suffix are reference-checked under direction, W=3, PPOP/PMASK, SRT, partial-word RMW, and memory stalls. Task 0167 connects those boundaries to maskable and NMIM=0 nonmaskable entry. The core stacks the rewound array opcode PC and an ST copy with PBX set, installs the ordinary clean handler ST, then lets RETI restore PBX and refetch the opcode. Four quiet B-file read states reconstruct all private FILL/PIXBLT state from the documented handler-preserved context; successful completion clears PBX. Every legal checkpoint in all eight forms is interrupted under three-cycle memory stalls, including repeated entries, with exact no-duplicate/no-skip traffic and ordinary/LINE PBX-negative controls. Task 0168 completes LINE interruption after every nonfinal pixel. Three quiet writeback states publish the next d/DADDR/COUNT in B0/B2/B10, after which a pending interrupt stacks the rewound opcode PC with PBX clear. RETI therefore uses the ordinary LINE setup to continue from the architectural image. All signed octants, horizontal/vertical/diagonal lines, repeated DI/NMI, W=3, RMW, SRT, exact traffic, and stalls are reference-checked; W=1/W=2 violations retain their complete-abort path without a restart checkpoint. The repository currently contains:

  • a multicycle 32-bit core with bit-addressed instruction and data access;
  • A/B register files, shared stack pointer, status register, ALU, shifter, multiplier, and multicycle divider;
  • the instruction families tracked in docs/instruction_coverage.md, including field-aware MOVE, stack/trap/interrupt operations, and conditional control;
  • PIXT, FILL, PIXBLT, DRAV, and LINE graphics datapaths with pixel processing, plane masking, transparency, and all window modes;
  • on-chip I/O-register storage plus every maskable pending-source boundary and maskable/nonmaskable entry path, including defined reserved-field and reserved-location behavior;
  • architectural reset and illegal-opcode vector entry, including HCS-selected host-present reset halt;
  • a synchronous direct-host HSTCTL boundary with complementary host/processor field ownership, active-low HINT, and instruction-boundary HLT;
  • an integrated synchronous four-register host engine with shared processor/host HSTADR/HSTDATA state, byte-order triggers, pre-read/post-write incrementing, stalled-request stability, and an exposed local-word client;
  • synchronized physical RUN/EMU sampling, exact Q1/Q2 EMUA pulse/halt indication, and resume;
  • a synthesizable field-to-word sequencer covering §4.1 alignment cases A–G, partial-word RMW locking/restart, and arbitrary word-side stalls;
  • a synthesizable fixed-priority HOLD/screen/DRAM/host/CPU local-cycle arbiter with held grants and a captured DRAM-refresh event;
  • an integrated functional-system wrapper that converges CPU/graphics, screen, DRAM-refresh, and host-indirect traffic on one abstract controller;
  • a standalone 8×-clock original-pin local-bus engine covering ordinary word, screen-transfer, program-controlled MTR/RTM, RAS-only, CAS-before-RAS, and I/O cycles, including LRDY waits and reset initialization;
  • an integrated core-clock-to-8× pin-system wrapper using a lossless MCP command/response CDC, including returned read data, IAQ, and screen ORG;
  • processor and host-indirect on-chip I/O access through dedicated RAS/LAL-only physical cycles, including internal read data and completion-qualified register writes;
  • active-low physical HOLD sampling and synchronized grant return, with early Q3/Q4 HOLDA indication and explicit Q2/Q3 LAD/control output-enable release/resume sequencing;
  • the original shared HLDA/EMUA output, with lossless EMU-event CDC and phase-exclusive EMUA/HLDA selection under simultaneous halt and HOLD;
  • the asynchronous original-pin host bus, with legal-access qualification, HCS-triggered HSTCTL delay, coherent register request/response capture, indirect-busy waits, latched read data, and per-byte HD output enables;
  • integrated dedicated-VCLK internal/external noninterlaced/interlaced video timing with live HCOUNT/VCOUNT, synchronized active-low sync inputs, DPYCTL.DXV/HSD direction control and output enables, DPYCTL.ENV blanking, field-aware DIP, live raw-decrement DPYADR, half-DUDATE field starts, and held screen-refresh scheduling; coherent MCP configuration/command/status, event, and completed-screen-transaction crossings; plus integrated REFCNT/refresh-request generation;
  • DPYCTL.SRT classification for every graphics engine, with explicit program-controlled VRAM MTR/RTM pin cycles and unaffected nonpixel traffic;
  • a Cyclone V top-level adapter with vendor-isolated PLLs, three reset conditioners, active-low video-clock/sync/blank phase mapping, and IOE-ready bidirectional host/local/video pads;
  • a Quartus Prime Lite 17 project with a complete SDC, deterministic implementation/report validator, 63 fitted 3.3-V LVTTL pins, and archived reproducible sign-off evidence;
  • 159 self-checking SystemVerilog testbenches, including non-integer-clock video CDC, cycle-by-cycle internal interlace and external-sync coverage, the generated 577-case display scheduler matrix, end-to-end SRT graphics-cycle coverage, direct FPGA pad/reset checks, and an exhaustive 65,536-opcode static status-policy sweep.

The Task 0160 processor/FPGA baseline and its reproducible Cyclone V implementation flow remain complete. Tasks 0162–0174 close every identified production functional, surviving-software, and final implementation gate, so the Task 0161 scoped GPU-complete claim is restored. A board-level system also needs external VRAM/DRAM or an equivalent memory/video subsystem, level translation, and signal-integrity validation to consume the landed transfer cycles and emit pixels; those surrounding-device responsibilities are not part of this processor.

Getting started

git submodule update --init --recursive
scripts/lint.sh
scripts/regress.sh
scripts/sim.sh tb_smoke
scripts/sim.sh tb_pixt_win
scripts/sim.sh tb_mame_graphics_replay
scripts/sim.sh tb_ti_workload_replay
# Explicit first-time network/build step for the optional live reference:
scripts/mame_graphics_diff.sh --setup
scripts/ti_workloads.sh --setup --build-rom
QUARTUS_SH=/path/to/quartus-17.0/quartus/bin/quartus_sh \
  scripts/synth_quartus.sh

The scripts prefer Questa/ModelSim and fall back to Verilator. Testbenches must print TEST_RESULT: PASS; the simulator exit code alone is not treated as a pass. scripts/regress.sh discovers and runs all testbenches; set REGRESS_JOBS to parallelize the Verilator flow. scripts/synth_quartus.sh requires Quartus Prime Lite 17.0.2 and rebuilds map, fit, assembly, and TimeQuest results before validating the exact warnings, pin fit, complete constraints, timing, synchronizers, and resource envelopes.

Before changing RTL, read AGENTS.md, tasks.md, architecture.md, and the relevant specification in the pinned third_party/TMS34010_Info submodule.

Repository map

  • rtl/ — synthesizable package, core, memory sequencing, I/O, video, and refresh RTL.
  • sim/models/ — nonsynthesizable behavioral memory model.
  • sim/tb/ — focused self-checking testbenches.
  • fpga/ — Quartus 17 project, SDC, DE10-Nano pinout, report script, and reproducible implementation evidence.
  • docs/ — architecture, assumptions, coverage, memory/timing notes, and the authoritative Cyclone V HDL coding-guideline bundle.
  • docs/completion_audit.md — primary-spec reconciliation and ordered exit gates for project completion.
  • docs/status_audit.md — complete individual-instruction N/C/Z/V policy, undefined-bit handling, and regression evidence.
  • docs/graphics_conformance.md — every production graphics instruction and defined graphics-register field, its primary source, RTL owner, side effects, undefined boundaries, and named regression evidence.
  • docs/srt_conformance.md — graphics-only SRT classification, zero-cycle conditions, continuation semantics, arbitration/CDC behavior, original-pin MTR/RTM phases, and processor/external-VRAM scope.
  • docs/mame_graphics_reference.md — exact MAME provenance/license boundary, deterministic differential corpus, opt-in setup, and the closed secondary-reference divergence ledger.
  • docs/ti_workloads.md — preserved TI media hashes, original-tool ROM rebuild, workload entry/memory maps, deterministic milestones, result locks, closed divergence classes, and optional legal arcade-ROM readiness.
  • scripts/ — lint, simulation, and Quartus entry points.
  • tasks.md / changelog.md — task-level design and implementation history.
  • third_party/TMS34010_Info/ — pinned primary/reference documentation.

About

TMS34010 reimplementation in SystemVerilog

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages