Support matrix
What Retrograd claims per backend, and what it does not. A capability is listed where it is verified, not where a code path merely exists. A row marked manual has no automatic coverage: it is run by hand on release, per Release criteria below.
Backends are chosen at build time with Cargo features (metal, vulkan, cuda; CPU is always included, and platform-gpu defaults to Metal on macOS for the root and Python packages, CPU elsewhere). metal is macOS-only and cuda is refused there.
The authoritative, per-build answer is GET /v1/capabilities: it reports what this binary actually registered, on the device it is actually running on. The tables below are the general shape; that endpoint is the specific one.
Training algorithms
| Capability | CPU | Metal | Vulkan | CUDA |
|---|---|---|---|---|
| LoRA training (SFT) | ✅ | ✅ | ✅ | ✅ |
| Base-weight training (full / partial / hybrid) | ✅ | ✅ | ✅ | ✅ |
| PPO / GRPO / agentic GRPO | ✅ | ✅ | ✅ | ✅ |
| Recurrent and hybrid model families | ✅ | ✅ | ✅ | ✅ |
Optimizers
| Optimizer | CPU | Metal | Vulkan | CUDA |
|---|---|---|---|---|
| AdamW | ✅ | ✅ | ✅ | ✅ |
| SGD | ✅ | ✅ | ✅ | ✅ |
| Muon | ✅ | ✅ | ✅ | ✅ |
| Gefen | ✅ | ✅ | 🟡 [a] | ✅ |
Gefen's two update phases are implemented on every backend. A run still asks the live device whether it has them (cap_opt_step_device in the capability report) and is refused at preflight where it does not, rather than silently falling back to a copy that would leave the real optimizer state stale.
- [a] Vulkan carries both phases, with one restriction under
variant = "quantized_m": the block size must be a multiple of four.shared_vhas no such restriction. The default (1024) satisfies it, and a run that asks for a non-multiple is refused at preflight.
Both optimizers are correct on every backend and neither is yet recommended over AdamW: see Optimizer cost and quality for what they were measured to cost and to approximate.
Base-weight storage precision
Base training reads the weights at the precision the GGUF stores them at and writes them back there. Nothing is converted, at load or at checkpoint. There are two update paths below F32, selected by training.master_weights.
In place ("off"). Each step is stochastically rounded back onto the weight's own precision. It costs nothing per parameter, but a step far smaller than one grid point is still a whole grid point taken at random; a run whose steps fall below that floor is refused.
Through an F32 master copy ("f32", and "auto" for half-precision base tensors). Each step runs in F32 and is rounded into the store once, so the floor above does not apply. It costs four bytes per trained element - 1.9 GiB for a 498M-parameter run, cheaper than converting the whole model to F32 because the store stays narrow in the forward. Measured on the GRPO full-finetuning example against a BF16 0.8B at lr = 1e-6: 8.47 GiB of device memory without the copy and 10.32 GiB with it; the run is refused without it, since at that rate every trained parameter takes a step under its own store's floor.
| Stored as | CPU | Metal | Vulkan | CUDA |
|---|---|---|---|---|
| F32 | ✅ | ✅ | ✅ | ✅ |
| F16, under AdamW or SGD, in place | ✅ | ✅ | ❌ [c] | ✅ |
| BF16, under AdamW or SGD, in place | ✅ | ✅ | ❌ [c] [e] | ✅ |
| F16 or BF16, under AdamW or SGD, master copy | ✅ | ✅ | ❌ [e] [f] | ✅ |
| F16 or BF16, under Muon or Gefen | ❌ [d] | ❌ [d] | ❌ [d] | ❌ [d] |
A cell is a measured combination of (dtype, optimizer, backend, master), run with tolerances read from the row; a combination with no row is refused before the graph is built. The live device is also asked each step whether it can run the update, and either the table or the device can refuse.
- [c] The update kernel exists and an exact CPU-against-device equality test holds it to the CPU's rounding, but no lane has run a model on it, so no row admits it.
- [d] Both steps write F32 parameters only: Muon's ends in a Newton-Schulz orthogonalization, Gefen's already approximates the first moment, and a rounded store would stack a second approximation on the first. The refusal says that, not an empty list of backends.
- [e] Vulkan does not decode a BF16 weight in
OUT_PROD, which an activation gradient needs, nor a BF16 head in the fused cross-entropy; the other three backends do. A BF16 model sends both to the CPU, whichRETRO_REQUIRE_GPU_RESIDENT=1refuses, and the master copy does not help: it changes the update, not the backward that feeds it. - [f] F16: measured, and refused on the resume lane: the loss of the step after a restore differs from the uninterrupted run's by one ulp of F32 while the restored weights are bit-identical - a difference in how the reported loss is reduced in a fresh context, not in the master copy. The other lanes pass.
BF16 keeps 8 significand bits against F16's 11, so its grid is eight times coarser and the measured gaps follow: ten times the F16 ones on the CPU, three times on CUDA, where BF16's own grid then dominates the reduction-order difference. Stochastic rounding keeps the difference a rounding rather than a bias.
The optimizer matters too. An AdamW step is wider than the grid, so it mostly lands on the same grid point whichever backend computed the gradient; an SGD step is the gradient itself, about one grid point wide, so a small difference between backends decides which side of a point each element lands on. That is why the SGD rows' outlier fractions move more between CPU and CUDA than AdamW's do. Metal is the limit case: on the fixture its F16 gradient equals the F32 twin's, so no element lands a grid point away under either optimizer.
Kernels and execution
| Capability | CPU | Metal | Vulkan | CUDA |
|---|---|---|---|---|
| Differentiable flash attention | ➖ [1] | 🟡 [2] | ✅ [3] | 🟡 [4] |
| Chunked cross-entropy | ✅ | ✅ [5] | ✅ [5] | ✅ [5] |
| Device-side sampling | ➖ | 🟡 [6] | 🟡 [6] | 🟡 [6] |
| Generated RIR kernels | ❌ [7] | 🟡 [8] | 🟡 [8] | 🟡 [9] |
| Multi-GPU | ➖ | ❌ | ❌ | ❌ [10] |
| Automatic CI lane | ✅ [11] | ⚠️ | ⚠️ | ⚠️ |
- [1] GPU only: a CPU run builds the materialized F32 backward graph, and so does any GPU shape the device declines. What follows is probed per model at load time (
supports_flash_attn_back), never read off a backend name. - [2] Metal stops at
head_dim256 (GGML_METAL_FA_BACK_MAX_D), K/V in F16 or F32, no attention sinks. - [3] Vulkan reaches
head_dim512 and is the only backend accepting attention sinks - the only one with a probe exercising them. It needs a 32-wide subgroup. - [4] CUDA reaches
head_dim512 (FA_BACK_MAX_D), K/V in F16 or F32, no attention sinks. - [5] On by default everywhere, and every backend ships both fused nodes - but they are probed against the model's head geometry and weight type (
cap_fused_sparse_ce). A device that declines either sends the whole loss tail to the CPU, which the backend report names (chunked_cross_entropy_status: cpu_fallback) andRETRO_REQUIRE_GPU_RESIDENT=1turns into a preflight failure. - [6] Advertised per
ggml_backend_dev_supports_op, never by a backend-name comparison. - [7] CPU is always native and serves as the Loop IR oracle; it never runs a generated kernel.
- [8] Per-kernel policy in the RIR registry (
crates/rir-kernels/src/integration.rs): a generated kernel may replace a native one, and two families (L2_NORM_BACK,RMS_NORM_BACK) have had theirs retired from the fork outright. - [9] Same registry, one gate more: a CUDA policy above native-only has to be a line in
CUDA_ADMITTED(crates/rir-gen/src/validate.rs), with its measurement. Two ops are promoted there today,L2_NORM_BACKandSCALE; the rest stay native or in observe. - [10] Single device (
CUDA0); no sharding, NCCL or peer-copy. - [11] CPU-only features, fast Rust and Python lanes. A manual row is run by hand on the reference hardware (§Release criteria).
Recurrent and hybrid model families train correctly everywhere; what varies per family is whether they can pack several sequences into one micro-batch row, which is a throughput property, not a correctness one, and is reported per model at load time rather than declared here.
Release criteria
An artifact is releasable when, for the backends it claims:
- the fast lanes and the lane of every touched subsystem pass (
docs/engineering/tests/notice.md); - a claimed CUDA build pins
RETRO_CUDA_ARCHITECTURESto an explicit list - the defaultnativeproduces a binary that only loads on the build machine's GPU generation, and the build warns about it in release; - every manual row above has been run by hand on the reference hardware - GitHub Actions runs the CPU fast lanes only - and the result is recorded (
docs/engineering/cuda/STATUS.mdfor the CUDA lane).
