Skip to content

RetrogradFine-tune GGUF models directly

Train a LoRA adapter or the model's own weights with SFT, preference optimization, PPO, GRPO, agentic GRPO or distillation. One TOML file, one binary, no Python stack.

Retrograd

How it works ​

Retrograd loads a GGUF model as is, with no conversion, and trains it with ggml on CPU, Metal, Vulkan or CUDA. By default it trains a LoRA adapter, written as a standard GGUF that llama.cpp loads with --lora. It can also train some or all of the model's weights.

text
retrograd inspect --model base.gguf    # check the model and pick LoRA targets
retrograd train run.toml               # train
retrograd bench run.toml --adapter adapter.gguf   # compare base and adapter
retrograd chat run.toml --compare      # try it interactively
retrograd serve run.toml               # serve it through the OpenAI chat API

Start with the quickstart.

Retrograd documentation