Skip to content

Configuration reference ​

A run is described by one TOML file. [run].algorithm selects the algorithm and its section: [sft], [ppo], [grpo], [distill], [preference], or [agent] for agent_grpo. Unknown keys, and sections that do not apply to the run, are errors. Relative paths resolve against the directory of the TOML file.

The same schema is published as JSON Schema, for an editor to complete and check a run file (with Taplo or Even Better TOML):

bash
retrograd-server config-schema > retrograd.schema.json

A running server serves it at GET /v1/config-schema.

Common sections ​

[run] ​

KeyDefaultDescription
algorithmrequiredsft, ppo, grpo, distill, preference or agent_grpo.
verbosefalseExtra runtime logging.

[model] ​

KeyDefaultDescription
pathrequiredBase GGUF model. Can be given with --model instead.
deviceautoauto (GPU if available, else CPU), cpu, or gpu (fail without a GPU).

[output] ​

KeyDefaultDescription
pathrequiredWhere the result is written.
kindsee belowadapter, trainable or model.
kindContentsAllowed with
adapterA standard LoRA GGUF, loadable by llama.cpp. Default for lora.lora
trainableThe trained base tensors (and the adapter for hybrid). Needs the base model and Retrograd to load. Default otherwise.full, partial, hybrid
modelA complete standalone GGUF with the trained weights.full, partial

[lora] ​

Required when training.trainable is lora (the default) or hybrid; refused otherwise.

KeyDefaultDescription
rank8Adapter rank.
alpha16.0Scale; the effective scale is alpha / rank.
seed42Initialization seed, also used to shuffle SFT data.
targetsq, k, v, o, ffn_up, ffn_down, ffn_gateAliases, tensor patterns (blk.*.attn_q.weight), or auto for architecture-specific targets. retrograd inspect lists the candidates.
dtypef16Adapter storage: f16 or f32.
init_adapternoneStart from an existing adapter GGUF (weights only, no optimizer state). Excludes the other keys and --resume.

[trainable] ​

Selects base tensors for training.trainable = "partial" or "hybrid". Refused for lora and full.

KeyDefaultDescription
layersallall, last:<n>, or <first>..<last> (inclusive).
modulesnoneattn, ffn, a single projection (attn_q), or a tensor pattern.
normsfalseTrain normalization weights.
biasesfalseTrain the biases of the selected modules.
output_headfalseTrain the output projection. Not possible when it shares storage with the input embedding.

At least one selector must match something. hybrid allows only norms and biases next to the adapter. Quantized tensors, the input embedding and rotary constants are never trained; the run lists what it excluded at start-up.

[training] ​

KeyDefaultDescription
trainableloralora, full, partial or hybrid. See SFT.
optimizeradamwadamw, sgd, muon or gefen.
lr0.0001Learning rate.
lr_schedulerconstantconstant, linear or cosine. linear and cosine decay to zero by the end of the run.
warmup_steps0Warm-up steps.
weight_decay0.0Decoupled weight decay, in [0, 1].
max_grad_norm1.0Global gradient-norm clipping.
epochs1Passes over the SFT dataset. Rollout algorithms use their own counters.
ctx128Context length in tokens.
micro_batch32Tokens per forward/backward pass. The main memory knob.
gradient_accumulation1Passes per optimizer step; micro_batch × gradient_accumulation must divide ctx. Fixed to ctx / micro_batch for rollout algorithms: leave it out.
threadsautomaticCPU threads. Leave it out for automatic; 0 is refused.
generation_concurrencyderivedAnswers generated at once (GRPO, agentic GRPO, distillation). Lower to save memory.
generation_batchderivedGeneration batch size.
shared_prefix_fanoutautoPacking of sequences that share a prompt (GRPO answers, preference pairs): auto, off, max, or an integer ≥ 2.
fast_sampling_contexttrueF16 KV cache and flash attention for generation. false for bit-exact sampling.
kv_dtypef16KV cache type for training: f16 or f32. Falls back to F32 when the device requires it.
gradient_checkpointingfalseRecompute activations to save memory.
checkpoint_every_n_layers4Checkpoint stride when gradient checkpointing is on.
checkpoint_dtypef32Stored activation precision: f32, f16 or bf16. Needs gradient checkpointing.
chunked_cross_entropytrueCompute the loss in vocabulary tiles to save memory (rollout algorithms).
chunked_ce_tiles8Number of vocabulary tiles.
chunked_ce_seq_chunk512Tokens per tile chunk; 0 for all at once.
master_weightsautoF32 master copy for F16/BF16 base weights: auto, f32 or off. Costs 4 bytes per trained parameter.
require_gpu_residentfalseFail if any training operation would run on the CPU.
max_gpu_duty_cycle1.0Cap on the share of time spent on the GPU, in (0, 1]. See Performance.

Optimizers. AdamW and SGD can update F16 and BF16 weights. Muon and Gefen update F32 weights only, so they need lora.dtype = "f32" for an adapter. The last two read their settings from [optimizer.<name>]; parameters they do not handle are updated with AdamW.

[optimizer.muon] ​

Muon applies to 2-D hidden weight matrices; everything else, including LoRA factors, uses AdamW at fallback_learning_rate.

KeyDefaultDescription
momentum0.95Momentum coefficient, in [0, 1).
nesterovtrueNesterov momentum.
ns_steps5Newton-Schulz iterations, from 1 to 32.
ns_epsilon1e-7Normalization epsilon.
fallback_learning_rate0.001AdamW learning rate for parameters Muon does not handle.

[optimizer.gefen] ​

Experimental. Block-wise second moments to reduce optimizer memory.

KeyDefaultDescription
variantshared_vshared_v, or quantized_m (8-bit first moment). Checkpoints are not interchangeable between variants.
block_size1024Elements per block; a power of two, at most 2^30.
min_numel4096Smaller parameters use AdamW.
beta10.9First-moment coefficient, in [0, 1).
beta20.999Second-moment coefficient, in [0, 1).
eps1e-8Update epsilon.

codebook = "uniform", codebook_levels = 256 and partition = "fixed" are the only accepted values of the remaining keys.

[evaluation] ​

KeyDefaultDescription
datarequiredHeld-out file, in the training data format.
every_iterations1Evaluate every N epochs (SFT) or updates (rollout algorithms).
max_examplesallRollout algorithms: cap on evaluated prompts, spread evenly across the file.
patiencenoneStop after N evaluations without improvement.
min_delta0.0Smallest change counted as an improvement.

[checkpoint] ​

KeyDefaultDescription
directoryrequiredCheckpoint directory.
moderequiredsteps, best_eval or steps_and_best_eval. The last two need [evaluation].
every_stepsrequired for step modesSave every N optimizer steps.
resume_fromnoneCheckpoint to resume from; same as train --resume.

See Checkpoints.

[metrics] ​

KeyDefaultDescription
tensorboard_dirnoneTensorBoard log directory.
wandb_export_dirnoneOffline export for Weights & Biases.

[observe] ​

PPO, GRPO and agentic GRPO. See Observing rollouts.

KeyDefaultDescription
directoryrequiredOutput directory.
every1Record rollouts every N updates.
max_text_chars0Truncate texts to N characters; 0 keeps them whole.
enabledtruefalse keeps the table and exports nothing.

[reference] ​

The frozen model used by the KL penalty (kl_coefficient > 0) in GRPO, agentic GRPO and on-policy distillation, and as the reference of a dpo or ipo preference run. Without it, the reference is the base model without its adapter, which only works while base weights are frozen: a run that trains base weights with a KL penalty requires this section.

KeyDefaultDescription
modelrequiredReference GGUF. Must share the model's tokenizer.
ctxtraining.ctxContext length; at least training.ctx.

[sft] ​

KeyDefaultDescription
datarequiredTraining file.
data_formatinferredtext or jsonl. See Datasets.
shuffletrueShuffle rows at each epoch.
template_variables{}Chat template variables, e.g. { enable_thinking = false }. Use the ones of the agentic run a warm-start precedes, or the format it teaches is not the one sampled.

[ppo] ​

KeyDefaultDescription
promptsrequiredChat JSONL prompts.
reward_commandrequiredReward program, as an argument list.
reward_modepersistentpersistent or oneshot. See reward program.
reward_timeout_seconds300Deadline for one batch of rewards.
updatesrequiredNumber of rollout batches.
rollout_batch_sizerequiredAnswers per update.
ppo_epochsrequiredOptimizer passes per batch.
clip_rangerequiredIn (0, 1), typically 0.2.
kl_coefficientrequiredPenalty toward the policy that generated the batch; ≥ 0.

[ppo.critic]:

KeyDefaultDescription
enabledtrueUse a value head and GAE advantages.
gamma1.0Discount, in (0, 1].
gae_lambda0.95GAE λ, in [0, 1].
value_lr0.01Value head learning rate.
value_epochs8Value head passes per update.
feature_dtypef32Stored feature precision: f32, f16 or bf16.

[ppo.sampling] (all required): temperature, top_p (in (0, 1]), max_new_tokens, seed.

[grpo] ​

KeyDefaultDescription
promptsrequiredChat JSONL prompts.
reward_commandrequiredReward program, as an argument list.
reward_modepersistentpersistent or oneshot.
reward_timeout_seconds300Deadline for one batch of rewards.
updatesrequiredNumber of updates.
prompts_per_updaterequiredPrompts per update.
group_sizerequiredAnswers per prompt, 2 to 256.
grpo_epochsrequiredOptimizer passes per update.
clip_range_lowrequiredLower clip, in (0, 1), e.g. 0.2.
clip_range_highrequiredUpper clip, at least the lower one, e.g. 0.28.
importance_sampling_leveltokentoken or sequence (GSPO).
kl_coefficientrequiredKL penalty toward the reference; 0 disables it.
mask_truncatedfalseIgnore answers cut at max_new_tokens.
baselinemeanmean or leave_one_out.
prompt_ordersequentialsequential or shuffled.
max_stalled_updates25Stop after N updates in a row without signal; 0 never stops.
judge_weightrequired with a judgeWeight of the judge's verdict.
judge_failuredrop_groupdrop_group or fail.
max_judge_dropped_fraction0.5Largest share of groups an update may lose to judge failures.

[grpo.sampling] (all required): temperature = 1.0, top_p = 1.0 (both enforced), max_new_tokens, seed.

Optional tables:

TableKeysDescription
[grpo.overlong_penalty]buffer_tokens, max_penaltyPenalty over the last buffer_tokens of the budget.
[grpo.kl_schedule]warmup_updates, targetKL warm-up and adaptive target; needs kl_coefficient > 0.
[grpo.dynamic_sampling]max_resample_factorReplace zero-signal groups, up to this multiple of prompts_per_update (≥ 2).
[grpo.judge]see JudgeLLM or command judge.

[distill] ​

KeyDefaultDescription
modeon_policyon_policy or topk_offline.
teacher_pathrequiredTeacher GGUF, with the student's tokenizer.

On-policy only:

KeyDefaultDescription
promptsrequiredChat JSONL prompts.
updatesrequiredNumber of updates.
prompts_per_updaterequiredPrompts per update.
samples_per_prompt1Answers per prompt, 1 to 256.
distill_epochs1Optimizer passes per update.
clip_range_low / clip_range_high0.2 / 0.28Used only when distill_epochs > 1.
weight_clip5.0Cap on a single token's weight, in nats.
kl_coefficient0.0Extra KL penalty toward the base model.
mask_truncatedfalseIgnore answers cut at max_new_tokens.
prompt_ordersequentialsequential or shuffled.
[distill.sampling]requiredAs [grpo.sampling].

Offline only:

KeyDefaultDescription
datarequiredChat JSONL corpus.
sidecarrequiredThe .topk file written by retrograd distill-teacher.
offline_epochs1Passes over the corpus.

Keys of one mode are refused in the other.

[preference] ​

KeyDefaultDescription
datarequiredPreference JSONL: prompt, chosen, rejected per line.
lossdpodpo, ipo, simpo or orpo.
betaper loss0.1, or 2.0 for simpo. For orpo, the weight of the odds-ratio term.
referenceinitialdpo and ipo only: initial or base. Omit it when [reference] names the model.
label_smoothing0.0dpo only, in [0, 0.5).
gamma_beta_ratio0.5simpo only: the target margin.
shuffletrueShuffle pairs at each epoch, seeded from lora.seed.
pairs_per_stepnoneMost pairs per optimizer step.
logps_drop_warn2.0Nats the chosen answers may lose, with the rejected ones, before a warning.

A key the chosen loss does not read is an error. training.epochs counts the passes over the pairs; training.gradient_accumulation is pinned to ctx / micro_batch. See Preference optimization.

Agentic GRPO ​

run.algorithm = "agent_grpo" reads [agent]. See the agentic GRPO guide. A run needs [agent.judge], [agent.environment], or both.

[agent] ​

KeyDefaultDescription
scenariosrequiredScenario JSONL file.
updates1Number of updates.
scenarios_per_update1Scenarios per update.
group_size8Trajectories per scenario (≥ 2).
epochs_per_update4Optimizer passes per update.
max_turns6Assistant turns per trajectory.
max_new_tokens_per_turn512Tokens per turn.
max_trajectory_tokensmodel contextToken budget for a whole trajectory.
max_rollout_secs300Time budget per trajectory; 0 disables it.
end_on_no_tool_calltrueEnd the trajectory on a turn without a tool call.
max_failed_turns0Cut a trajectory after N turns in a row without a valid call; 0 never cuts.
truncationdropdrop or min_reward for trajectories that hit a limit.
max_dropped_fraction0.5Stop when more than this share of an update is lost.
skip_empty_updatesfalseSkip, rather than fail, an update with fewer than two usable trajectories.
judge_failuredrop_groupdrop_group or fail.
drop_degenerate_groupsfalseDrop groups the judge scored identically.
clip_range_low / clip_range_high0.2 / 0.28Clip range.
importance_sampling_leveltokentoken or sequence (GSPO).
kl_coefficient0.0KL penalty toward the reference.
system_suffix""Text appended to every system message.
template_variables{}Chat template variables, e.g. { enable_thinking = false }.
seed42Seed.
mcp_confignonePath, or list of paths, to mcp.json files.

[[agent.mcp_servers]] ​

KeyDefaultDescription
namerequiredServer name.
command / urlone requiredLocal server command (with optional env), or remote URL (with optional headers).
allowed_tools / denied_toolsall / noneTool name filters; * is a wildcard and deny wins.
requiredtrueFail the run if the server cannot connect.
statelessfalseMust be true when an environment is declared.
tool_timeout_secs30Per-call timeout.
max_tool_result_bytes65536Cap on a tool result.
cwdinheritedWorking directory of a local server.
env_passthroughallNames of the environment variables a local server inherits; PATH and HOME always are.

[agent.environment] ​

type is container, http or local.

KeyDefaultDescription
containerRequires the container build feature.
profilepythonDefault image and package cache: python, typescript or custom.
imageprofile'sContainer image; prefer a digest.
toolsrequiredToolsets, see below.
allow_networkfalseAllow outbound network.
cache_volumenoneVolume mounted read-only on the profile's package cache, so reuse = "workspace" installs once per run.
setup_timeout_secs300Timeout for a scenario's setup.
verify_timeout_secsnoneTimeout for a scenario's verify.
limitscpus = 1.0, memory_mb = 1024, pids = 256, exec_timeout_secs = 30, max_output_bytes = 65536.
poolmax_live = 8, min_idle = 0, reuse = "never" (or "workspace"), max_leases_per_container = 32.
http
base_urlrequiredEnvironment server URL.
request_timeout_secs60Timeout per request.
connect_timeout_secs10Connection timeout.
pool_size16Concurrent connections; keep at least group_size.
max_result_bytes65536Cap on an observation.
headersnoneHeaders added to every request.
localRuns tools on the host, without isolation.
allow_unsandboxedfalseMust be true.
tools, setup_timeout_secs, verify_timeout_secsAs for containers.

[agent.environment.tools] ​

For container and local. See Tools and toolsets.

KeyDefaultDescription
defaultrequiredToolset of a scenario that names none (base, python, typescript, or one you define).
scenario_toolsetsnoneOther toolsets a scenario may select with metadata.env.toolset.
filesnoneDefinition files ([[tool]], [toolset.NAME]), relative to this config.
[[...tools.tool]]noneInline tool definitions: id, version, name, description, input_schema, and builtin or exec.
[...tools.toolset.NAME]noneInline toolsets: include, tools (id or id@version), deny.

[agent.collect_api] ​

The OpenAI-compatible endpoint retrograd collect --api generates traces with. A training run ignores it.

KeyDefaultDescription
base_urlrequiredEndpoint, e.g. https://api.openai.com/v1.
modelrequiredModel name sent with each request.
api_key_envrequiredEnvironment variable holding the key; the key itself never goes in the file.
timeout_secs120Timeout per request.
temperatureprovider'sSampling temperature.

[agent.scenario_generation] ​

Used by retrograd scenarios generate. model, base_url and api_key_env are required; count (24), batch_size (12), min_difficulty / max_difficulty (1 / 5, between 1 and 5), custom_instructions, seed, shuffle (true), timeout_secs (120), max_retries (2) and max_catalog_bytes (262144) are optional.

Judge ​

[grpo.judge] and [agent.judge] share the same format.

type = "command": command (required, argument list), timeout_secs (30).

type = "ruler", an OpenAI-compatible LLM judge:

KeyDefaultDescription
base_url, modelrequiredEndpoint and model.
api_key_envOPENAI_API_KEYEnvironment variable holding the API key.
rubricbuilt-inGrading criteria. A prompt or scenario can set its own.
pairwise_rubricbuilt-inCriteria for pairwise comparisons.
temperatureendpoint default0.0 for the most repeatable verdicts.
max_concurrency4Parallel requests.
timeout_secs120Request timeout.
max_retries2Retries on transient errors.
cache_pathnoneVerdict cache file.

[<section>.judge.strategy]:

KeyDefaultDescription
modeautoauto (one request per group, split into chunks if too long), listwise (always one request), chunked (always split), pairwise (two answers per request, most reliable, most requests).
anchortrueChunked: repeat one answer in every chunk to keep scores on one scale.
max_pairsall pairsPairwise: comparisons per group.
both_orderstruePairwise: judge each pair in both orders to cancel position bias (doubles requests).
aggregationwin_ratePairwise: win_rate or bradley_terry (better when max_pairs limits comparisons).

[<section>.judge.context], budgets in characters: max_request_chars (60000), max_trajectory_chars (8000), max_message_chars (2000), head_ratio (0.4, share of an elided message kept from its start), include_env_state (true, show the environment's final state to the judge).

[<section>.judge.compaction], optional: summarize the middle of long transcripts with the judge model. trigger_chars, target_chars and keep_last are required.

Retrograd documentation