dragon_reinforce() trains a model with Group Relative
Policy Optimization (GRPO), the method behind the recent reasoning
models, in its plain on-policy form. For every prompt the model writes
several answers, each is scored by rewards you define, and the model is
nudged towards the answers that beat their group’s average. A KL penalty
against the model it started from keeps it from drifting.
Reinforcement learning on a small model works when the reward is verifiable: something a function can check without opinion.
It works poorly as a substitute for preference data on goals such as
“be more helpful” or “sound friendlier”. A reward that is really a
model’s opinion is noisy, and a small policy will find the noise. For
those goals use dragon_prefer() with judge-ranked pairs
instead; see the post-training loop vignette.
Rows carry a prompt and, optionally, a reference answer for rewards
that compare against one. Every other column travels along as
fields for custom rewards.
library(dragonfarm)
a <- sample(10:999, 400, replace = TRUE)
b <- sample(10:999, 400, replace = TRUE)
math <- data.frame(
question = sprintf("What is %d + %d? Reply with just the number.", a, b),
answer = as.character(a + b)
)
prompts <- dragon_map_prompts(dragon_dataset(math), prompt = "question", reference = "answer")
dragon_preview(prompts, n = 1)dragon_reward() builds one reward; a list of them is
summed by weight.
rewards <- list(
dragon_reward("numeric"), # last number equals the reference
dragon_reward("length", max_chars = 12, weight = 0.3) # keep it terse
)The built-ins: "exact", "contains",
"numeric", "regex" (with
pattern), "json" (with optional
keys, partial credit per key), "length" (with
min_chars and max_chars), and
"keyword" (with words and mode).
For anything else, "custom" points at a Python file:
writeLines(c(
"def reward(prompt, completion, reference, row):",
" # 1.0 when the reply mentions the product named in the row",
" return 1.0 if row.get('product', '') and row['product'] in completion else 0.0"
), "product_reward.py")
mention <- dragon_reward("custom", file = "product_reward.py", name = "mentions_product")The file is copied into the run, so the run stays self-contained and travels with the cloud bundle.
Start from a fine-tuned run when you have one; the RL stage folds its adapters in before adding its own.
rl <- dragon_reinforce(
prompts, sft, # or a model id such as "Qwen/Qwen2.5-0.5B-Instruct"
rewards = rewards,
group_size = 6, # answers sampled per prompt
beta = 0.04, # KL penalty towards the starting model
temperature = 1.0, # sampling temperature; must be above 0
max_new_tokens = 16,
args = dragon_train_args(learning_rate = 1e-5, epochs = 1, batch_size = 4, grad_accum = 1, save_steps = 20),
wait = TRUE
)A step samples batch_size * group_size completions,
scores them, and takes one optimizer step. Sampling dominates the cost,
so short max_new_tokens and modest group_size
keep steps fast. Progress rows carry reward,
reward_std, kl, and
completion_len, plus one column per reward:
dragon_progress(rl)[, c("step", "reward", "reward_numeric", "reward_length", "kl", "completion_len")]The Monitor panel plots mean reward on a second axis next to the loss.
dragon_evaluate(rl) # mean held-out reward, per reward
dragon_generate(rl, "What is 417 + 285? Reply with just the number.", temperature = 0)
dragon_compare(sft, rl)reward in the comparison table is the mean total reward
on held-out prompts. Compare it with the same rewards computed on the
starting model to see what the stage bought:
dragon_evaluate() on the starting run does not know these
rewards, so the quickest check is dragon_generate() on both
with base = TRUE on the RL run and a few prompts.
beta trades speed of change for
stability. Raise it if replies degrade in ways the rewards do not see;
lower it if nothing moves.temperature needs to be high enough
that the group varies. 0.8 to 1.2 is typical.Reinforcement learning is a stage like the others. A typical order is fine-tune, then preference optimization for style, then RL for a verifiable target, and each starts from the previous run. In a pipeline: