Reinforcement learning with rewards you can check

dragon_reinforce() trains a model with Group Relative Policy Optimization (GRPO), the method behind the recent reasoning models, in its plain on-policy form. For every prompt the model writes several answers, each is scored by rewards you define, and the model is nudged towards the answers that beat their group’s average. A KL penalty against the model it started from keeps it from drifting.

When it helps, and when it does not

Reinforcement learning on a small model works when the reward is verifiable: something a function can check without opinion.

It works poorly as a substitute for preference data on goals such as “be more helpful” or “sound friendlier”. A reward that is really a model’s opinion is noisy, and a small policy will find the noise. For those goals use dragon_prefer() with judge-ranked pairs instead; see the post-training loop vignette.

Data: prompts and references

Rows carry a prompt and, optionally, a reference answer for rewards that compare against one. Every other column travels along as fields for custom rewards.

library(dragonfarm)

a <- sample(10:999, 400, replace = TRUE)
b <- sample(10:999, 400, replace = TRUE)
math <- data.frame(
  question = sprintf("What is %d + %d? Reply with just the number.", a, b),
  answer = as.character(a + b)
)
prompts <- dragon_map_prompts(dragon_dataset(math), prompt = "question", reference = "answer")
dragon_preview(prompts, n = 1)

Rewards

dragon_reward() builds one reward; a list of them is summed by weight.

rewards <- list(
  dragon_reward("numeric"),                            # last number equals the reference
  dragon_reward("length", max_chars = 12, weight = 0.3) # keep it terse
)

The built-ins: "exact", "contains", "numeric", "regex" (with pattern), "json" (with optional keys, partial credit per key), "length" (with min_chars and max_chars), and "keyword" (with words and mode). For anything else, "custom" points at a Python file:

writeLines(c(
  "def reward(prompt, completion, reference, row):",
  "    # 1.0 when the reply mentions the product named in the row",
  "    return 1.0 if row.get('product', '') and row['product'] in completion else 0.0"
), "product_reward.py")
mention <- dragon_reward("custom", file = "product_reward.py", name = "mentions_product")

The file is copied into the run, so the run stays self-contained and travels with the cloud bundle.

Train

Start from a fine-tuned run when you have one; the RL stage folds its adapters in before adding its own.

rl <- dragon_reinforce(
  prompts, sft,                    # or a model id such as "Qwen/Qwen2.5-0.5B-Instruct"
  rewards = rewards,
  group_size = 6,                  # answers sampled per prompt
  beta = 0.04,                     # KL penalty towards the starting model
  temperature = 1.0,               # sampling temperature; must be above 0
  max_new_tokens = 16,
  args = dragon_train_args(learning_rate = 1e-5, epochs = 1, batch_size = 4, grad_accum = 1, save_steps = 20),
  wait = TRUE
)

A step samples batch_size * group_size completions, scores them, and takes one optimizer step. Sampling dominates the cost, so short max_new_tokens and modest group_size keep steps fast. Progress rows carry reward, reward_std, kl, and completion_len, plus one column per reward:

dragon_progress(rl)[, c("step", "reward", "reward_numeric", "reward_length", "kl", "completion_len")]

The Monitor panel plots mean reward on a second axis next to the loss.

Read the result

dragon_evaluate(rl)   # mean held-out reward, per reward
dragon_generate(rl, "What is 417 + 285? Reply with just the number.", temperature = 0)
dragon_compare(sft, rl)

reward in the comparison table is the mean total reward on held-out prompts. Compare it with the same rewards computed on the starting model to see what the stage bought: dragon_evaluate() on the starting run does not know these rewards, so the quickest check is dragon_generate() on both with base = TRUE on the RL run and a few prompts.

Settings that matter

Chaining with the other stages

Reinforcement learning is a stage like the others. A typical order is fine-tune, then preference optimization for style, then RL for a verifiable target, and each starts from the previous run. In a pipeline:

dragon_pipeline("Qwen/Qwen2.5-0.5B-Instruct", list(
  dragon_step_train(tickets),
  dragon_step_reinforce(prompts, rewards = rewards, group_size = 6, max_new_tokens = 16),
  dragon_step_evaluate()
))