Definition
Reinforcement learning with verifiable rewards trains a policy using rewards computed by a task verifier. The verifier might compare a mathematical answer with a known result, run generated code against tests, or check an instruction constraint. Training updates the policy to make higher-reward responses more likely.
Verifiable describes the reward procedure. A checker can establish only the properties it tests. Passing a finite test suite gives evidence about those tests; it cannot establish every property of an arbitrary program.
Origin and attribution
Ai2's Tülu 3 team introduced the named method on November 21, 2024. Nathan Lambert and coauthors documented it in the Tülu 3 report, replacing a learned reward model with a verification function in their training recipe. They acknowledge earlier work using checked outcomes. The documented introduction concerns this named recipe; it does not establish a sole inventor of learning from automatically verified rewards.
DeepSeek's January 2025 R1 report provides another implementation with rule-based accuracy and format rewards. The reward source and the optimization algorithm are separate choices.
Disagreement about capability gains
Yang Yue and coauthors' April 2025 study found that the RLVR setups they tested improved success with few samples but did not exceed the base model's answer coverage at large sample budgets. Mingjie Liu and coauthors' May 2025 ProRL study reported gains beyond base-model coverage after prolonged training with different tasks and controls.
The disagreement concerns capability gains under different experimental conditions. A capability claim needs the base model, training conditions, evaluation tasks, and sampling budget used in the comparison.
Operational significance
Design the verifier alongside the reward. Weak checks can reward a response that exploits the test, satisfies a formatting rule, or passes exposed examples while missing the intended goal. Evaluate held-out cases and inspect errors that the training verifier cannot detect.
A factory using an RLVR-trained model still needs runtime validation and permission checks. The training procedure supplies no authorization for a deployment, payment, or other external action.
Distinguish it from nearby terms
- Reinforcement learning is the broader family of policy optimization methods. RLVR identifies a source of reward within that family.
- Reinforcement learning from human feedback uses human judgments or preferences, often through a learned reward model. RLVR computes its task rewards through verification.
- An optimizer specifies how parameters change. PPO and GRPO are optimization choices; RLVR specifies where the reward comes from.
- A verification gate checks whether work may proceed during execution. Without policy training, that gate is not RLVR.
Check your understanding
A coding model trains against ten public tests, then passes those tests in production. What additional evidence would show that it solves the underlying task, and which failures could the reward procedure have missed?