Context and knowledge

Recursive language model (RLM)

working definition
Machine-readable Download Markdown

Definition

A recursive language model is an inference framework that places a large input in an external environment and lets a language model inspect and decompose it programmatically. The model can select passages, compute over the input, and call language models on smaller questions. Those subcalls can themselves use the framework.

In the original implementation, a Python REPL holds the input as a variable. The model operates on that variable instead of receiving the entire input in one neural context window. Which passages it selects and how it combines sub-results remain part of the inference process.

Origin and attribution

Alex L. Zhang, Tim Kraska, and Omar Khattab introduced Recursive Language Models in December 2025. Credit for the named framework belongs to these authors. Their paper acknowledges earlier work on recursive decomposition.

Version 3, published in May 2026, studies recursion depths from zero through three and includes a small model post-trained for the framework. Using an RLM therefore does not require changing a base model's weights, while training a model specifically for RLM execution remains possible.

Scope and naming

The word model in the name can refer to the inference system around a neural model. Zhang explains this model-versus-scaffold framing in his February 2026 essay. The framework specifies how inference runs around the model; its name alone does not establish a new neural architecture.

External context can exceed the base model's input capacity, but each model call still has a finite context window. Access to a larger document does not establish that the system examined every relevant passage or combined the results correctly.

Operational significance

Budget the whole call tree. The current paper identifies runaway sub-call costs, underexplored guardrails, and limited evaluation on more difficult natural tasks. Its results do not show that deeper recursion always improves performance.

Keep the source passages used by each subcall and the steps that combine their answers. Enforce limits on recursion depth, total calls, elapsed time, and returned data. Code that inspects external context also needs an execution boundary appropriate to its tools and permissions.

Distinguish it from nearby terms

  • A context window limits the input of one model call. An RLM manages a larger input through an external environment and multiple operations.
  • Retrieval-augmented generation supplies selected material to a model. Retrieval can be one operation within an RLM, whose controller can also inspect and compute over its external context.
  • Prompt compression creates a shorter representation. An RLM can revisit the retained external input, subject to its selection procedure.
  • Recursive self-improvement changes the system or its capability over rounds. Recursive inference calls alone do not establish self-improvement.
  • Programmatic tool calling coordinates tools through code. An RLM uses programmatic access specifically to manage context and model subqueries.

Check your understanding

An RLM answers a question about a large repository after inspecting only three files. What would establish that those files cover the relevant evidence, and which records would let you check the subcalls and final synthesis?