The bottleneck

Agentic reinforcement learning splits training from rollout: the model trains on one cluster and acts on another, and every policy update has to reach the rollout cluster before the next batch can start. That transfer is dead time.

A paper published on 7 October by researchers at Nvidia and Aalto University puts a number on it. Sending a full checkpoint of a one-trillion-parameter model between two AWS regions takes 87.5 minutes. Nothing trains while that happens.

The observation the method rests on

The authors measured how much of the model actually changes, and report that only about 1% of BF16 weights differ from one training step to the next. The other 99% is being copied across a continent for no reason.

Their system, NeMo-DCR, sends only the values that changed, and does it bit-exactly: the receiving cluster ends up with the same parameter and buffer bits it would have had from a full dense refit. That matters more than it sounds, because an approximate refit introduces drift between the policy being trained and the policy being sampled, which is precisely the thing reinforcement learning is sensitive to.

Shipping containers stacked at a port
The paper reports transport accounts for most of the remaining refit time. Illustrative image. Wolfgang Weiser · pexels · Pexels License

The machinery includes fixed index mappings that place each change at a canonical checkpoint coordinate, a choice between XOR and overwrite encoding, in-place application with retry after a failure, and a joint commit binding each policy version to the baseline it was built from. Transfers go over object storage or a relay tree, with no cross-cluster collective operation.

The figures

All of these are the authors’ own measurements and have not been independently reproduced.

At change rates of 3% and 5%, refits of models from 30 billion to a trillion parameters run 12 to 40 times faster than a transport-only full-checkpoint reference. At 120 billion parameters a refit takes 22.6 to 49.7 seconds against a 750-second reference for a 247.2 GB checkpoint. At a trillion parameters with a 3% change rate, the relay tree takes about 150 seconds instead of 87.5 minutes.

The encoding choices carry their own numbers: XOR masks compress 1.7 to 2.2 times smaller than overwrite values, and a mixed XOR and overwrite scheme cuts payload bytes on Qwen3 by 38 to 40% against overwrites alone.

A close-up of a graphics card circuit board
The evaluation covers models from 30 billion to a trillion parameters. Illustrative image. Armando Are · pexels · Pexels License

What it says about where the time goes

One figure in the paper is worth more than the speedups. In the stress cases, transport accounts for 77 to 94% of relay-tree refit latency — that is, after the compression work is done, almost all the remaining time is just moving bytes.

That tells you the method is close to the floor set by the network rather than by the encoding, and that further gains have to come from sending less or sending it differently, not from compressing harder.

What to watch

Whether the 1% change rate holds for other training regimes, whether anyone reproduces the figures outside Nvidia’s own stack, and whether the method lands in a released version of NeMo. A technique that only works on the hardware and framework of the company that wrote the paper is a product feature rather than a result.