What the model does

Black Forest Labs, the German company best known for its image generators, has released FLUX 3 Action, a seven-billion-parameter model for controlling robots, with open weights and a model card published on Thursday. The weights are on Hugging Face.

The company calls it a world action model. It takes camera frames, the robot’s current state and a text instruction, and returns the next chunk of actions — denoised jointly with the next video frames, so the model is predicting what it will see as well as what it will do. Its action space is fifty-dimensional for robots, covering end-effector pose, hand articulation and gripper state, with a separate sixty-four-dimensional space for mouse and keyboard.

The benchmark claim, and who made it

Black Forest Labs reports 42.92 per cent overall success on RoboLab-120, Nvidia’s simulated manipulation benchmark. On the company’s own table that is 6.1 percentage points ahead of Cosmos 3 Nano, the previous best open model on the public leaderboard, which has 16 billion parameters against FLUX 3 Action’s seven. The vision-language-action model π0.5 is listed at 28.0 per cent.

A worker moving boxes on a warehouse shelf
RoboLab-120 measures simulated manipulation tasks. Stock photograph, not from the benchmark. Alexander Isreb · pexels · Pexels License

These are vendor-reported figures on a vendor-selected benchmark, and the model card is not consistent to the decimal: the guidance-distilled variant is quoted at 42.2 and 42.24 per cent in the text while the comparison table carries 42.92. The size and speed claims are easier to check than the score — the company puts latency at 41.06 milliseconds per action chunk at FP8 on an Nvidia B200, and says the model runs between 1.34 and 3.95 times faster than the models it compares itself to.

One number that did not come from the vendor

The most useful figure in the release is a third-party one. Positronic Robotics tested the model on a Franka arm across ten tasks from the DROID dataset and reported 28 successes in 30 attempts, or 93.3 per cent, against 90 per cent for Cosmos 3 Nano, 66.7 per cent for DreamZero and 43.3 per cent for π0.5. That is a small sample on one arm, but it is a physical one rather than a simulated one.

A rack of servers with green status lights in a machine room
The model is small enough to run on the robot rather than in a data centre. Stock photograph. panumas nikhomkhai · pexels · Pexels License

The licence, and what to watch

The release is open weights rather than open source: Black Forest Labs publishes the weights and offers separate commercial terms, which is the same structure it has used for its image models. That distinction decides who can actually deploy this, and it is the first thing a robotics team should read.

The number to watch next is independent evaluation at scale. A seven-billion-parameter model that runs at forty milliseconds a chunk is small enough to sit on a robot rather than in a data centre, and that, more than the leaderboard position, is what would change how these systems get built.