One network, two kinds of token
Reka AI released Rho-1 on 5 October as a research preview. It is a 19-billion-parameter model that takes in and produces text, images, video and robot actions inside a single network, rather than routing each job to a specialist system behind a shared interface.
The design Reka describes is deliberately symmetric: everything entering or leaving the network is one of two formats. Text and reasoning are discrete tokens. Images, video, robot actions and proprioception — a robot’s own sense of where its joints are — are continuous tokens. The same weights that predict the next frame of video also produce the next action for a robot arm.
Robot training data is the usual bottleneck for this kind of work, because teleoperated footage is expensive to collect. Reka’s answer is an inverse dynamics model that recovers control signals from ordinary internet video, which turns a large and cheap corpus into something an action model can learn from.

What it does, and what it does not do yet
Reka published performance figures rather than benchmark scores. The base model generates video at a median of 0.79 times real time, with roughly six seconds to the first frame. A distilled variant cuts the denoising process from 99 steps to eight. Prompted edits to a running video land in between 1.11 and 2.01 seconds. Native resolution is 672 by 384.
The company is unusually direct about the limits. Structure drifts in rollouts past about 30 seconds. Temporal grounding is weak, so the model loses track of objects across a video. Targeted visual editing is brittle. The native resolution is below what specialised video systems produce. None of these figures has been checked by anyone outside Reka.

The compute number is the argument
Rho-1 was trained on 320 Nvidia H100 GPUs over about three months. That is a small fraction of what a frontier text model consumes, and it is the part of the announcement aimed at anyone deciding whether unified multimodal architectures are a research curiosity or a practical option.
Reka is not releasing weights. Access is a research preview, and the company has asked anyone working on embodied robotics, interactive simulation or vision-action systems to get in touch directly.
The useful test will be whether the single-network claim survives contact with a real robot. A model that drives an arm from the same weights that render a video is a clean idea; whether it is better than a stack of specialised models wired together is a question the preview does not answer.