No learned control layer between the model and the robot
Researchers at Stanford and Caltech have run a Unitree G1 humanoid through multi-step household tasks in rooms it had never seen, with a general-purpose vision-language model issuing the commands directly. The system, called HomeBody, tidies spaces, fetches objects and opens drawers to retrieve what is inside them.
What makes it worth attention is what the researchers left out. Most humanoid work of the past two years routes perception and language through a trained vision-language-action layer — a model taught, on robot demonstrations, to turn instructions into motion. HomeBody has no such layer. OpenAI’s GPT-6 Astra is plugged in as what the team calls a plug-and-play VLM, and it calls into a library of composable skills without ever being trained on the robot or the room.
The skill library is small: navigate, pick, place, open a drawer, and pick from a drawer. Each is a conventional controller. Astra’s job is to look at the current view and decide which one to call and on what.

How it knows where things are
An unfamiliar room is the hard part, and the team solves it before the task starts rather than during it. The robot explores, and the run is reconstructed into a digital twin in Nvidia’s Isaac Sim. That gives the system a persistent spatial memory, so it can be asked for an object that is not currently in view and go to where it saw one.
Around that sit conventional components: SAM 2.1 for tracking, Fast-FoundationStereo for depth, Super Odometry with ICP registration for localisation, and an AMO policy for coordinating the lower body with whatever the arms are doing. On each step the model picks a skill and a target from the current first-person view, the map, the gripper state and what it recalls seeing. When a grasp misses, it retries locally with visual servoing rather than restarting the task.
The local stack runs on one RTX 4090 laptop GPU. Astra runs remotely.
What the team says does not work yet
The paper is unusually direct about its costs, and they are the interesting part.
Reconstruction has to happen before a new room is usable, which takes time and API spend. Task length is capped by the hardware rather than the software — the robot’s endurance runs out first, and the-decoder notes servos overheating during runs. And because Astra is called remotely for every decision, its latency shows up as visible pauses between actions.

Why it matters
If a frontier model can drive a humanoid through unseen rooms using only hand-written skills, then the trained action layer is a convenience rather than a requirement, and robot-specific demonstration data becomes less of a moat than the field has assumed. That is a claim about architecture, made on a project page from two university labs, not a benchmarked result against competing systems — and it should be read that way until someone reproduces it.
The costs also point at what would change it. Latency and reconstruction are both engineering problems with obvious directions: a smaller model on the robot for the fast loop, a faster reconstruction pass, or an Astra-class model served closer to the hardware. None of those requires a new idea, which is why the result is worth watching rather than filing away.