What was released
Alibaba’s Qwen team has published Qwen-Drive-1.0-4B, an open vision-language model for autonomous driving that combines understanding a driving scene with planning where the vehicle goes next. Code, weights and demo data are available under the Apache 2.0 licence, and TechNode reported the release on 7 September. The project was developed with Huazhong University of Science and Technology.
The accompanying paper, posted to arXiv on 31 August, describes it as an initial step toward a vision-language foundation model for driving rather than a finished product.
One model, three jobs
The design choice the authors emphasise is what they did not change. Qwen-Drive is built on Qwen3.5-4B, a natively multimodal vision-language model, and the pretrained VLM architecture is left untouched. The driving capabilities are attached around it.
A bird’s-eye-view perception head handles 3D object detection, semantic occupancy prediction and BEV map segmentation. A separate Planning Expert conditions on the same shared representations to generate the vehicle’s future trajectory. The team ships two planning variants: one trained by imitation from driving examples, one refined with reinforcement learning.

Why keeping the VLM intact matters
Driving stacks have historically been built as pipelines: a perception module produces boxes and lanes, a planner consumes them, and the two share nothing but an interface. That works, and it also means the planner never sees what the perception module discarded.
Unifying them in one representation is the current bet across the field. The risk is that specialising a general model on driving data destroys the general capability that made it useful — the thing that lets it answer “why did you slow down here” in words. The authors’ claim is that their staged training recipe, which mixes driving supervision with general-purpose vision-language data, avoids that: they report strong 3D perception and driving scene understanding while largely preserving general vision-language ability.
The benchmark claims are the authors’ own
The paper reports what it calls highly competitive motion-planning performance across open-loop, pseudo-closed-loop and closed-loop evaluations. That is a claim by the model’s authors about their own model, and it has not yet been independently reproduced. Open weights make that reproduction possible, which is more than can be said for most driving models.

The size is the point
Four billion parameters is small. That is not a limitation being apologised for; it is the reason this release is interesting. A driving model has to run in a vehicle, on a power and latency budget that a data centre model never faces, and 4B under Apache 2.0 is a size that a university lab or a mid-size supplier can actually fine-tune and deploy.
It also continues a pattern. Alibaba has spent two years releasing Qwen weights openly while its Western counterparts publish papers about models nobody else can run, and the effect on who does downstream research has been visible.
What to watch
Independent evaluation on the closed-loop benchmarks, which is where driving models usually stop looking good. And whether any vehicle programme adopts it — an open model that no one puts in a car is a research artefact, not a product.