What was released

Alibaba’s DAMO Academy published a general-purpose medical imaging model called RADAR in Science on Thursday and released the code and weights on GitHub and Hugging Face under an Apache 2.0 licence at the same time.

RADAR reads contrast-enhanced computed tomography scans of the abdomen. It covers 18 organs and reports on 146 distinct clinical findings, including cancers. DAMO describes it as the first expert-level generalist model for medical imaging — the company’s own characterisation, not an independent assessment.

The reported results

Across nearly 40,000 real-world examinations, the researchers report a mean area under the curve of 0.913 over the 146 findings. AUC measures how well a test separates positive from negative cases; 1.0 would be perfect discrimination.

Two clinicians discussing results in a hospital office
Illustration: in a reader study the model's mean accuracy exceeded that of 23 of 26 radiologists. Los Muertos Crew · pexels · Pexels License

In a reader study against 26 radiologists drawn from several hospitals, the model’s average accuracy exceeded that of 23 of them. The researchers also report that when radiologists were given the model’s prompts, their sensitivity rose by about 10% and their reading time fell by more than 30%.

The training set was 424,911 contrast-enhanced abdominal CT examinations and more than 15 million anatomy-aware image-text pairs.

What the numbers do and do not establish

These are figures reported by the team that built the model, in a peer-reviewed paper. Peer review is a meaningful filter, but it is not the same as external replication on a different population, and a reader study is not a clinical trial. Radiology models have a long history of performing worse on scanners, protocols and patient populations they were not trained on.

A hospital corridor with medical equipment
Illustration: the code and weights are published under an Apache 2.0 licence. RDNE Stock project · pexels · Pexels License

The assistive result — faster reads with higher sensitivity — is arguably more interesting than the head-to-head comparison, because it describes the way such a model would actually be used. A radiologist reading with a prompt is a different clinical arrangement from a model reading alone, and it is the one regulators are more willing to approve.

Why the licence matters

Apache 2.0 with published weights means any hospital, national health system or competing lab can download RADAR, run it on its own scans, and check the claims without asking Alibaba for anything. That is the part of this release with the most consequence: it makes independent verification possible rather than promised, and it puts a capable diagnostic model into the hands of health systems that could not commission one.

What happens next is external evaluation on populations outside the training distribution. Until that exists, the strongest claim available is that a model reported these results on the data its authors used.