A frontier-scale model with no full attention anywhere in it
NaiveAI, a Beijing lab founded in February that had not previously released a model, has published Naive-N0.5-Flash under an MIT licence. It is a mixture-of-experts model with 309 billion total parameters and 15.5 billion active, and it claims a native context window of one million tokens. What makes it unusual is what is missing: according to the model card, there is no full-attention layer anywhere in the network.
Transformers normally keep at least a few layers that attend to the entire context. Those layers are what preserve long-range information, and at million-token lengths they are also where most of the decoding cost goes. Naive-N0.5-Flash removes them entirely. Its 48 layers are split into eight six-layer modules: five sliding-window attention layers followed by one DeepSeek Sparse Attention layer, with the first layer of the first module also replaced by DSA. That gives 39 sliding-window layers and 9 sparse ones.
The sliding window is 128 tokens. The sparse layers run a lightweight 16-head indexer across the full history, then compute attention over only the top 2,048 tokens it selects. The full key-value cache is still kept and the indexer still scans everything, so the saving is in attention computation and memory traffic rather than in storage.

Built on Xiaomi’s open base model
Naive-N0.5-Flash is not trained from scratch. It builds on Xiaomi’s open-weight MiMo-V2.5 base, which NaiveAI credits in its acknowledgements alongside DeepSeek for the sparse attention design and the SGLang project for inference infrastructure. Adapting the base model to the new attention structure was, the company says, one purpose of continued pretraining: 3.25 trillion tokens in three stages, comprising 50 billion tokens of indexer warm-up, 3 trillion of sparse attention training and 200 billion of learning-rate decay.
That lineage matters for reading the release. This is a sizeable architectural retrofit of somebody else’s base model, released openly, rather than a new pretraining run — and it is being offered as evidence that hybrid sparse-plus-sliding stacks can hold quality out to a million tokens.
NaiveAI also published NaiveRT, its own inference system, which it says combines mega-kernel fusion, Programmatic Dependent Launch and speculative decoding to deliver 50 tokens per second per user in a standard mode and up to 2,000 tokens per second in an “Ultrafast” mode. The weights and inference code are MIT-licensed; a hosted API is priced at $0.10 per million input tokens, $0.40 per million output and $0.01 per million cache reads.
The benchmark numbers are the company’s own
The model card reports results across coding and agentic tasks — SWE-Bench Pro, DeepSWE v1.1, Terminal-Bench 2.1, ALE-CLI, ProgramBench — and a set of AI research-and-development benchmarks including PostTrainBench, MLE-bench-30 and PaperBench. These are NaiveAI’s own evaluations, not independent results, and the company is explicit about the harness: unless noted, it ran the tests through Claude Code 2.1.207 with a one-million-token context, temperature 1.0 and top-p 0.95, exposing only basic file I/O and Bash tools. Competing scores are quoted from other vendors’ published blogs, model cards and public leaderboards rather than re-run.

A lab valued before it had a product
NaiveAI was founded in February 2026 by Dai Jifeng, an associate professor at Tsinghua University who previously worked at Microsoft Research Asia and at the computer-vision company SenseTime. Backed by Tencent, the lab raised about $400m and reached a reported valuation of roughly $1.42bn — a figure that drew attention last week precisely because the company had shipped nothing at the time.
It has now shipped, and the pitch in the model card is the company’s slogan rather than a benchmark: “Building Frontier AI with AI.” NaiveAI says NaiveRT itself was built and optimised through what it calls AI-centred research and development, and points readers to a case study in its technical blog for the process.
The thing to watch is whether the architecture survives contact with other people’s workloads. A model with no full-attention layer is a genuine test of whether a 128-token window plus 2,048 sparsely selected tokens can substitute for global attention at a million-token context, and the MIT licence means outside groups can run that test themselves rather than take the company’s figures for it. The weights require FP8-capable Nvidia hardware and occupy roughly 315GB, so the test will not be a cheap one.