By Evan Vega
Twenty-four gigabytes. 2.24x. 2.1 points. Those three numbers came out of the local AI world inside five weeks, and together they describe something that used to need a data centre. Qwen 3.8 27B, a native vision-language model that holds a quarter of a million tokens, now runs in the memory of a laptop. It got twice as fast, with its output distribution mathematically unchanged. And its ability to refuse was removed with linear algebra.
Two of those numbers are quoted everywhere. The third is what the removal cost, and exactly one builder published it. This is a report on all three, with every figure attributed to whoever measured it.
In five weeks, Qwen 3.8 27B moved onto a desk, got twice as fast and had its refusals removed. 24 GB: the native vision-language model runs a 262K-token window in laptop memory because 48 of its 64 la
Abliteration identifies the direction a model’s internal state moves along when it refuses and orthogonalises that direction out of the residual stream. Every abliterated build of this model reports the same result: zero refusals, or zero out of a hundred. One build, OBLITERATED V3, published on 25 August by the builder working as OBLITERATUS, also measured what the removal cost. MMLU fell from 84.5% to 82.3%. STEM fell from 81.8% to 78.5%. Humanities moved by one point.
That is a real, bounded cost: not catastrophic, not free, and measurable by anyone willing to run the harness. Four builds publish their refusal number to the decimal; one publishes the price. The table is at What Removing Refusals Costs, the mechanism at One Vector, and the four model cards compared at Reading the Model Cards.
Qwen 3.8 27B has sixty-four layers. Forty-eight are Gated DeltaNet linear-attention layers holding a fixed recurrent state of about 72 MiB in total; only sixteen are full-attention layers that keep a key-value cache growing with every token. At 64 KiB per token, the full 262,144-token window costs 16 GiB of cache. A conventional all-attention model of the same depth would need 64.
At six-bit weights and a normal working context the model runs in about 24 GB; eight-bit is about 30, and the gap is a flat six gigabytes at every context length, all of it weights. On Apple Silicon, decode speed is set by memory bandwidth, so the larger build is also the slower one. The arithmetic, with the runtime tables, is at Why 24 GB Runs a 262K Window.
The model was trained with a multi-token prediction head: it drafts several of its own next tokens and a runtime verifies the whole block in one pass. MTPLX, built for Apple Silicon, commits those tokens through exact rejection sampling with residual correction, so the fast output and the normal output come from the same distribution. It measures 1.6x on a 16 GB M4 Mac mini and 2.24x on an M5 Max. oMLX 0.6.1, released 17 August, measured +34% decode throughput at 16K context with its own Lightning MTP.
The field is far enough into this that an open mlx-vlm issue is hunting a 2.19 ms (+10.6%) regression in MTP verify cycles that nobody has yet localised. Details at Twice as Fast, Same Distribution.
Novel Cognition did not re-run MMLU and did not benchmark on its own hardware; every capability and speed figure here is the named builder’s own, reproduced as published. oMLX 0.6.3’s feature list comes from a user post. Whether the fast speculative-decoding paths work with a third-party abliterated drafter has not been shown by anyone. That page is load-bearing and is at What We Did Not Test.
The full file is at threenumbers.novcog.us.com, including the install instruction that does not work. Primary sources: the Qwen3.8-27B model card, MTPLX, the oMLX 0.6.1 release, and OBLITERATED V3.
Part of the Frontier Watch Series: Read the previous investigation
More Coverage:
→ Read this investigation on North Denver Tribune
→ Coverage from Daily Colorado News