Black Sheep AI
Per-tensor sensitivity
← All research
RAM Research

“Protect Attention” Is the Wrong Rule for Mixed-Precision Quantization

October 2026 · Black Sheep AI Research

A common rule when quantizing a model is to keep attention at high precision and squeeze the MLP and expert weights. We tested that rule against per-tensor allocation from a single probe pass. The rule produced a bigger model with worse perplexity. Then we tried the opposite rule, and it was worse again. Sensitivity turns out to belong to individual tensors, not to whole classes of them.

The head-to-head

On Qwen3.5-35B-A3B we built a GPTQ-style “protect attention” model: every attention and shared-expert tensor left at 16-bit, everything else at 4-bit, group size 128. All builds use plain round-to-nearest, so the only thing that differs is where the bits go. WikiText-2 test, 128 sequences of 2,048 tokens, median perplexity:

Build Size (GB) Median PPL vs BF16
BF1669.36.494–
RAM (probe allocation)19.56.585+1.4%
Protect attention (GPTQ-style)20.56.697+3.1%
Uniform 4-bit, group 6417.46.764+4.2%
Uniform 4-bit, group 12817.26.928+6.7%

The protect-attention build is a gigabyte larger and loses more than twice as much quality relative to BF16. It pays for every attention tensor, including the many that don’t need the help, and leaves all the routed experts at a flat 4-bit.

Where the probes actually put the bits

The RAM manifest for the same model looks nothing like a class rule. It keeps 93 of 260 attention tensors at 16-bit, but it also puts 22 attention tensors at 2-bit. All 40 MoE routers stay at 16-bit. Most shared-expert tensors get 8-bit. The routed expert stacks, which hold 32.2 of the model’s 33.6 billion scored parameters, are spread across 2, 3, 4 and 5 bits.

The sensitivity ranking behind that is consistent across the MoE models we tested. On Qwen3.5-35B-A3B the 20 most sensitive tensors are 16 attention projections (mainly the output and value projections of the late full-attention layers) and 4 routers. By median percentile, routers sit at the 90th, attention at the 65th, shared experts at the 46th and routed expert stacks at the 34th. GLM-4.7-Flash and Qwen3-30B-A3B show the same ordering.

On the dense Qwen3-8B the order flips: 14 of the 20 most sensitive tensors are MLP projections, and MLP sits at the 61st percentile against the 42nd for attention. A rule tuned on one architecture is backwards on the other.

The opposite rule loses too

If attention isn’t special, maybe the MLP and experts are. We built two forced variants at the same 19.5 GB and added MMLU (5-shot, full test set), because perplexity alone hides a lot:

Build at 19.5 GB Median PPL MMLU
RAM allocation6.5850.657
G: every attention tensor at 8-bit, rest reduced to fit6.5980.645
H: every MLP/expert tensor at 4-bit, attention upgraded6.6220.589

Config H is the worst on both, and its MMLU drop is large: almost seven points below RAM. Taking bytes away from the expert tensors costs far more factual recall than taking them from attention. Config G is closer, but it is still worse on both metrics, because adding bytes to every attention tensor helps less than adding them to the specific tensors the probes flag.

One probe pass, any memory budget

RAM’s per-tensor scores don’t depend on the target size. A knapsack solver turns them into an allocation for whatever byte budget you give it, so one 188-second probe pass on Qwen3.5-35B-A3B produced all of these builds:

Build Size (GB) Median PPL MMLU
BF1669.36.4940.721
Uniform 8-bit34.36.5170.720
RAM, 34 GB target33.66.5040.731
RAM, 25 GB target22.96.5840.713
RAM, 23 GB target21.26.5250.683
RAM, 21 GB target19.56.5850.657
Uniform 4-bit, group 6417.46.7640.704

At 33.6 GB, under half the BF16 size, the RAM build scores 0.731 on MMLU, a point above BF16 (about two standard errors) and slightly better on perplexity than uniform 8-bit at a larger size. GPTQ and AWQ would need a separate calibration run for each of those target sizes.

The lower rows show the catch. Every RAM build beats uniform 4-bit on perplexity, but the 21.2 and 19.5 GB builds lose to it on MMLU. The 22.9 GB build, about a third of the BF16 size at 4.86 average bits, is the smallest that wins on both. The manifests explain why. At tight budgets the perplexity-optimal allocation pushes routed expert stacks down to 3-bit. Next-token prediction on ordinary text tolerates that because any one expert fires rarely. MMLU asks about every knowledge domain, so it hits every expert. Above roughly 23 GB the budget keeps the experts at 4 to 5 bits and the problem goes away.

So if you care about knowledge-heavy tasks and you’re below about a third of the BF16 size, uniform 4-bit’s even spread of precision across experts may serve you better. Above that, the per-tensor build is ahead on both metrics.

Across seven models

Median WikiText-2 perplexity against uniform 4-bit builds, all on one Mac Studio. Uniform 4-bit comes in one fixed size per model, so the sizes don’t always line up; the paper flags the rows that aren’t size-matched, and we flag them here too:

Model Architecture RAM (GB) Uniform 4-bit (GB) Change Probe time
Qwen3.5-35B-A3BMoE, 256 experts19.517.2−4.9%188 s
MiniMax-M2.5MoE, 256 experts, FP8112.3119.8−3.5%301 s
Llama-4-ScoutMoE, 16 experts56.056.9−5.0%214 s
GLM-4.7-FlashMoE + MLA16.2community build−13.6%99 s
Qwen3-30B-A3BMoE, 128 experts16.8community build−3.1%*53 s
Qwen3.5-122B-A10BMoE, 128 experts107.660.4−4.9%†345 s
Qwen3-8Bdense7.34.3−3.5%†36 s

* At this tight budget the allocator admitted no upgrades, so the “RAM” build is uniform 4-bit at group size 64. The gain is a group-size effect, not mixed precision. † The RAM build is much larger than the uniform one, so this compares the budget a practitioner chose rather than the allocator.

The biggest gain, on GLM-4.7-Flash, comes from handling its Multi-head Latent Attention properly. Random probes in the compressed latent space spend their energy on directions the compression path never produces. Passing them through the block’s own compression weights first fixed that: +2.7% against BF16 instead of +10.4% for the best plain-probe build. Part of that gap is size (16.2 GB against 14.6 GB).

The analysis is fast because it never runs text through the model. The probe pass took 36 seconds on the 8B model, under six minutes on the 250 GB Qwen3.5-122B-A10B, and 539 seconds on Llama-4-Maverick at 803 GB, loading one layer at a time with a peak of about 109 GB on a 192 GB machine.

Where it lost

On the small dense model the allocator has much less room to work. We asked for a 4 GB Qwen3-8B. Both builds came out at 5.7 to 5.8 GB, because our 10% size reserve badly under-estimated the unquantized embedding and output head (two 151,936 × 4,096 matrices, 2.5 GB together at 16-bit). Worse, both were beaten by plain uniform 4-bit at 4.3 GB, which reached a median of 9.982 against 10.299 and 10.494. On this model the 3-bit MLP tensors the solver liked cost more perplexity than the bytes they saved, which is why RAM now raises the minimum to 4 bits at tight budgets on dense and smaller-MoE models.

At a 6 GB target (7.3 GB actual, 48% of BF16) the same model matches BF16 within 0.2%. But calibration-based AWQ reaches 9.746 at 4.4 GB, a size RAM can’t match on this model. When a small dense model has no tolerant class of tensors to take bytes from, calibration data captures something random probes don’t.

What we’d tell someone quantizing an MoE model

Don’t allocate by tensor class. Measure each tensor, because the sensitive ones are a scattered minority: a third of the attention tensors, all the routers, and a handful of late-layer projections, not “attention” as a block. Check MMLU or another knowledge-heavy task before trusting a perplexity win at aggressive budgets, because a third of the BF16 size is roughly where the two metrics part ways. And report median perplexity alongside the mean: on GLM-4.7-Flash five sequences on which even the BF16 model is near-random lift its mean from 8.47 to 11.52.


All results from one Mac Studio (M2 Ultra, 192 GB) with MLX. Perplexity: WikiText-2 test, 128 sequences of 2,048 tokens, seed 42, median of per-sequence perplexity. MMLU: 5-shot, full test set, lm-evaluation-harness, standard errors about 0.35 points. Probes: 50 propagated sequences of 8 positions, RTN at group size 64, candidate widths 2/3/4/5/6/8. Uniform baselines for MiniMax-M2.5 and Llama-4-Scout were evaluated on 256 sequences.

Read the Full Paper

The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

arxiv.org/abs/2609.33923

I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0

Related articles

SEE ALL