Black Sheep AI
Two sensitivity signals
← All research
RAM Research

Two Sensitivity Signals Disagreed on Every Tensor. They Built the Same Model

October 2026 · Black Sheep AI Research

RAM can score a tensor’s sensitivity to quantization in two ways. On a 35B MoE model the two rankings are uncorrelated: Spearman −0.01 across 390 tensors, and only 7 tensors in common between their top-30 lists. We fed each one to the same allocator at the same budget, expecting two quite different models. Perplexity couldn’t tell them apart. On a dense model, it could.

The two signals

The isolated estimator quantizes a tensor and measures its rounding error relative to its own weight energy. You can compute it exactly, and its behaviour is well understood: one Gaussian probe estimates it to within a few percent. The propagated signal pushes random probes through the network, so each layer gets the inputs the earlier layers produce, and measures how far quantizing one tensor moves the block output, after the nonlinearity and the residual path.

They measure different things, and it shows. On Qwen3.5-35B-A3B at 4-bit:

Tensor set Tensors Rank correlation between signals
All matched tensors390−0.01
Attention projections190−0.43
Shared-expert projections1200.06
Routed expert stacks400.59
Routers400.53

Within attention they disagree in opposite directions. The isolated estimator’s 30 most sensitive tensors are 29 routers and one attention projection. The propagated signal’s are 22 attention projections and 8 routers. Both are right about what they measure. A 256 × 2,048 router really is the worst-rounded tensor per unit of weight energy. Its rounding error just barely moves the layer output under the inputs the network produces, while an attention output projection’s does.

Same allocator, same budget, two signals

We computed the exact isolated score for every tensor at all six candidate bit-widths and gave it to the same greedy solver, at the same 21 GB budget, with the same candidate sets. The only difference between the two builds is the cost signal. Average bits by tensor class:

Signal Attention Routers Shared experts Routed experts Other
Propagated8.516.08.03.811.8
Isolated5.78.08.03.910.6

The isolated allocation spends less on attention (30 attention tensors at 2-bit against 22) and halves the routers to 8-bit, and puts those bytes into the routed expert stacks: 109 of 120 stacks at 4-bit against 74. Half the tensors got the same bit-width in both.

Build at 19.5 GB Median PPL Mean PPL MMLU
Propagated signal6.5446.6220.668
Isolated estimator6.5506.6250.679

On perplexity they are indistinguishable. For scale, simply upgrading the inference runtime (mlx_lm 0.30.4 to 0.31.3) moved the same propagated build from 6.585 to 6.544, about ten times the gap between the two signals. On MMLU the isolated allocation is 1.1 points ahead, about two standard errors, which fits it keeping more expert stacks at 4-bit rather than 3-bit.

Why the disagreement didn’t matter here

Two things absorb it. First, on this model the routed expert stacks hold 96% of the scored parameters, and both signals put them low. Whatever the signals do with the small tensors, the byte budget is decided by the experts. Second, the greedy solver upgrades whichever tensor buys the biggest drop in cost per added byte. That size-normalized rule pushes small tensors up and huge ones down almost regardless of which signal supplies the costs, so the coarse shape of the allocation comes out the same.

On a dense model, one signal wins

Qwen3-8B has no dominant class of huge, tolerant tensors. Its flexible tensors are comparable in size, so the allocation really is decided tensor by tensor. Same experiment, 6 GB target, byte totals within 0.02%:

Signal Median PPL Beats propagated on MMLU
Propagated probe9.630–0.749
HAWQ-V2 (needs gradients and data)9.65663 of 128 sequences0.750
Isolated estimator9.73720 of 128 sequences0.749

Here the isolated estimator is clearly worse, by 0.6% in perplexity and on 108 of 128 sequences. The propagated probe ties HAWQ-V2 (paired p = 0.96), which uses gradients and calibration text. The two data-free signals agree on only 50% of tensor bit-widths, and their 4-bit costs have a rank correlation of 0.04. MMLU doesn’t separate any of them at this budget.

The catch: the propagated signal is noisy

The isolated estimator is precise but narrow. The propagated one is the opposite. We re-ran the propagated pass with 200 probe sequences and measured how much a tensor’s score varies from one sequence to the next:

Model Per-sequence CV Spread across tensors Rank agreement at 50 seq. Median error at 50 seq.
Qwen3.5-9B0.161.06 dex1.0001.3%
Qwen3.5-35B-A3B0.240.86 dex0.9982.6%
Qwen3-8B2.450.64 dex0.96719%

For comparison, the isolated estimator’s per-probe CV is 0.04 to 0.07, but its scores span only 0.05 to 0.07 dex across tensors. The propagated scores spread ten to twenty times wider, so per-tensor noise that would scramble a narrow ranking leaves a wide one intact. At 50 sequences the rankings are stable on all three models.

Qwen3-8B is the odd one out, and the reason is visible in the reference outputs. From the third decoder layer on, its residual-stream norm varies by a factor of about 1.7 across Gaussian probe sequences, against 0.03 to 0.04 on the two Qwen3.5 models. A few random sequences excite very large activations, and every tensor’s score inherits that heavy tail. The propagated signal on this model is not calibrated per sequence at all, even though its ranking is still good enough to tie HAWQ-V2. More sequences, or taking a median instead of a mean, are the obvious remedies. We haven’t tested either.

What we took from it

A signal you can prove things about and a signal that builds good models aren’t necessarily the same signal. On an MoE model the allocator’s structure does most of the work and the choice barely shows, with the calibrated isolated estimator slightly ahead on MMLU. On a dense model the network-aware signal is the one that matches gradient-based sensitivity. If you only test sensitivity methods on one kind of architecture, you can draw the wrong conclusion about which one is better.


Models: Qwen3.5-35B-A3B (21 GB target, 19.5 GB builds) and Qwen3-8B (6 GB target, 252 flexible tensors, 72 norms pinned). Same greedy solver, budget, quantizer and candidate sets for every signal; 2-bit veto decisions copied from the propagated manifest. Perplexity: WikiText-2 test, 128 × 2,048 tokens. MMLU: 5-shot, full test set. Calibration study: 200 probe sequences of 8 positions, estimates from the first 100 scored against the mean of the other 100.

Read the Full Paper

The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

arxiv.org/abs/2609.33923

I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0

Related articles

SEE ALL