Black Sheep AI
Spectral flatness of quantization error
← All research
RAM Research

Quantization Error Is Spectrally Flat, So One Random Probe Is Enough

October 2026 · Black Sheep AI Research

If you push one random Gaussian vector through a quantized layer and measure how far the output moved, you get an unbiased estimate of that layer’s quantization error. That part is textbook. What surprised us is how good one vector turns out to be. Across 1,683 weight tensors, round-to-nearest error behaves almost exactly like random noise, and that one fact turns a noisy Monte-Carlo sample into a measurement with a 4 to 7% error bar you can compute before loading the weights.

The estimator

Take a linear layer with weights W and its quantized version Ŵ. The quantization error is the matrix Δ = W − Ŵ, and for an input x the output moves by Δx. Draw x from a standard Gaussian and the expected squared size of that movement is exactly the squared Frobenius norm of Δ:

E[ ||Δx||² ] = Tr(ΔᵀΔ) = ||Δ||F²   for x ~ N(0, I)

This is Hutchinson’s trace estimator with the matrix set to ΔᵀΔ. Nothing new in the identity. The interesting question is the variance, because an unbiased estimator that swings by 100% per sample is useless for ranking a few thousand tensors against each other.

For a Gaussian probe the per-probe coefficient of variation works out to √(2 / deff), where deff is the effective dimensionality of the error spectrum: roughly, how many directions the error energy is spread over. If all the error sat in one direction, a single probe would be off by about 141%. If it were spread evenly over a thousand directions, a single probe would be off by about 4.5%. So everything hinges on what the spectrum of real quantization error looks like.

We measured the spectrum instead of assuming it

For an isolated tensor you don’t need probes to get the answer. ||Δ||F² and the quantity that controls the variance can both be computed in closed form from the smaller Gram matrix. So we did that for every weight tensor of two models, at four bit-widths, with round-to-nearest at group size 64:

  • Qwen3.5-35B-A3B, a 256-expert MoE: 1,317 tensors covering attention, linear attention, dense MLP, fused expert stacks and the vision blocks.
  • Qwen3.5-9B, dense: 366 tensors.
  • Bit-widths 2, 3, 4 and 8.

The reference point is what pure i.i.d. noise of the same shape would give. For an m × n noise matrix the expected effective dimensionality is mn/(m+n). Real quantization error lands very close to it:

Model Tensors Median deff vs i.i.d. noise IQR
Qwen3.5-35B-A3B (MoE)1,3170.93×0.86 to 0.97
Qwen3.5-9B (dense)3660.96×0.90 to 0.98

The biggest tensors sit closest to the noise ceiling. The fused expert stacks, 262,144 × 2,048 each, predict a deff of 2,032.1 and measure 2,031.7. That is not an approximation we chose to make. It is what the rounding residuals actually look like: inside each group of 64 weights they are small, nearly independent fluctuations below one quantization step, so Δ behaves like a noise matrix rather than a low-rank perturbation.

It doesn’t depend on the bit-width

We expected the shape of the error to change as bits went down. It didn’t. On the 35B model the median effective dimensionality at 2, 3, 4 and 8 bits was:

Bits 2 3 4 8
Median deff (Qwen3.5-35B-A3B)386.1386.0386.1386.2

The size of the error shrinks by a factor of 64 between 2-bit and 8-bit. Its spectral shape stays put to three significant figures. On the 9B model the four medians fall within 3.5% of each other (1,424 to 1,476).

This matters for anyone building a mixed-precision allocator, because the allocator has to compare a tensor at 2-bit against the same tensor at 8-bit. If the estimator got noisier at low bit-widths, the comparisons that matter most would be the least reliable. They aren’t.

The error bars are predictable from shape alone

Since deff ≈ 0.93 · mn/(m+n), you can work out the probe noise for any tensor from its dimensions. We checked the prediction against measurement at 4-bit. The ratio of measured to predicted per-probe CV had a median of 0.999 on the 35B model and 1.000 on the 9B model, and the 200-probe average matched the exact closed-form value with median ratios of 0.9998 and 1.0000. So the estimator is not only unbiased, it is calibrated: it tells you how wrong it is likely to be, and it is right about that.

To make that concrete, here is the formula applied to two shapes you’ll find in most current models (our arithmetic from the formula, not additional measurements):

Tensor shape Predicted deff 1 probe 20 probes
4,096 × 4,096 projection≈ 1,905≈ 3.2%≈ 0.7%
256 × 2,048 MoE router≈ 212≈ 9.7%≈ 2.2%

Small, skinny tensors like routers are where a single probe is weakest, and you know that before you start.

How fast it converges, and a number we got wrong

Because the exact answer is available in closed form, we could measure convergence against ground truth rather than against a big probe average. We drew 200 independent probes per tensor and scored estimates built from disjoint blocks of them:

Probes Rel. error (35B) Rank ρ (35B) Rel. error (9B) Rank ρ (9B)
16.2%0.6715.5%0.627
52.8%0.8552.5%0.864
201.4%0.9381.3%0.957
1000.64%0.9810.53%0.991

The error falls with a log-log slope of −0.50 on the 35B model and −0.51 on the 9B model. That is the ordinary 1/√p rate you get from averaging independent samples.

An earlier version of this work reported a slope of −1.43, which would have meant the estimate converges far faster than averaging should allow. That number was wrong. We had measured convergence against a 500-probe reference whose probe bank included the probes being evaluated, so the estimate and the reference shared samples and the error was forced to zero at p = 500. Measured against exact values with disjoint probes, there is no super-convergence. There doesn’t need to be: the starting point is already around 6% because the spectrum is flat.

The rank correlation looks less impressive than the error numbers, and that has a mundane cause. At 4-bit the true sensitivities of all 1,317 tensors on the 35B model span only 0.41 dex (standard deviation 0.065 dex), so many tensors are nearly tied. With 20 probes the leftover rank disagreements are swaps between tensors whose true values differ by less than the 1.4% estimate error. Whether swaps that small change a finished model is a separate question, and the paper doesn’t claim to answer it.

Why probe at all, then?

Fair question. For a single tensor viewed in isolation, ||Δ||F² is cheap to compute directly, and an isotropic probe tells you nothing more than that number. If that were the whole job, the probe would be a roundabout way of computing a Frobenius norm.

The probe earns its place once you stop treating tensors in isolation. In RAM the probes are pushed through the network, so each layer receives the inputs the earlier layers actually produce and the damage is measured after the nonlinearity and residual path. That is a different quantity, and it ranks tensors very differently: on Qwen3.5-35B-A3B the isolated and propagated scores have a Spearman correlation of −0.01 across 390 matched tensors. The isolated measurement here is the part we can prove things about. The propagated version is what the allocator uses, and it gets its own article.

There is a second, smaller benefit. Running several probes gives you a sample variance for free, and a tensor whose variance is far above its shape-predicted value is one whose error is not flat. Those are rare, and worth knowing about.

What to take from it

If you are building a data-free sensitivity estimator for quantization, you can stop worrying about whether random probes are too noisy. For round-to-nearest error they behave close to the best case, the noise is predictable from tensor shape, and twenty probes get you to about 1.4%. Budget your probes from the shape formula, expect skinny tensors to need more, and measure convergence against something that doesn’t share samples with your estimate. We learned that last one the hard way.


Measurements: Qwen3.5-35B-A3B (1,317 tensors) and Qwen3.5-9B (366 tensors), RTN at group size 64, bit-widths 2/3/4/8, exact ||Δ||F² and ||ΔᵀΔ||F² from the small-side Gram matrix. Convergence: 200 independent Gaussian probes per tensor, disjoint blocks, 4-bit. Hardware: one Mac Studio (M2 Ultra, 192 GB), MLX. The 4,096 × 4,096 and router rows in the shape table are computed from the paper’s formula, not measured.

Read the Full Paper

The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

arxiv.org/abs/2609.33923

I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0

Related articles

SEE ALL