Black Sheep AI
Probing Multi-head Latent Attention
← All research
RAM Research

Random Probes Are Blind Inside Multi-Head Latent Attention

October 2026 · Black Sheep AI Research

RAM scores every tensor in a model by pushing random Gaussian probes through it and measuring how much quantizing that tensor disturbs the output. On most architectures that works with no special handling. On models with Multi-head Latent Attention it didn’t, and the allocation it produced for GLM-4.7-Flash was markedly worse than it should have been. The cause was a mismatch between the probes and the subspace those weights actually see. The fix took one extra step per attention block.

What MLA changes

Multi-head Latent Attention, used in DeepSeek-V2 and GLM-4.7-Flash, doesn’t project the hidden state straight into keys and values. It first compresses it into a much smaller latent vector, normalizes that, and only then expands it back out:

latent = Norm( W_compress · x )        # d_c dimensions, d_c « d
K      = W_K · latent                   # up-projection

The point of the design is a smaller KV cache. The side effect, for anyone measuring sensitivity, is that the up-projection weights never see an arbitrary input. They only ever see the normalized image of the compression weights, which occupies a particular region of the latent space.

Why isotropic probes get it wrong

A standard Gaussian probe in the latent space spreads its energy evenly over all dc directions. The up-projection weights, at inference, get energy only in the directions the compression path produces. So a probe-based sensitivity score for those weights is averaging over a lot of directions that never occur in practice. Some weights look more sensitive than they are, others less, and the allocator spends bytes accordingly.

This is a specific case of a general point the paper makes about probes. On an isolated tensor, a standard Gaussian probe estimates the squared Frobenius norm of the quantization error and nothing else. Probes only start to say something about how the network uses a tensor when their distribution matches the inputs that tensor really receives. RAM’s normal propagated probes get that by passing through the earlier layers. Inside an MLA block, the compression step is one more layer they need to pass through.

Two-stage probing

The fix probes each MLA block in two stages:

  • Compression weights see the full hidden dimension, so they are probed with the ordinary propagated probes, the same as any other tensor.
  • Up-projection weights are probed with latent probes built by passing the propagated probes through that block’s own compression weight and latent normalization. Each candidate quantization of WK is scored by how far WK·latent moves.

Weights whose input is the full hidden size, such as the output projection, keep the full-dimensional probes. Once every tensor in the block is scored, the unquantized block output propagates to the next layer as usual. Detection is automatic: the pipeline checks the model config for a KV compression rank (kv_lora_rank), so nobody has to flag a model as MLA by hand.

The result on GLM-4.7-Flash

All builds below use the same allocator and target 16 GB. Median perplexity on WikiText-2 test, 128 sequences of 2,048 tokens:

Probe strategy Size (GB) Median PPL vs BF16
BF1658.28.470–
MLA-aware (latent stage)16.28.700+2.7%
Propagated, with adaptive monitor14.69.354+10.4%
Propagated, no latent stage14.69.486+12.0%
Uniform 4-bit (community build)–10.075+18.9%

One caveat belongs right next to that table. The MLA-aware manifest came out at 16.2 GB and the two plain-probe manifests at 14.6 GB, so part of the gap is size. The bigger build had more bytes to spend. Even so, going from +10.4% to +2.7% is far larger than the difference in size would suggest on its own, and the 94 post-bottleneck weights are exactly where the plain probes were misplacing bytes. This model also gave the largest gain over uniform 4-bit of the seven architectures in the paper: 13.6% lower median perplexity.

Where the bytes ended up

The allocation the MLA-aware probes produced for GLM-4.7-Flash, read from the build manifest (embeddings, output head and norms stay unquantized and aren’t counted):

Tensor class Tensors Parameters Mean bits
Attention3291.02B9.0
Routers460.01B16.0
Shared experts1380.43B11.5
Routed expert stacks13827.8B3.6

The routed experts hold almost all of the parameters and take most of the compression: 58 of the 138 stacks at 3-bit and 80 at 4-bit. All 46 routers stay at 16-bit. Attention is spread out rather than protected as a block: of the 329 attention tensors, 138 sit at 16-bit and 131 at 8-bit, and the rest are split between 3, 4, 5 and 6 bits.

A note on how we measured

GLM-4.7-Flash is also the model where the choice of statistic matters most. Its BF16 build has five sequences, out of 128, with per-sequence perplexity above 25,000. Those five lift its mean perplexity to 11.52 while the median is 8.470. Uniform 4-bit has six such sequences and a mean of 14.75. The ordering is the same either way, but the size of every gap reported in means is set by a handful of sequences on which even the full-precision model is close to random. That is why the table above uses medians.

If you quantize MLA models

Check how your sensitivity method feeds inputs to the weights after the bottleneck. Anything that assumes isotropic inputs in the latent space, whether random probes or a weight-only statistic, is measuring those weights under conditions they never see. Passing the probes through the block’s own compression and normalization is cheap and needs no data. The broader lesson is the one that keeps coming up with probes: the closer their distribution is to what the tensor really receives, the more the score means.


Model: GLM-4.7-Flash (MoE with MLA). Allocator target 16 GB for every probe strategy. Perplexity: WikiText-2 test, 128 × 2,048 tokens, seed 42, median of per-sequence perplexity. Probe pass: 99 seconds on one Mac Studio (M2 Ultra, 192 GB) with MLX. Allocation counts are from the paper’s Appendix C.

Read the Full Paper

The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

arxiv.org/abs/2609.33923

I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0

Related articles

SEE ALL