Deciding how many bits each tensor of a model deserves usually means running real text through the full model, often for hours, and often on hardware most teams don’t have. RAM’s probe analysis scored every tensor of Llama-4-Maverick, 400 billion parameters and 803 GB in BF16, in 539 seconds on a single Mac Studio with 192 GB of memory. The model is four times bigger than the machine’s memory. Here is how that works, and what it does and doesn’t tell you.
Why this step is normally expensive
Calibration-based quantizers like GPTQ load the whole model and run forward passes over hundreds of sequences of real text. On a 400B-parameter MoE that takes hours on high-end hardware, and it has to be repeated for every target size, because the result depends on the budget. Gradient-based sensitivity is heavier again. When we ran HAWQ-V2, which estimates Hessian traces by double back-propagation, its peak memory already exceeded 64 GB on a 14B model, and it would not fit a 35B MoE on our 192 GB machine at all.
What RAM does instead
The analysis never runs text through the model. It draws 50 random Gaussian probe sequences of 8 positions and works through the decoder one layer at a time:
- Run the layer on its probe inputs to get a reference output.
- For every weight tensor in the layer with at least 1,024 elements, and for each candidate bit-width (2, 3, 4, 5, 6 and 8), quantize that one tensor, re-run the layer, record how far the output moved, then put the original weights back.
- Pass the layer’s reference output on as the next layer’s input, free the layer’s weights, and move on.
In lazy mode only one decoder layer is in memory at any time. Peak memory is set by the largest layer plus the probe batch, not by the model. The run reports show about 25 GB peak for the 217 to 250 GB models and about 109 GB for the 803 GB Maverick.
Analysis time by model
Wall-clock time for the full probe pass on one Mac Studio (M2 Ultra, 192 GB). The asterisk marks lazy loading:
| Model | BF16 size | Probe pass |
|---|---|---|
| Qwen3-8B | 15 GB | 36 s |
| Qwen3-30B-A3B | 61 GB | 53 s* |
| GLM-4.7-Flash | 60 GB | 99 s |
| Qwen3.5-35B-A3B | 69 GB | 188 s |
| Llama-4-Scout | 217 GB | 214 s* |
| MiniMax-M2.5 | 230 GB | 301 s* |
| Qwen3.5-122B-A10B | 250 GB | 345 s* |
| Llama-4-Maverick | 803 GB | 539 s* |
Time doesn’t track byte count closely. The 61 GB Qwen3-30B-A3B, run in lazy mode, finished in about half the time of the 60 GB GLM-4.7-Flash, and the 803 GB Maverick took less than three times as long as the 69 GB Qwen3.5-35B-A3B.
Against the obvious alternative
The straightforward data-free way to score tensors is a rate-distortion scan: quantize each tensor in isolation at a grid of settings and measure the error. We ran one at 13 (bit-width, group-size) configurations on three models:
| Model | Probe pass | Rate-distortion scan | Speed-up |
|---|---|---|---|
| Qwen3-8B | 36 s | 195 s | 5.4× |
| GLM-4.7-Flash | 99 s | 1,847 s | 19× |
| Qwen3-30B-A3B | 53 s | 2,634 s | 50× |
Speed isn’t the only difference. The scan scores each tensor on its own, while the probe pass measures each tensor inside the layer, with the inputs the earlier layers produce. That turns out to matter a great deal for which tensors get bits, which our other articles from this paper go into.
One pass, every budget
The scores don’t depend on the target size. A knapsack solver turns them into an allocation for whatever byte budget you ask for, so the probe pass is done once per model, not once per build. On Qwen3.5-35B-A3B a single 188-second pass produced builds at 19.5, 21.2, 22.9 and 33.6 GB. A calibration-based method would need a calibration run for each of those.
The workflow from the paper’s reproduction notes looks like this:
# score every tensor once (50 probes, six bit-widths, one layer in memory at a time)
python experiments/probe_allocator_v2.py \
--model /path/to/bf16-model --budget-gb 21 \
--num-probes 50 --lazy --output manifest.json
# build the quantized model from the manifest
python experiments/convert_probe_model.py \
--model /path/to/bf16-model --manifest manifest.json \
--output /path/to/quantized
For dense models, and MoE models with 128 or fewer experts, below about 30% of the BF16 size, the paper adds --min-bits 4. On those models 3-bit tensors cost more quality than the bytes they save.
What the 539 seconds doesn’t cover
To be clear about what we measured on Maverick: the probe analysis, not a finished, evaluated model. Quality results in the paper cover seven models from 8B to 122B parameters. Maverick appears only in the timing table. The analysis is also only the first step: converting the weights and evaluating the result take their own time, and the evaluation in particular is not a nine-minute job.
The data-free approach has limits too. On a small dense model, calibration-based AWQ reached 9.746 median perplexity at 4.4 GB on Qwen3-8B, a size RAM couldn’t match on that model. Real activations capture things random probes don’t, and the paper suggests a hybrid, with RAM for the initial allocation and calibration-based refinement for critical tensors, as the natural next step.
What the timing does show is that the sensitivity analysis is no longer the step that needs a cluster. If you can store a model and convert it, you can work out where its bits should go on the same machine, in minutes, without any data.
Hardware: one Apple Mac Studio, M2 Ultra (24-core CPU, 76-core GPU, 192 GB unified memory), macOS 15.5, MLX. Probe settings: 50 propagated sequences of 8 positions, RTN at group size 64, candidate widths 2/3/4/5/6/8, tensors with at least 1,024 elements. Peak-memory figures are from the run reports. The rate-distortion scan used 13 (bit, group-size) configurations per tensor.
Read the Full Paper
The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:
Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
arxiv.org/abs/2609.33923I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0



