Black Sheep AI
Analysing a large model on one workstation
← All research
RAM Research

Analysing an 803 GB Model in Nine Minutes on One Mac

October 2026 · Black Sheep AI Research

Deciding how many bits each tensor of a model deserves usually means running real text through the full model, often for hours, and often on hardware most teams don’t have. RAM’s probe analysis scored every tensor of Llama-4-Maverick, 400 billion parameters and 803 GB in BF16, in 539 seconds on a single Mac Studio with 192 GB of memory. The model is four times bigger than the machine’s memory. Here is how that works, and what it does and doesn’t tell you.

Why this step is normally expensive

Calibration-based quantizers like GPTQ load the whole model and run forward passes over hundreds of sequences of real text. On a 400B-parameter MoE that takes hours on high-end hardware, and it has to be repeated for every target size, because the result depends on the budget. Gradient-based sensitivity is heavier again. When we ran HAWQ-V2, which estimates Hessian traces by double back-propagation, its peak memory already exceeded 64 GB on a 14B model, and it would not fit a 35B MoE on our 192 GB machine at all.

What RAM does instead

The analysis never runs text through the model. It draws 50 random Gaussian probe sequences of 8 positions and works through the decoder one layer at a time:

  • Run the layer on its probe inputs to get a reference output.
  • For every weight tensor in the layer with at least 1,024 elements, and for each candidate bit-width (2, 3, 4, 5, 6 and 8), quantize that one tensor, re-run the layer, record how far the output moved, then put the original weights back.
  • Pass the layer’s reference output on as the next layer’s input, free the layer’s weights, and move on.

In lazy mode only one decoder layer is in memory at any time. Peak memory is set by the largest layer plus the probe batch, not by the model. The run reports show about 25 GB peak for the 217 to 250 GB models and about 109 GB for the 803 GB Maverick.

Analysis time by model

Wall-clock time for the full probe pass on one Mac Studio (M2 Ultra, 192 GB). The asterisk marks lazy loading:

Model BF16 size Probe pass
Qwen3-8B15 GB36 s
Qwen3-30B-A3B61 GB53 s*
GLM-4.7-Flash60 GB99 s
Qwen3.5-35B-A3B69 GB188 s
Llama-4-Scout217 GB214 s*
MiniMax-M2.5230 GB301 s*
Qwen3.5-122B-A10B250 GB345 s*
Llama-4-Maverick803 GB539 s*

Time doesn’t track byte count closely. The 61 GB Qwen3-30B-A3B, run in lazy mode, finished in about half the time of the 60 GB GLM-4.7-Flash, and the 803 GB Maverick took less than three times as long as the 69 GB Qwen3.5-35B-A3B.

Against the obvious alternative

The straightforward data-free way to score tensors is a rate-distortion scan: quantize each tensor in isolation at a grid of settings and measure the error. We ran one at 13 (bit-width, group-size) configurations on three models:

Model Probe pass Rate-distortion scan Speed-up
Qwen3-8B36 s195 s5.4×
GLM-4.7-Flash99 s1,847 s19×
Qwen3-30B-A3B53 s2,634 s50×

Speed isn’t the only difference. The scan scores each tensor on its own, while the probe pass measures each tensor inside the layer, with the inputs the earlier layers produce. That turns out to matter a great deal for which tensors get bits, which our other articles from this paper go into.

One pass, every budget

The scores don’t depend on the target size. A knapsack solver turns them into an allocation for whatever byte budget you ask for, so the probe pass is done once per model, not once per build. On Qwen3.5-35B-A3B a single 188-second pass produced builds at 19.5, 21.2, 22.9 and 33.6 GB. A calibration-based method would need a calibration run for each of those.

The workflow from the paper’s reproduction notes looks like this:

# score every tensor once (50 probes, six bit-widths, one layer in memory at a time)
python experiments/probe_allocator_v2.py \
  --model /path/to/bf16-model --budget-gb 21 \
  --num-probes 50 --lazy --output manifest.json

# build the quantized model from the manifest
python experiments/convert_probe_model.py \
  --model /path/to/bf16-model --manifest manifest.json \
  --output /path/to/quantized

For dense models, and MoE models with 128 or fewer experts, below about 30% of the BF16 size, the paper adds --min-bits 4. On those models 3-bit tensors cost more quality than the bytes they save.

What the 539 seconds doesn’t cover

To be clear about what we measured on Maverick: the probe analysis, not a finished, evaluated model. Quality results in the paper cover seven models from 8B to 122B parameters. Maverick appears only in the timing table. The analysis is also only the first step: converting the weights and evaluating the result take their own time, and the evaluation in particular is not a nine-minute job.

The data-free approach has limits too. On a small dense model, calibration-based AWQ reached 9.746 median perplexity at 4.4 GB on Qwen3-8B, a size RAM couldn’t match on that model. Real activations capture things random probes don’t, and the paper suggests a hybrid, with RAM for the initial allocation and calibration-based refinement for critical tensors, as the natural next step.

What the timing does show is that the sensitivity analysis is no longer the step that needs a cluster. If you can store a model and convert it, you can work out where its bits should go on the same machine, in minutes, without any data.


Hardware: one Apple Mac Studio, M2 Ultra (24-core CPU, 76-core GPU, 192 GB unified memory), macOS 15.5, MLX. Probe settings: 50 propagated sequences of 8 positions, RTN at group size 64, candidate widths 2/3/4/5/6/8, tensors with at least 1,024 elements. Peak-memory figures are from the run reports. The rate-distortion scan used 13 (bit, group-size) configurations per tensor.

Read the Full Paper

The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

arxiv.org/abs/2609.33923

I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0

Related articles

SEE ALL