Black Sheep AI
The calibration objective is the wrong target
← All research
RAM Research

We Estimated GPTQ’s Objective Without Data. Then It Lost

October 2026 · Black Sheep AI Research

Calibration-based quantizers like GPTQ minimize a layer-wise objective computed from real activations. We showed that random probes pushed through the network estimate that same objective without any data, and estimate it well. Then we allocated bits from it, and from the exact version computed on real text, and both lost to a cruder signal measured at the block output. The estimate was fine. The objective was the wrong thing to allocate from.

Two ways to probe a tensor

RAM scores tensors with random Gaussian probes and needs no calibration text. There are two ways to do that. The isolated way feeds standard Gaussian vectors straight into one tensor. It is clean, and it measures little more than how badly the weights round (the previous article covers it). The propagated way, which RAM’s allocator uses, draws 50 Gaussian probe sequences of 8 positions, feeds them into the first decoder layer, and then gives every later layer the reference output of the layer before it. Each tensor therefore sees inputs shaped by everything upstream of it.

For each tensor and each candidate bit-width (2, 3, 4, 5, 6 and 8), the pipeline quantizes that one tensor, re-runs the layer, and records the cosine distance between the perturbed block output and the reference. That cosine distance is the allocator’s cost signal. The whole pass takes 36 seconds on an 8B model.

The propagated probe tracks the calibration objective

The paper’s Proposition 3 says something checkable: if the probes carry the network’s own input statistics, the quadratic form they estimate is the layer objective GPTQ and OBQ minimize, ||ΔX||², without needing the X. So we computed that objective from real activations (32 sequences of 512 tokens from the WikiText-2 training split, never the test split) on every linear tensor of three models and compared rankings at 4-bit:

Model Propagated probe Isolated estimator Block cosine Real text, other domain
Qwen3-8B0.83−0.070.500.99
Qwen3.5-9B0.81−0.060.110.99
Qwen3.8-27B (hybrid)0.83−0.15––

The last column is a ceiling: the same objective computed on Tulu-3 dialogue instead of WikiText-2, which shows how stable the real ranking is across domains. Against that 0.99, a data-free estimate at 0.81 to 0.83 is good. The isolated estimator, the one we can prove the nice variance results for, is uncorrelated with the objective. Which tensors matter is decided by the second moment of their inputs, not by how badly their weights round, and only the propagated probe carries those inputs. For comparison, HAWQ-V2’s Hessian trace, which needs gradients and calibration text, tracks the same objective at 0.25 on Qwen3-8B.

In absolute terms the probe over-estimates the objective by 12 to 29% at the median. Gaussian probes over-excite the few high-energy channels of the residual stream: on Qwen3-8B the top eight channels carry 29% of probe energy against 17% of real energy. The ranking survives that.

A safety switch that broke it

The first time we ran this on the 27B hybrid, the correlation was 0.27. Split by depth, the first 32 layers sat at 0.79 and the last 32 at −0.22. The cause was our own code. The pipeline has an adaptation rule that watches the block outputs and, if three consecutive layers look unstable, swaps the propagated probes for fresh isotropic ones. On Qwen3.8-27B it fired at layer 32, and from there on every tensor was being probed with inputs that carry no network statistics at all, which is exactly what the isolated estimator does.

We disabled the rule and changed nothing else. The overall correlation went to 0.83, the same as Qwen3-8B, and layers 32 to 63 moved from −0.22 to 0.85. The rule had been added to stabilize the block-output signal on hybrid stacks. For estimating the objective it does precisely the wrong thing, because it throws away the input statistics the estimate depends on. One projection type still isn’t tracked even with the rule off: the β-gate input of the DeltaNet layers (0.06 over 48 small tensors, 0.05% of the linear parameters).

Then we allocated from it

A good estimate of the objective GPTQ optimizes sounds like exactly what an allocator wants, and it is cheaper to compute than the block cosine. So we tested it on Qwen3.8-27B in a setup that reuses none of our own build path. Each signal drives the same greedy knapsack over llama.cpp’s k-quant types (Q2_K to Q8_0), the budget is bisected until the file matches llama.cpp’s own IQ3_M mix, and llama-quantize builds every arm from the vendor BF16 file with one shared importance matrix. One arm is an oracle that allocates from the real-activation objective itself, computed on WikiText-2 training text. Perplexity is measured on 40 chunks of 2,048 tokens of WikiText-2 test:

Signal Uses text? File (GB) Perplexity Beats IQ3_M on
IQ3_M (llama.cpp mix)no12.586.336–
Q3_K_M (llama.cpp mix)no13.306.20731 / 40
Block cosineno12.726.04136 / 40
Block cosine × isolated CKAno12.795.99137 / 40
Quadratic form, rule onno12.616.20434 / 40
Quadratic form, rule offno12.666.23830 / 40
Real-activation objective (oracle)yes12.806.08737 / 40

The block-output signals win. The cosine-and-CKA arm reaches 5.991 against 6.336 for IQ3_M at the same bytes, about 5.5% lower, and llama.cpp’s Q3_K_M needs 5.7% more bytes to get to 6.207. The oracle, which allocates from the exact quantity calibration-based methods optimize, reaches 6.087 and loses to the cosine-and-CKA arm on 35 of 40 chunks (Wilcoxon p = 10−5) and to the plain cosine on 29 of 40 (p = 0.004). Our data-free estimate of the objective does worse again, at 6.204 to 6.238. A rank correlation of 0.83 with the objective would suggest the opposite ordering.

Why the objective is the wrong target

The two signals normalize by different things. The layer objective divides a tensor’s output perturbation by that tensor’s own output energy. It tells you how badly the tensor rounds relative to itself and nothing about how much its output matters to the rest of the block. The block cosine measures the perturbation against the whole block output, residual stream included, so a tensor whose contribution is small next to the residual scores low however badly it rounds.

You can see it in where the bits went. The cosine-and-CKA arm spent 2.7 to 2.8 bits per weight on layers 0 to 31, 4.3 to 4.5 on layers 32 to 63, and 4.1 bits on the sixteen full-attention blocks. The oracle was nearly flat by depth (3.2 to 3.9 bits), gave full attention 3.4 bits, and spent its extra bytes on the DeltaNet output projections instead. The two agreed on the quant type for 53% of tensors.

The signal an allocator actually needs is the downstream-weighted version: the same quadratic form, but measured through the rest of the block. The block cosine approximates it at the cost of re-running each layer. The quadratic form could compute it with one matrix product if the downstream Jacobian were supplied. We haven’t built that yet.

Against gradients: a tie

The closest calibration-based comparison is HAWQ-V2, which scores tensors by Hessian trace estimated with double back-propagation on real text. We ran all three signals (propagated probe, isolated estimator, HAWQ-V2) through one allocator on Qwen3-8B at a 6 GB target, with byte totals within 0.02% of each other, and tested per sequence:

Signal Needs data? Median PPL Mean PPL MMLU
Propagated probeno9.6309.6410.749
HAWQ-V2 (Hessian trace)yes9.6569.6410.750
Isolated estimatorno9.7379.6980.749

The propagated probe and HAWQ-V2 are indistinguishable: identical mean perplexity, the probe build lower on 65 of 128 sequences and HAWQ-V2 lower on 63, paired p = 0.96. The isolated estimator trails both and is worse than the propagated build on 108 of 128 sequences. On MMLU all three sit at 0.749 to 0.750. One caveat in HAWQ-V2’s favour: we used a single Hutchinson seed for it, so this shows parity with one draw of the gradient estimator, not superiority. One caveat in the probe’s favour: the double back-propagation didn’t fit the 35B MoE on our 192 GB machine at all.

Limits

The allocation test covers one model and one quantizer family. The objective measurement covers three non-MoE models, two of them hybrids. MoE expert stacks, whose inputs are routed, weren’t measured. And the most interesting signal here, the downstream-weighted quadratic form, is untested.

The practical lesson is narrower and more useful: if you are choosing what to allocate bits from, “the objective GPTQ minimizes” is not automatically the right answer, even when you can compute it exactly. Measure what quantizing a tensor does to the block output, and check any adaptive switch in your pipeline for whether it is quietly discarding the statistics your signal depends on.


Objective comparison: every linear tensor of Qwen3-8B (252), Qwen3.5-9B (248) and Qwen3.8-27B (496), 4-bit, real activations from 32 × 512 tokens of WikiText-2 train, probes 200 sequences × 8 positions. Allocation test: Qwen3.8-27B, llama.cpp k-quants with a shared importance matrix (100 chunks × 512 tokens), files matched to IQ3_M, perplexity on 40 × 2,048-token chunks of WikiText-2 test. During these runs a tensor-naming gap in our bridge left the 192 DeltaNet input projections (1.9 GB) on llama.cpp’s default types in every arm; with it fixed and the file matched to the IQ3_M byte count exactly, the ensemble arm reaches 6.018 whether the signal controls those projections or llama.cpp does (6.020). HAWQ-V2: 16 × 256 tokens, 8 Rademacher probes each, one seed.

Read the Full Paper

The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

arxiv.org/abs/2609.33923

I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0

Related articles

SEE ALL