RAM quantizes every tensor with plain round-to-nearest. It decides how many bits each tensor gets, but not how the rounding is done. That is a deliberate simplification, and an obvious place to look for more quality, so we tested four alternatives on the same tensors. Two of them helped on paper. Two did nothing, for reasons that turn out to be instructive. And the most practical gain disappeared the moment the weights were packed for inference.
The setup
We took 30 tensors from Qwen3.5-35B-A3B, quantized each one at 4-bit with group size 64, and measured the normalized reconstruction error (NRMSE) for each method against plain RTN. None of these methods were allowed calibration data, because the whole point of RAM is not to need any.
| Method | Change in NRMSE vs RTN | Tensors improved |
|---|---|---|
| Hadamard rotation | −7.92% | 30 of 30 |
| HQQ (half-quadratic) | −2.96% | 30 of 30 |
| Data-free OBQ (H ≈ WᵀW) | +0.50% | 0 of 18 |
| Data-free AdaRound (weight-only MSE) | 0.000% | – |
Hadamard rotation: the biggest win, and the hardest to ship
A random orthogonal rotation applied before quantizing spreads outliers across the group, so no single large weight dictates the scale. It lowered per-tensor error by 7.92% on average and improved every one of the 30 tensors, the largest gain of anything we tried.
The catch is deployment. A rotated weight has to be un-rotated at inference, which means computing y = Wq(Hᵀx) with a custom dequantization kernel. The runtime we use doesn’t provide one, so we never evaluated rotation at the model level. The number above is a per-tensor result only.
HQQ: real at the model level, until we packed it
HQQ optimizes each group’s scale and zero-point rather than taking them straight from the min and max. Per tensor it cut error by 2.96%, again on all 30. Unlike rotation, we could test it end to end. On Qwen3-8B at uniform 4-bit, group size 64:
| Build | Median PPL |
|---|---|
| RTN | 9.987 |
| HQQ weights, evaluated in BF16 | 9.768 |
| HQQ weights, re-packed into the runtime’s 4-bit format | 10.338 |
In BF16 the HQQ-optimized weights are 2.2% better than RTN. Then we packed them into the runtime’s 4-bit storage, which re-rounds them with its own scale and offset per group, and the build ended up worse than plain RTN. The optimization only survives if the conversion keeps HQQ’s exact quantization parameters. That needs a format-aware converter, and we haven’t built one.
If you evaluate a rounding method by dequantizing to BF16 and measuring perplexity, check that the gain survives the format you will actually ship. Ours didn’t.
Data-free OBQ: the wrong Hessian is worse than none
Optimal Brain Quantization rounds weights one at a time and compensates the error on the remaining weights, using the Hessian of the layer’s loss. That Hessian comes from real activations. Without data, the obvious stand-in is H ≈ WᵀW, built from the weights themselves.
It made reconstruction worse: +0.50% NRMSE, and not one of 18 tensors improved. A weight-only Hessian is a poor proxy for the activation Hessian, so the error compensation pushes the weights in directions that don’t correspond to how the layer is used. OBQ needs calibration data, and pretending otherwise costs quality.
Data-free AdaRound: exactly zero, and that’s the point
AdaRound learns whether to round each weight up or down, minimizing the error of the layer output on real inputs. Replace that activation-weighted objective with plain weight-space MSE, which is all you can do without data, and the NRMSE change is exactly 0.000%.
That isn’t a failed experiment. Without activation data, the rounding direction that minimizes weight error for each element is simply the nearest grid point, which is precisely what RTN already computes. There is nothing left for the optimizer to find. Any data-free method that claims to beat RTN by choosing rounding directions has to be getting its signal from somewhere other than the weights alone.
Rounding versus allocation
How big are these effects next to the decision RAM actually makes, where to put the bits? Some reference points from the same paper, all median perplexity on WikiText-2:
| Change | Model | Effect |
|---|---|---|
| Group size 128 to 64, uniform 4-bit | Qwen3.5-35B-A3B | 6.928 → 6.764 (−2.4%) |
| HQQ instead of RTN, in BF16 | Qwen3-8B | 9.987 → 9.768 (−2.2%) |
| Probe allocation vs uniform 4-bit g128 | Qwen3.5-35B-A3B | 6.928 → 6.585 (−4.9%) |
| Probe allocation vs uniform 4-bit | GLM-4.7-Flash | 10.075 → 8.700 (−13.6%) |
The allocation rows aren’t perfectly size-matched (the RAM builds are somewhat larger), so read them as orders of magnitude rather than a precise ranking. In our experiments, where the bits go mattered more than how each tensor is rounded. But a percent or two from a better rounder is not nothing, and in principle it could add to what allocation gives. We haven’t tested the combination.
Why RTN is still in the pipeline
RTN is fast, needs no data, and every runtime supports it. The two methods that genuinely beat it each come with a deployment cost: a custom kernel for rotation, a format-aware converter for HQQ. The two data-free imitations of calibration-based methods don’t beat it at all. Combining a better rounder with mixed-precision allocation is open work, and we make no claim that RTN is optimal. It is simply the best option we can currently ship without data.
Per-tensor tests: 30 tensors of Qwen3.5-35B-A3B (18 for data-free OBQ), 4-bit, group size 64, normalized RMSE against RTN. Model-level HQQ test: Qwen3-8B, uniform 4-bit, group size 64, WikiText-2 test, 128 × 2,048 tokens, median perplexity. The rounding-versus-allocation table collects numbers from different sections of the paper; the allocation builds are not exactly size-matched to the uniform ones.
Read the Full Paper
The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:
Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
arxiv.org/abs/2609.33923I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0



