Black Sheep AI
Alternatives to round-to-nearest
← All research
RAM Research

Better Rounding Than Round-to-Nearest: What Helped and What Didn’t

October 2026 · Black Sheep AI Research

RAM quantizes every tensor with plain round-to-nearest. It decides how many bits each tensor gets, but not how the rounding is done. That is a deliberate simplification, and an obvious place to look for more quality, so we tested four alternatives on the same tensors. Two of them helped on paper. Two did nothing, for reasons that turn out to be instructive. And the most practical gain disappeared the moment the weights were packed for inference.

The setup

We took 30 tensors from Qwen3.5-35B-A3B, quantized each one at 4-bit with group size 64, and measured the normalized reconstruction error (NRMSE) for each method against plain RTN. None of these methods were allowed calibration data, because the whole point of RAM is not to need any.

Method Change in NRMSE vs RTN Tensors improved
Hadamard rotation−7.92%30 of 30
HQQ (half-quadratic)−2.96%30 of 30
Data-free OBQ (H ≈ WᵀW)+0.50%0 of 18
Data-free AdaRound (weight-only MSE)0.000%–

Hadamard rotation: the biggest win, and the hardest to ship

A random orthogonal rotation applied before quantizing spreads outliers across the group, so no single large weight dictates the scale. It lowered per-tensor error by 7.92% on average and improved every one of the 30 tensors, the largest gain of anything we tried.

The catch is deployment. A rotated weight has to be un-rotated at inference, which means computing y = Wq(Hᵀx) with a custom dequantization kernel. The runtime we use doesn’t provide one, so we never evaluated rotation at the model level. The number above is a per-tensor result only.

HQQ: real at the model level, until we packed it

HQQ optimizes each group’s scale and zero-point rather than taking them straight from the min and max. Per tensor it cut error by 2.96%, again on all 30. Unlike rotation, we could test it end to end. On Qwen3-8B at uniform 4-bit, group size 64:

Build Median PPL
RTN9.987
HQQ weights, evaluated in BF169.768
HQQ weights, re-packed into the runtime’s 4-bit format10.338

In BF16 the HQQ-optimized weights are 2.2% better than RTN. Then we packed them into the runtime’s 4-bit storage, which re-rounds them with its own scale and offset per group, and the build ended up worse than plain RTN. The optimization only survives if the conversion keeps HQQ’s exact quantization parameters. That needs a format-aware converter, and we haven’t built one.

If you evaluate a rounding method by dequantizing to BF16 and measuring perplexity, check that the gain survives the format you will actually ship. Ours didn’t.

Data-free OBQ: the wrong Hessian is worse than none

Optimal Brain Quantization rounds weights one at a time and compensates the error on the remaining weights, using the Hessian of the layer’s loss. That Hessian comes from real activations. Without data, the obvious stand-in is H ≈ WᵀW, built from the weights themselves.

It made reconstruction worse: +0.50% NRMSE, and not one of 18 tensors improved. A weight-only Hessian is a poor proxy for the activation Hessian, so the error compensation pushes the weights in directions that don’t correspond to how the layer is used. OBQ needs calibration data, and pretending otherwise costs quality.

Data-free AdaRound: exactly zero, and that’s the point

AdaRound learns whether to round each weight up or down, minimizing the error of the layer output on real inputs. Replace that activation-weighted objective with plain weight-space MSE, which is all you can do without data, and the NRMSE change is exactly 0.000%.

That isn’t a failed experiment. Without activation data, the rounding direction that minimizes weight error for each element is simply the nearest grid point, which is precisely what RTN already computes. There is nothing left for the optimizer to find. Any data-free method that claims to beat RTN by choosing rounding directions has to be getting its signal from somewhere other than the weights alone.

Rounding versus allocation

How big are these effects next to the decision RAM actually makes, where to put the bits? Some reference points from the same paper, all median perplexity on WikiText-2:

Change Model Effect
Group size 128 to 64, uniform 4-bitQwen3.5-35B-A3B6.928 → 6.764 (−2.4%)
HQQ instead of RTN, in BF16Qwen3-8B9.987 → 9.768 (−2.2%)
Probe allocation vs uniform 4-bit g128Qwen3.5-35B-A3B6.928 → 6.585 (−4.9%)
Probe allocation vs uniform 4-bitGLM-4.7-Flash10.075 → 8.700 (−13.6%)

The allocation rows aren’t perfectly size-matched (the RAM builds are somewhat larger), so read them as orders of magnitude rather than a precise ranking. In our experiments, where the bits go mattered more than how each tensor is rounded. But a percent or two from a better rounder is not nothing, and in principle it could add to what allocation gives. We haven’t tested the combination.

Why RTN is still in the pipeline

RTN is fast, needs no data, and every runtime supports it. The two methods that genuinely beat it each come with a deployment cost: a custom kernel for rotation, a format-aware converter for HQQ. The two data-free imitations of calibration-based methods don’t beat it at all. Combining a better rounder with mixed-precision allocation is open work, and we make no claim that RTN is optimal. It is simply the best option we can currently ship without data.


Per-tensor tests: 30 tensors of Qwen3.5-35B-A3B (18 for data-free OBQ), 4-bit, group size 64, normalized RMSE against RTN. Model-level HQQ test: Qwen3-8B, uniform 4-bit, group size 64, WikiText-2 test, 128 × 2,048 tokens, median perplexity. The rounding-versus-allocation table collects numbers from different sections of the paper; the allocation builds are not exactly size-matched to the uniform ones.

Read the Full Paper

The proofs, the variance bounds, the per-tensor tables and every experiment behind this article are in the paper:

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

arxiv.org/abs/2609.33923

I. Kennedy and T. Kennedy, September 2026. Licensed under CC BY-NC-ND 4.0

Related articles

SEE ALL