Black Sheep AI

Ram · The Reasoning Layer

Frontier reasoningat the sizeyou name

Ram is the reasoning layer of the Bounded AI Architecture (BAA). Name a budget in gigabytes and the allocator solves for the precision of every tensor so the file lands on it, with the model's reasoning kept and measured. Bind a Ram model to Paddock and it reasons over the modules it is allowed to see. When a better model ships, we compress that one and you bind it instead.

The Reasoning Layer

Bound, not baked in

Ram holds no facts of its own. Knowledge stays in Paddock's Knowledge Modules, and the model bound on top only reads what its modules allow. That is what makes the reasoning layer swappable.
Stateless

The model holds no facts

Your knowledge lives in modules, permissioned one by one. Ram reads them at answer time and remembers nothing between questions, so there is nothing in the weights to delete, date, or permission.
Swappable

Swap it in one operation

When a better frontier model ships, we compress that one to the same budget and you bind it in place of the old one. Same modules, same permissions, same citations. Nothing in the plane changes.
Certified

Reasoning kept, and certified

Compression that keeps the size and loses the reasoning is worthless at this layer. Every Ram build carries Watchman's five certificates, and the reasoning gate is a hard one: fail it and the build does not ship.

Three properties, and the third is the one that surprises people.

How it decides

It measures your checkpoint

An oscilloscope tracing a signal

It measures your checkpoint

Random probe sequences are pushed through the network, carrying its own propagated statistics, to measure how much each individual tensor’s output moves when that tensor alone is compressed. Uniform quantization applies a rule; this measures the specific model in front of it.

It solves, rather than guesses

Boxes of different sizes packed into a crate

It solves, rather than guesses

Choosing one bit-width per tensor under a hard byte ceiling is a knapsack problem, and it is solved as one. Against an exact solver on a 27B model the production allocator landed within 0.1% of the provably optimal allocation at two different budgets.

It needs no data

Static noise on a screen

It needs no data

No calibration corpus. Nothing to license, nothing to leak, and no quiet dependence on text that resembles your users’. The probe is noise, and that turns out to be enough, which is the result our paper is about.

Measurements

What we measured

Same model, same file size, same evaluation. The only honest way to compare two compressions.
  • Against the reference implementation

    Bisected to within 1.7% of the file size of llama.cpp’s own 3-bit mix on a 27B model, our build scored 5.5% lower perplexity, ahead on 37 of 40 evaluation chunks, which at that sample size is not a coin toss.
  • On a 16 GB card

    That 27B model at 12.6 GB holds a 32,000-token context in 14 GB of VRAM and answers retrieval questions across it perfectly: 32 of 32, including two-hop. At 64K and 128K it still fits, with the cache compressed.
  • Against gradient-based methods

    Matched byte-for-byte against a Hessian method that needs backward passes, the data-free signal is a statistical tie. We do not claim to beat it. We claim to equal it without the gradients, which is what lets it run on models where that method will not fit.

TARGET : 12.6 GB

0 GB16 GB

12.6 GB

MODEL ON DISK

32K

CONTEXT HELD

16 GB

GPU MEMORY

HARDWARE : ONE 16 GB GPU, THE CLASS EVERY CLOUD RENTS BY THE HOUR

Hardware

The hardware you already rent

A 16 GB accelerator is the cheapest useful GPU on every major cloud, and it is the class Ram targets.

AWS

g4dn.xlarge: 16 GB NVIDIA T4. The entry GPU instance, available in every region and on spot pricing.

Azure

NC4as_T4_v3: 16 GB NVIDIA T4, the smallest GPU size in the NC family.

Google Cloud

g2-standard-4: 24 GB NVIDIA L4, with headroom above what the model needs.

The output is an ordinary GGUF file. It runs under unmodified llama.cpp, Ollama or LM Studio, on any CUDA GPU, with no runtime of ours anywhere in the path. Our published measurements were taken on a 16 GB NVIDIA card on our own hardware; the cloud instances above are the same class of accelerator.

Scale

It does not stop at small models

Eighty compressed models are published and downloadable, and the largest is not a demonstration.
A full rack of storage drives
01

To 447 GB

Trillion-parameter-class mixture-of-experts models, compressed and published: Kimi K2.6, GLM-5.x, DeepSeek-V3.2, a 397B Qwen. These are files people download, not slideware.
A compact desktop workstation
02

Memory that does not scale with the model

The probe streams layer by layer, so peak memory tracks the largest layer rather than the checkpoint. An 803 GB model was probed in nine minutes.
A camera lens and film strip
03

Not only language models

The same allocator ships image and video diffusion models: the method is about where error goes, not about transformers specifically. We publish those builds; we have not yet published quality measurements for them, so we are not claiming any.

Limits

Where it does not help

Every report we publish states its limits in full. So does this page.
  1. 01

    We are not the smallest

    On one head-to-head against a good third-party 4-bit build, at a larger size, we scored lower on a graduate-level benchmark. Compression is a field of trade-offs and anyone claiming to win all of them is selling you something.
  2. 02

    Very tight budgets cost real quality

    The benefit of measuring per tensor is largest when the budget is tight enough to force hard choices, and at those budgets some capability does go. The certificates exist to tell you how much, before you deploy it.
  3. 03

    Format conversion has its own losses

    Converting between runtime formats can cost more quality than the compression did, particularly on newer hybrid architectures. We measure the artifact you will actually run, not the one that flatters the pipeline.