
01
To 447 GB
Trillion-parameter-class mixture-of-experts models, compressed and published: Kimi K2.6, GLM-5.x, DeepSeek-V3.2, a 397B Qwen. These are files people download, not slideware.
Ram · The Reasoning Layer
The Reasoning Layer
Three properties, and the third is the one that surprises people.
HOW IT DECIDES
It measures your checkpoint

It solves, rather than guesses

It needs no data

Measurements
TARGET : 12.6 GB
12.6 GB
MODEL ON DISK
32K
CONTEXT HELD
16 GB
GPU MEMORY
HARDWARE : ONE 16 GB GPU, THE CLASS EVERY CLOUD RENTS BY THE HOUR
Hardware
g4dn.xlarge: 16 GB NVIDIA T4. The entry GPU instance, available in every region and on spot pricing.NC4as_T4_v3: 16 GB NVIDIA T4, the smallest GPU size in the NC family.g2-standard-4: 24 GB NVIDIA L4, with headroom above what the model needs.The output is an ordinary GGUF file. It runs under unmodified llama.cpp, Ollama or LM Studio, on any CUDA GPU, with no runtime of ours anywhere in the path. Our published measurements were taken on a 16 GB NVIDIA card on our own hardware; the cloud instances above are the same class of accelerator.
Scale



Limits