Local AI

What Quantization Actually Does to a Local LLM

September 28, 2026

If you have looked at a model card and seen labels like Q4_K_M or FP16 and wondered what they mean for your hardware, quantization is the answer. It is the single biggest lever for fitting a large model into a GPU that would otherwise be too small, and understanding it changes how you shop for a local AI machine.

Quantization is compression, not magic

Every parameter in a neural network is a number, and the precision used to store that number determines how much memory the model takes up. Full precision, usually FP16 or BF16, stores each parameter using 16 bits. Quantization reduces that to fewer bits per parameter, most commonly 8, 5, or 4 bits, which shrinks the file size and the VRAM footprint roughly in proportion to the bit reduction.

This is why a 7B model at FP16 needs roughly 14GB of VRAM just to load the weights, while the same model quantized to 4-bit needs somewhere around 4 to 5GB. The parameter count has not changed. The precision used to represent each weight has, and that precision loss is the tradeoff you are making in exchange for fitting on smaller hardware.

Why 4-bit became the default for local setups

Among the common quantization levels, 4-bit formats such as Q4_K_M have become the de facto standard for local inference because they hit a sweet spot between size and usable output quality. Going from FP16 down to 8-bit roughly halves memory use with minimal perceptible quality loss for most tasks. Going further to 4-bit roughly halves it again, and for most general-purpose chat, coding assistance, and retrieval-augmented tasks, the difference in output quality compared to 8-bit is small enough that most people would not notice without a side-by-side comparison.

Going below 4-bit, into 3-bit or 2-bit territory, is where things get noticeably worse. Reasoning tasks, long-context coherence, and code generation are usually the first things to degrade, so most people who care about output quality treat sub-4-bit quantization as a last resort for squeezing a model onto hardware that is otherwise too small, rather than a normal operating mode.

What this means for VRAM planning

Because quantization changes the memory math, the useful way to plan hardware is to think in terms of a model's parameter count at your target quantization level, not the raw parameter count alone. A 14B model at 4-bit needs roughly 8 to 10GB of VRAM including context overhead, which comfortably fits on a 16GB card. A 32B model at 4-bit needs closer to 20GB, which is why that class of model is generally the ceiling for a single 16GB or even 24GB-class card once you account for context length and any additional overhead from the runtime.

This is the reasoning behind how Spectre's pre-configured builds are tiered. The 16GB cards in the Starter, Pro, and Ultra builds are sized for models up to roughly 14B at practical quantization levels, while the 32GB RTX 5090 in the Apex, and the combined 64GB of VRAM in the Apex Dual, open up room for 32B-class models with more headroom for longer context windows. None of these are marketed as running 70B models, because that class of model needs substantially more memory even at 4-bit, which is a job for the Apex Pro's 96GB card rather than anything smaller.

Quality tradeoffs are task-dependent, not universal

A common mistake is assuming quantization loss applies evenly across every use case, when in practice it depends heavily on what you are asking the model to do. Casual conversation, summarization, and general Q&A tend to hold up well even at 4-bit. Tasks that require precise multi-step reasoning, exact code syntax, or careful adherence to complex instructions are more sensitive, and that sensitivity tends to get worse as you go to lower bit depths or as the underlying model itself gets smaller.

This is also why picking a larger model at a lower quantization level is not automatically better than picking a smaller model at a higher one. A 32B model at 4-bit will often outperform a 14B model at 8-bit on tasks that benefit from more parameters, but the reverse can be true for tasks where precision in each individual weight matters more than raw model size. There is no universal answer here, which is why testing a candidate model at your intended quantization level on your actual workload matters more than trusting any general rule of thumb.

Matching quantization strategy to hardware you actually own

Once you understand the memory math, hardware selection becomes a straightforward exercise: decide what model size and quantization level covers your actual workload, then buy enough VRAM to run it with room to spare for context length, which also consumes memory as conversations or documents get longer. Buying a card that just barely fits a model at 4-bit leaves no margin for longer context windows or for running a second model alongside it, which is a common frustration for people who sized their hardware too tightly.

If your workload does not map cleanly onto one of the standard tiers, for example if you need a specific combination of VRAM and CPU core count for a mixed inference and fine-tuning workflow, it is worth looking at a custom build rather than forcing a fit into a pre-configured option. Getting the VRAM-to-quantization ratio right up front is cheaper than discovering months later that your model barely fits and every context-heavy task is running against a hard ceiling.

Want this installed in your office?

See the levels of local AI, pick a machine, or browse our custom-built towers.