Buying Guides
How Much System RAM Do You Actually Need for Local AI
October 5, 2026
When people plan a local AI build they obsess over VRAM and forget that system RAM has its own job to do, and getting it wrong either wastes money or quietly bottlenecks the machine.
VRAM and system RAM do different jobs
The GPU's VRAM is where the model weights and the active context actually live while inference happens, and that is the number that determines whether a given model fits and runs fast. System RAM is a different pool entirely, managed by the CPU, and it handles everything around the model: the operating system, your browser and IDE, background services, the inference server's own overhead, and in some setups, any part of the model that spills out of VRAM.
People sometimes assume more system RAM makes a model run faster the way more VRAM does, and that is only true in a narrow sense. If a model fits entirely in VRAM, extra system RAM does nothing for inference speed. It matters for everything happening alongside inference, and for how gracefully the machine handles larger models or multiple things running at once.
A practical baseline: 32GB
For a single-GPU machine running models that comfortably fit in 16GB of VRAM, 32GB of system RAM is a reasonable baseline. That covers the OS, a modern browser with a dozen tabs, an IDE, a local inference server, and some headroom without constantly hitting swap. This is the configuration on Spectre's Starter and Pro tiers, both built around 16GB cards, and for the workloads those cards are meant to handle, 32GB keeps pace without being wasteful.
Where 32GB starts to feel tight is when you are doing more than running one model. If you are running an inference server, a vector database, a coding agent, and a few browser tabs with large contexts open all at once, 32GB can get consumed quickly, and you will notice it in system responsiveness rather than in model speed.
Why 64GB is the sensible default for serious work
Once you move to larger GPUs or start running agent workflows with multiple processes, 64GB becomes the more sensible default, which is why it shows up on the Ultra, Apex, and Apex Dual tiers. An AI coding agent, for instance, is not just an inference call, it is a long-running process that holds context, manages tool calls, and often keeps a local index of your codebase in memory alongside whatever model server it is talking to. Stack a few of those alongside normal development work and 64GB gives you room to breathe instead of running close to the edge constantly.
64GB also matters if you ever want to run smaller models directly on CPU for testing or run multiple models side by side for comparison. None of that requires exotic amounts of RAM, but it does require enough that the system is not swapping to disk, which causes the kind of stutter that is easy to blame on the GPU when the real cause is memory pressure elsewhere.
When 128GB actually earns its place
128GB is overkill for a single model running inference, and it is not there to make a 32B model run faster. It earns its place on dual-GPU and very large single-GPU machines for two reasons: those builds tend to run heavier, more parallel workloads, and the kind of buyer choosing a machine like the Apex Dual or Apex Pro is usually running several things simultaneously rather than one chat interface. Multiple agents, a database, containerized services, and a model server all running at once can add up fast, and 128GB means none of that competes for memory with the model itself.
It is also a sensible match for the RTX PRO 6000 Blackwell in the Apex Pro, which is the only configuration built around 70B-class models. A model operating at that scale is usually part of a more involved pipeline rather than a simple one-off query, and the system RAM needs to match the seriousness of everything running around it.
Matching RAM to how you actually plan to use the machine
The right amount of RAM depends less on the model size alone and more on how many things will be running at once, and that is worth thinking through honestly before you spec a machine. Someone running a single 14B model through a chat interface in the evenings has very different needs than someone running an agent pipeline with several concurrent processes during the workday, even if both are using similarly sized GPUs.
If you are unsure where you fall, it is usually safer to go one tier up on RAM than to go one tier up on GPU, since RAM is the cheaper way to buy yourself headroom. You can look through the pre-configured builds to see how RAM is paired with each GPU tier, and if your workload does not match any of the standard configurations, a custom build lets you adjust the RAM independently of everything else.