RAM, cache and memory pressure: why free memory is not the metric
Why free memory is not the metric: pages, cache, reclaim and memory pressure.

A host with 60 GB of 64 GB “in use” is not automatically in trouble, and a host with 50 GB “free” is not automatically healthy. Memory is a hierarchy that ends in a reclaim path, and the number that decides whether a workload is comfortable is not how much memory is free, but how much memory can be made available quickly enough.
The model: a hierarchy with a reclaim path
application
↓ asks for pages
L1/L2/L3 cache → RAM → swap / page file → disk
↓ ↑
fastest, smallest overflow, slowest
↑
page cache (files read from disk stay here)
When a process asks for memory the kernel hands it pages. Pages that are not actively needed can be taken back — reclaimed — and given to somebody else. The hierarchy matters because every step away from the CPU costs latency: caches close to the processor exist precisely to avoid the cost of going to memory, and a cache miss is what the memory subsystem pays for. Performance problems that look like “the disk is slow” are very often “the working set no longer fits”.
Terms used here
- Page — the unit the kernel manages memory in; processes see virtual addresses, the kernel maps them onto physical frames.
- Page cache — the kernel’s copy of file data (and some filesystem metadata) kept in RAM to avoid going to disk again.
- Anonymous memory — memory a process allocated that is not backed by a file (a heap, a stack).
- Reclaim — freeing reclaimable physical pages and repurposing them.
- Swap / page file — the place anonymous pages go when RAM cannot hold them.
- Working set — the memory a process is actively using, as seen by the OS.
- Committed memory — virtual memory the system has promised to processes, which may be larger than the physical memory installed.
The numbers, and which one answers the question
On Linux, /proc/meminfo is the classic source, and its field names are worth knowing because every
monitoring tool is a view of them:
MemTotal— total usable RAM (physical RAM minus a few reserved regions and the kernel itself).Cached— in-memory cache for files read from disk, not counting memory that was swapped out and swapped back in (SwapCached).MemAvailable— an estimate of how much memory is available for starting new applications without swapping. This, notMemFree, is the field that answers “can I start something else?”.
On Windows the counters have different names for the same ideas: Physical Total and Physical Available for RAM, Committed Bytes and Commit Limit for the virtual memory the system has promised and the ceiling it can promise against RAM plus page file, and Available MBytes for remaining free RAM. The working set is the part of a process that is resident.
The practical rule is therefore not “keep memory free” but “know what is reclaimable”: the kernel’s own documentation classifies pages as reclaimable (page cache, anonymous memory) or unreclaimable, and only the second category is a hard commitment.
What memory pressure actually is
Memory pressure is not a percentage; it is a behaviour. It appears when demand for pages exceeds what can be freed fast enough:
- Reclaim runs in the background. A kernel thread reclaims pages asynchronously so that normal allocation does not have to pay for it.
- Direct reclaim appears. If the allocator cannot find a page, it reclaims inside the allocation request: the allocation is stalled until enough pages are freed. This is where an otherwise healthy-looking host starts showing long tail latency.
- Swap or the page file is used. Anonymous pages that are not needed right now are written out. When they are needed again they must be read back — the cost shows up as latency, not as a percentage.
- The OOM killer arrives. If the kernel cannot reclaim enough memory to continue, it kills a process. That is the failure mode at the end of the path.
A large file cache is not waste: it is work already avoided. A host whose “free” memory is near zero
while MemAvailable (or Physical Available) stays high is doing its job.
A common misconception: free memory is good memory
Reading only the “free” figure inverts the picture, because unused RAM produces nothing. The two
figures that matter together are available memory (what can be handed to a new workload without
swapping) and the reclaim behaviour under load (whether allocations start to stall or the swap
path lights up). Sizing, overcommit policy, cgroup memory limits and NUMA placement are separate
subjects, and they belong to the operational material, not to this definition.
What to remember
- RAM is a hierarchy ending in a reclaim path, not a single number.
- Free memory is memory nothing has needed yet; cached memory is memory already working for you.
MemAvailable(Linux) and Physical Available (Windows) answer “can I start something else?”;MemFreealone does not.- Memory pressure is a behaviour: background reclaim, then stalled allocations, then swap, then the OOM killer.
- Committed memory can exceed installed RAM; that is why the commit limit exists.
Level and prerequisites
L1 — fundamentals. Prerequisites: none beyond the idea that an operating system manages memory for
processes. The sheet covers the vocabulary and the mechanism; sizing, tuning, cgroup limits, NUMA
and swap design are L2–L4 material.
Where to go next
- Server & Virtualization — the area this sheet belongs to.
References
- Linux kernel documentation — The /proc Filesystem (
/proc/meminfofield definitions). - Linux kernel documentation — Memory Management concepts (page cache, reclaimable pages, reclaim, direct reclaim, anonymous memory, OOM).
- Intel — Intel Hyper-Threading Technology Technical User’s Guide (2003): cache hierarchy and the cost of a cache miss.
- Microsoft Learn — Memory Performance Information (physical, committed and cache counters).
- Microsoft Learn — Troubleshoot memory consumption between identical Windows Server environments (Available MBytes, working set, paging counters).