This post was originally published on go2linux.org. The domain is no longer mine, but I am the original author. I am republishing it here on garron.me with corrections and improvements.

vmstat (virtual memory statistics) prints a compact, one-line-per-sample view of processes, memory, swap, disk I/O and CPU. I started using it when a small VPS of mine kept running out of memory: the CPU usage climbed until the server had to be restarted, and vmstat showed that the real problem was swapping.

It is part of the procps package and is installed on practically every Linux system.

Basic usage

vmstat [options] [delay [count]]

delay is the number of seconds between samples and count is how many samples to take. With a delay and no count it runs until you press Ctrl+C.

vmstat 5 6
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 0  1      0 986900 103532 492096    0    0    24    31  320  943 10  2 87  1  0
 0  0      0 986760 103536 492212    0    0    21     0  714 1706  7  2 91  0  0
 0  0      0 986760 103544 492212    0    0     0     4  877 3647 13  3 85  0  0
 1  0      0 983528 103576 492232    0    0     0    64  667 1816  6  1 93  0  0
 1  0      0 984768 103588 492232    0    0     0     9  767 1979 12  3 85  0  0
 0  0      0 984396 103612 492388    0    0     0    90  942 3215 19  3 77  0  0

The first line is different from the rest. It shows averages since the last boot, not the current state. Ignore it when you are looking for a problem happening now, and never run plain vmstat with no delay to judge the current load.

The columns

procs

  • r — processes running or waiting for a CPU.
  • b — processes blocked, normally waiting for I/O.

memory

Values are in kibibytes by default.

  • swpd — swap space in use. A non-zero value here is not a problem by itself: the kernel may have moved pages that nobody touched for days out to swap and left them there.
  • free — memory not used for anything.
  • buff — buffers, mostly file system metadata.
  • cache — the page cache: contents of files read from or written to disk, kept in memory in case they are needed again. That is why opening a large program the second time is faster.

A low free value is normal. Linux uses spare memory for cache and gives it back as soon as a program needs it. To know how much memory is really available to applications, look at the available column of free -h.

swap

  • si — memory swapped in from disk, per second.
  • so — memory swapped out to disk, per second.

These two are the columns that tell you whether the machine is short of memory.

io

  • bi — blocks read from block devices, per second.
  • bo — blocks written to block devices, per second.

system

  • in — interrupts per second.
  • cs — context switches per second.

cpu

Percentages of total CPU time.

  • us — user code.
  • sy — kernel code.
  • id — idle.
  • wa — idle while waiting for I/O to finish.
  • st — steal: time the hypervisor gave to other virtual machines while this one wanted to run.

Recent versions add a gu column for time spent running guest virtual machines.

How to read it

Memory shortage. si and so are zero on a healthy system. An occasional burst is fine. If they stay above zero sample after sample, the system is moving pages to disk and back because physical memory is not enough:

procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 2  3 812340  11208    904  23116  620  940  1480  1920  980 2100 12  9 21 58  0
 1  4 818992  10876    872  22640  710 1105  1730  2210 1040 2350 10 11 17 62  0
 3  3 825120  11340    860  21988  655  980  1520  2004 1010 2290 14 10 19 57  0

Swap in use keeps growing, free and cache are squeezed to almost nothing, and wa is high because the CPU sits waiting for the disk. The server feels frozen although us is low. The fix is to find what is using the memory (see how to find which process is eating RAM), reduce it, or add memory.

CPU saturation. Compare r with the number of CPUs (nproc). If r stays above that number and id is near zero, there is more work than processors.

Disk bottleneck. High b and high wa with si/so at zero mean processes are waiting on disk I/O that is not swapping. Continue with iotop to find the process.

Noisy neighbors. On a VPS, a st value that stays in the double digits means the host is oversold or your plan is CPU-limited. Nothing you change inside the virtual machine will fix it.

Useful options

Wide output, so the columns do not run into each other on machines with a lot of memory:

vmstat -w 5

Show memory in mebibytes instead of kibibytes:

vmstat -S M 5

Add a timestamp to every line, useful when you redirect the output to a file:

vmstat -t 5

Show active and inactive memory instead of buffers and cache:

vmstat -a 5

A table of counters since boot (total memory, pages swapped, forks, and so on):

vmstat -s

Per-disk statistics, or the summary for one partition:

vmstat -d
vmstat -p /dev/sda1

Record it while you wait for the problem

When the server slows down at unpredictable times, leave vmstat logging and read the file afterwards:

nohup vmstat -t -w 10 > vmstat.log &

One line every ten seconds is about 8,600 lines a day, small enough to keep for a week.

See also

man vmstat — full reference. free -h for a simple memory summary, and top or htop to see which processes are responsible.