This post was originally published on go2linux.org. The domain is no longer mine, but I am the original author. I am republishing it here on garron.me with corrections and improvements.
vmstat (virtual memory statistics) prints a compact, one-line-per-sample view of processes, memory, swap, disk I/O and CPU. I started using it when a small VPS of mine kept running out of memory: the CPU usage climbed until the server had to be restarted, and vmstat showed that the real problem was swapping.
It is part of the procps package and is installed on practically every Linux system.
Basic usage
vmstat [options] [delay [count]]
delay is the number of seconds between samples and count is how many samples to take. With a delay and no count it runs until you press Ctrl+C.
vmstat 5 6
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
0 1 0 986900 103532 492096 0 0 24 31 320 943 10 2 87 1 0
0 0 0 986760 103536 492212 0 0 21 0 714 1706 7 2 91 0 0
0 0 0 986760 103544 492212 0 0 0 4 877 3647 13 3 85 0 0
1 0 0 983528 103576 492232 0 0 0 64 667 1816 6 1 93 0 0
1 0 0 984768 103588 492232 0 0 0 9 767 1979 12 3 85 0 0
0 0 0 984396 103612 492388 0 0 0 90 942 3215 19 3 77 0 0
The first line is different from the rest. It shows averages since the last boot, not the current state. Ignore it when you are looking for a problem happening now, and never run plain vmstat with no delay to judge the current load.
The columns
procs
r— processes running or waiting for a CPU.b— processes blocked, normally waiting for I/O.
memory
Values are in kibibytes by default.
swpd— swap space in use. A non-zero value here is not a problem by itself: the kernel may have moved pages that nobody touched for days out to swap and left them there.free— memory not used for anything.buff— buffers, mostly file system metadata.cache— the page cache: contents of files read from or written to disk, kept in memory in case they are needed again. That is why opening a large program the second time is faster.
A low free value is normal. Linux uses spare memory for cache and gives it back as soon as a program needs it. To know how much memory is really available to applications, look at the available column of free -h.
swap
si— memory swapped in from disk, per second.so— memory swapped out to disk, per second.
These two are the columns that tell you whether the machine is short of memory.
io
bi— blocks read from block devices, per second.bo— blocks written to block devices, per second.
system
in— interrupts per second.cs— context switches per second.
cpu
Percentages of total CPU time.
us— user code.sy— kernel code.id— idle.wa— idle while waiting for I/O to finish.st— steal: time the hypervisor gave to other virtual machines while this one wanted to run.
Recent versions add a gu column for time spent running guest virtual machines.
How to read it
Memory shortage. si and so are zero on a healthy system. An occasional burst is fine. If they stay above zero sample after sample, the system is moving pages to disk and back because physical memory is not enough:
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
2 3 812340 11208 904 23116 620 940 1480 1920 980 2100 12 9 21 58 0
1 4 818992 10876 872 22640 710 1105 1730 2210 1040 2350 10 11 17 62 0
3 3 825120 11340 860 21988 655 980 1520 2004 1010 2290 14 10 19 57 0
Swap in use keeps growing, free and cache are squeezed to almost nothing, and wa is high because the CPU sits waiting for the disk. The server feels frozen although us is low. The fix is to find what is using the memory (see how to find which process is eating RAM), reduce it, or add memory.
CPU saturation. Compare r with the number of CPUs (nproc). If r stays above that number and id is near zero, there is more work than processors.
Disk bottleneck. High b and high wa with si/so at zero mean processes are waiting on disk I/O that is not swapping. Continue with iotop to find the process.
Noisy neighbors. On a VPS, a st value that stays in the double digits means the host is oversold or your plan is CPU-limited. Nothing you change inside the virtual machine will fix it.
Useful options
Wide output, so the columns do not run into each other on machines with a lot of memory:
vmstat -w 5
Show memory in mebibytes instead of kibibytes:
vmstat -S M 5
Add a timestamp to every line, useful when you redirect the output to a file:
vmstat -t 5
Show active and inactive memory instead of buffers and cache:
vmstat -a 5
A table of counters since boot (total memory, pages swapped, forks, and so on):
vmstat -s
Per-disk statistics, or the summary for one partition:
vmstat -d
vmstat -p /dev/sda1
Record it while you wait for the problem
When the server slows down at unpredictable times, leave vmstat logging and read the file afterwards:
nohup vmstat -t -w 10 > vmstat.log &
One line every ten seconds is about 8,600 lines a day, small enough to keep for a week.
See also
man vmstat — full reference. free -h for a simple memory summary, and top or htop to see which processes are responsible.