Low-Latency Software · Technology
13Linux Tuning
Run a thread that does nothing but read the time-stamp counter in a loop, and record every gap between two consecutive readings longer than , when one pass of the loop takes about fifteen. Each such gap is a moment when the thread was not running: something else had the CPU. On the laptop of this book, pinned to a CPU that no other program of ours uses, the thread loses about 1.3% of its time: some 750 times a second for five microseconds or more, a few times a second for more than a hundred, and every few seconds for a millisecond or more. A market-data message that arrives during a gap waits for the rest of it. None of this is the program’s fault, and none of it can be fixed in the program: it is the operating system doing its own work on the CPU the program thought it had. This chapter is about taking the CPU back. It moves the kernel’s work, the interrupts and the timer tick off the cores that run the hot path, replaces sleeping with polling, keeps the program’s memory in RAM, and ends with a build that audits a host against the placement plan of chapter 4 before a trading process is allowed to start on it.
13.1 Core isolation and CPU affinity
Definition 13.1 (CPU affinity, core isolation, housekeeping core)
The CPU affinity of a thread is the set of CPUs the scheduler may run it on; it is set with sched_setaffinity (or taskset from a shell) and inherited by the threads a thread creates. Core isolation removes CPUs from the scheduler’s load balancing, so that no thread runs on them unless its affinity names them explicitly, and moves the kernel’s own unbound work away from them. A housekeeping core is a CPU left out of the isolated set, which runs everything else: the kernel’s work, the interrupts of devices the hot path does not use, and the process’s cold threads.
Affinity alone fixes where a thread runs, not what else runs there: a pinned thread still shares its CPU with any other runnable task the scheduler places on it, with kernel threads, and with interrupts. The kernel documentation calls all of these “noise”, asynchronous (interrupts, timers, work queues, kernel threads) or synchronous (system calls and page faults), and notes that it usually goes unnoticed: the timer interrupt can run 1 024 times a second “without a significant and measurable impact most of the time”. The exceptions it names are workloads such as very low latency network processing. Isolation is how the noise is moved: the boot parameter isolcpus (with its default domain flag) takes CPUs out of “the general SMP balancing and scheduling algorithms”, irreversibly until the next boot, and control-group cpusets do the same at run time. A thread reaches an isolated CPU only through its affinity, which is exactly what chapter 9’s firm.affinity sets from chapter 4’s plan.
Figure 13.1 shows what the hook’s thread sees in three situations on this untuned laptop. Pinned alone or left free to move, it loses about one percent of its time in gaps of microseconds, and a few gaps a second reach a hundred microseconds or more. Sharing its CPU with a second spinning thread, it loses about half of it, about 120 times a second for about four milliseconds each: the scheduler hands the CPU to the rival for a slice, and a slice of that length matches the period of this kernel’s tick, which runs 250 times a second (its configuration, CONFIG_HZ). The shared case is the one that isolation removes outright; the small gaps of the other two need the rest of this chapter, and some of them, on a virtual machine, belong to the hypervisor, which the guest cannot see or tune.
bench_tuning.py.On a tuned host the plan of chapter 4 becomes a boot configuration. Figure 13.2 draws it for the two-socket fixture of that chapter: the three hot threads on three cores of the network card’s socket, their hyper-threading siblings taken offline (the kernel’s own checklist says to avoid SMT, “to prevent your hardware thread from being ‘preempted’ by another one”), the card’s interrupts on housekeeping CPUs of the same socket, and everything else on the rest. The documentation asks for at least one housekeeping CPU, preferably one per NUMA node; a trading host keeps several, since the logger, the monitoring agents and the kernel all live there, and a housekeeping CPU that is itself overloaded delays the work (log draining, timers) that the hot path hands to it.
isolcpus=21-23,29-31 nohz_full=21-23 rcu_nocbs=21-23 irqaffinity=0-15 with the siblings taken offline, the card’s interrupt queues bound to CPUs 16–20; the build’s audit accepts this host (the tuned fixture).13.2 Interrupts and their affinity
Definition 13.2 (Interrupt affinity)
The interrupt affinity of an interrupt line is the set of CPUs allowed to handle it, read and written in /proc/irq/N/smp_affinity_list; the boot parameter irqaffinity sets the default for lines that no one configures.
An interrupt runs its handler on whichever CPU receives it, in the middle of whatever that CPU was doing: the network card announcing a packet, a disk finishing a write, another CPU asking this one to flush its translation buffer. On an untuned machine every line may go to every CPU: the audit of this laptop in the build below finds the default mask and all of its active lines allowed on the CPUs that chapter 4 would give the hot path. The kernel documentation’s instruction is short: “isolate the IRQs whenever possible, so that they don’t fire on the target CPUs”. Two details make it harder than writing a mask. Many distributions run irqbalance, a daemon that periodically spreads interrupts across CPUs and would undo a hand-made placement; the Red Hat real-time tuning guide says it is not needed where applications are bound to CPUs, and disables it. And some interrupts belong on a hot core’s socket but not on the core: the network card’s receive queues should interrupt a housekeeping CPU of the card’s node (Figure 13.2), so that the kernel’s packet processing leaves the data in the socket’s shared cache, where the hot threads read it, rather than across the interconnect (chapter 4).
As of September 2026 — Distribution tuning profiles
Red Hat’s guide to optimising RHEL 9 for Real Time for low-latency operation (consulted September 2026) walks through the same steps as this chapter: isolating CPUs with the TuneD real-time profile’s isolated_cores variable, disabling irqbalance and binding interrupts by hand, locking memory with mlock and mlockall, and the real-time throttling parameters of Section 13.5. Profiles and their variable names change between releases; the kernel parameters they set are the stable part.
13.3 Tickless kernels and housekeeping
Definition 13.3 (Tickless kernel)
A tickless kernel (adaptive ticks, nohz_full) stops the periodic scheduling-clock interrupt on listed CPUs while each runs a single task, and moves the work that interrupt did for them (timekeeping, a residual once-a-second tick, the callbacks of the read-copy-update mechanism) to housekeeping CPUs.
The scheduling-clock interrupt, the tick, lets the kernel account time, run timers and decide whether to switch tasks, 100 to 1 000 times a second depending on the build. A CPU with a single runnable task has nothing to switch to, and the kernel’s documentation on adaptive ticks puts the gain precisely: real-time applications “improve their worst-case response times by the maximum duration of a scheduling-clock interrupt”. The same document lists the conditions: the kernel must be built with CONFIG_NO_HZ_FULL (this laptop’s is not), the CPU must run one task only, the thread must stay out of the kernel (system calls and page faults bring the noise back), and at least one CPU must keep ticking for timekeeping. Read-copy-update callbacks, deferred work the kernel queues on each CPU, go with the tick: nohz_full offloads them from the listed CPUs, and rcu_nocbs does so explicitly. The price is paid elsewhere: entering and leaving the kernel costs more on a tickless CPU, and the housekeeping CPUs take the offloaded work.
What does a gap cost a message? The hook’s thread spends its time either running or stopped, and a message arriving at a random moment waits only if it arrives during a stop.
Proposition 13.4 (The wait added by gaps)
A thread that must handle each message immediately is stopped during gaps of lengths in an observation of length . A message arriving at a uniformly random moment of waits, beyond its own processing,
and the fraction of time during which it would wait at all is .
Proof. The message falls inside gap with probability , and then its position in the gap is uniform, so it waits a remaining time uniform on , with mean and . Outside the gaps it does not wait. Summing over the gaps gives and . ∎
The squares matter: the wait is dominated by the long gaps, not the frequent ones. A tick of at 250 Hz steals 0.05% of the time and adds half a nanosecond on average; a single gap of a second steals 0.1% but adds on average and makes one message in two thousand wait more than half a millisecond. On the laptop, pinned alone, the meter’s gaps give an average added wait of about three microseconds (it moves by a factor of two or more from run to run, with the run’s longest gap: the squares at work), a wait beyond for about 0.45% of messages and beyond for 0.1 to 0.2%: the host adds a 99.9th percentile of over a hundred microseconds to a path whose own work takes a few microseconds or less (chapter 1). The proposition assumes messages arrive independently of the gaps; bursts that coincide with the kernel’s work (an interrupt from the same card that brings the data) make it worse.
13.4 Busy polling and kernel bypass
Definition 13.5 (Context switch, busy polling, kernel bypass)
A context switch is the kernel saving the state of the thread running on a CPU and restoring another’s; a thread that blocks waiting for data gives its CPU away and is switched back in when the data arrives. Busy polling checks for new data in a loop that never blocks, keeping the CPU. Kernel bypass moves the network card’s receive and transmit queues into the application’s address space, so that packets reach the program without a system call, an interrupt or the kernel’s network stack.
A blocking receive is the default and the right choice for almost every program: the thread sleeps, the CPU does other work or saves power, and the kernel wakes the thread when the packet has gone through the interrupt handler, the protocol stack and the socket’s buffer (Figure 13.3, left). Each step costs microseconds, and the wake-up is the worst of them: the sleeping thread’s CPU may be idle in a low-power state, and the thread must be scheduled back in. Busy polling removes the sleep: the thread asks for data in a loop, with a non-blocking call, and finds it as soon as the kernel has put it in the socket (middle). The socket option SO_BUSY_POLL makes the kernel itself spin for up to a set number of microseconds inside a blocking receive (socket(7)); the tutorial polls from user space instead, which works on any socket. Kernel bypass removes the kernel: in the open-source DPDK, a poll-mode driver “accesses the RX and TX descriptors directly without any interrupts”, in a run-to-completion loop on a dedicated core (right). One Quant Book 14 covers the stacks and the cards; what matters here is that each step to the right trades a CPU spinning at 100% for microseconds.
Figure 13.4 measures the first two on the laptop’s loopback interface, where no card is involved and the difference is the sleep alone. A round trip of 64 bytes between two pinned threads takes about at the median with blocking receives and about busy-polled, over ten times less. Two threads bouncing a byte through a pair of pipes on the same CPU complete a round trip in about : two context switches with their system calls, each well under a microsecond when the CPU is awake. The same ping-pong between two CPUs takes about : the cost is not the switch but waking a CPU that went idle while waiting, which is also why the blocking receive is slow. Kernel bypass, which removes the system calls as well, needs a card that supports it, which this laptop does not have.
bench_tuning.py.The tails tell the rest: busy polling cuts the median, not the p99.9, which on this laptop is still tens of microseconds, because the spinning thread meets the gaps of Figure 13.1 like any other. Polling and isolation go together: a thread that polls on a CPU where the kernel still interrupts it is fast on average and slow in the tail, and an isolated CPU whose thread sleeps in a blocking call wakes up slowly. The idle states are the third part: a CPU that is never idle never enters a deep one, but the housekeeping CPUs that wake the hot path’s helpers can, and the kernel documentation lists intel_idle.max_cstate= as the way to limit their depth, and idle=poll as a stronger one that “can cause your CPU to overheat” and disables turbo frequencies.
13.5 Real-time priorities and memory locking
Definition 13.6 (Real-time scheduling policy, memory locking)
A real-time scheduling policy (SCHED_FIFO or SCHED_RR in Linux) gives a thread a static priority above every ordinary thread: a runnable real-time thread preempts any ordinary one and, under SCHED_FIFO, runs without time slicing until it blocks, yields or is preempted by a higher priority. Memory locking (mlock, mlockall) keeps a process’s pages in physical memory: they are never paged out, and with MCL_FUTURE every page of every later mapping is made present when the mapping is created.
A real-time policy answers the shared-CPU curve of Figure 13.1: the rival no longer gets a slice. On an isolated core with a single thread there is no rival to preempt, and the policy brings a risk instead. A real-time thread that spins forever would starve the CPU’s kernel threads, so Linux throttles real-time threads by default: they may run of every second, and the remaining go to ordinary tasks (the kernel’s documentation of real-time group scheduling). A spinning SCHED_FIFO thread on a CPU that also has an ordinary task runnable is therefore stopped for up to once a second, a hiccup a thousand times longer than any on the untuned laptop. Throttling can be disabled (sched_rt_runtime_us set to ), which the Red Hat guide calls adequate only when the real-time tasks are well engineered. An ordinary process may not ask for the policy at all: on this laptop the request fails with EPERM. The firm’s design therefore spins ordinary threads on isolated cores, where there is nothing to preempt, and keeps real-time priorities for threads that must preempt others on shared CPUs.
Page faults are the synchronous noise of the kernel documentation, and chapter 6 measured them: the first write to a page costs a microsecond or more. The mlockall manual page names the reason to lock: “paging is one major cause of unexpected program execution delays” for real-time algorithms. On the laptop, writing one byte to each page of a fresh 32 MiB mapping takes 8 192 page faults and about one to two microseconds a page. After mlockall(MCL_CURRENT | MCL_FUTURE), the same writes take no fault and about a page: the cost has moved to the mmap call, which now populates every page up front, at start-up, where it belongs. Locking is limited by RLIMIT_MEMLOCK, 64 MiB for an ordinary user on this machine: a process whose book and rings need a gigabyte must be given the limit (the build checks it), or mlockall fails with ENOMEM.
13.6 Tutorial: measure an untuned machine
Goal. Measure what this machine does to a thread that wants its CPU, and audit it against a placement plan. End state: Figures 13.1 and 13.4, the page-fault counts with and without locking, and the audit’s report for this machine next to the two fixtures’.
The hiccup meter. A loop reads the time-stamp counter and records every gap above a threshold in a buffer reserved before the loop; it allocates nothing and makes no system call, so every gap it sees was imposed on it.
template <class Now> Hiccups hiccup_meter(Now now, std::uint64_t duration, std::uint64_t threshold, std::size_t max_gaps) { Hiccups h; h.gaps.reserve(max_gaps); const std::uint64_t start = now(); std::uint64_t prev = start; while (prev - start < duration) { const std::uint64_t t = now(); const std::uint64_t gap = t - prev; if (gap > threshold) { h.stolen += gap; if (gap > h.worst) h.worst = gap; if (h.gaps.size() < max_gaps) h.gaps.push_back(gap); } prev = t; ++h.loops; } h.elapsed = prev - start; return h; }Listing 13.1. The hiccup meter’s loop, generic over its clock so that the test can script one. code/low-latency/13-linux-tuning/cpp/ll_tuning.hpp - Run it with
python bench_tuning.py: five seconds pinned alone withtaskset, five seconds unpinned, and five seconds pinned beside a second spinning thread; the script converts the gaps into Figure 13.1 and applies Proposition 13.4 to them. Poll instead of sleeping. The same receive, blocking or in a non-blocking loop:
// Receive one datagram: a blocking call that sleeps until data arrives, or a loop of non-blocking calls that never // gives up the CPU (busy polling from user space). static ssize_t receive(int s, char* buf, bool spin) { if (!spin) return recv(s, buf, 64, 0); for (;;) { const ssize_t n = recv(s, buf, 64, MSG_DONTWAIT); if (n >= 0 || errno != EAGAIN) return n; } }Listing 13.2. A receive that sleeps and one that spins. code/low-latency/13-linux-tuning/cpp/ll_tuning_bench.cpp Lock the memory and count the faults. The C++ test maps 8 MiB, touches it (2 048 faults), then calls
mlockalland maps and touches another 8 MiB (none); if the process may not lock, it says so and skips that half.// Write one byte per 4 KiB page; returns the minor page faults the writes took. inline long touch_faults(volatile char* p, std::size_t bytes) { const long f0 = minor_faults(); for (std::size_t i = 0; i < bytes; i += 4096) p[i] = 1; return minor_faults() - f0; } // Lock every present and future page of the process in memory: 0, or the errno (ENOMEM above RLIMIT_MEMLOCK, // EPERM without the privilege). inline int lock_all() { return mlockall(MCL_CURRENT | MCL_FUTURE) == 0 ? 0 : errno; }Listing 13.3. Counting page faults, and locking every present and future page. code/low-latency/13-linux-tuning/cpp/ll_tuning.hpp - Ask for a real-time policy (
ll_tuning_bench privileges): an ordinary process getsEPERM. - Audit the machine:
python firm_tuneaudit.py /plans the hot threads withfirm.coreplanand checks the host against the plan (next section).
What to change next. Run the meter on a CPU while another program (a compile, a test suite) runs on its hyper-threading sibling: the gaps do not change, but count the loop’s passes a second. On a Linux host you administer, boot with isolcpus and nohz_full on two CPUs, move their interrupts away, and run the meter there.
13.7 Build: the host audit
firm.tuneaudit reads the same files an engineer would, from /proc and /sys or from a fixture tree with their layout, and checks each rule of this chapter against the plan of firm.coreplan: the hot CPUs isolated, their siblings offline or isolated, the hot CPUs tickless and their read-copy-update callbacks offloaded, no interrupt line (nor the default mask) allowed on a hot CPU, irqbalance not running, the performance frequency governor on the hot CPUs, the idle states capped on the boot line, transparent huge pages only where asked for (madvise, with no stalling always defragmentation, which the kernel documentation says makes an allocation “stall on allocation failure and directly reclaim pages and compact memory”), a locked-memory limit large enough, and real-time throttling off if the hot threads are real-time.
isolated = set(parse_cpulist(_text(root, "sys/devices/system/cpu/isolated") or ""))
miss = hot_cpus - isolated
add("isolated", not miss, f"hot CPUs not isolated: {_fmt(miss)}" if miss else f"isolated {_fmt(isolated)}")
online = {c for c in idle if _text(root, f"sys/devices/system/cpu/cpu{c}/online") != "0"}
busy = online - isolated
add("siblings", not busy, f"idle siblings online and not isolated: {_fmt(busy)}" if busy else "siblings idle")
nohz = _text(root, "sys/devices/system/cpu/nohz_full")
if nohz is None:
add("nohz_full", False, "kernel built without adaptive ticks (no nohz_full file)")
else:
miss = hot_cpus - set(parse_cpulist(nohz))
add("nohz_full", not miss, f"ticking hot CPUs: {_fmt(miss)}" if miss else f"nohz_full {nohz}")
offloaded = set(parse_cpulist(cmd.get("rcu_nocbs", ""))) | set(parse_cpulist(nohz or ""))
miss = hot_cpus - offloaded
add("rcu_nocbs", not miss, f"RCU callbacks run on {_fmt(miss)}" if miss else "RCU callbacks offloaded")
bad = []
dflt = _text(root, "proc/irq/default_smp_affinity")
if dflt is not None and mask_cpus(dflt) & hot_cpus:
bad.append("default")
for d in sorted((root / "proc/irq").glob("[0-9]*"), key=lambda p: int(p.name)):
cpus = _text(root, f"proc/irq/{d.name}/smp_affinity_list")
if cpus is not None and set(parse_cpulist(cpus)) & hot_cpus:
bad.append(d.name)
shown = ", ".join(bad[:5]) + (f" and {len(bad) - 5} more" if len(bad) > 5 else "")
add("irq_affinity", not bad, f"interrupts allowed on hot CPUs: {shown}" if bad else "no interrupt on hot CPUs")
A setting the host does not expose is reported unknown, not passed: this laptop has no frequency governor to read, as virtual machines often do not. Its report fails seven of the eleven rules (the hot CPUs neither isolated nor tickless, their siblings busy, every interrupt allowed on them, no idle-state cap, a 64 MiB locked-memory limit), passes three and cannot judge one. The tuned fixture, which describes Figure 13.2, passes all eleven; the mistuned fixture fails nine, one per usual mistake.
Purpose. The check that a host is fit to run the hot path before a process starts on it: the deployment of chapter 25 runs it and refuses a host that fails, and its report goes with every latency measurement, so that a change of host tuning is visible when the numbers move.
Interface. Python firm.tuneaudit: audit(root, plan, hot, memlock_bytes, fifo) returns a list of Finding(rule, status, detail) with status ok, fail or unknown, one per rule of RULES in order; failures(findings); report(findings); the command line audits a root and an interface with firm.coreplan’s plan and exits non-zero on any failure.
Rules. Read-only: the audit changes nothing on the host. Every rule is derived from the plan, not from a fixed CPU list; an unreadable setting is unknown, never ok; details name the offending CPUs or interrupt lines.
Acceptance tests. code/firm/tuneaudit/: the tuned fixture passes every rule and the mistuned fixture fails exactly the nine expected ones, with the offending siblings, interrupt lines and governor named; real-time threads with default throttling fail; a larger locking requirement fails only the mistuned limit; an empty tree reports unknown where nothing can be read. The fixtures are written by make_tuneaudit_fixture.py for firm.coreplan’s two-socket fixture.
Stretch. Check that the card’s interrupt queues are on housekeeping CPUs of the card’s node; check the kernel’s own configuration (CONFIG_NO_HZ_FULL) when it is exposed; emit the boot line and interrupt masks that would fix a failing host.
Sources and further reading
- Linux kernel documentation: “The kernel’s command-line parameters” (isolcpus, nohz_full, rcu_nocbs, irqaffinity); “CPU Isolation”; “NO_HZ: Reducing Scheduling-Clock Ticks”; “Real-Time group scheduling”; “Transparent Hugepage Support”.
- Linux man pages
sched(7),socket(7)(SO_BUSY_POLL),mlockall(2). - Red Hat, Optimizing RHEL 9 for Real Time for low latency operation.
- DPDK Programmer’s Guide, “Poll Mode Driver” (release 24.07).
13.8 Exercises
Exercise 13.1 ★
A kernel ticks 1 000 times a second and each tick takes on the CPU it interrupts. What fraction of a spinning thread’s time does it take, what average wait does it add to a message arriving at random (Proposition 13.4), and what does nohz_full remove from the worst case?
Solution
Solution of Exercise 13.1.
of the time. The mean added wait is . On a tickless CPU running one task, the tick no longer interrupts the thread: the worst case loses the tick’s full duration, here.
Exercise 13.2 ★
A process whose rings and books occupy 3 GiB calls mlockall(MCL_CURRENT | MCL_FUTURE) as an ordinary user on this laptop. What happens, and what must change?
Solution
Solution of Exercise 13.2.
The call fails with ENOMEM: an ordinary user’s locked-memory limit is 64 MiB on this laptop, and 3 GiB is far above it. Raise the process’s RLIMIT_MEMLOCK (in the service’s configuration or the system’s limits), or give it the CAP_IPC_LOCK capability, for which the limit is not enforced; the audit’s memlock rule checks the first.
Exercise 13.3 ★
A host has 32 cores with two hardware threads each; the sibling of CPU is . The plan puts four hot threads on CPUs 12–15. Write the isolcpus and nohz_full lists, and say what to do with the siblings.
Solution
Solution of Exercise 13.3.
isolcpus=12-15,44-47 and nohz_full=12-15 (with rcu_nocbs=12-15). The siblings 44–47 share the hot cores’ execution units: take them offline, or leave them isolated with nothing pinned to them, so that no other thread competes with the hot threads for their cores.
Exercise 13.4 ★★
A thread loses 1 000 gaps a second of each. What fraction of its time is stolen, what is the mean added wait, and what fraction of messages wait more than ? Compare with one gap a second of , which steals the same time.
Solution
Solution of Exercise 13.4.
Both steal 1%. The thousand short gaps add a mean wait of , and 0.5% of messages wait more than . The single gap adds seconds, on average, a thousand times more, and almost 1% of messages wait more than , some of them ten milliseconds.
Exercise 13.5 ★★
Busy polling keeps a core at 100% whether or not data arrives. When is that a price worth paying, and what does it cost besides the core?
Solution
Solution of Exercise 13.5.
When the microseconds saved are worth more than a core, its power and its heat: on the few threads of the hot path, not on the rest. Besides the core, it costs power and heat (the kernel documentation warns that polling in the idle loop can overheat the processor and disables turbo frequencies), and on a hyper-threaded core it slows whatever runs on the sibling.
Exercise 13.6 ★★
A team gives its spinning feed thread SCHED_FIFO priority on a CPU where a monitoring agent also runs, with the default throttling. What gap does the feed thread suffer, how often, and what share of its time? What are the two fixes?
Solution
Solution of Exercise 13.6.
With default throttling, real-time threads may use of each second; when the agent is runnable, it gets the other : a gap of up to once a second, 5% of the time. Fixes: move the agent off the CPU (isolate it), after which the throttle has nothing to give the time to; or disable throttling, which risks locking up the CPU if the thread misbehaves. Better, run the feed thread as an ordinary thread on an isolated core.
Exercise 13.7 ★★★
Coding. Add a rule nic_irqs to firm.tuneaudit: the interrupt lines of the plan’s network card (read their names from a fixture /proc/interrupts) must be allowed only on housekeeping CPUs of the card’s node. Extend both fixtures and the tests.
Solution
Solution of Exercise 13.7.
Parse the fixture’s /proc/interrupts for lines whose name contains the interface (for example ens1f0-TxRx-0), and require each such line’s smp_affinity_list to be a subset of the housekeeping CPUs of the plan’s node. The tuned fixture binds them to CPUs 16–20 and passes; add a line to the mistuned fixture bound to CPU 3 (node 0) and assert the failure names it.
Exercise 13.8 ★★★
Find the flaw. “We isolated CPUs 2–5 with isolcpus and started the feed handler with taskset -c 2-5 so that its four threads would share them. Its latency got worse.”
Solution
Solution of Exercise 13.8.
Isolated CPUs are taken out of the scheduler’s load balancing: threads whose affinity covers several isolated CPUs are not spread over them and can all stay on the one where they started, time-sharing it (the shared curve of Figure 13.1). Pin each thread to its own isolated CPU, as firm.affinity does from the plan.
13.9 Problem: The Hiccup
Problem 13.1
Weekend problem — what the host takes from a spinning thread
The hiccup meter ran for five seconds on this laptop in three situations (measured_hiccup_summary.csv). A tuned host is modelled with stated assumptions, not measurements: with the tick running, 250 interruptions a second of ; with the hot CPU isolated and tickless, one interruption of a second.
Part I — The meter.
- Why must the meter’s loop never allocate memory or call the kernel?
- From the summary, how long does one pass of the loop take, and why is the threshold safe?
- What fraction of its time does the thread lose pinned alone, and what is its worst gap?
- What fraction does it lose sharing its CPU with a second spinner, and why about that?
Part II — What a message sees.
- Using Proposition 13.4, what does the summary give for the mean added wait pinned alone?
- What fraction of messages wait more than , and more than ?
- In the shared case, what is the median added wait of a message?
- Why do long gaps matter more than frequent ones?
Part III — The tuned host.
- With the 250 Hz tick of : stolen fraction and mean added wait.
- With the tick: what fraction of messages waits more than ?
- Isolated and tickless: stolen fraction and mean added wait.
- Where does the residual once-a-second tick of a tickless CPU run?
- If the feed thread were
SCHED_FIFOwith default throttling and a runnable ordinary task on its CPU, what would the worst gap and stolen fraction be? - Why can the laptop not be tuned to the model?
Part IV — The verdict.
- State the named result: the fraction of wall time a spinning thread is not running, and its worst gap, on this laptop against the tuned host.
- Which of the audit’s rules does this laptop fail (
measured_audit.csv)? - How would you run the meter on a production host without disturbing trading?
- What does the meter not tell you about a gap, and how would you find it?
- What should be recorded with every latency measurement, following this chapter?
- In one sentence: what does tuning buy, and what does it cost?
Solution
Solution of Problem 13.1.
- Either would bring the kernel onto the CPU (a page fault, a system call), and the meter would measure itself instead of what is imposed on it.
- On the committed run, 5 seconds for about 323 million passes: about a pass, so a gap is more than ten passes lost, not a slow iteration.
- About 1.3%, with a worst gap of about .
- About 51%: two equal threads that both always want the CPU each get half of it.
- About on the committed run; it varies with the run’s longest gap.
- About 0.45% beyond , and about 0.15% beyond .
- Between and : 50.2% of messages wait more than the first and 49.0% more than the second. About half wait nothing, and those that wait, wait milliseconds.
- The wait grows with the square of a gap: a message is more likely to land in a long gap, and then waits longer.
- and .
- .
- of the time and on average.
- On a housekeeping CPU, which runs it on behalf of the tickless ones.
- Up to once a second, 5% of the time: worse than the untuned laptop.
- It is a virtual machine: its kernel has no adaptive ticks, the hypervisor’s own work steals time the guest cannot see or move, and an ordinary user cannot change the boot line.
- Named result. Pinned alone on this laptop, a spinning thread is not running about 1.3% of the time, with a worst gap of about ; on the modelled tuned host, of the time, with a worst gap of : some six thousand times less time, a worst gap some 2 700 times shorter.
isolated,siblings,nohz_full,rcu_nocbs,irq_affinity,cstateandmemlock: seven of eleven.- On an isolated spare CPU of the same host, configured like the hot ones, or on a hot CPU before the open; never beside a hot thread.
- Its cause: the meter sees that the CPU was taken, not by whom. Kernel tracing (interrupt counts per CPU in
/proc/interruptsbefore and after, the scheduler’s trace events) attributes it. - The audit’s report of the host, next to the placement plan, the kernel’s version and its boot line.
- It gives the hot path whole CPUs that the kernel leaves alone, at the price of cores, power and a host that must be configured and checked for it.
13.10 Interview questions
Interview question 13.1 ★ developer
What do isolcpus, nohz_full and rcu_nocbs do, and why are they used together?
Solution
Solution of Interview question 13.1.
isolcpus removes CPUs from load balancing so that only explicitly pinned threads run there; nohz_full stops the scheduling-clock tick on them while they run one task; rcu_nocbs moves the kernel’s deferred callbacks off them. Each removes a different source of interruption, and the tick cannot stop while other tasks or callbacks still need the CPU.
What the interviewer is looking for: three sources of noise and why each needs its own setting.
Interview question 13.2 ★★ developer
Your pinned, spinning thread shows a latency spike every few milliseconds. How do you find the cause?
Solution
Solution of Interview question 13.2.
A regular period suggests the tick or a timer: check the kernel’s tick rate and whether the CPU is tickless; compare interrupt counts per CPU before and after a run; look for other runnable tasks on the CPU and for the kernel’s per-CPU work; trace scheduler events on that CPU. Confirm with a hiccup meter on the same CPU.
What the interviewer is looking for: measure first, then attribute with counters and tracing.
Interview question 13.3 ★★ developer
Blocking receive, busy polling, or kernel bypass: what does each cost, and when would you choose each?
Solution
Solution of Interview question 13.3.
Blocking costs nothing when idle but adds the wake-up, tens of microseconds on the chapter’s machine. Busy polling costs a spinning core and saves the wake-up ( against here). Kernel bypass costs a supported card, a separate network stack and a spinning core, and removes the kernel from the path. Choose by the latency budget, per thread.
What the interviewer is looking for: costs on each side and the budget as the deciding rule.
Interview question 13.4 ★★ developer
Should the hot threads of a trading process run under SCHED_FIFO? What can go wrong?
Solution
Solution of Interview question 13.4.
Usually not, if they run alone on isolated cores: there is nothing to preempt. The risks are real-time throttling (a gap a second by default when another task is runnable), a runaway thread locking up a CPU, and priority inversion with lower-priority threads it shares locks with (chapter 11).
What the interviewer is looking for: throttling and starvation, not just “higher priority is faster”.
Interview question 13.5 ★★ developer
Where should the network card’s interrupts go on a tuned host, and why not simply away from every CPU of its socket?
Solution
Solution of Interview question 13.5.
On housekeeping CPUs of the card’s socket: off the hot cores, so they are not interrupted, but on the card’s node, so that the kernel’s work on the packet leaves it in the socket’s shared cache near the threads that read it, not across the interconnect.
What the interviewer is looking for: isolation from the hot core, locality to the card.
Interview question 13.6 ★★★ developer, researcher
A colleague reports that tuning halved the median tick-to-trade but did nothing to the p99.9. What would you look at?
Solution
Solution of Interview question 13.6.
The rare, long gaps that tuning did not remove: interrupts still allowed on the hot CPUs, a tick that is still running (a second runnable task, a kernel without adaptive ticks), page faults (memory not locked or pre-faulted), a hyper-threading sibling in use, the hypervisor on a virtual machine. Measure the gaps on the hot CPU and audit the host.
What the interviewer is looking for: the tail comes from rare events; the proposition’s squares.