---
title: "Trading in a Public Cloud"
book: "Networks, Hardware and Trading Infrastructure"
subject: quant
language: en
chapter: 16
exercises: 8
source: https://one-course.com/books/quant/14/en/chapter/16-trading-in-a-public-cloud
---

# Chapter 16 — Trading in a Public Cloud

In one cloud provider’s documentation, the [availability zone](#def-nw-trading-in-a-public-cloud-azid) called us-east-1a in one account is not necessarily the zone called us-east-1a in another: “AWS maps the physical Availability Zones randomly to the Availability Zone names for each AWS account”. Only the zone identifier, a code such as use1-az1, names the same physical place in every account. Two firms that agree to meet in “zone a” may be in different buildings; a trading firm that places its machines in the zone its venue calls “a” may be a zone boundary away from it, and pay for it in every round trip.

Part IV of this book is about venues that run in a [public cloud](#def-nw-trading-in-a-public-cloud-cloud), which is where most crypto trading happens (chapter 17), and about the cloud itself as a place to trade from. This chapter describes the cloud’s geography (regions, zones and the identifiers that locate them), the tools that place machines close together, the network and the virtualisation between a program and the wire, the noise that shared hardware adds, and the clocks and prices. Nothing in it is measured by this book: the provider documentation is cited and dated, one published measurement study supplies the round trips, and the rest is a labelled simulation.

## 16.1 Regions, availability zones and zone identifiers

**Definition 16.1 (Public cloud).**

A *public cloud* is computing, storage and networking that a provider runs in its own data centres and rents to any customer on demand, by the hour or the unit, through a programming interface, on hardware shared between customers unless they pay for dedicated machines.

**Definition 16.2 (Availability zone, availability-zone identifier).**

An *availability zone* is one of the isolated groups of data centres that make up a cloud region, with its own power, cooling and networking, connected to the region’s other zones by low-latency links. Its *availability-zone identifier* is the provider’s code for the physical zone, the same in every account, as distinct from the zone’s name, which the provider may assign differently in each account.

A cloud region (One Quant Book 3) is a geographic area; its zones are buildings or campuses within it, kilometres apart, chosen so that one zone’s failure does not take the others down. For a trading firm the consequence is chapter 9’s in a new form: two machines in the same zone talk faster than two in different zones, and the only way to know which zone a venue’s servers are in is by identifier.

![Zone names against zone identifiers for two accounts, from the chapter’s fixtures (the documented response shape of the provider’s describe-availability-zones call, with synthetic values; four of the six zones shown). The firm reads the venue’s zone identifier, not its name.](https://one-course.com/images/onecourse/chapters/quant-14/nw-trading-in-a-public-cloud/fig-a2ed9af79a86.svg)

***Figure 16.1.** Zone names against zone identifiers for two accounts, from the chapter’s fixtures (the documented response shape of the provider’s describe-availability-zones call, with synthetic values; four of the six zones shown). The firm reads the venue’s zone identifier, not its name.*

**Proposition 16.3 (Trusting names).**

If a region has $n$ zones and each account’s names are an independent uniform shuffle of them, two accounts’ zones with the same name are the same physical zone with probability $1/n$. If a round trip inside a zone takes $t_s$ and across zones $t_c$, choosing by name rather than by identifier costs $(1 - 1/n)(t_c - t_s)$ per round trip on average.

**Proof.** Fix the venue’s zone. The firm’s zone with the same name is a uniformly random physical zone, so it is the venue’s with probability $1/n$; the rest of the time the round trip crosses zones and takes $t_c$ instead of $t_s$. ∎

```python
def load_zones(path):
    with open(path) as f:
        doc = json.load(f)
    zones = [z for z in doc["AvailabilityZones"] if z.get("ZoneType") == "availability-zone"]
    return {z["ZoneName"]: z["ZoneId"] for z in zones}


def same_physical(zones_a, zones_b, name):
    return zones_a[name] == zones_b[name]


def locate(zones, zone_id):
    """The name under which this account sees a physical zone."""
    return next(n for n, i in zones.items() if i == zone_id)


def p_same_name_same_zone(n):
    """Names shuffled independently per account over n zones: P(same name, same zone)."""
    return 1.0 / n


def expected_penalty_us(n, same_us, cross_us):
    return (1 - p_same_name_same_zone(n)) * (cross_us - same_us)
```

***Listing 16.1.** Zone names resolved to identifiers from a describe-availability-zones response, and [Proposition 16.3](#prop-nw-trading-in-a-public-cloud-names). code/firm/cloudplan/firm_cloudplan.py*

With the published round trips of this chapter’s fourth section (median $270\,\text{µ}\mathrm{s}$ within a zone, $555\,\text{µ}\mathrm{s}$ across), trusting names costs $190\,\text{µ}\mathrm{s}$ per round trip on average in a region of three zones and $238\,\text{µ}\mathrm{s}$ in one of six ([Figure 16.2](#fig-nw-trading-in-a-public-cloud-penalty)): far more than any tuning of the machine would recover.

![The expected round-trip penalty of choosing a zone by name instead of identifier (), against the number of zones in the region, with the published median round trips within and across zones. Data: fig_cloud.py, nw_cloud.penalty_rows().](https://one-course.com/images/onecourse/chapters/quant-14/nw-trading-in-a-public-cloud/fig-03148dc0fbea.svg)

***Figure 16.2.** The expected round-trip penalty of choosing a zone by name instead of identifier ([Proposition 16.3](#prop-nw-trading-in-a-public-cloud-names)), against the number of zones in the region, with the published median round trips within and across zones. Data: `fig_cloud.py`, `nw_cloud.penalty_rows()`.*

The same regional logic reaches the traditional venues: CME and Google Cloud are building a private cloud region and co-location facility in Aurora (chapter 10), so that zones and identifiers may one day matter for futures as much as they already do for crypto.

## 16.2 Placement groups and instance selection

**Definition 16.4 (Placement group).**

A *placement group* is a customer’s instruction to a cloud provider about where to put a set of its machines relative to each other: close together for low latency (a cluster), apart for resilience (a spread), or on hardware with particular capabilities.

The providers name the tool differently and promise different things. AWS’s cluster [placement groups](#def-nw-trading-in-a-public-cloud-pg) keep instances in one [availability zone](#def-nw-trading-in-a-public-cloud-azid), in the same high-bisection-bandwidth segment of its network, and recommend them for low latency; its precision time [placement groups](#def-nw-trading-in-a-public-cloud-pg) put instances on hardware with direct access to high-precision time sources. Azure’s proximity [placement groups](#def-nw-trading-in-a-public-cloud-pg) make resources “physically located close to each other”. Google Cloud’s compact placement policies place instances close together in a zone “on a best-effort basis”, optionally within a maximum distance. None of them publishes a latency figure, which is why the firm measures (chapter 17).

**As of September 2026 — Placement and time services, from provider documentation.**

- AWS: zone names mapped randomly per account, zone identifiers consistent across accounts; placement strategies cluster, partition, spread and precision time; precision time groups give an enhanced local NTP source, a PTP hardware clock for microsecond-accurate synchronisation (Linux), and hardware packet timestamping with nanosecond resolution, at no extra charge.
- AWS networking: enhanced networking through SR-IOV; ENA Express raising a single flow’s bandwidth from 5 to 25 Gb/s and reducing tail latency within a zone; [bare-metal instances](#def-nw-trading-in-a-public-cloud-bare) without virtualisation overhead.
- Azure: proximity [placement groups](#def-nw-trading-in-a-public-cloud-pg) . Google Cloud: compact placement policies, best effort, with an optional maximum distance.

Instance selection follows the same logic as chapter 8’s servers: the clock and the core count of the machine, its network bandwidth, and whether the network adapter and its offloads reach the program without a hypervisor in the way. In the cloud the choice is a line in a catalogue with an hourly price, and the firm can try several and keep the best (chapter 17’s instance lottery).

## 16.3 Network adapters, virtualisation and bare metal

**Definition 16.5 (Single-root I/O virtualisation).**

*Single-root I/O virtualisation* (SR-IOV) lets one physical network adapter present several virtual functions, each assigned directly to a virtual machine, so that the machine’s driver talks to the hardware without the hypervisor copying its packets.

**Definition 16.6 (Bare-metal instance).**

A *bare-metal instance* is a whole physical server rented through the cloud’s interface, with no hypervisor between the customer’s operating system and the hardware, and no other customer on the machine.

Between a trading program on a cloud instance and the wire there are, in order: the program, the kernel or a kernel-bypass library (One Quant Book 13), the virtual function of an SR-IOV adapter, the provider’s own network card and its software, the host’s physical network, and the region’s fabric. The provider documents the middle of this stack only in outline: enhanced networking through SR-IOV, ENA Express on a proprietary transport for higher single-flow bandwidth and lower tail latency within a zone, bare metal for customers who need the hardware to themselves. The firm can control the ends: its program and its kernel settings (chapter 3), and which instance, placement and zone it rents.

## 16.4 Jitter and noisy neighbours

**Definition 16.7 (CPU steal time, noisy neighbour).**

*CPU steal time* is the time a virtual CPU was ready to run but the hypervisor ran something else on the physical CPU; Linux reports it as “involuntary wait”. A *noisy neighbour* is another tenant of the same physical server or network whose load takes resources (CPU, cache, memory bandwidth, network queues) from a customer’s machine and so adds delay to it.

The one published measurement this chapter relies on, by Hilyard and co-authors, recorded round trips between small virtual machines in one US East region of three providers for six hours. On AWS, machines in the same subnet of one zone had a median round trip of $270\,\text{µ}\mathrm{s}$, a 99th percentile of $395\,\text{µ}\mathrm{s}$ and a 99.99th of $760\,\text{µ}\mathrm{s}$; machines in different zones 555, 755 and $1905\,\text{µ}\mathrm{s}$; and the worst round trips reached 2 900 times the average. The model fits a lognormal to each median and 99th percentile and adds rare tail events calibrated to the 99.99th ([Table 16.1](#tab-nw-trading-in-a-public-cloud-rtt)); a third case, a shared host with a [noisy neighbour](#def-nw-trading-in-a-public-cloud-steal), keeps the same-zone body and adds spikes of 5 to 50 times one round trip in 500, an assumption.

| Case | p50 | p90 | p99 | p99.9 | p99.99 |
| --- | --- | --- | --- | --- | --- |
| Same zone (published) | 270 | – | 395 | – | 760 |
| Same zone (model) | 270 | 333 | 396 | 463 | 743 |
| Cross zone (published) | 555 | – | 755 | – | 1 905 |
| Cross zone (model) | 555 | 657 | 755 | 851 | 1 908 |
| Shared host, [noisy neighbour](#def-nw-trading-in-a-public-cloud-steal) (model) | 270 | 333 | 400 | 7 474 | 13 940 |

***Table 16.1.** Round-trip percentiles in microseconds: the published measurement (Hilyard et al., AWS, same subnet and cross AZ) and the simulation that `firm.cloudplan` fits to it; the noisy-neighbour case is an assumption. Data: `nw_cloud.percentiles()`.*

![Simulation fitted to a published measurement: round-trip percentiles within a zone and across zones, and on a shared host with an assumed noisy neighbour. Crossing a zone doubles the body of the distribution; a noisy neighbour leaves the body alone and multiplies the tail. Data: fig_cloud.py.](https://one-course.com/images/onecourse/chapters/quant-14/nw-trading-in-a-public-cloud/fig-fb894e4eab37.svg)

***Figure 16.3.** Simulation fitted to a published measurement: round-trip percentiles within a zone and across zones, and on a shared host with an assumed [noisy neighbour](#def-nw-trading-in-a-public-cloud-steal). Crossing a zone doubles the body of the distribution; a [noisy neighbour](#def-nw-trading-in-a-public-cloud-steal) leaves the body alone and multiplies the tail. Data: `fig_cloud.py`.*

```python
def sample_rtt(rtt, n, seed=0):
    rng = np.random.default_rng(seed)
    sigma = math.log(rtt.p99_us / rtt.median_us) / Z99
    x = rtt.median_us * np.exp(rng.normal(0.0, sigma, n))
    hit = rng.random(n) < rtt.spike_p
    x[hit] *= rng.uniform(rtt.spike_lo, rtt.spike_hi, hit.sum())
    return x
```

***Listing 16.2.** The round-trip model: a lognormal fitted to a median and a 99th percentile, with rare spikes. code/firm/cloudplan/firm_cloudplan.py*

The lesson is the one of One Quant Book 13 in a harsher setting: the median is not the problem. A strategy that races in the cloud races in the tail, and the tail belongs to other tenants unless the firm rents the whole machine; the 99.9th percentile in the noisy-neighbour case is 16 times the same-zone figure. Bare metal removes the neighbours on the host, not those on the network.

## 16.5 Clocks and costs in the cloud

Chapter 4’s clocks have a cloud version. AWS documents a local time service available to every instance and, in precision time [placement groups](#def-nw-trading-in-a-public-cloud-pg), a PTP hardware clock for “microsecond-accurate” synchronisation and hardware timestamps of packets with nanosecond resolution; that is enough to timestamp the chapter’s own round trips, which is what a firm must do to see its tail.

**Definition 16.8 (Data-transfer charge).**

A *data-transfer charge* is a cloud provider’s fee for data moved across a boundary of its network (between zones, between regions, or out to the internet), per gigabyte and per direction, on top of the price of the machines.

**As of September 2026 — Published cloud prices.**

AWS on-demand Linux prices in US East (N. Virginia), from the price list published on 25 September 2026: c7i.xlarge (4 vCPUs) 0.1785 USD an hour; c7i.4xlarge (16 vCPUs, up to 12.5 Gb/s) 0.714; c6in.4xlarge (16 vCPUs, up to 50 Gb/s) 0.9072; c7i.24xlarge and the bare-metal c7i.metal-24xl (96 vCPUs, 37.5 Gb/s) 4.284 each; c7i.metal-48xl 8.568. Data within one zone is free; data crossing zones is charged 0.01 USD per gigabyte in each direction.

![Three monthly plans priced from the dated catalogue of : four 16-vCPU instances in one zone, the same spread over two zones exchanging 20 terabytes a month, and two bare-metal servers in one zone. Data: fig_cloud.py, nw_cloud.plan_costs().](https://one-course.com/images/onecourse/chapters/quant-14/nw-trading-in-a-public-cloud/fig-a72d2bcd3480.svg)

***Figure 16.4.** Three monthly plans priced from the dated catalogue of [Box 16.2](#dat-nw-trading-in-a-public-cloud-prices): four 16-vCPU instances in one zone, the same spread over two zones exchanging 20 terabytes a month, and two bare-metal servers in one zone. Data: `fig_cloud.py`, `nw_cloud.plan_costs()`.*

The prices are small next to chapter 9’s [colocation](https://one-course.com/books/quant/14/en/chapter/9-colocation-products-and-how-they-are-sold#def-nw-colocation-products-and-how-they-are-sold-colo): four 16-core instances cost 2 085 USD a month, two whole 96-core servers 6 255 USD, where one Nasdaq cabinet costs 7 230 USD before its power and connections. Spreading the four instances over two zones and exchanging 20 terabytes a month adds 400 USD in [data-transfer charges](#def-nw-trading-in-a-public-cloud-dtc), and $285\,\text{µ}\mathrm{s}$ to every median round trip between them. In the cloud the expensive thing is not the machine but the distance.

**Method 16.9 (Placing a trading system in a cloud region).**

1. Find the venue’s region and, if it publishes it, the zone identifier of its servers (chapters 17 and 18); otherwise find it by measuring.
2. Resolve your own zone names to identifiers in every account you use, and place by identifier.
3. Put the latency-critical machines in one zone, in a cluster (and, if you need the clock, a precision time) [placement group](#def-nw-trading-in-a-public-cloud-pg) ; consider bare metal for the ones that race.
4. Timestamp your own round trips with the provider’s time service and measure the tail, not the median; re-measure after every maintenance or relaunch.
5. Price machines and data transfer together; keep cross-zone traffic for what needs resilience.

## 16.6 Tutorial: zones, tails and a bill

**Goal.** Resolve zone names, simulate round trips from a published measurement, and price three plans; no cloud account is needed. **End state:** Figures [16.2](#fig-nw-trading-in-a-public-cloud-penalty), [16.3](#fig-nw-trading-in-a-public-cloud-rtt) and [16.4](#fig-nw-trading-in-a-public-cloud-plans) and [Table 16.1](#tab-nw-trading-in-a-public-cloud-rtt).

1. **Zones.** `firm_cloudplan.load_zones` reads the two fixtures in `data/` (the documented response shape, synthetic values); `same_physical` and `locate` compare them ( [Listing 16.1](#lst-nw-trading-in-a-public-cloud-zones) ).
2. **Penalty.** `nw_cloud.penalty_rows()` evaluates [Proposition 16.3](#prop-nw-trading-in-a-public-cloud-names) for 2 to 6 zones.
3. **Round trips.** `sample_rtt` ( [Listing 16.2](#lst-nw-trading-in-a-public-cloud-rtt) ) and the three `CASES` .
4. **Bill.** `load_catalogue` and `monthly_cost` price `PLANS` .

**What to change next.** Run the chapter’s code on real instances in two zones with hardware timestamps, replace the fitted case with your own percentiles, and compare.

## 16.7 Build: the cloud placement plan

**Purpose.** Where the firm’s machines are in a cloud region, what they cost, and what their round trips look like: the base of chapters 17 to 19 and of the cloud rows of chapter 29’s plan.

**Interface.** `firm_cloudplan`: `load_zones`, `same_physical`, `locate`, `p_same_name_same_zone`, `expected_penalty_us`, `RTT_PUBLISHED`, `Rtt`, `sample_rtt`, `Instance`, `load_catalogue`, `monthly_cost`.

**Rules.** Place by zone identifier, never by name; prices come from the dated catalogue with their source; simulated round trips say which measurement they are fitted to.

**Acceptance tests.** `code/firm/cloudplan/tests/`: the fixtures’ names and identifiers, the $1/n$ probability against a shuffle simulation, the fit’s median and 99th percentile, and a bill by hand.

**Stretch.** Read live responses and price lists through the provider’s interfaces; add regions and inter-region round trips.

Sources and further reading

- AWS documentation: Availability Zone IDs; placement strategies and precision time placement groups; enhanced networking, ENA Express and Nitro bare-metal instances; Amazon VPC pricing and the public price list (September 2026).
- Microsoft Learn, proximity placement groups; Google Cloud, placement policies.
- O. Hilyard, B. Cui, M. Webster, A. Bangalore Muralikrishna and A. Charapko, “Cloudy Forecast: How Predictable is Communication Latency in the Cloud?”, arXiv:2309.13169 (2023).

## 16.8 Exercises

**Exercise 16.1 ★.**

In a region of four zones, what is the probability that your “zone b” is the venue’s “zone b”?

**Solution of Exercise 16.1.**

$1/4$, if names are shuffled independently and uniformly in each account.

**Exercise 16.2 ★.**

What does it cost a month to move 5 terabytes each way between two zones?

**Solution of Exercise 16.2.**

$5\,000~\mathrm{GB} \times 0.01~\mathrm{USD} \times 2$ directions $= 100$ USD a month.

**Exercise 16.3 ★.**

What does [CPU steal time](#def-nw-trading-in-a-public-cloud-steal) measure, and why does a trading program care?

**Solution of Exercise 16.3.**

The time a virtual CPU wanted to run but the hypervisor gave the physical CPU to something else. For a trading program it is time during which a market event waits unprocessed: a delay invisible to the program’s own clocks except as a gap, and a sign that the host is shared and busy.

**Exercise 16.4 ★★.**

Why does crossing a zone double the median round trip in the published measurement, while a [noisy neighbour](#def-nw-trading-in-a-public-cloud-steal) leaves it unchanged?

**Solution of Exercise 16.4.**

Crossing zones adds distance and switching hops to every packet: a fixed cost that moves the whole distribution. A [noisy neighbour](#def-nw-trading-in-a-public-cloud-steal) takes resources only now and then, so most round trips are untouched and a few are much longer: it moves the tail, not the body.

**Exercise 16.5 ★★.**

What does a [bare-metal instance](#def-nw-trading-in-a-public-cloud-bare) remove, and what does it not?

**Solution of Exercise 16.5.**

It removes the hypervisor and the other tenants on the host (steal time, shared caches, shared network card queues). It does not remove the shared network beyond the host, the provider’s own card and software, or the distance to the venue’s zone.

**Exercise 16.6 ★★.**

How would you find the zone identifier of a venue’s servers if the venue does not publish it?

**Solution of Exercise 16.6.**

Launch machines in every zone of the region, measure round trips to the venue’s endpoints from each with hardware timestamps, and keep the zone with the lowest and most stable figure (chapter 17); read the zone identifiers of the machines, which are physical, and record the result with its date.

**Exercise 16.7 ★★★.**

*Coding.* With `firm.cloudplan`, find the spike probability at which the noisy-neighbour case’s 99.9th percentile first exceeds the cross-zone case’s.

**Solution of Exercise 16.7.**

Above a spike probability of about 0.092% (one round trip in about 1 090) the noisy-neighbour case’s 99.9th percentile jumps into the spikes and exceeds the cross-zone case’s $851\,\text{µ}\mathrm{s}$: the tail percentile is decided by how often the neighbour strikes, not by how hard.

**Exercise 16.8 ★★★.**

*Find the flaw.* “We launched in us-east-1a, like the venue says in its documentation, so we are in its zone.”

**Solution of Exercise 16.8.**

The venue’s documentation names the zone in its own account; the same name in the firm’s account is a different physical zone with probability $1 - 1/n$. Only the zone identifier locates the venue; the firm must compare identifiers, or measure.

## 16.9 Problem: us-east-1a Is Not Your us-east-1a

**Problem 16.1.**

Weekend problem — the price of trusting a name

A venue runs its matching engine in one zone of a six-zone region and tells its users the zone’s name in its own account. A firm launches its machines in the zone of the same name in its account. Use the published round trips: $270\,\text{µ}\mathrm{s}$ within a zone, $555\,\text{µ}\mathrm{s}$ across.

**Part I — The names.**

1. With independent uniform shuffles, what is the probability that the firm is in the venue’s zone?
2. What is the expected round-trip penalty, and what if the region had three zones?
3. In the chapter’s fixtures, which of account A’s names match account B’s physically?
4. Under which name does account B see account A’s us-east-1a?

**Part II — The tail.**

5. What are the modelled 99th and 99.99th percentiles within and across zones?
6. How close is the model to the published 99.99th percentiles?
7. What does a [noisy neighbour](#def-nw-trading-in-a-public-cloud-steal) do to the 99.9th percentile?
8. Which matters more for a racing strategy, the zone or the neighbour?

**Part III — The bill.**

9. What do four c7i.4xlarge instances cost a month?
10. What do 20 terabytes a month each way across zones add?
11. What do two bare-metal c7i.metal-24xl cost, and against what [colocation](https://one-course.com/books/quant/14/en/chapter/9-colocation-products-and-how-they-are-sold#def-nw-colocation-products-and-how-they-are-sold-colo) figure should that be compared?
12. What is cheaper, a zone mistake or a bare-metal upgrade, and why is that the wrong question?

**Part IV — The verdict.**

13. State the *named result* : the probability that two accounts’ same-named zone is the same physical zone when names are shuffled over $n$ zones, and the expected round-trip penalty of trusting names instead of identifiers.
14. How would the firm check its placement without the venue’s help?
15. Why do providers shuffle names at all?
16. What should a venue publish, and what should a firm ask for?
17. Which parts of this problem are documentation, which measurement, which assumption?
18. How often should the placement be re-checked?
19. Does the same problem exist in a [colocation](https://one-course.com/books/quant/14/en/chapter/9-colocation-products-and-how-they-are-sold#def-nw-colocation-products-and-how-they-are-sold-colo) hall?
20. In one sentence: what does a zone name tell you?

**Solution of Problem 16.1.**

**Part I.**

1. $1/6$ .
2. $(1 - 1/6)(555 - 270) = 237.5\,\text{µ}\mathrm{s}$ on average; with three zones, $190\,\text{µ}\mathrm{s}$ .
3. us-east-1e and us-east-1f: the only names that point to the same identifiers (use1-az3 and use1-az5) in both accounts.
4. As us-east-1b: A’s us-east-1a is use1-az4, which is B’s us-east-1b.

**Part II.**

1. Within a zone $396\,\text{µ}\mathrm{s}$ and $743\,\text{µ}\mathrm{s}$ ; across zones $755\,\text{µ}\mathrm{s}$ and $1908\,\text{µ}\mathrm{s}$ .
2. Within 2% (743 against 760) and 0.2% (1 908 against 1 905).
3. It takes it from $463\,\text{µ}\mathrm{s}$ to $7474\,\text{µ}\mathrm{s}$ , sixteen times more.
4. For the median race, the zone; for the tail (the races lost to a stall), the neighbour. A racing strategy needs both right.

**Part III.**

1. $4 \times 0.714 \times 730 = 2\,084.88$ USD.
2. $20\,000 \times 0.01 \times 2 = 400$ USD.
3. $2 \times 4.284 \times 730 = 6\,254.64$ USD a month: less than one Nasdaq Ultra High Density Cabinet’s monthly fee (7 230 USD, chapter 9) before its power and connections.
4. The zone mistake costs nothing on the bill and $285\,\text{µ}\mathrm{s}$ on the median round trip; bare metal costs thousands a month and removes only the neighbours. The question is wrong because the zone is free to fix and the neighbours are not.

**Part IV.**

1. *Named result* : two accounts’ same-named zones are the same physical zone with probability $1/n$ ( $1/6$ for six zones); trusting names costs $(1 - 1/n)(t_c - t_s)$ per round trip on average, $237.5\,\text{µ}\mathrm{s}$ with six zones and the published medians.
2. By reading its own zone identifiers and measuring round trips to the venue from every zone, keeping the fastest.
3. So that customers who all choose “a” do not all land in the same physical zone, which would concentrate load and failures.
4. The venue should publish the zone identifier of its servers (or offer endpoints per identifier); the firm should ask for it and verify it by measurement.
5. Documentation: the random mapping, the identifiers, the placement tools, the prices. Measurement: the published round trips. Assumptions: the uniform shuffle, the [noisy neighbour](#def-nw-trading-in-a-public-cloud-steal) , the traffic volume.
6. After every launch and relaunch, after the venue’s maintenance, and continuously for the round trips (chapter 17).
7. In a different form: the hall’s position of a cabinet (chapter 9), which equalisation hides; a [colocation](https://one-course.com/books/quant/14/en/chapter/9-colocation-products-and-how-they-are-sold#def-nw-colocation-products-and-how-they-are-sold-colo) customer knows its building, while a cloud customer must discover its zone.
8. Nothing about where you are, unless you translate it into an identifier.

## 16.10 Interview questions

**Interview question 16.1 ★ developer.**

What is an [availability zone](#def-nw-trading-in-a-public-cloud-azid), and why does it matter for latency?

**Solution of Interview question 16.1.**

One of the isolated groups of data centres that make up a region, kilometres from the others. Machines in one zone talk faster than machines in two: in the published measurement, $270\,\text{µ}\mathrm{s}$ against $555\,\text{µ}\mathrm{s}$ of median round trip.

*What the interviewer is looking for: Zones as buildings; names against identifiers; the order of magnitude.*

**Interview question 16.2 ★★ developer.**

Your cloud trading system’s median latency is fine but its 99.9th percentile is ten times worse. Where do you look?

**Solution of Interview question 16.2.**

At the tail’s causes: steal time and [noisy neighbours](#def-nw-trading-in-a-public-cloud-steal) on the host (check steal, move to dedicated or bare metal), the network (timestamp packets to separate network from host), cross-zone paths, the program’s own pauses (garbage collection, page faults, interrupts; One Quant Book 13).

*What the interviewer is looking for: Measure before guessing; separate host, network and program; hardware timestamps.*

**Interview question 16.3 ★★ developer.**

How do you place machines as close as possible to a venue that runs in the same cloud region?

**Solution of Interview question 16.3.**

Find the venue’s zone identifier (published or measured), launch in that zone by identifier, use a cluster [placement group](#def-nw-trading-in-a-public-cloud-pg), choose instances with enhanced networking or bare metal, and keep the few that measure best.

*What the interviewer is looking for: Identifier, [placement group](#def-nw-trading-in-a-public-cloud-pg), instance type, measurement; chapter 17’s lottery.*

**Interview question 16.4 ★★ developer, researcher.**

Can you timestamp packets accurately in a [public cloud](#def-nw-trading-in-a-public-cloud-cloud)? How?

**Solution of Interview question 16.4.**

With a provider that exposes a PTP hardware clock and hardware packet timestamps (AWS documents both in its precision time [placement groups](#def-nw-trading-in-a-public-cloud-pg)): synchronise the clock with the hardware clock, read packet timestamps from the network card, and check the error budget as in chapter 4.

*What the interviewer is looking for: Hardware clock and timestamps; documented accuracy; verification.*

**Interview question 16.5 ★★ developer.**

What does it cost to run a trading system in the cloud compared with [colocation](https://one-course.com/books/quant/14/en/chapter/9-colocation-products-and-how-they-are-sold#def-nw-colocation-products-and-how-they-are-sold-colo), and where do the hidden costs come from?

**Solution of Interview question 16.5.**

Machines are cheap by the hour; the hidden costs are data transfer across zones and regions and out to the internet, bare metal or dedicated hosts for predictability, the engineering to measure and re-place, and the latency tail of shared infrastructure.

*What the interviewer is looking for: [Data-transfer charges](#def-nw-trading-in-a-public-cloud-dtc); predictability as a cost; comparison with [colocation](https://one-course.com/books/quant/14/en/chapter/9-colocation-products-and-how-they-are-sold#def-nw-colocation-products-and-how-they-are-sold-colo) fees.*

**Interview question 16.6 ★★★ developer, trader.**

Would you run a latency-sensitive equities strategy from a [public cloud](#def-nw-trading-in-a-public-cloud-cloud)? What would have to be true?

**Solution of Interview question 16.6.**

Only if the venue itself ran in that region (so that the cloud is the [colocation](https://one-course.com/books/quant/14/en/chapter/9-colocation-products-and-how-they-are-sold#def-nw-colocation-products-and-how-they-are-sold-colo)), the firm could place by zone identifier next to it, and the tail could be controlled (bare metal, measurement). Against venues in their own data centres, a [public cloud](#def-nw-trading-in-a-public-cloud-cloud) is a remote site: fine for research and slow strategies, not for races.

*What the interviewer is looking for: Where the venue is decides; the tail; honesty about the use case.*
