Quantitative Finance · Book 14 · Technology

Networks, Hardware and Trading Infrastructure

Networks, Hardware and Trading Infrastructure · Technology

19Cross-Region Crypto Networks

Crypto liquidity is spread over a few cloud regions, and a firm that trades across them needs its messages to cross oceans. In April 2026 a round trip from Tokyo to Frankfurt over the cloud provider’s own backbone, by VPC peering, took a median 224.3 ms224.3\,\mathrm{m}\mathrm{s}; over the same provider’s Direct Connect joined to a specialised network, 136.5 ms136.5\,\mathrm{m}\mathrm{s}. Light in straight glass along the geodesic would need 91.3 ms91.3\,\mathrm{m}\mathrm{s}. The provider’s backbone is built for “resilience and scale”, its blog says, and it takes the long way; the specialist takes a shorter one; and the firm that cannot choose sends the same message down both and keeps the first to arrive.

Chapters 16 to 18 worked inside one region. This chapter connects the regions: the hubs and their distances, the backbone and the internet against private lines and specialised providers, and route racing, with the duplicate orders it creates and what a venue’s rules do with them. The published figures are cited and dated; the racing is a labelled simulation.

19.1 The hubs: Tokyo, Singapore, Hong Kong, London, Virginia

Definition 19.1 (Path inflation)

The path inflation of a route between two places is its measured round trip divided by the round trip light would need along the geodesic in the route’s medium (for long-haul networks, fibre): chapter 10’s route factor, for round trips over networks whose medium and path are not published.

The hubs are chapter 17’s: Tokyo, where AWS’s ap-northeast-1 region hosts Binance and BitMEX and Hyperliquid’s validators; Singapore, Bybit’s region; Hong Kong; London and Frankfurt in Europe; Northern Virginia in the United States. The provider publishes only the regions’ names and their numbers of zones (four in Tokyo, six in Northern Virginia, three in the others here), so the chapter places each at its city, and Northern Virginia at Ashburn, an assumption flagged in the site table.

From Tokyo toGeodesicVacuumFibreBackboneSpecialistPath inflation
(km)RTT (ms)RTT (ms)(ms)(ms)backbonespecialist
Hong Kong2 88319.228.1––––
Singapore5 31135.451.868.065.41.311.26
N. Virginia10 89572.7106.3150.2135.41.411.27
Frankfurt9 35962.491.3224.3136.52.461.50
London9 58664.093.5210.0139.62.251.49
Stockholm8 19454.779.9245.3124.33.071.56
Table 19.1. Tokyo’s hubs: geodesic distances between the region cities and the round trips light needs along them, against the median round trips AWS published for April 2026 over its backbone (VPC peering) and over Direct Connect with a specialised network, and the path inflation of each against the fibre floor. Data: nw_race.hub_rows().

The table’s pattern is the chapter’s subject. To Singapore and Virginia both networks run 26 to 41% above straight glass. To Europe the backbone runs 2.2 to 3.1 times the floor, the specialist 1.5 times. Neither publishes its path; the inflation says the backbone goes a long way round and the specialist a shorter one. The best European route from Tokyo, 124.3 ms124.3\,\mathrm{m}\mathrm{s} to Stockholm, is still more than twice the vacuum floor.

Round trips from Tokyo: the fibre floor along the geodesic, and the published medians over a specialised network and over the provider’s backbone (April 2026). Data: fig_race.py, .
Figure 19.1. Round trips from Tokyo: the fibre floor along the geodesic, and the published medians over a specialised network and over the provider’s backbone (April 2026). Data: fig_race.py, Table 19.1.

19.2 Cloud backbone, public internet and private lines

Definition 19.2 (Border Gateway Protocol, cloud backbone)

The Border Gateway Protocol (BGP) is the internet’s routing protocol between autonomous systems: each exchanges reachability information, including the list of autonomous systems a route traverses, and chooses among routes by policy, not by latency. A cloud backbone is a provider’s private network between its regions, which carries traffic between them (peering, replication, its own services) under the provider’s routing.

Three kinds of path join two regions. The public internet, where each hop’s operator chooses the next by business policy and BGP carries reachability, not delay: short paths are a side effect. The provider’s backbone, engineered for capacity and resilience, and in AWS’s own blog not the lowest latency. And private lines: dark fibre or wavelengths (chapter 13) and radio (chapter 14), bought directly or from a network that specialises in low latency and connects to the provider’s regions through its interconnection product, as the blog’s specialist does.

As of September 2026 — Cross-region routes and duplicate rules, from public sources

  • AWS (June 2026): median TCP round trips from Tokyo on 28 April 2026, over VPC peering and over Direct Connect with Avelacom: Singapore 68.0 and 65.4 ms; Virginia 150.2 and 135.4; Frankfurt 224.3 and 136.5; London 210.0 and 139.6; Stockholm 245.3 and 124.3.
  • Binance spot API: a client order identifier is unique among open orders; an order reusing one can be accepted only once the earlier order is filled, otherwise it is rejected (“Duplicate order sent.”).

19.3 Specialised low-latency providers and virtual-server hosts

Between Tokyo and Europe the specialist of the blog saves 70 to 121 milliseconds of round trip over the backbone. For a strategy that trades a Tokyo venue against a European one, that is the difference between seeing the other market’s prices a tenth of a second late and seeing them in time to matter. The specialist is a financial extranet (chapter 15) for cloud regions: private lines between the hubs, delivered into the provider’s regions through its interconnection product.

The instance lottery of chapter 17 applies at both ends: the fastest cross-region route is wasted on a machine that sits far from the venue in its own region.

19.4 Racing the same message down several routes

Definition 19.3 (Route racing, idempotency key)

Route racing sends the same message on several routes at once and uses the first copy to arrive. An idempotency key is an identifier the sender attaches to every copy of one message (for orders, the client order identifier) so that the receiver can recognise the copies and act on only one.

Proposition 19.4 (What racing buys)

The first of kk copies arrives at the minimum of the routes’ latencies. If the routes’ delays were perfectly correlated, the minimum would be the fastest route’s delay and racing would buy nothing; the less correlated they are, the more the minimum’s tail falls below every single route’s tail, since the minimum exceeds a level only when every route does.

Proof. P(min⁡iXi>t)=P(X1>t,…,Xk>t)P(\min_i X_i > t) = P(X_1 > t, \dots, X_k > t), which equals the smallest single survival probability when the routes move together and falls to their product when they are independent. ∎

The model races three one-way routes from Tokyo to Singapore: the specialist and the backbone at half their published round trips (32.7 and 34.0 ms34.0\,\mathrm{m}\mathrm{s}), each with an assumed 99th percentile 15% above its median, and an internet route at 35 ms35\,\mathrm{m}\mathrm{s} with a heavier tail (an assumption). A Gaussian copula ties them with correlation ρ\rho, a common cause such as a shared cable or congestion.

Simulation: the 99th-percentile one-way latency saved by racing, Tokyo to Singapore, against the best single route, as the routes’ correlation rises. At = 0 two routes save 1.6 and three 2.1\, m s; at = 0.9 both save about 50\, µ s. Data: fig_race.py, nw_race.race_rows().
Figure 19.2. Simulation: the 99th-percentile one-way latency saved by racing, Tokyo to Singapore, against the best single route, as the routes’ correlation rises. At ρ=0\rho = 0 two routes save 1.6 and three 2.1 ms2.1\,\mathrm{m}\mathrm{s}; at ρ=0.9\rho = 0.9 both save about 50 µs50\,\text{µ}\mathrm{s}. Data: fig_race.py, nw_race.race_rows().
def first_arrival(routes, rho, n=200_000, seed=0):
    """Correlated lognormal draws (one column per route) and the row-wise first arrival."""
    rng = np.random.default_rng(seed)
    k = len(routes)
    common = rng.normal(size=(n, 1))
    own = rng.normal(size=(k, n)).T          # route i keeps its draws whatever k
    z = math.sqrt(rho) * common + math.sqrt(1 - rho) * own
    draws = np.column_stack([r.median_us * np.exp(z[:, i] * math.log(r.p99_us / r.median_us) / Z)
                             for i, r in enumerate(routes)])
    return draws.min(axis=1), draws


def percentile_gain(routes, rho, q=99, n=200_000, seed=0):
    first, draws = first_arrival(routes, rho, n, seed)
    best_single = min(float(np.percentile(draws[:, i], q)) for i in range(draws.shape[1]))
    return best_single - float(np.percentile(first, q))
Listing 19.1. Correlated routes through a Gaussian copula, the first arrival, and the percentile saved against the best single route. code/firm/routerace/firm_routerace.py

Racing pays in the tail, and only while the routes fail independently. With independent routes the second route saves 1623 µs1623\,\text{µ}\mathrm{s} at the 99th percentile and the third another 518 µs518\,\text{µ}\mathrm{s}; at ρ=0.5\rho = 0.5, 771 and 90; at 0.9, 53 and nothing. At an assumed 10 USD a month per microsecond of 99th percentile, the third route is worth 5 182 USD a month if the routes are independent, 902 at ρ=0.5\rho = 0.5 and nothing at 0.9: the price above which racing it does not pay.

What arrives second is a problem. The venue receives the same order several times; the idempotency key must make it act once. Binance’s spot interface is typical of the rule to check: a client order identifier must be unique among open orders, and a copy is rejected as a duplicate only while the first order is open. Once the first copy has filled, a late copy with the same identifier is a new order.

Simulation: with three raced copies (= 0.5) and a venue that rejects a duplicate identifier only while the first order is open, the expected number of late copies accepted as new orders, against the time the first copy takes to fill. An order that fills on arrival is executed three times. Data: fig_race.py, nw_race.duplicate_rows().
Figure 19.3. Simulation: with three raced copies (ρ=0.5\rho = 0.5) and a venue that rejects a duplicate identifier only while the first order is open, the expected number of late copies accepted as new orders, against the time the first copy takes to fill. An order that fills on arrival is executed three times. Data: fig_race.py, nw_race.duplicate_rows().

Proposition 19.5 (Racing orders that fill)

Under a rule that rejects a repeated identifier only while the first order is open, a raced order that fills within time ff of the first copy’s arrival is executed again by every later copy that arrives after ff. For orders that fill on arrival (f=0f = 0), kk raced copies are kk orders.

Proof. A late copy that finds the first order filled finds no open order with its identifier and is accepted. With f=0f = 0 every later copy is late. ∎

The model makes the danger concrete: three copies of an order that fills on arrival are three executions; one that takes a millisecond to fill still produces 1.57 unwanted executions on average; eight milliseconds, 0.07. Racing is therefore safe for messages that are idempotent at the venue (a cancel, a query, market data) or under a rule that remembers identifiers after the order has ended, and dangerous for taking orders under an open-only rule. One Quant Book 13’s idempotent processing is what the venue must do; this chapter’s lesson is to read the venue’s rule before racing.

Route racing, schematically: three copies of one order with one idempotency key; the venue keeps the first and must recognise the others. Under an open-only rule it rejects them while the order is open and accepts them as new orders once it has filled.
Figure 19.4. Route racing, schematically: three copies of one order with one idempotency key; the venue keeps the first and must recognise the others. Under an open-only rule it rejects them while the order is open and accepts them as new orders once it has filled.

Method 19.6 (Racing routes between regions)

  1. Measure each candidate route’s latency distribution, not only its median, and the correlation of their delays (do they share a cable, a landing, a peering point?).
  2. Race only routes whose tails are not too correlated, and price each extra route against the tail it saves.
  3. Read the venue’s duplicate rule. Race cancels, amendments and queries freely; race taking orders only if the venue remembers identifiers after an order ends, or if the firm’s own logic tolerates a duplicate.
  4. Log every copy’s arrival (from the venue’s acknowledgements) and re-measure the routes’ correlation; it changes when a cable is cut or repaired.

19.5 Tutorial: hubs and races

Goal. Measure path inflation between the hubs and simulate route racing with its duplicates. End state: Figures 19.1, 19.2 and 19.3 and Table 19.1.

  1. Hubs. The region cities are rows of data/networks/sites.csv; nw_race.hub_rows() computes geodesics, floors and path inflation against the published round trips.
  2. Routes. ROUTES holds three routes from published medians and assumed tails; first_arrival (Listing 19.1) draws them correlated.
  3. Gain and price. race_rows() and break_even().
  4. Duplicates. duplicates applies the open-only rule; duplicate_rows() sweeps the fill time.

What to change next. Give the backbone route to Frankfurt its published 224 ms and race it against the specialist; set the rule to “ever” and see the duplicates disappear.

19.6 Build: the route racer

Purpose. The latency, cost and risk of sending messages on several routes between regions: the input of the cross-region rows of chapter 29’s plan.

Interface. firm_routerace: Route, first_arrival, percentile_gain, break_even_price, duplicates, path_inflation.

Rules. Correlation is explicit; route draws use common random numbers across route counts; duplicate handling follows a stated venue rule; simulations say so.

Acceptance tests. code/firm/routerace/tests/: the marginals’ median and 99th percentile, the first arrival, no gain at perfect correlation, the duplicate rules by hand, path inflation of a perfect route.

Stretch. Fit route delays and their correlation from logged acknowledgements; a sender that races cancels and sends orders once.

Sources and further reading

  • AWS for Industries blog, “Ultra-low-latency cross-Region crypto trading with Avelacom and AWS” (June 2026); AWS Regions documentation.
  • RFC 4271 (BGP-4); Binance spot API documentation (client order identifiers and duplicate orders).

19.7 Exercises

Exercise 19.1 ★

What is the path inflation of the backbone and of the specialist from Tokyo to Frankfurt?

Solution

Solution of Exercise 19.1.

Against a fibre round-trip floor of 91.3 ms91.3\,\mathrm{m}\mathrm{s}: 224.3/91.3=2.46224.3/91.3 = 2.46 for the backbone and 136.5/91.3=1.50136.5/91.3 = 1.50 for the specialist.

Exercise 19.2 ★

Why does BGP not choose the fastest path?

Solution

Solution of Exercise 19.2.

BGP chooses among routes by the policies of the networks that announce them (business relationships, preferences, path length in autonomous systems), not by measured delay; it does not know the latency of a path at all.

Exercise 19.3 ★

How much round trip does the specialist save over the backbone from Tokyo to each European hub?

Solution

Solution of Exercise 19.3.

Frankfurt 224.3−136.5=87.8224.3 - 136.5 = 87.8 ms; London 210.0−139.6=70.4210.0 - 139.6 = 70.4 ms; Stockholm 245.3−124.3=121.0245.3 - 124.3 = 121.0 ms of round trip.

Exercise 19.4 ★★

Two routes share a subsea cable for half their length. What does that do to the value of racing them?

Solution

Solution of Exercise 19.4.

It correlates their delays and their failures: congestion or a fault on the shared cable hits both, so the minimum of the two is often no better than either, and racing saves little in exactly the events it was meant to cover.

Exercise 19.5 ★★

Which messages can be raced safely under an open-only duplicate rule, and which cannot?

Solution

Solution of Exercise 19.5.

Safe: messages whose repetition has no effect (cancels, which fail harmlessly once the order is gone; queries; heartbeats), and orders that rest long enough that every copy arrives while the first is open. Unsafe: orders that can fill on arrival (taking orders, immediate-or-cancel), whose late copies are accepted as new orders once the first has filled.

Exercise 19.6 ★★

A strategy trades Binance in Tokyo against a European venue. Where should its engine run?

Solution

Solution of Exercise 19.6.

Next to the venue whose prices it must react to fastest, usually the one it quotes or takes on, with the other market’s data carried to it on the fastest route; if it hedges on both, an engine in each region with a fast link between them. The 100 ms100\,\mathrm{m}\mathrm{s}-scale round trip means the engine must be where the race is.

Exercise 19.7 ★★★

Coding. With firm.routerace, find the correlation above which the third route saves less than 100 µs100\,\text{µ}\mathrm{s} at the 99th percentile.

Solution

Solution of Exercise 19.7.

Between ρ=0.46\rho = 0.46 (110 µs110\,\text{µ}\mathrm{s}) and ρ=0.48\rho = 0.48 (98 µs98\,\text{µ}\mathrm{s}): above about 0.48 the third route saves less than 100 µs100\,\text{µ}\mathrm{s} at the 99th percentile.

Exercise 19.8 ★★★

Find the flaw. “We race every order on three routes; the venue rejects duplicates, so it is free latency.”

Solution

Solution of Exercise 19.8.

A venue that rejects duplicates only while the first order is open accepts a late copy once the first has filled: for taking orders, racing on three routes can mean three executions. It is free latency only for messages the venue treats idempotently.

19.8 Problem: Three Routes to Singapore

Problem 19.1

Weekend problem — what racing is worth

A firm in Tokyo sends orders to a venue in Singapore on up to three routes: a specialised network (32.7 ms32.7\,\mathrm{m}\mathrm{s} one way, median), the provider’s backbone (34.0 ms34.0\,\mathrm{m}\mathrm{s}) and the internet (35 ms35\,\mathrm{m}\mathrm{s}); the two private routes’ 99th percentiles are 15% above their medians, the internet’s 30%. A microsecond of 99th percentile is worth 10 USD a month to it.

Part I — The floors.

  1. What are the geodesic, the vacuum and the fibre round-trip floors between Tokyo and Singapore?
  2. What is each published route’s path inflation?
  3. How much of the specialist’s round trip is above the fibre floor?
  4. Why is the internet route’s median an assumption rather than a published figure?

Part II — The race.

  1. What does racing two routes save at the 99th percentile for ρ=0\rho = 0, 0.5 and 0.9?
  2. And three routes?
  3. What is the third route worth a month at each correlation?
  4. What does racing do to the median?

Part III — The duplicates.

  1. Under an open-only rule, how many late copies of an order that fills on arrival are executed?
  2. And of an order that takes a millisecond, four, eight to fill?
  3. Which of the firm’s messages would you race?
  4. What rule would you ask the venue for?

Part IV — The verdict.

  1. State the named result: the 99th-percentile gain of racing two or three routes with correlation ρ\rho, and the price per route above which racing does not pay.
  2. Which input moves the answer most?
  3. How would you estimate ρ\rho?
  4. What happens to ρ\rho when a cable is cut?
  5. Would you race to Frankfurt, where the backbone is 88 ms slower than the specialist?
  6. What does the firm’s position inside each region (chapter 17) do to all of this?
  7. Where does this go in the plan of chapter 29?
  8. In one sentence: when does racing pay?
Solution

Solution of Problem 19.1.

Part I.

  1. 5311 km5311\,\mathrm{k}\mathrm{m}; round-trip floors of 35.4 ms35.4\,\mathrm{m}\mathrm{s} in vacuum and 51.8 ms51.8\,\mathrm{m}\mathrm{s} in fibre.
  2. 1.26 for the specialist, 1.31 for the backbone.
  3. 65.4−51.8=13.6 ms65.4 - 51.8 = 13.6\,\mathrm{m}\mathrm{s} of round trip.
  4. Because nobody published it: it is the chapter’s assumption for a path chosen by BGP, slower and with a heavier tail.

Part II.

  1. 1 623, 771 and 53 µs53\,\text{µ}\mathrm{s}.
  2. 2 141, 861 and 53 µs53\,\text{µ}\mathrm{s}.
  3. At 10 USD per microsecond: 5 182, 902 and 0 USD a month.
  4. It lowers it too, by 536 µs536\,\text{µ}\mathrm{s} at ρ=0\rho = 0 with two routes and 290 at 0.5, because the backbone sometimes beats the specialist; but the gain is largest in the tail.

Part III.

  1. Two: every late copy.
  2. 1.57, 0.50 and 0.07 on average.
  3. Cancels, amendments where the venue treats them idempotently, and queries; not taking orders.
  4. That a client order identifier be rejected for the life of the session (or a day), not only while the order is open.

Part IV.

  1. Named result: racing two routes saves 1623 µs1623\,\text{µ}\mathrm{s} of 99th percentile when they are independent, 771 at ρ=0.5\rho = 0.5 and 53 at 0.9; a third saves 518, 90 and 0 more, so at 10 USD per microsecond a month racing it stops paying above 5 182, 902 and 0 USD a month.
  2. The correlation ρ\rho.
  3. From logged copies: send probes (or every message) on all routes with timestamps and compute the correlation of the routes’ delays, with care for the tails.
  4. It jumps, if both routes share the cut cable’s replacement path, or falls, if one route moves to a separate path; re-measure after every event.
  5. Racing the backbone against a specialist 88 ms88\,\mathrm{m}\mathrm{s} faster saves nothing unless the specialist fails; use the backbone as a fall-back, not a racer.
  6. Everything is added to it: a slow instance at either end wastes the fastest route, so the lottery comes first.
  7. In the cross-region rows: routes, prices, the racing rule per message type, and the venues’ duplicate rules.
  8. When the routes fail independently, the tail matters to the strategy, and the venue will not execute a late copy twice.

19.9 Interview questions

Interview question 19.1 ★ developer

Why is the round trip from Tokyo to Frankfurt over a cloud backbone more than twice the physical floor?

Solution

Solution of Interview question 19.1.

The backbone’s path is chosen for capacity and resilience, not latency, and may go a long way round; add the equipment and the routing between the region’s zones and the backbone. The published figures put it at 2.5 times the fibre floor, where a specialist runs at 1.5.

What the interviewer is looking for: Path inflation; routing by policy; the floor as reference.

Interview question 19.2 ★★ developer

When would you pay for a specialised network between two cloud regions?

Solution

Solution of Interview question 19.2.

When the strategy’s edge depends on seeing another region’s prices sooner than competitors (cross-venue arbitrage, hedging), when the saving is tens of milliseconds as between Tokyo and Europe, and when the traded volume pays for it; not when the backbone is already close to the floor, as to Singapore.

What the interviewer is looking for: Value of latency against cost; where the backbone is bad; measurement.

Interview question 19.3 ★★ developer, researcher

You send the same message on two routes. What does it buy you, and what does it depend on?

Solution

Solution of Interview question 19.3.

The minimum of the two delays: a lower median and a much lower tail if the routes fail independently, little if they are correlated. It depends on the correlation, on each route’s tail, and on whether the receiver can discard the second copy safely.

What the interviewer is looking for: Order statistics; correlation; idempotency.

Interview question 19.4 ★★ developer

What is an idempotency key, and how can it fail to protect you?

Solution

Solution of Interview question 19.4.

An identifier attached to every copy of one request so that the receiver acts once. It fails if the receiver forgets the identifier too soon (for example, once the order has filled), if the sender reuses identifiers, or if different copies carry different keys.

What the interviewer is looking for: Receiver-side memory; lifetime of keys; sender discipline.

Interview question 19.5 ★★ trader

How does a hundred milliseconds between Tokyo and Europe shape a cross-venue crypto strategy?

Solution

Solution of Interview question 19.5.

It makes cross-venue strategies local: an engine must sit next to the venue it trades, see the other market a tenth of a second late, and trade only signals that survive that delay; fast arbitrage between regions is a contest of routes, won by those on the best private line.

What the interviewer is looking for: Engine placement; signal horizon; the value of private lines.

Interview question 19.6 ★★★ developer, researcher

Design an order-sending layer that races routes without ever executing an order twice.

Solution

Solution of Interview question 19.6.

Race only idempotent messages; for new orders, send once on the best route with a fast fail-over; if a venue remembers keys for the session, race orders too with one key; never generate a new key for a copy; reconcile every acknowledgement against the key, and cancel any unexpected duplicate fill’s exposure at once.

What the interviewer is looking for: Message classes; venue rule; reconciliation; fail-over rather than racing.

Terms defined in this chapter

See all 2333 terms in the glossary