Research, Data and Risk Platforms · Technology
26Cloud Against On-Premises
On 4 November 2021 CME Group and Google Cloud announced a ten-year partnership: CME would move its technology to Google Cloud, starting the next year with data and clearing services and eventually all of its markets. The same year most trading firms kept every order path in their own racks, as Book 14 describes. Both decisions can be right, because they are about different workloads. This chapter prices the one where the question is most often asked and least often computed — the research cluster — with every cloud price taken from a dated, published list, every assumption about owned hardware stated, and chapter 13’s simulator run on interruptible capacity.
26.1 What moves to the cloud and what cannot
A public cloud (Book 14, chapter 16) rents machines by the hour in regions and availability zones the firm does not control. Four questions sort a workload. How sensitive is it to latency and jitter — the order path, Book 14 shows, is still measured in the firm’s own racks next to the matching engine. How variable is its demand — a sweep that needs thousands of cores for a weekend and none on Monday is the cloud’s case. Where is its data — the next section’s question. And who must be able to inspect it — the last section’s.
Definition 26.1 (Infrastructure as code)
Infrastructure as code describes servers, networks, storage and their permissions in version-controlled files from which a tool creates, changes and destroys them, so that an environment can be rebuilt, reviewed and compared like software.
Infrastructure as code is what makes elastic capacity usable: a cluster that exists only while a sweep runs must be created and destroyed by a program, identically every time, and an owned cluster described the same way can be rebuilt after a failure. It applies to both sides of this chapter’s comparison.
26.2 The economics of compute: owned, reserved, on demand, interruptible
Definition 26.2 (On-demand, reserved and spot capacity)
On-demand capacity is rented by the second or hour with no commitment, at the provider’s list price. Reserved capacity is committed for one or three years at a lower hourly price, paid whether used or not. Spot capacity is the provider’s unused capacity rented at a lower, changing price, which the provider can reclaim at short notice, interrupting whatever runs on it.
The chapter prices one cloud instance of 96 hardware threads in one US region from the provider’s public price list of 25 September 2026: 4.284 dollars an hour on demand, 2.834 reserved for one year and 1.886 for three years without upfront payment, and 1.737 on the spot market on 28 September. Per thread-hour that is 4.46, 2.95, 1.96 and 1.81 cents. Owned hardware has no such list. The model builds a two-socket server of 512 threads from the list price of its CPU (10 931 dollars each), assumes the CPUs are half of the server’s price, spreads it over four years, powers it at the US average commercial price of electricity (14.53 cents a kilowatt-hour in July 2026) with a 1.3 kW draw and a building overhead of 1.4, and adds $300 a month of rack space and a two-hundredth of an engineer’s $250 000 a year. Those last figures are assumptions, not facts, and the results below are shown for server prices up to three times the assumption.
@dataclass(frozen=True)
class OwnedServer:
price: float # purchase, USD
life_months: int # straight-line depreciation
kw: float # average draw at load, kW
pue: float # facility overhead (power usage effectiveness)
power_price: float # USD per kWh
space: float # USD per server-month
staff: float # USD per server-month
@property
def monthly(self) -> float:
power = self.kw * self.pue * HOURS_PER_MONTH * self.power_price
return self.price / self.life_months + power + self.space + self.staff
def per_thread_hour(self, threads: int) -> float:
return self.monthly / (threads * HOURS_PER_MONTH)
The owned server then costs $1 508 a month, 0.40 cents per thread-hour when fully used. The comparison turns on utilisation: capacity that is paid for every hour — owned or reserved — costs its hourly price divided by the share of hours it is used, while on-demand and spot capacity cost their price only when used (Figure 26.1).
fig_cloudcost.py.| server price | per month | cents per thread-hour | break-even: on demand | spot |
|---|---|---|---|---|
| as assumed ($43 724) | $1 508 | 0.40 | 9.0% | 22.3% |
| twice | $2 419 | 0.65 | 14.5% | 35.8% |
| three times | $3 330 | 0.89 | 20.0% | 49.2% |
A thread is not a unit of work: the owned CPU and the cloud instance’s CPU run a given backtest at different speeds, and a firm should compare them with chapter 14’s measurements of its own workload before trusting any break-even. A factor of two in speed divides or multiplies the owned server’s whole cost per unit of work by two, and the break-even with it: 4.5% or 18% against on demand.
26.3 Buying for a year of demand
A research cluster’s demand is not a utilisation but a profile. The model’s year has an interactive base of 1 500 threads, 1 000 more in weekday working hours, a nightly batch of 4 000 more on weekdays, a weekend sweep of 12 000 more for eight hours every Saturday, and twenty unplanned bursts of 8 000 for four hours: a mean of 3 364 threads against a peak of 13 500. Sorting the hours by demand gives the load-duration curve, and each level of capacity is used in the share of hours the demand reaches it.
def plan_capacity(demand, owned: float, reserved: float, on_demand: float,
spot: float | None = None, spot_share: float = 1.0) -> Plan:
"""Buy each unit of capacity the cheapest way: the k-th unit is needed in the share
f(k) of hours with demand >= k; fixed options cost their price every hour, metered
ones only f(k) of hours. `spot_share` of metered work may run on spot."""
d = np.sort(np.asarray(demand, dtype=float))
hours = len(d)
metered = on_demand if spot is None else (
spot_share * spot + (1 - spot_share) * on_demand)
levels = np.arange(1, int(np.ceil(d[-1])) + 1)
f = 1.0 - np.searchsorted(d, levels - 1e-9, side="left") / hours # hours >= k
n_own = int(np.sum(owned < np.minimum(reserved, metered * f)))
n_res = int(np.sum((reserved <= owned) & (reserved < metered * f)))
over = np.clip(d - (n_own + n_res), 0, None).sum() # unit-hours bought hourly
cost = {"owned": n_own * owned * hours if n_own else 0.0,
"reserved": n_res * reserved * hours if n_res else 0.0,
"metered": over * metered}
return Plan(n_own, n_res, float(over), cost)
fig_cloudcost.py.fig_cloudcost.py.Definition 26.3 (Cloud bursting)
Cloud bursting runs a workload on the firm’s own capacity up to its size and sends the excess to rented cloud capacity, so that the owned cluster is sized for the frequent load and the cloud absorbs the rare peaks.
The optimum in Figure 26.3 is a hybrid, and it is one because of the shape of the demand, not because of any single price: the base and the nightly batch are used more than 9% of the time and are cheaper owned; the weekend sweeps and bursts, 496 hours of the year, are cheaper rented. A firm whose research runs flat out every night owns; a firm whose research is a few enormous sweeps a year rents; most are in between.
26.4 Spot capacity and the cost of interruption
Spot capacity is 59% cheaper than on demand for this instance, and the provider’s own advisor places it in the highest band of interruption frequency, above 20% of instances reclaimed in the trailing month — at least 0.03% an hour. Chapter 13’s sweep of 10 000 backtests, forty of them seven and a half hours long, runs on 24 instances, each billed until its last task ends.
def checkpointed(tasks: list, chunk: float, overhead: float) -> list:
"""Split each task into chunks of `chunk` seconds (plus `overhead` per checkpoint),
chained so that a preempted chunk restarts from the last checkpoint."""
out, nid = [], max(t.tid for t in tasks) + 1
for t in tasks:
k = max(1, int(np.ceil(t.duration / chunk)))
prev = None
for i in range(k):
d = t.duration / k + (overhead if k > 1 else 0.0)
tid = t.tid if i == 0 else nid
if i > 0:
nid += 1
deps = (prev,) if prev is not None else ()
out.append(J.Task(tid, d, t.estimate / k, t.team, deps=deps))
prev = tid
return out
fig_cloudcost.py.On demand, the sweep costs $203 and takes 7.5 hours, the length of its longest backtest. On spot without interruptions it costs $82. What interruption costs is mostly time, and it falls on the long tasks (Figure 26.4): a 7.5-hour backtest interrupted after six hours starts again, so a single interruption of one of the forty long tasks nearly doubles the sweep’s length. Cutting each backtest into 15-minute chunks that write a checkpoint — the training checkpoint of Book 12, chapter 23, applied to backtests — bounds the loss at one chunk, and the sweep costs the same at every interruption rate the chart shows. Spot capacity is for work that can be interrupted, and checkpointing is what makes work interruptible.
26.5 Data gravity and data-transfer charges
Definition 26.4 (Data gravity)
Data gravity is the tendency of computation to move to where large data already is, because moving the data costs more in time and data-transfer charges than moving the computation.
The provider’s data-transfer charge (Book 14, chapter 16) is asymmetric: data coming into the region is free, data leaving it to the internet costs 9 cents a gigabyte for the first 10 terabytes of a month, falling in tiers to 5 cents above 150 terabytes. Keeping a 500-terabyte tick store in the cloud’s standard object storage (chapter 2’s price) costs $11 776 a month; taking a full copy out once costs $29 491, about twenty months of an owned server. So the tick store’s location decides where research runs: once it is in the cloud, research follows it, and results — a few gigabytes a sweep — are what should cross the boundary. A hybrid design keeps one authoritative copy and sends the computation to it, or keeps two copies and pays once to fill the second.
26.6 Security, control and regulators
Most cloud incidents at financial firms are on the customer’s side of that line: a storage bucket left public, a key committed to a repository, a permission granted too widely — the subjects of chapter 29. Regulators add a second line. In the European Union the EBA’s guidelines on outsourcing, which apply since 30 September 2019 and absorbed its earlier recommendation on cloud outsourcing, require a documented exit strategy for every critical or important function outsourced: a plan, tested, for moving the function elsewhere if the provider fails or the relationship ends. For a research cluster that is easy — the owned half of a hybrid is the exit. For an order path or a clearing system, it is a second full implementation, which is part of why a move like CME’s is a ten-year programme. Concentration — many firms relying on the same few providers — is the operational risk (Book 6, chapter 28) that supervisors watch most.
As of September 2026 — Cloud prices and the rules that apply
AWS’s public price list of 25 September 2026 gives, for a c7i.24xlarge instance (96 vCPUs) in US East (N. Virginia) running Linux, 4.284 dollars an hour on demand, 2.83387 reserved for one year and 1.88608 for three years without upfront payment; its spot price feed gave 1.7372 on 28 September, and its Spot Instance Advisor put the instance’s interruption frequency in the highest band (above 20% in the trailing month) with 65% savings over on demand. Data transfer out of the region to the internet costs 0.09 dollars per gigabyte for the first 10 TB a month, 0.085 for the next 40, 0.07 for the next 100 and 0.05 above 150 TB; inbound is free. AMD lists its 128-core EPYC 9755 at 10 931 dollars (1 000-unit price); the EIA reports an average US commercial electricity price of 14.53 cents per kWh for July 2026. CME Group and Google Cloud announced their ten-year partnership, with a one-billion-dollar equity investment by Google, on 4 November 2021. The EBA’s outsourcing guidelines (EBA/GL/2019/02) apply from 30 September 2019.
26.7 Tutorial: own, rent or both
Goal. Price a research cluster’s year every way, find the break-even utilisation, and run a sweep on spot capacity. End state: Figures 26.1, 26.3 and 26.4.
- Prices:
pl_cloudcost.PRICESandSERVER;per_thread()andbreak_evens(k)for . - Demand:
demand(); plot its load-duration curve. - Plans:
plans(), the four ways of buying the year. - Burst:
burst(rate, price, checkpoint)over the interruption rates, with and without checkpoints. - Data:
firm_cloudcost.transfer_outfor the tick store and for a sweep’s results.
What to change next. Replace the thread as the unit by measured backtests per hour on each machine; let reserved capacity cover the nightly batch when the firm cannot buy servers in time.
26.8 Build: the cloud cost model
Purpose. Price any compute plan from dated, cited lists and stated assumptions, and choose the mix of owned, reserved, on-demand and spot capacity for a demand profile.
Interface. PriceItem, OwnedServer, cost_per_used_hour, break_even, plan_capacity, checkpointed, billed_node_hours, transfer_out.
Rules. Every price carries its source and date; every assumption is named and shown with a sensitivity; plans are compared on the same demand; interruptible capacity only for checkpointed work; data-transfer charges counted both ways.
Acceptance tests. code/firm/cloudcost/tests/: the owned server’s monthly and hourly cost; fixed and metered costs and the break-even; the planner on a two-level demand, with reserved and spot capacity; checkpoint chunks, their chaining and the billing; the tiers of the transfer charge.
Stretch. Savings plans and committed-use discounts; interruption rates from measured history; storage tiers and retrieval charges in the data-gravity comparison.
Sources and further reading
- CME Group and Google Cloud, press release of 4 November 2021.
- AWS public price lists (EC2 on demand and reserved, data transfer), spot price feed and Spot Instance Advisor; AWS, Shared Responsibility Model.
- AMD, EPYC 9755 product page; U.S. Energy Information Administration, Electric Power Monthly, Table 5.6.A.
- European Banking Authority, Guidelines on outsourcing arrangements, EBA/GL/2019/02.
- One Quant Book 14, chapter 16 (public clouds, zones, data-transfer charges); Book 16, chapter 20 (total cost of ownership).
26.9 Exercises
Exercise 26.1 ★
What does one thread-hour cost on demand, reserved for three years, and on spot, at the cited prices?
Solution
Solution of Exercise 26.1.
The instance’s hourly price divided by its 96 threads: 4.46 cents on demand, 1.96 cents reserved for three years without upfront payment, and 1.81 cents on spot (on 28 September 2026).
Exercise 26.2 ★
Above what utilisation is three-year reserved capacity cheaper than on-demand?
Solution
Solution of Exercise 26.2.
Above 1.886/4.284, 44%: below that, paying every hour for reserved capacity costs more than paying on demand only for the hours used.
Exercise 26.3 ★
What does it cost to take a 500-terabyte tick store out of the region once?
Solution
Solution of Exercise 26.3.
$29 491: 10 TB at 9 cents a gigabyte, 40 TB at 8.5, 100 TB at 7 and the remaining 350 TB at 5 (with 1 024 GB to the TB). Bringing it in was free.
Exercise 26.4 ★★
Why does the planner never choose reserved capacity when owning is allowed?
Solution
Solution of Exercise 26.4.
Both are paid every hour whether used or not, so for any unit of capacity the cheaper hourly price wins at every utilisation: 0.40 cents owned against 1.96 reserved. Reserved capacity wins only when owning is not possible — no data centre, no time to buy, a need that lasts less than the server’s life.
Exercise 26.5 ★★
Why does a single interruption of a long backtest nearly double the sweep’s length?
Solution
Solution of Exercise 26.5.
The sweep’s length is set by its longest backtests, 7.5 hours each. An interruption restarts one from the beginning, so if it strikes six hours in, that backtest ends after 13.5 hours and the sweep with it: 14.4 hours at the advisor’s band.
Exercise 26.6 ★★
The firm’s CPUs run a backtest twice as fast as the cloud instance’s. What happens to the break-even?
Solution
Solution of Exercise 26.6.
The owned server’s cost per unit of work halves, and so does the break-even: 4.5% against on demand, 11% against spot. Speed matters as much as price, which is why the comparison should be made in backtests per dollar, measured.
Exercise 26.7 ★★★
Coding. Rerun plans() on demand(spread_sweep=True), where the weekend sweep’s thread-hours run inside the five weeknight batch windows instead. How much capacity does the planner own, and what does the year cost?
Solution
Solution of Exercise 26.7.
The sweep adds 2 400 threads to every weeknight’s batch, a level used 24% of hours, so the planner owns 7 900 threads instead of 5 500 and the year costs $291 003 with on-demand bursts, against $361 215: moving flexible work into the hours owned capacity already covers raises its utilisation and saves more than any price negotiation.
Exercise 26.8 ★★★
Find the flaw. "Spot is 59% cheaper, so we will run the whole research cluster on spot."
Solution
Solution of Exercise 26.8.
Spot capacity can be reclaimed at any moment: interactive work and anything without checkpoints loses its progress, and the chapter’s sweep without checkpoints took twice as long at the advisor’s band. And for the base load, used all the time, owned capacity is more than four times cheaper per thread-hour than spot at the model’s prices.
26.10 Problem: Own, Rent or Both
Problem 26.1
Weekend problem — own, rent or both
The chapter’s prices, owned server, demand profile and sweep.
Part I — Prices.
- What are the four cloud prices per thread-hour, and where do they come from?
- What is cited and what is assumed in the owned server’s cost?
- What does the owned server cost a month and per thread-hour?
- How does utilisation enter the comparison?
- What are the break-evens against on-demand and spot, and how do they move with the server’s price?
Part II — The year.
- What is the model’s demand profile, its mean and its peak?
- How does the planner decide each unit of capacity?
- How much does it own, and how many hours are bought by the hour?
- What does each of the four plans cost?
- Why is the optimum a hybrid?
Part III — Spot and data.
- What does the provider say about this instance’s interruptions?
- What do the sweep’s cost and length become on spot without checkpoints?
- What do checkpoints change?
- What does it cost to store the tick store in the cloud, and to take it out?
- What does a regulator require before a critical function moves to a cloud?
Part IV — The verdict.
- State the named result: the utilisation above which owning a research cluster costs less than renting it at cited prices, and the cost of the chapter 13 sweep on spot capacity with preemption against on demand.
- Which of the model’s assumptions would you check first, and how?
- Where should the tick store live?
- What would make you move the whole cluster to the cloud?
- In one sentence: what decides between owning and renting compute?
Solution
Solution of Problem 26.1.
- 4.46 cents on demand, 2.95 reserved one year, 1.96 reserved three years, 1.81 spot: the provider’s published price list of 25 September 2026 and its spot feed of 28 September, divided by 96 threads.
- Cited: the CPU’s list price and the average commercial electricity price. Assumed: the rest of the server (as much again as the CPUs), four-year life, 1.3 kW, overhead 1.4, $300 of space and a two-hundredth of an engineer a month.
- $1 508 a month, 0.40 cents per thread-hour fully used.
- Fixed capacity (owned, reserved) costs its hourly price divided by utilisation; metered capacity (on demand, spot) costs its price per hour used.
- 9.0% against on demand and 22.3% against spot; 14.5% and 35.8% at twice the server price, 20.0% and 49.2% at three times.
- Base 1 500 threads, 1 000 more in weekday hours, 4 000 more on weeknights, 12 000 more for eight hours on Saturdays, twenty bursts of 8 000 for four hours; mean 3 364, peak 13 500.
- From the load-duration curve: the -th unit is used in the share of hours, and it is owned if owning costs less than times the hourly price, reserved on the same test when owning is excluded, otherwise bought by the hour.
- 5 500 threads owned; 496 hours have demand above it.
- $1 315 233 on demand only, $940 460 reserved and on demand, $361 215 owned and on demand, $281 883 owned with spot bursts.
- Because the demand has a base used most of the time and peaks used rarely, on opposite sides of the break-even.
- Savings of 65% over on demand over the last 30 days, and an interruption frequency in the highest band, above 20% in the trailing month.
- At the advisor’s band, $94 and 14.4 hours; at 0.2 an hour, $185 and 59 hours; against $203 and 7.5 hours on demand.
- The loss is bounded at one chunk: $80–82 and under 9 hours at every rate.
- $11 776 a month in standard object storage; $29 491 to take a copy out once.
- In the EU, under the EBA’s outsourcing guidelines, among other things a documented exit strategy for a critical or important function.
- Named result. At the cited prices and the model’s owned server, owning costs less than renting on demand above 9% utilisation (22% against spot; 20% and 49% even at three times the server price). Chapter 13’s sweep costs $203 on demand and $82 on spot without interruptions; at the provider’s highest interruption band it costs $94 and takes 14.4 hours instead of 7.5, and with 15-minute checkpoints $80 in under 9 hours.
- The relative speed of the two machines on the firm’s own backtests, measured; then the server’s full price from a quote.
- Where most of the computation that reads it runs; with one authoritative copy, and computation sent to it.
- Demand dominated by rare large peaks, no data centre or staff, or data already in the cloud.
- The shape of the demand against the price of fixed and metered capacity, where the data is, and how the work can be interrupted.
26.11 Interview questions
Interview question 26.1 ★ developer
What are on-demand, reserved and spot instances, and when would you use each?
Solution
Solution of Interview question 26.1.
On demand: no commitment, highest price, for irregular work. Reserved: one or three years committed at a lower price, for steady load. Spot: spare capacity at the lowest price that can be reclaimed, for interruptible, checkpointed batch work.
What the interviewer is looking for: Commitment, price, interruptibility.
Interview question 26.2 ★★ developer
How would you decide whether to buy a research cluster or rent one?
Solution
Solution of Interview question 26.2.
Build the demand profile and its load-duration curve, price owned capacity per hour with its assumptions stated, and buy each level of capacity the cheapest way given the share of hours it is used; check data location and interruptibility; compare in work per dollar, measured.
What the interviewer is looking for: Utilisation and break-even, not list prices.
Interview question 26.3 ★★ developer
How do you run a large backtest sweep on capacity that can be taken away at any moment?
Solution
Solution of Interview question 26.3.
Make every task restartable: checkpoint long tasks at intervals, keep tasks idempotent with results written atomically, let the scheduler resubmit interrupted work, and spread over instance types and zones.
What the interviewer is looking for: Checkpoints and idempotent tasks.
Interview question 26.4 ★★ developer
Your tick store is 500 terabytes on premises. A team wants to run research in the cloud. What do you tell them?
Solution
Solution of Interview question 26.4.
Moving it in is free but it must then be stored there (about $12 000 a month in standard storage) and taking it out costs about $29 000 a copy; decide where the authoritative copy lives and send computation to it, or copy only the subset the research needs.
What the interviewer is looking for: Data gravity and data-transfer charges.
Interview question 26.5 ★★★ developer
Design a hybrid research platform that uses owned capacity for the base load and the cloud for peaks.
Solution
Solution of Interview question 26.5.
Owned cluster sized at the level used above the break-even share of hours; one scheduler over owned and cloud nodes; bursts on spot for checkpointed work and on demand for the rest; data near the owned cluster with cached subsets in the cloud; everything as code; an exit plan.
What the interviewer is looking for: Size owned for the base, burst the peaks.