Machine Learning for Markets · Machine learning
28The Machine-Learning Team
A model that took a researcher three weeks took the firm five months to trade. Nothing in it was hard. Every handoff between the people who built it was a queue: the review committee met weekly, the engineers who rewrote the features for production had a backlog, the validators sent it back twice because the rewrite did not match the research, and paper trading could not be hurried. This chapter is about the organisation around the models of the previous twenty-seven chapters: who does what, what passes between them, and how to measure the result. It builds two things. The first is a handoff specification that a continuous-integration job can check, so that the package a researcher hands over either runs in production as it ran in research or fails a named test. The second is a simulator of the research-to-production pipeline as a network of queues. On the chapter’s pipeline, the median model reaches production 20 weeks after research ends, and 65% of models are lost on the way, most of them at the first review; replacing the rewrite by a validated package cuts the median to 11 weeks, one of which is the researcher’s time spent building the package.
28.1 Roles: researcher, engineer, platform
Definition 28.1 (Machine-learning engineer, model owner)
A machine-learning engineer builds and operates the systems that train, serve and monitor models in production, combining the skills of data scientists, data engineers and software engineers. The model owner is the named person accountable for a model in production: its performance, its monitoring, its retraining and its retirement, and the answers to the questions validators, risk managers and regulators ask about it.
A trading firm’s machine-learning work involves at least three kinds of people. Researchers (quantitative researchers or data scientists) find the signal and fit the model, with the methods of this book and the craft of Book 7. Machine-learning engineers turn the model into a service that runs every day without its author: features computed online as they were offline (chapter 24), a compiled model inside the latency budget (chapter 26), monitors (chapter 27). A platform team builds what every model shares. Kreuzberger, Kühl and Hirschl, from a literature review, a tool review and expert interviews, list seven roles around machine-learning operations (business stakeholder, solution architect, data scientist, data engineer, software engineer, DevOps engineer, and the machine-learning or MLOps engineer who combines several of them); a trading firm adds the model validator of Book 6 (chapter 26), independent of the others, and the trader or portfolio manager who uses the model.
The model owner is a role, not a team. It is usually the researcher who proposed the model or the manager of the desk that uses it; what matters is that it is one named person, recorded in the registry (chapter 25) and on the model card (chapter 21), who answers the page of a performance alarm and signs the decision to retrain or retire. A model without an owner is the one whose monitors page nobody. Amershi and co-authors, observing teams at Microsoft, found that model customisation and reuse require skills different from those of software teams, and that models are harder to treat as distinct modules than ordinary software components; both are reasons the division of work between researchers and engineers is a design choice with a cost, not a given.
28.2 How work moves: the handoff
Definition 28.2 (Handoff specification)
A handoff specification is the contract between the people who build a model and the people who run it: the list of what a model package must contain (the artefact and its content hash, the feature definitions, test vectors of inputs with the outputs the model must reproduce, the latency budget with a measurement, the monitoring specification, the model card) and the checks the package must pass before production accepts it.
In the hook’s firm the handoff was a rewrite. The researcher handed over a notebook; an engineer reimplemented the features and the model in production code; the validator compared the two and, when they differed, sent the model back. Every difference was a round trip through two queues. Zinkevich’s rules for machine-learning engineering say it directly: “Re-use code between your training pipeline and your serving pipeline whenever possible”. The handoff specification is the other half of that advice: when the same code runs in both places, what passes between the teams is a package that can be checked mechanically.
The chapter’s package is for chapter 26’s compiled forest (Listing 28.1). Its manifest names the forest file and its content hash (Book 7’s firm.workflow); sixteen feature definitions in firm.featstore’s format; research’s values of those features on a reference session of the simulated order book; 200 test vectors with the forest’s outputs, to be reproduced to ; a latency budget of 5 microseconds with the measured 2.6; four monitors with thresholds, a false-page rate and an owner each; and the model card’s fifteen fields. The validator runs nine checks, and the clean package passes all of them. Then the chapter plants eight defects a handoff typically carries, one at a time (Table 28.1).
| defect planted | checks that fail |
|---|---|
| none (the package as handed over) | none of the nine |
| artefact rebuilt after it was hashed | artefact hash; test vectors |
| a feature reads an input the feed does not carry | feature specification |
| a 5-second window written as 5 000 (milliseconds) | feature vectors |
| one feature missing (15 for a model reading 16) | feature specification |
| test vectors holding another model’s outputs | test vectors |
| latency of the library call, not the compiled model | latency |
| a monitor without an owner | monitoring |
| a model card without limitations | model card |
ml_team.handoff_table.Every defect is caught, and by the check written for it, with one lesson in the third row. A window in the wrong unit is still a valid feature definition, and the offline and online computations agree with each other, since both use the same wrong window: the specification and parity checks pass. Only the comparison with research’s own feature values on a reference session finds it. Test vectors must cover the features as well as the model.
28.3 What the platform owns
Definition 28.3 (Machine-learning platform)
A machine-learning platform is the shared infrastructure on which every model is built and run: data and feature stores, training infrastructure, experiment tracking and the model registry, compilation and serving, monitoring, and the checks a model package passes on its way to production, owned by one team and used by all.
The platform owns everything that is the same for every model. In this book’s terms: the feature store with its parity tests (chapter 24), the training infrastructure (chapter 23), the tracker, registry and audit trail (chapter 25), the compilers and serving kernels (chapter 26), the monitors (chapter 27) and the handoff validator. Researchers own the model and its features; the platform owns the guarantee that a feature means the same thing in research and in production. The line matters because each side tends to push work across it: researchers ask for one-off pipelines, and platforms grow features no model uses. A useful test is Zinkevich’s fourth rule, “Keep the first model simple and get the infrastructure right”: a platform is justified by the second model, not the first. Its other justification is visibility: among the risks specific to machine-learning systems, Sculley and co-authors list undeclared consumers and data dependencies, which a shared registry and feature store make explicit.
The validator is the platform’s most leveraged piece, because it turns a conversation into a test. In the chapter’s pipeline below, it is what lets the integration of a model take a week instead of a rewrite of four.
28.4 Measuring the pipeline
The pipeline is a network of queues (Listing 28.5); its parameters are illustrative, not measured at any firm. Models arrive when research ends, at 0.35 a week (18 a year), and pass five stations. The review committee (one slot, 0.6 weeks a model) kills 35% of the models it reviews and sends 15% back for two weeks of research fixes. Two engineers reimplement each surviving model in about four weeks; 30% of rewrites go back to the researcher for a week and a half of clarification and are then rewritten, and 5% are abandoned. One validator (1.5 weeks) kills 10% and returns 25% to the engineers. Paper trading lasts four weeks and kills 20%; a two-week canary kills 5%. Service times are exponential except for the two fixed trading periods; the first 500 of 4 000 simulated models are discarded while the pipeline fills.
The simulator is checked on the one case with a closed form. A single station of the M/M/1 queue of Book 4 (chapter 8), arrivals at 0.8 and service at 1 a week, gives a mean time in the system of 5.05 weeks against the theoretical ; the stationary law of the corresponding birth–death chain, computed with Book 4’s firm.queues, gives a mean of 4.00 models in the system and an empty system 20% of the time, and the simulation’s arrival rate times its mean time gives 4.04, as Little’s law requires. On the whole pipeline, the time-average number of models in it is 4.87 and the arrival rate times their mean time in it (14.0 weeks, counting killed models until they die) is 4.90.
| rewrite (handoff) | validated package | |
|---|---|---|
| median lead time to production (weeks) | 20.4 | 11.1 |
| quartiles (weeks) | 14.8 to 30.2 | 9.6 to 13.7 |
| 90th percentile (weeks) | 42.9 | 16.5 |
| share of models reaching production | 35.1% | 37.5% |
| lost at review | 41.1% | 39.5% |
| lost at the build (reimplementation / integration) | 5.3% | 3.5% |
| lost at validation | 7.9% | 6.9% |
| lost in paper trading | 8.7% | 10.3% |
| lost in the canary | 1.9% | 2.2% |
| passes through the build, per model past review | 1.79 | 1.14 |
| engineers’ utilisation | 74% | 12% |
ml_team.pipeline.Table 28.2 is the chapter’s result. With the rewrite, the median model reaches production 20.4 weeks after research ends, about five months, and one in ten takes more than 42.9 weeks; 35% of models get there. The largest loss is at review, 41% of all models: that is where most models should die, cheaply and early. Of the time, service without any waiting would account for 14.9 weeks of the median; the rest is queueing, concentrated at the engineers, who are busy 74% of the time because each model passes through them 1.79 times on average. Replacing the rewrite by a validated package (integration in about a week, rework at the build falling from 30% to 5% and at validation from 25% to 10%, plus a week of the researcher’s time) cuts the median to 11.1 weeks, a saving of 9.3 weeks (45%), and the 90th percentile from 42.9 to 16.5. The shares lost at each gate barely move, which is right: removing a handoff should make decisions faster, not different.
ml_team.rate_curve.The table holds at one arrival rate; Figure 28.1 shows why it matters more than it seems. The engineers’ station is a queue near saturation, and a queue’s waiting time grows without bound as its utilisation approaches one: from 0.35 to 0.44 models a week, a quarter more research, the median lead time with the rewrite rises from about 22 to 70 weeks, while with the package it stays near 11. A pipeline is measured by its lead time, the share lost at each gate and the utilisation of its busiest station, and a manager who adds researchers without looking at the last number makes every model slower.
Method 28.4 (Running a research-to-production pipeline)
- Name a model owner for every model, in the registry and on the model card.
- Replace rewrites by packages: the same feature and model code in research and production, and a handoff specification checked in continuous integration, with test vectors for the features as well as the model.
- Kill models early: the cheapest gate should reject the most.
- Measure lead time, losses per gate and the utilisation of each team; keep the busiest well below saturation.
- Put what every model shares on a platform owned by one team, and justify it by the second model.
28.5 Tutorial: who owns the model?
Goal. Validate a handoff package in the way continuous integration would, plant defects and see which check catches each; simulate the pipeline with and without the rewrite. End state: Table 28.1, Table 28.2, Figure 28.1.
The validator.
def validate(pkg, events=None, times=None): """Every check, in order; a check that cannot run because an earlier one failed is reported as failed.""" out = [] missing = [k for k in REQUIRED if k not in pkg] out.append(("schema", not missing, f"missing: {missing}" if missing else "all fields present")) if missing: return out raw = pathlib.Path(pkg["artefact"]).read_bytes() ok = content_hash(raw) == pkg["artefact_hash"] out.append(("artefact hash", ok, "matches" if ok else "the artefact is not the one the hash names")) forest = read_forest(pkg["artefact"]) n_in = int(forest["feature"].max()) + 1 try: defs = _features(pkg) except TypeError as e: defs, err = [], str(e) else: err = "" bad = [d.name for d in defs if d.input not in INPUTS or d.agg not in AGGS or d.window < 0] names = [d.name for d in defs] ok = not err and not bad and len(set(names)) == len(names) and len(defs) == n_in detail = err or (f"unknown input or aggregation: {bad}" if bad else f"{len(defs)} features for a model reading {n_in}" if len(defs) != n_in else "duplicate names" if len(set(names)) != len(names) else f"{len(defs)} features") out.append(("feature specification", ok, detail)) if events is not None and ok: off = offline(events, defs, times) p = parity(off, stream(events, defs, times)) out.append(("feature parity", p["mismatch share"] == 0, f"mismatch share {p['mismatch share']:.3f}")) if pkg.get("feature_vectors") is not None: want = np.loadtxt(pkg["feature_vectors"], delimiter=",", ndmin=2) e = float(np.max(np.abs(off - want))) out.append(("feature vectors", e <= 1e-9, f"largest error {e:.3g} against research's values"))Listing 28.1. Schema, artefact hash, feature specification, parity and feature vectors. code/firm/modelpkg/firm_modelpkg.py The rest of the checks.
V = np.loadtxt(pkg["test_vectors"], delimiter=",", ndmin=2) got, want = forest_predict(forest, V[:, :n_in]), V[:, n_in] err = float(np.max(np.abs(got - want))) out.append(("test vectors", err <= pkg["tolerance"], f"{len(V)} vectors, largest error {err:.3g}")) ok = 0 < pkg["latency_measured_us"] <= pkg["latency_budget_us"] out.append(("latency", ok, f"{pkg['latency_measured_us']} us measured, budget {pkg['latency_budget_us']} us")) mons = pkg["monitoring"] bad = [m.get("statistic") for m in mons if m.get("statistic") not in MONITORS or not all(k in m for k in ("threshold", "pages_per_month", "owner"))] out.append(("monitoring", bool(mons) and not bad, f"incomplete: {bad}" if bad else f"{len(mons)} monitors")) miss = ModelCard(pkg["card"]).missing() out.append(("model card", not miss, f"missing: {miss}" if miss else "complete")) return outListing 28.2. Test vectors through the forest kernel, latency, monitoring and the model card. code/firm/modelpkg/firm_modelpkg.py The planted defects (
ml_team._defects).def unknown_input(p, d): p["features"][3] = {"name": "microprice minus mid", "input": "microprice", "agg": "last", "window": 0.0} def milliseconds(p, d): p["features"][5]["window"] = 5000.0 # 5 s written as 5,000 ms def missing_feature(p, d): p["features"].pop() def stale_vectors(p, d): # expected outputs of another model V = np.loadtxt(p["test_vectors"], delimiter=",") V[:, 16] = V[:, 17] f = pathlib.Path(d) / "vectors.csv" np.savetxt(f, V, delimiter=",", fmt="%.17g") p["test_vectors"] = str(f) def over_budget(p, d): p["latency_measured_us"] = 265.0 # the library call, not the compiled model def unowned_monitor(p, d): del p["monitoring"][3]["owner"] def no_limitations(p, d): p["card"]["limitations"] = ""Listing 28.3. Seven of the eight defects, each an edit of a loaded package. code/ml/28-the-machine-learning-team/python/ml_team.py The pipeline.
def stations(handoff=True): """Weeks. The main line: review, reimplementation (or integration of a package), validation, paper trading, canary; side loops (names starting '~') for rework with the researcher.""" build = (Station("reimplementation", 2, 4.0, kill=0.05, rework=0.30, rework_to=6) if handoff else Station("integration", 2, 1.0, kill=0.05, rework=0.05, rework_to=6)) return (Station("review", 1, 0.6, kill=0.35, rework=0.15, rework_to=5), build, Station("validation", 1, 1.5, kill=0.10, rework=0.25 if handoff else 0.10, rework_to=1), Station("paper trading", 0, 4.0, kill=0.20, dist="fixed"), Station("canary", 0, 2.0, kill=0.05, dist="fixed"), Station("~research fixes", 0, 2.0, next_to=0), Station("~clarification", 0, 1.5, next_to=1))Listing 28.4. Stations, with and without the rewrite. code/ml/28-the-machine-learning-team/python/ml_team.py The simulator.
while ev: t, kind, i, k = heapq.heappop(ev) if kind == 0: visits[i, k] += 1 if stations[k].servers == 0: start(t, i, k) elif free[k] > 0: free[k] -= 1 start(t, i, k) else: queues[k].append(i) continue if stations[k].servers > 0: if queues[k]: start(t, queues[k].pop(0), k) else: free[k] += 1 s, u = stations[k], rng.random() if u < s.kill: killed[i], exit_t[i] = k, t elif u < s.kill + s.rework: heapq.heappush(ev, (t, 0, i, s.rework_to)) elif s.next_to >= 0: heapq.heappush(ev, (t, 0, i, s.next_to)) elif k + 1 < len(stations) and not stations[k + 1].name.startswith("~"): heapq.heappush(ev, (t, 0, i, k + 1)) else: done[i] = exit_t[i] = t return {"arrival": arrival, "done": done, "exit": exit_t, "killed_at": killed, "visits": visits}Listing 28.5. The event loop: queues, service, and the kill, rework and pass decisions. code/firm/modelpkg/firm_modelpkg.py - Run
ml_team.handoff_table(),pipeline(True),pipeline(False, 1.0),mm1_check()andfig_team.py(about 20 seconds on one core).
What to change next. Let the review committee meet twice as often, and find which station is then the busiest.
28.6 Build: the model package
Purpose. A contract between research and production that a machine can check, and a model of the pipeline it flows through.
Interface. REQUIRED, INPUTS, MONITORS; load(path); validate(pkg, events, times) returning (check, passed, detail) triples, on firm.workflow, firm.featstore, firm.mlinfer and firm.modelcard; Station(name, servers, mean, kill, rework, rework_to, dist, next_to); simulate_pipeline(stations, rate, n, seed).
Rules. Every field of the specification is required; a check that cannot run fails; feature values, not only model outputs, are compared with research’s; the simulator is deterministic given its seed.
Acceptance tests. code/firm/modelpkg/tests/: a complete package passes; a missing field, a wrong hash, an unknown input and a stale test vector each fail their check; one M/M/1 station matches ; a pure delay line returns its fixed time; kill and rework probabilities are respected.
Stretch. Run the validator as a pre-merge job on the registry, with the package’s manifest as the registered artefact; priorities at the busiest station.
Sources and further reading
- S. Amershi and co-authors, “Software engineering for machine learning: a case study”, ICSE Software Engineering in Practice, 2019.
- D. Kreuzberger, N. Kühl and S. Hirschl, “Machine learning operations (MLOps): overview, definition, and architecture”, IEEE Access, 2023.
- M. Zinkevich, Rules of Machine Learning: Best Practices for ML Engineering, Google for Developers.
- J. D. C. Little, “A proof for the queuing formula ”, Operations Research, 1961.
- D. Sculley and co-authors, “Hidden technical debt in machine learning systems”, NeurIPS, 2015.
28.7 Exercises
Exercise 28.1 ★
Name the three kinds of people in a trading firm’s machine-learning work and one thing each owns.
Solution
Solution of Exercise 28.1.
Researchers own the model and its features (the signal, the fit, the validation evidence); machine-learning engineers own the service that runs it (online features, compiled model, monitors); the platform team owns what every model shares (feature store, training, registry, serving, the handoff validator).
Exercise 28.2 ★
Why must the model owner be one named person? What goes wrong when it is a team?
Solution
Solution of Exercise 28.2.
Because someone must answer a performance alarm, decide on retraining or retirement and explain the model to validators and regulators. A team as owner diffuses all three: the page goes to a list, each member assumes another is acting, and the model runs unwatched.
Exercise 28.3 ★
Check Little’s law on the chapter’s single station: arrivals at 0.8 a week, a mean time in the system of 5.05 weeks. What is the mean number of models in the system?
Solution
Solution of Exercise 28.3.
models, against the chain’s 4.00 (theory: with ).
Exercise 28.4 ★★
Why do the specification and parity checks pass a window written in milliseconds, and what check catches it?
Solution
Solution of Exercise 28.4.
A window of 5 000 is a valid definition (a known input, a known aggregation, a positive window), and the offline and online computations both use it, so they agree with each other. Only the comparison of the computed feature values with research’s own values on the reference session shows that the feature is not the one the model was trained on.
Exercise 28.5 ★★
The engineers are busy 74% of the time with the rewrite. Using the M/M/1 formula as a rough guide, by how much does the mean wait change if their utilisation rises to 90%?
Solution
Solution of Exercise 28.5.
In an M/M/1 queue the mean wait before service is service times: 2.8 at 74% and 9 at 90%, a little over three times longer. The chapter’s station has two servers and rework, so the numbers differ, but the shape (a wait that explodes near saturation) is the same, as Figure 28.1 shows.
Exercise 28.6 ★★
Find the flaw. “Our pipeline is slow because validation kills too many models; we should relax the validators.”
Solution
Solution of Exercise 28.6.
Validation kills 7.9% of models; review kills 41%, and the time is lost in queues and rework at the engineers, not in validation’s verdicts. Relaxing validators would let more bad models into paper trading and production without shortening the lead time much; removing the rewrite (the rework loop between engineers and validators) is what shortens it.
Exercise 28.7 ★★★
Coding. Keep the rewrite and give the engineers a third member. Compare the median lead time with the package’s, and say which change you would make first.
Solution
Solution of Exercise 28.7.
ml_team.third_engineer: with three engineers the engineers’ utilisation falls to 50% and the median lead time to 16.6 weeks (means over three seeds), against 22.3 with two engineers and 11.1 with the package, over the same seeds. Headcount removes part of the queueing but none of the rework; the package removes both, and frees engineers’ time instead of adding to it. The package first; a third engineer only if the queue returns as research grows.
Exercise 28.8 ★★★
Write the handoff specification for chapter 27’s monitored linear model: the fields, the test vectors, and the checks.
Solution
Solution of Exercise 28.8.
Artefact: the coefficient vector and intercept, with their content hash. Features: five definitions with units, sources and the standardisation fitted on the 60 training days (means and standard deviations as part of the artefact). Feature vectors: raw inputs and standardised values on reference days. Test vectors: 200 feature rows with predictions to . Latency: trivial, but stated. Monitoring: the five monitors of chapter 27 with their thresholds (0.115, 0.115, 0.089, 0.07 and the CUSUM’s and ), one false page a month, owners. Model card: purpose, target (next-day return), horizon, the trading threshold, limitations (linear, sensitive to units). Checks: as in the chapter, plus a range check on each raw input, which would have caught the Tuesday of chapter 27 before the model saw it.
28.8 Problem: Who Owns the Model?
Problem 28.1
Weekend problem — who owns the model?
The chapter’s handoff package and pipeline.
Part I — Roles.
- Name the roles in a trading firm’s machine-learning work.
- What does the model owner answer for?
- What does the platform own, and what does it not?
- Why is the division between researchers and engineers a design choice?
Part II — The handoff.
- What does the chapter’s package contain?
- Which checks does the validator run?
- Which check catches each planted defect?
- Why must test vectors cover the features?
Part III — The pipeline.
- Describe the simulated pipeline.
- How is the simulator checked?
- Where do models die, and where does the time go?
- What happens as more models arrive?
Part IV — The verdict.
- State the named result: the pipeline’s median lead time, the share of models lost at each gate, and the saving from removing one handoff.
- Why should the losses per gate barely change when the handoff is removed?
- What does the researcher pay for the package, and is it worth it?
- Would a third engineer do as well?
- What should the firm measure every month?
- Who should be the owner of the chapter’s forest?
- What should happen when a package fails a check?
- In one sentence: what makes a handoff cheap?
Solution
Solution of Problem 28.1.
Part I.
- Researchers, machine-learning engineers, platform engineers, the model validator, the trader or portfolio manager, and the model owner (a role held by one of them); Kreuzberger and co-authors list seven roles around MLOps.
- Performance, monitoring, retraining and retirement, and explaining the model to validation, risk and regulators.
- Everything every model shares (stores, training, tracking, registry, serving, monitors, the handoff validator); not the models or their features.
- Because the split determines what passes between teams: a rewrite creates queues and rework, a shared codebase with a checked package does not; models are harder to modularise than software (Amershi and co-authors).
Part II.
- The forest and its hash, 16 feature definitions, research’s feature values on a reference session, 200 test vectors to , a 5-microsecond budget with 2.6 measured, four monitors with owners, a fifteen-field model card.
- Schema, artefact hash, feature specification, feature parity, feature vectors, test vectors, latency, monitoring, model card.
- Table 28.1: rebuilt artefact, hash and test vectors; unknown input and missing feature, the specification; millisecond window, feature vectors; another model’s outputs, test vectors; library latency, latency; ownerless monitor, monitoring; card without limitations, the card.
- A wrong feature can be computed identically offline and online; only a comparison with research’s values catches it.
Part III.
- Poisson arrivals at 0.35 a week; review (kills 35%, 15% back to research), reimplementation by two engineers (30% back for clarification), validation (kills 10%, 25% back to the engineers), four weeks of paper trading (kills 20%), a two-week canary (kills 5%).
- Against the M/M/1 queue: 5.05 weeks simulated against 5, 4.00 models in the system from Book 4’s chain, 4.04 by Little’s law; on the whole pipeline 4.87 against 4.90.
- Most at review (41%); the time goes to rework loops and to the queue at the engineers, busy 74% of the time, each model passing through them 1.79 times.
- With the rewrite the median climbs from about 16 weeks at 0.2 models a week to 70 at 0.44; with the package it stays near 11.
Part IV.
- Who owns the model? With a rewrite at the handoff, the median model reaches production 20.4 weeks after research (90th percentile 42.9), and 35% get there: 41% are lost at review, 5% at reimplementation, 8% at validation, 9% in paper trading and 2% in the canary. A validated package instead of the rewrite cuts the median to 11.1 weeks, a saving of 9.3 weeks.
- Because the gates judge the models, and the models have not changed; only the time spent between gates should.
- A week of research time per model; it buys 9.3 weeks of lead time and cuts the engineers’ load from 74% to 12%.
- No: 16.6 weeks against 11.1, since it shortens the queue but keeps the rework.
- Median and 90th-percentile lead time, losses per gate, utilisation per team, rework rates per handoff.
- The researcher who built it or the head of the desk that trades it, named in the registry, not the platform team.
- It goes back to its author with the failing check’s detail, and nothing reaches production until it passes.
- One codebase for research and production and a package a machine can check.
28.9 Interview questions
Interview question 28.1 ★ mle
What does a machine-learning engineer do that a researcher does not?
Solution
Solution of Interview question 28.1.
Builds and runs the production path: online features equal to offline ones, a model inside its latency budget, deployment, monitoring, retraining pipelines and rollbacks. The researcher finds and validates the signal; the engineer makes it run without its author.
What the interviewer is looking for: production concerns: parity, latency, monitoring, operation.
Interview question 28.2 ★★ mle, developer
What should a research team hand to production with a model?
Solution
Solution of Interview question 28.2.
The artefact and its hash, the feature definitions in the shared format with reference values, test vectors of inputs and outputs, a latency budget and measurement, the monitoring specification, a model card and an owner, all checkable automatically.
What the interviewer is looking for: a checkable package, not a notebook.
Interview question 28.3 ★★ mle
Your features differ between research and production. How do you find out, and how do you stop it happening again?
Solution
Solution of Interview question 28.3.
Compare offline and online values at the same decision times on a sample (parity) and against research’s reference values; find the first diverging feature and input. Prevent it with one feature definition executed by both paths (a feature store), parity checks in continuous integration and in production, and feature test vectors in every handoff.
What the interviewer is looking for: parity testing and a single definition.
Interview question 28.4 ★★ researcher, mle
Models take five months to reach production. How would you find out why?
Solution
Solution of Interview question 28.4.
Record every model’s timestamps at each stage, and measure waiting against working time, rework loops and the utilisation of each team; model the pipeline as queues to test changes. Usually the time is in queues before a saturated team and in rework caused by the handoff.
What the interviewer is looking for: measurement per stage, queues and rework.
Interview question 28.5 ★★ mle
What belongs on a machine-learning platform, and what should stay with the model’s team?
Solution
Solution of Interview question 28.5.
The platform: storage and feature computation, training infrastructure, tracking and registry, compilation and serving, monitoring, handoff checks. The model’s team: the model, its features and labels, its validation evidence, its monitoring thresholds and its ownership.
What the interviewer is looking for: shared infrastructure against model-specific choices.
Interview question 28.6 ★★★ mle
Your platform team is asked to double the number of models it supports next year. What do you measure first?
Solution
Solution of Interview question 28.6.
The utilisation of each team and the rework rates today: doubling arrivals at a team already busy three-quarters of the time makes lead times explode long before it runs out of hours. Then decide what to automate (checks, packaging) before hiring.
What the interviewer is looking for: utilisation and queueing near saturation.