---
title: "Native Extensions"
book: "Research, Data and Risk Platforms"
subject: quant
language: en
chapter: 9
exercises: 8
source: https://one-course.com/books/quant/15/en/chapter/9-native-extensions
---

# Chapter 9 — Native Extensions

A researcher rewrote the inner step of a signal in C++ and called it from Python once per tick, a million times. It was slower than the Python function it replaced: 118 nanoseconds a call against 60. Called once on the whole array of a million values, the same C++ loop took 1.4 nanoseconds an element, more than twice as fast as the library routine it was meant to beat. Everything about a [native extension](#def-pl-native-extensions-extension) is decided at the boundary between the interpreter and the compiled code: how often it is crossed, what has to be converted on the way, and who owns the memory on each side. This chapter binds one kernel written in C++20 and in Rust into Python two ways, with pybind11 and through a C interface loaded with `ctypes`, passes numpy arrays across without copying them, and measures the boundary (`firm.natext`).

## 9.1 Why and when to leave Python

**Definition 9.1 (Native extension, language binding).**

A *native extension* is compiled code (C, C++, Rust) loaded into the Python process and called like a Python function. A *language binding* is the layer that makes it callable: it converts arguments and results between the two languages’ representations, reports errors across the boundary, and states who owns each piece of memory.

Chapter 8 gave the order of preference: first [array programming](https://one-course.com/books/quant/15/en/chapter/8-python-at-scale#def-pl-python-at-scale-array), which is already compiled code written by someone else; then a [native extension](#def-pl-native-extensions-extension) for what arrays cannot express — a loop whose state changes with each element (a book, a queue position, a state machine), a merge of two sorted sequences, a recursion with branches. The chapter’s two kernels are one of each kind: an exponentially weighted average (which numpy can express as a filter, so the native version competes with a library) and an as-of lookup of sorted times by one merge pass (which numpy does with a binary search per element, $O(n \log m)$ instead of $O(n + m)$).

```cpp
namespace firm::natext {

// y_0 = x_0; y_t = (1 - alpha) y_{t-1} + alpha x_t. Writes n outputs to out.
inline void ewma(const double* x, std::size_t n, double alpha, double* out) {
    if (n == 0) return;
    double y = x[0];
    out[0] = y;
    for (std::size_t i = 1; i < n; ++i) {
        y = (1.0 - alpha) * y + alpha * x[i];
        out[i] = y;
    }
}

// One step of the average: the unit the per-call benchmark calls a million times.
inline double ewma_step(double y, double x, double alpha) {
    return (1.0 - alpha) * y + alpha * x;
}

// For each left time, the index of the last right time at or before it (-1 if none).
// Both inputs sorted ascending: one merge pass, O(nl + nr).
inline void asof_index(const std::int64_t* left, std::size_t nl,
                       const std::int64_t* right, std::size_t nr, std::int64_t* out) {
    std::size_t j = 0;
    for (std::size_t i = 0; i < nl; ++i) {
        while (j < nr && right[j] <= left[i]) ++j;
        out[i] = static_cast<std::int64_t>(j) - 1;
    }
}

}  // namespace firm::natext
```

***Listing 9.1.** The two kernels in C++20, free of Python: the average, one step of it, and the as-of lookup by one merge pass. code/firm/natext/cpp/firm_natext.hpp*

## 9.2 The boundary: calls, types and the interpreter lock

**Definition 9.2 (Foreign function interface, application binary interface).**

A *foreign function interface* lets code in one language call functions compiled from another. It relies on an *application binary interface*: the machine-level convention (how arguments are passed in registers and on the stack, the layout of structures, the names of symbols) that the caller and the callee compiled separately both obey; the C convention is the one every language can speak.

**Definition 9.3 (Boundary-crossing cost).**

The *boundary-crossing cost* of a native call is the fixed work done on each call, independently of the size of its data: finding the function, checking and converting each argument from a Python object, releasing and reacquiring the interpreter lock, converting the result back and handling errors.

![Two routes across the boundary. pybind11 generates a binding that checks each argument’s type and memory layout and can release the interpreter lock while C++ runs; ctypes loads any library with a C interface and passes what it is told, so the Python wrapper must check. Both hand the kernel a pointer into the same numpy buffer.](https://one-course.com/images/onecourse/chapters/quant-15/pl-native-extensions/fig-57e953f341d4.svg)

***Figure 9.1.** Two routes across the boundary. pybind11 generates a binding that checks each argument’s type and memory layout and can release the interpreter lock while C++ runs; ctypes loads any library with a C interface and passes what it is told, so the Python wrapper must check. Both hand the kernel a pointer into the same numpy buffer.*

[Figure 9.2](#fig-pl-native-extensions-calls) measures the crossing. A one-step function called a million times from a Python loop costs 60 nanoseconds a call when it is a Python function, 118 through pybind11 and 337 through `ctypes`: each native call converts three Python floats to C doubles and one back, and `ctypes` does it dynamically, by the argument types declared at run time. The arithmetic, one multiplication and one addition, is less than a nanosecond. A native function called per element is slower than the Python it replaces, because the loop around it is still Python.

![The cost of calling a one-step function once per element from a Python loop, a million times: as a Python function, as C++ through pybind11, and as Rust through ctypes. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, one thread, machine otherwise idle. Data: bench_natext.py.](https://one-course.com/images/onecourse/chapters/quant-15/pl-native-extensions/fig-04b443175bca.svg)

***Figure 9.2.** The cost of calling a one-step function once per element from a Python loop, a million times: as a Python function, as C++ through pybind11, and as Rust through `ctypes`. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, one thread, machine otherwise idle. Data: `bench_natext.py`.*

**Proposition 9.4 (Move the loop, not the body).**

If a native call costs $c_0$ per crossing and $c_1$ per element, and the Python it replaces costs $p$ per element, then calling it once per element costs $n(c_0 + c_1)$ and is faster than Python only if $c_0 + c_1 < p$; calling it once on the whole array costs $c_0 + n c_1$, which is faster than Python for every $n > c_0/(p - c_1)$ and approaches $n c_1$.

**Proof.** Sum the costs of the calls in each arrangement and compare with $np$. ∎

With $c_0 \approx 118$ ns and $p \approx 60$ ns, the first arrangement never wins; the second wins beyond a handful of elements. The interpreter lock is the other half of the boundary: while the kernel runs it touches no Python object, so the binding releases the lock ([Listing 9.2](#lst-pl-native-extensions-pybind)) and other Python threads run meanwhile (chapter 8’s sort did the same).

## 9.3 Binding C++ with a header-only binder

pybind11 is a header-only C++ library: the binding is a C++ file that includes it, describes each function’s arguments, and is compiled with the Python headers into a shared library that Python imports as a module. `firm.natext` compiles it on first use with the compiler of the series (`g++ -std=c++20 -O2 -shared -fPIC` with the include paths from `python3-config` and pybind11).

```cpp
using Doubles = py::array_t<double, py::array::c_style>;
using Ints = py::array_t<std::int64_t, py::array::c_style>;

static Doubles ewma(const Doubles& x, double alpha) {
    const auto n = static_cast<std::size_t>(x.size());
    Doubles out(static_cast<py::ssize_t>(n));
    const double* in = x.data();
    double* o = out.mutable_data();
    {
        py::gil_scoped_release unlocked;  // pure C++ from here: other Python threads may run
        firm::natext::ewma(in, n, alpha, o);
    }
    return out;
}
```

***Listing 9.2.** The pybind11 binding of the average: a numpy array taken as a C-contiguous float64 buffer, a result array allocated once, and the interpreter lock released while C++ runs. code/firm/natext/cpp/firm_natext_py.cpp*

**Definition 9.5 (Buffer protocol).**

The *buffer protocol* is Python’s interface by which an object exposes its underlying memory — a pointer, the element type, the shape and the strides — to other objects and to compiled code, so that they can read and write it without copying.

A numpy array exposes its memory this way, and pybind11’s `array_t` reads it. The declaration decides what happens when the array is not what the kernel expects. With `forcecast`, pybind11 silently converts: a float32 array or a strided view is copied into a new contiguous float64 buffer, and a function that was meant to write into the caller’s array writes into the copy instead. The chapter’s binding declares `c_style` without `forcecast` and marks its arguments `noconvert`, so both are refused with an error: a copy should be a decision, not a surprise.

## 9.4 Binding Rust through the C ABI

Rust has a binding generator for Python too; this series uses no external crates, so the chapter takes the route that needs none: a library of `extern "C"` functions compiled as a C-compatible dynamic library, loaded with Python’s `ctypes`.

```rust
/// # Safety
/// `x` must point to `n` readable f64 values and `out` to `n` writable f64 values
/// that do not overlap `x`.
#[no_mangle]
pub unsafe extern "C" fn natext_ewma(x: *const f64, n: usize, alpha: f64, out: *mut f64) {
    if n == 0 {
        return;
    }
    let x = std::slice::from_raw_parts(x, n);
    let out = std::slice::from_raw_parts_mut(out, n);
    ewma(x, alpha, out);
}
```

***Listing 9.3.** A Rust function with a C interface: raw pointers and a length in, a slice built from them under stated conditions, and the safe kernel called on the slices. code/firm/natext/rust/src/lib.rs*

**As of September 2026 — Rust’s Python bindings.**

PyO3, described in its user guide as “Rust bindings for Python, including tools for creating native Python extension modules” (consulted September 2026), generates the equivalent of pybind11’s argument checks and conversions for Rust. It is an external crate, which the series does not build; the C interface of this chapter is what such a binding produces underneath. A `cdylib` is, in the Rust Reference’s words, a dynamic system library used when compiling a library to be loaded from another language.

`ctypes` is the other extreme from pybind11: it calls any C function with the arguments it is told to pass and checks nothing about the memory behind a pointer. The wrapper ([Listing 9.4](#lst-pl-native-extensions-ctypes)) declares each function’s argument types once, checks that each array is C-contiguous and of the right type, and passes the pointers; without the check, a strided view would be read as if it were contiguous, silently.

```python
@functools.lru_cache(maxsize=1)
def rust_library() -> ctypes.CDLL:
    crate = HERE / "rust"
    subprocess.run(["cargo", "build", "--release", "-q"], cwd=crate, check=True)
    lib = ctypes.CDLL(str(crate / "target" / "release" / "libnatext_rs.so"))
    lib.natext_ewma.argtypes = [_D, ctypes.c_size_t, ctypes.c_double, _D]
    lib.natext_ewma.restype = None
    lib.natext_asof.argtypes = [_I, ctypes.c_size_t, _I, ctypes.c_size_t, _I]
    lib.natext_asof.restype = None
    lib.natext_ewma_step.argtypes = [ctypes.c_double, ctypes.c_double, ctypes.c_double]
    lib.natext_ewma_step.restype = ctypes.c_double
    return lib


def _check(a, dtype):
    """ctypes passes a pointer and trusts it: the wrapper must check what the C++ binding
    checks for us."""
    if not isinstance(a, np.ndarray) or a.dtype != dtype or not a.flags.c_contiguous:
        raise TypeError(f"need a C-contiguous numpy array of {np.dtype(dtype)}")
    return a


def ewma_rust(x, alpha: float) -> np.ndarray:
    x = _check(x, np.float64)
    out = np.empty_like(x)
    rust_library().natext_ewma(x.ctypes.data_as(_D), len(x), alpha, out.ctypes.data_as(_D))
    return out
```

***Listing 9.4.** Loading the Rust library with `ctypes`: declared argument types, and a wrapper that checks what `ctypes` does not before passing pointers into numpy buffers. code/firm/natext/firm_natext.py*

## 9.5 Zero-copy across the boundary

**Definition 9.6 (Zero-copy transfer).**

A *zero-copy transfer* of data across a boundary hands the receiver a reference to the sender’s memory (a pointer and a layout) instead of a copy, so that its cost does not depend on the size of the data; it requires the two sides to agree on the layout and on who may write and free the memory, and for how long.

Both bindings are zero-copy for their inputs; `ewma_into` is zero-copy for its output as well, writing into an array the caller allocated (a test checks that the array’s address is unchanged). The ownership rules are what makes it safe: the kernel reads and writes only during the call, keeps no pointer afterwards, and never frees what Python allocated; Python keeps the arrays alive for the call’s duration. Chapter 4’s [memory-mapped files](https://one-course.com/books/quant/15/en/chapter/4-a-tick-store-on-open-formats#def-pl-a-tick-store-on-open-formats-mmap) extend the same idea to disk and to other processes, and Arrow’s layout of chapter 3 to other libraries and languages.

![Time per element of the whole-array average against the array’s length. For short arrays every route is dominated by its fixed cost per call (c_0/n); for long arrays the native loops settle at about 1.3 to 1.9 ns an element and numpy’s filter at 3 to 3.6, while pure Python stays at hundreds. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, one thread, machine otherwise idle. Data: bench_natext.py.](https://one-course.com/images/onecourse/chapters/quant-15/pl-native-extensions/fig-8b7ceaee105a.svg)

***Figure 9.3.** Time per element of the whole-array average against the array’s length. For short arrays every route is dominated by its fixed cost per call ($c_0/n$); for long arrays the native loops settle at about 1.3 to 1.9 ns an element and numpy’s filter at 3 to 3.6, while pure Python stays at hundreds. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, one thread, machine otherwise idle. Data: `bench_natext.py`.*

Figures [9.3](#fig-pl-native-extensions-sweep) and [9.4](#fig-pl-native-extensions-asof) are the whole-array versions. For the average, both native routes beat numpy’s filter at every length measured: the filter routine has a fixed cost of several microseconds, larger than pybind11’s; at a million elements C++ takes 1.4 nanoseconds an element against numpy’s 3.3. For the as-of lookup the algorithm matters more than the language: numpy’s binary search per element costs more per element as the arrays grow, 21 ns at a million, while the merge pass stays at about 6, three times less. The `ctypes` route pays a larger fixed cost (the pointer conversions of the wrapper) and overtakes numpy only from about a thousand elements for the lookup.

![Time per element of the as-of lookup of n sorted times in n others. numpy’s binary search costs more per element as the arrays grow; the native merge pass costs a constant few nanoseconds once the fixed cost is amortised. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, one thread, machine otherwise idle. Data: bench_natext.py.](https://one-course.com/images/onecourse/chapters/quant-15/pl-native-extensions/fig-a43afe680c1c.svg)

***Figure 9.4.** Time per element of the as-of lookup of $n$ sorted times in $n$ others. numpy’s binary search costs more per element as the arrays grow; the native merge pass costs a constant few nanoseconds once the fixed cost is amortised. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, one thread, machine otherwise idle. Data: `bench_natext.py`.*

**Method 9.7 (Writing a native kernel for the research platform).**

(1) Write the kernel in C++ or Rust with no knowledge of Python: pointers, lengths, plain types. (2) Test it alone in its own language. (3) Bind it once per array, never once per element: arrays in, arrays out, the loop inside. (4) Refuse layouts and types the kernel does not read in place, rather than copying silently. (5) Release the interpreter lock around the pure computation. (6) Test the binding against the Python reference on the same data, including empty and edge inputs.

## 9.6 Tutorial: one kernel, four implementations

**Goal.** Build the C++ and Rust kernels, bind them, show that four implementations agree, and measure the boundary. **End state:** Figures [9.2](#fig-pl-native-extensions-calls), [9.3](#fig-pl-native-extensions-sweep) and [9.4](#fig-pl-native-extensions-asof).

1. **Kernels** : read [Listing 9.1](#lst-pl-native-extensions-kernels) ; compile and run `firm_natext_test.cpp` ; run `cargo test` in `code/firm/natext/rust` .
2. **Bind** : `firm_natext.cpp_module()` compiles and imports the pybind11 module; `rust_library()` builds the `cdylib` and declares its functions to `ctypes` .
3. **Agree** : the acceptance tests compare the four averages and the four lookups, on random and edge inputs.
4. **Refuse** : pass a strided view and a float32 array to each binding and read the errors.
5. **Measure** : `bench_natext.py` (per-call cost and the size sweep).

**What to change next.** Add `forcecast` to the binding and watch `ewma_into` write into a copy; call the Rust kernel from two Python threads through `ctypes` (which releases the lock around foreign calls) and measure the speed-up.

## 9.7 Build: native kernels for the research platform

**Purpose.** The pattern and the build machinery every native kernel of the platform follows: the [tick store](https://one-course.com/books/quant/15/en/chapter/4-a-tick-store-on-open-formats#def-pl-a-tick-store-on-open-formats-store)’s readers of chapter 4, the backtest engine’s inner loops (chapter 11) and the payoff compiler’s evaluation (chapter 19) can all be moved across the boundary this way.

**Interface.** `ewma_python`, `ewma_numpy`, `ewma_cpp`, `ewma_cpp_into`, `ewma_rust`; `asof_python`, `asof_numpy`, `asof_cpp`, `asof_rust`; `step_python`, `step_cpp`, `step_rust`; `cpp_module()`, `rust_library()`. C++20 header `firm::natext` (`ewma`, `ewma_step`, `asof_index`); Rust crate `natext_rs` with `natext_ewma`, `natext_asof`, `natext_ewma_step`.

**Rules.** Kernels know nothing of Python; one call per array; inputs read in place or refused; the lock released around the computation; no pointer kept after a call; builds are part of the tests and a missing toolchain fails them.

**Acceptance tests.** `code/firm/natext/tests/`, `cpp/firm_natext_test.cpp`, `rust/`: four averages and four lookups equal, edges included; three steps equal; output written in place; strided and wrongly typed arrays refused by both bindings.

**Stretch.** A kernel over the [tick store](https://one-course.com/books/quant/15/en/chapter/4-a-tick-store-on-open-formats#def-pl-a-tick-store-on-open-formats-store)’s flat records (a structured buffer, chapter 4); a second-order filter in all four languages; the bindings compared under two threads.

Sources and further reading

- pybind11 documentation, *NumPy* (buffer protocol, `array_t` , `c_style` , `forcecast` , `noconvert` ).
- T. Oliphant and C. Banks, *PEP 3118 – Revising the buffer protocol* (Final).
- The Rust Reference, *Linkage* ( `cdylib` ); PyO3 user guide.

## 9.8 Exercises

**Exercise 9.1 ★.**

At 118 ns a call through pybind11 and 60 ns a step in Python, how long does a day of two hundred million steps take each way, called per element? And computed by one call at 1.4 ns an element?

**Solution of Exercise 9.1.**

Per element through pybind11: $118 \times 2\times10^8$ ns $= 23.6$ s; in Python: 12.0 s; one call on the whole array at 1.4 ns: 0.28 s.

**Exercise 9.2 ★.**

Why is the `ctypes` call slower than the pybind11 call, although both end in the same machine instructions?

**Solution of Exercise 9.2.**

`ctypes` converts each argument at run time by the declared types, through generic machinery, and builds a C call frame dynamically; pybind11’s conversions are compiled into the binding for the exact signature. The machine instructions of the kernel are the same; the work around them is not.

**Exercise 9.3 ★.**

What does `forcecast` do with a float32 array passed to a function declared for float64, and why is that dangerous for `ewma_into`?

**Solution of Exercise 9.3.**

It converts the array into a new float64 array and passes the copy. For `ewma_into` the kernel then writes into the copy, which is discarded, and the caller’s array is unchanged: a silent wrong result.

**Exercise 9.4 ★★.**

With [Proposition 9.4](#prop-pl-native-extensions-loop), $c_0 = 118$ ns, $c_1 = 1.4$ ns and $p = 60$ ns, from how many elements is one call on the whole array faster than the Python loop?

**Solution of Exercise 9.4.**

For $n > c_0/(p - c_1) = 118/58.6 \approx 2.0$: from three elements, one call on the whole array beats the Python loop.

**Exercise 9.5 ★★.**

Why does numpy’s as-of lookup cost more per element as the arrays grow, and the merge pass not?

**Solution of Exercise 9.5.**

numpy’s lookup is a binary search for each left time, $O(\log m)$ steps each, and each step touches a different part of a larger array (more cache misses as it grows); the merge walks both sorted arrays once, forward, a constant cost per element with sequential memory access.

**Exercise 9.6 ★★.**

A kernel keeps a pointer to a numpy array between calls to reuse it. Describe how that goes wrong, and what the ownership rule forbids.

**Solution of Exercise 9.6.**

The array can be freed or reallocated by Python after the call (garbage-collected, resized, replaced), leaving the kernel with a dangling pointer: reads of freed memory or writes into someone else’s. The rule: a kernel reads and writes the buffers only during the call and keeps no pointer afterwards; if data must persist, the caller keeps the array alive and passes it again.

**Exercise 9.7 ★★★.**

*Coding.* With the chapter’s functions, compute the as-of lookup of 100 000 random sorted times in 50 000 others by all four implementations, and check that they agree, including a left time before the first right time and one after the last.

**Solution of Exercise 9.7.**

All four agree on every one of the 100 000 indices: $-1$ for the time before the first right time, and 49 999 (the last index) for the time after the last.

**Exercise 9.8 ★★★.**

*Find the flaw.* “Our C++ extension takes a list of Python floats, converts it to a `std::vector<double>`, computes, and returns a new list. It is faster than numpy on our benchmark of ten numbers, so we use it everywhere.”

**Solution of Exercise 9.8.**

Converting a list of Python floats into a vector and a new list back is a copy and an object creation per element, the cost native code was meant to avoid; a benchmark of ten numbers measures the call’s fixed cost, not the per-element cost that matters on real data. Take and return numpy arrays through the [buffer protocol](#def-pl-native-extensions-buffer), and benchmark on realistic lengths.

## 9.9 Problem: The Million Calls

**Problem 9.1.**

Weekend problem — where to cross the boundary

The chapter’s kernels, their four implementations, the two bindings and the measurements of `bench_natext.py`.

**Part I — The crossing.**

1. What does a native call do besides the arithmetic?
2. What are the measured costs per call of the three routes?
3. Why is a per-element native call slower than the Python function it replaces?
4. What does releasing the interpreter lock allow, and when is it safe?
5. What is an [application binary interface](#def-pl-native-extensions-ffi) , and why is the C one the common ground?

**Part II — The two bindings.**

6. What does pybind11 check that `ctypes` does not?
7. What must the `ctypes` wrapper check itself?
8. What happens to a strided view passed to each binding?
9. Why is `forcecast` not used?
10. What would PyO3 add to the Rust route?

**Part III — Arrays.**

11. What are the costs per element of the whole-array average at a million elements, by route?
12. Why do the native routes beat numpy’s filter at every length measured?
13. What are the costs per element of the as-of lookup at a million, by route?
14. From what length does the Rust lookup beat numpy’s?
15. Which of the two kernels gains more from going native, and why?

**Part IV — The verdict.**

16. State the *named result* : the per-call cost of each binding, the lengths above which each native version beats numpy, and the saving of moving the loop, not the body, across the boundary.
17. How much time does a day of two hundred million steps save by moving the loop into C++?
18. When is a [native extension](#def-pl-native-extensions-extension) not worth writing?
19. What does the ownership rule require of the kernel and of the caller?
20. In one sentence: where should the boundary between Python and native code be?

**Solution of Problem 9.1.**

1. It locates the function, checks and converts each argument from a Python object, may release and reacquire the lock, converts the result back and checks for errors.
2. About 60 ns (a Python function), 118 ns (C++ through pybind11) and 337 ns (Rust through `ctypes` ), loop included.
3. The Python loop around it remains, and each call adds conversions that cost more than the step itself.
4. Other Python threads run while the native code computes; it is safe only if the native code touches no Python object while the lock is released.
5. The machine-level calling and layout convention; C’s is simple, stable and spoken by every language’s [foreign function interface](#def-pl-native-extensions-ffi) .
6. The argument’s type, contiguity and dtype, and it converts or refuses accordingly.
7. That each array is a numpy array of the right dtype and C-contiguous, and pass its length.
8. Both refuse it: pybind11 because the argument is `c_style` without `forcecast` and `noconvert` ; the wrapper because it checks contiguity.
9. It would copy silently: an output written into a copy is lost, and a copy of a large array costs what the extension was meant to save.
10. Generated argument checks and conversions for Rust, like pybind11’s for C++, and Rust types for Python objects; underneath, the same C interface.
11. About 1.4 ns (C++), 1.5 ns (Rust) and 3.3 ns (numpy’s filter).
12. The filter routine has a fixed cost of several microseconds per call, larger than the bindings’, and a slower inner loop.
13. About 21 ns (numpy), 6 ns (C++) and 6 ns (Rust).
14. From about a thousand elements.
15. The lookup: its native version has a better algorithm (a merge instead of a binary search per element), not only a faster loop.
16. **Named result.** A call costs about 60 ns as a Python function, 118 ns through pybind11 and 337 ns through `ctypes` ; on whole arrays both native averages beat numpy at every length measured and the Rust lookup from about a thousand elements, and moving the loop instead of the body turns a 24-second day of per-element calls into 0.3 s.
17. About 12 s against the Python loop (12.0 s to 0.28 s), 23 s against per-element native calls, and 0.4 s against numpy.
18. When the computation is already an array operation of a library, when arrays are short and calls are few, or when the gain does not pay for building, testing and distributing compiled code on every platform the researchers use.
19. The kernel keeps no pointer after the call and frees nothing it did not allocate; the caller keeps its arrays alive and unchanged during the call.
20. Around whole arrays: data cross once, loops stay native.

## 9.10 Interview questions

**Interview question 9.1 ★ developer, researcher.**

You wrote a function in C++ to speed up a Python loop, and it got slower. What happened?

**Solution of Interview question 9.1.**

It was called once per element from the Python loop: each call pays a conversion and dispatch cost larger than the arithmetic, and the loop stayed in Python. Move the loop into the native function and pass whole arrays.

*What the interviewer is looking for: Per-call cost versus per-element cost.*

**Interview question 9.2 ★★ developer.**

How do you pass a numpy array to C++ or Rust without copying it, and what can go wrong?

**Solution of Interview question 9.2.**

Through the [buffer protocol](#def-pl-native-extensions-buffer): pass the array’s pointer, dtype and shape (pybind11’s `array_t`, or `ctypes` pointers from `arr.ctypes`). Wrong dtype or strides (read as if contiguous), silent copies (forcecast), writing to a read-only array, and pointers kept after the call.

*What the interviewer is looking for: Layout checks and ownership.*

**Interview question 9.3 ★★ developer.**

Compare pybind11, ctypes and a Rust binding generator for exposing a numerical kernel to Python.

**Solution of Interview question 9.3.**

pybind11: compiled C++ bindings with type checks, conversions and lock handling, the norm for C++. `ctypes`: no compilation of a binding, works with any C interface, but checks nothing and costs more per call. A Rust generator (PyO3): pybind11’s comfort for Rust; the C interface with `ctypes` works without it.

*What the interviewer is looking for: Safety and per-call cost against convenience and build complexity.*

**Interview question 9.4 ★★ developer.**

Why should a [native extension](#def-pl-native-extensions-extension) release the interpreter lock, and what must it not do while it is released?

**Solution of Interview question 9.4.**

So that other Python threads run while it computes. It must not touch any Python object, including reference counts, raise Python exceptions, or call back into Python until it has reacquired the lock.

*What the interviewer is looking for: No Python API without the lock.*

**Interview question 9.5 ★★★ developer, researcher.**

Design the native layer of a research platform: which kernels go native, how they are bound, built, tested and distributed.

**Solution of Interview question 9.5.**

Kernels chosen by profiling (loops that arrays cannot express, merges, state machines); written in C++ or Rust against plain pointers, tested alone; bound once per array with pybind11 (or a C interface), refusing layouts they cannot read in place and releasing the lock; built in the continuous-integration pipeline for each supported platform and Python version and distributed as wheels; tested against a Python reference on the same data.

*What the interviewer is looking for: Selection by measurement, separation of kernel and binding, reproducible builds.*
