All books

Professional

Apps About Coach Log in Start reading

Quantitative Finance · Glossary

What is Process-based parallelism, serialisation cost?

Also known as: process-based parallelism · serialisation cost

Definition 8.8 Research, Data and Risk Platforms · Chapter 8 — Python at Scale

Process-based parallelism runs work in several operating-system processes, each with its own interpreter and lock, so that Python code runs on several cores at once. Its serialisation cost is the time spent converting data to bytes and back to move them between processes (pickling, copying through a pipe), which shared memory or files mapped by both processes avoid.

Moving a 64 MB array to a second process (pickled through a pipe, or placed in shared memory), and running a pure-Python loop and a numpy sort on one thread and on two (each thread doing the whole task). Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, machine otherwise idle. Data: bench_pyscale.py.
Figure 8.4. Moving a 64 MB array to a second process (pickled through a pipe, or placed in shared memory), and running a pure-Python loop and a numpy sort on one thread and on two (each thread doing the whole task). Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, machine otherwise idle. Data: bench_pyscale.py.
Read in context →