A kernel whose intensity is below the machine’s ridge point is memory-bound: its rate is limited by bandwidth, and more arithmetic units do not help it. A kernel above the ridge is compute-bound.
Exemples
Example 14.6 (Faster overnight, not intraday)
The same device, the same code style, two verdicts. The exposure job is large and memory-bound, moves little across the link, and runs 140 times faster than on one core (17.6 times for the whole process by Amdahl’s law). The price request is tiny, and the latency of a launch is thirty times its whole cost on the CPU. The break-even batch depends on the latency assumed: 4 trades at 1 microsecond, 16 at 5, 31 at 10, 155 at 50. A pricing service that can batch its requests to hundreds per call would gain; one that answers each request as it comes cannot.