جميع الكتب

مهني

1 Markets I: The Ecosystem and Exchange-Traded Marketsالأسواق عبر الإنترنت 2 Markets II: Rates, FX and Creditالأسواق عبر الإنترنت 3 Markets III: Commodities, Energy and Cryptoالأسواق عبر الإنترنت 4 Quantitative Methodsالأساليب عبر الإنترنت 5 Derivatives and Volatilityالمشتقات عبر الإنترنت 6 Rates, Credit, XVA and Riskالفائدة والائتمان والمخاطر عبر الإنترنت 7 Research Craft: Predictors, Backtests, Measurement, Portfoliosالبحث عبر الإنترنت 8 Strategies I: Equities and Futuresالاستراتيجيات عبر الإنترنت 9 Strategies II: Volatility, Relative Value, Macro and the Bank Desksالاستراتيجيات عبر الإنترنت 10 Microstructure and Executionالتنفيذ عبر الإنترنت 11 Market Making and High-Frequency Tradingصناعة السوق عبر الإنترنت 12 Machine Learning for Marketsتعلم الآلة عبر الإنترنت 13 Low-Latency Softwareالتكنولوجيا عبر الإنترنت 14 Networks, Hardware and Trading Infrastructureالتكنولوجيا عبر الإنترنت 15 Research, Data and Risk Platformsالتكنولوجيا عبر الإنترنت 16 The Desk and the Firmالشركة عبر الإنترنت 17 The Industry: Firms, Roles and Careersالمسارات المهنية عبر الإنترنت 18 The Interview Bookالمسارات المهنية عبر الإنترنت
التطبيقات حول المدرب تسجيل الدخول ابدأ القراءة

Quantitative Finance · المسرد

ما معنى Model compilation؟

Definition 26.2 Machine Learning for Markets · الفصل 26 — Low-Latency Inference

Model compilation translates a trained model into code or data structures specialised for inference on a target (generated source, a flat array layout, a hardware description), removing the training framework from the prediction path (Asadi, Lin and de Vries, 2014; Lucchese and co-authors, 2015, for ranking forests).

Single-call latency of the same two models by execution path; 200 000 calls in C++ and Rust, 5 000 in Python (300 for the pure-Python forest), timed one by one. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2 at low priority, no isolated cores; other jobs may have run. Data: bench_infer.py, measured_latency.csv.
Figure 26.1. Single-call latency of the same two models by execution path; 200 000 calls in C++ and Rust, 5 000 in Python (300 for the pure-Python forest), timed one by one. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2 at low priority, no isolated cores; other jobs may have run. Data: bench_infer.py, measured_latency.csv.
اقرأ في الفصل →