جميع الكتب

مهني

1 Markets I: The Ecosystem and Exchange-Traded Marketsالأسواق عبر الإنترنت 2 Markets II: Rates, FX and Creditالأسواق عبر الإنترنت 3 Markets III: Commodities, Energy and Cryptoالأسواق عبر الإنترنت 4 Quantitative Methodsالأساليب عبر الإنترنت 5 Derivatives and Volatilityالمشتقات عبر الإنترنت 6 Rates, Credit, XVA and Riskالفائدة والائتمان والمخاطر عبر الإنترنت 7 Research Craft: Predictors, Backtests, Measurement, Portfoliosالبحث عبر الإنترنت 8 Strategies I: Equities and Futuresالاستراتيجيات عبر الإنترنت 9 Strategies II: Volatility, Relative Value, Macro and the Bank Desksالاستراتيجيات عبر الإنترنت 10 Microstructure and Executionالتنفيذ عبر الإنترنت 11 Market Making and High-Frequency Tradingصناعة السوق عبر الإنترنت 12 Machine Learning for Marketsتعلم الآلة عبر الإنترنت 13 Low-Latency Softwareالتكنولوجيا عبر الإنترنت 14 Networks, Hardware and Trading Infrastructureالتكنولوجيا عبر الإنترنت 15 Research, Data and Risk Platformsالتكنولوجيا عبر الإنترنت 16 The Desk and the Firmالشركة عبر الإنترنت 17 The Industry: Firms, Roles and Careersالمسارات المهنية عبر الإنترنت 18 The Interview Bookالمسارات المهنية عبر الإنترنت
التطبيقات حول المدرب تسجيل الدخول ابدأ القراءة

Quantitative Finance · المسرد

ما معنى Off-policy evaluation, doubly robust estimator؟

يُعرف أيضًا باسم: off-policy evaluation · doubly robust estimator

Definition 17.6 Machine Learning for Markets · الفصل 17 — Reinforcement Learning Foundations

Off-policy evaluation estimates the value of a target policy from episodes generated by another, the behaviour policy, typically by reweighting each episode by the ratio of the two policies’ probabilities of the actions taken (importance sampling, Book 4, chapter 26). The doubly robust estimator adds to a model’s estimate of the target’s value the importance-weighted errors of the model on the logged rewards; it is unbiased if either the probabilities or the model are right, and has lower variance than importance sampling when the model is close (Dudík, Langford and Li, 2011).

estimatorbiasstandard deviationroot mean square error
per-decision importance sampling0.201.751.76
weighted importance sampling0.490.630.80
model alone0.3000.30
doubly robust−0.03-0.030.290.29
Table 17.1. Estimates of the target policy’s cost (true value 2.77 bp per lot) from 300 logged episodes of the desk’s behaviour policy, over 100 repetitions (bp per lot). Data: ml_rl.off_policy.
اقرأ في الفصل →