Off-policy evaluation estimates the value of a target policy from episodes generated by another, the behaviour policy, typically by reweighting each episode by the ratio of the two policies’ probabilities of the actions taken (importance sampling, Book 4, chapter 26). The doubly robust estimator adds to a model’s estimate of the target’s value the importance-weighted errors of the model on the logged rewards; it is unbiased if either the probabilities or the model are right, and has lower variance than importance sampling when the model is close (Dudík, Langford and Li, 2011).
| estimator | bias | standard deviation | root mean square error |
|---|---|---|---|
| per-decision importance sampling | 0.20 | 1.75 | 1.76 |
| weighted importance sampling | 0.49 | 0.63 | 0.80 |
| model alone | 0.30 | 0 | 0.30 |
| doubly robust | 0.29 | 0.29 |
ml_rl.off_policy.