Domain randomisation trains a policy on many versions of a simulator whose uncertain parameters are drawn at random for each episode, so that the policy works across them instead of exploiting one (Tobin and co-authors, 2017).
ml_rltrade.execution.ml_rltrade.agents.