Reward shaping changes the rewards an agent learns from without changing the policy that is optimal, for instance by removing a term whose expectation is known to be zero, or by adding a potential difference, to make learning faster or more stable. Action masking removes the actions that are not allowed in a state (selling more than is held, quoting a side that would breach an inventory limit) from the agent’s choice and from the maximum in its learning target.
Quantitative Finance · Begrippenlijst
Wat is Reward shaping, action masking?
Ook bekend als: reward shaping · action masking