Randomizing simulator parameters across training episodes so the real world falls inside the range the policy has already seen.
Continue to AI University →