2 Comments
User's avatar
Nidhin Gangisetty's avatar

Great article! Given the need for novelty, what are your thoughts on actively training policies designed to target noisy/difficult distributions as supplementary data generators for the main policy (ex: seeking variance-seeking behavior)?

Nano Gregory's avatar

Interesting article...