AI Stack Exchange
2022-11-28 10:54 UTC
By Antonis Karvelas
AI-110-20221128-social-media-7be9ec6e
PPO: dealing with variable episodic length
I'm dealing with a project that has episodes of variable length raging from just 3 steps to 20 steps. Now, I'm guessing that this may cause problems with GAE, as actions in large episodes will have much larger advantages than actions in smaller episodes simply because of the cascading addition of future rewards/costs. Is there some smart way of dealing with discounted future returns in such scenarios? Thank you.
I'm dealing with a project that has episodes of variable length raging from just 3 steps to 20 steps. Now, I'm guessing that this may cause problems with GAE, as actions in large episodes will have much larger advantages than actions in smaller episodes simply because of the cascading addition of future rewards/costs. Is there some smart way of dealing with discounted future returns in such scenarios? Thank you.
Full article content could not be extracted automatically. Read the original below.
Source:
AI Stack Exchange
· ai.stackexchange.com