Chaper 14 Deterministic policy gradients results are quite noisy.

Question

Chaper 14 Deterministic policy gradients results are quite noisy.

isu10503054a opened this issue 4 years ago · comments

In the results of Chapter 14 Deterministic policy gradients in the book,
why the training is not very stable and noisy?

I read the content repeatedly, but I still don’t understand why.

Max Lapan · Answer 1 · Tue Oct 27 2020 20:34:27 GMT+0800 (China Standard Time)

Random weights initialization adds randomness to initial starting point. Usage if different parallel environments also might add stochastisity вт, 27 окт. 2020 г., 12:01 isu10503054a <notifications@github.com>:

…

In the results of Chapter 14 Deterministic policy gradients in the book, why the training is not very stable and noisy? I read the content repeatedly, but I still don’t understand why. — You are receiving this because you are subscribed to this thread. Reply to this email directly, view it on GitHub <#86>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/AAAQE2WTJOWPGQGYY3MOTRLSM2D5XANCNFSM4TAQL7BQ> .

isu10503054a · Answer 2 · Wed Oct 28 2020 16:41:16 GMT+0800 (China Standard Time)

Random weights initialization adds randomness to initial starting point. Usage if different parallel environments also might add stochastisity вт, 27 окт. 2020 г., 12:01 isu10503054a notifications@github.com:
…
In the results of Chapter 14 Deterministic policy gradients in the book, why the training is not very stable and noisy? I read the content repeatedly, but I still don’t understand why. — You are receiving this because you are subscribed to this thread. Reply to this email directly, view it on GitHub <#86>, or unsubscribe https://github.com/notifications/unsubscribe-auth/AAAQE2WTJOWPGQGYY3MOTRLSM2D5XANCNFSM4TAQL7BQ .

Is there any hyperparameter in the source code that can modification to improve this situation?
thx

Max Lapan · Answer 3 · Wed Oct 28 2020 18:23:47 GMT+0800 (China Standard Time)

Tons of :). In fact any constant in the code could be seen as hyperparameter: * learning rate * gamma * amount of environments * optimisation method etc, etc, etc

…

On Wed, Oct 28, 2020 at 11:41 AM isu10503054a ***@***.***> wrote: Random weights initialization adds randomness to initial starting point. Usage if different parallel environments also might add stochastisity вт, 27 окт. 2020 г., 12:01 isu10503054a ***@***.***: … <#m_7201119268102051534_> In the results of Chapter 14 Deterministic policy gradients in the book, why the training is not very stable and noisy? I read the content repeatedly, but I still don’t understand why. — You are receiving this because you are subscribed to this thread. Reply to this email directly, view it on GitHub <#86 <#86>>, or unsubscribe https://github.com/notifications/unsubscribe-auth/AAAQE2WTJOWPGQGYY3MOTRLSM2D5XANCNFSM4TAQL7BQ . Is there any Hyperparameter in the source code that can modification to improve this situation? thx — You are receiving this because you commented. Reply to this email directly, view it on GitHub <#86 (comment)>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/AAAQE2WD2H3KPPJI7OQAZQLSM7KLXANCNFSM4TAQL7BQ> .

-- wbr, Max Lapan