History

Add warnings for duplicate usage of action-bounded actor and action scaling method (#850 )

- Fix the current bug discussed in #844 in `test_ppo.py`.
- Add warning for `ActorProb ` if both `max_action ` and
`unbounded=True` are used for model initializations.
- Add warning for PGpolicy and DDPGpolicy if they find duplicate usage
of action-bounded actor and action scaling method.

2023-04-23 16:03:31 -07:00

results

Add BranchingDQN for large discrete action spaces (#618 )

2022-05-15 21:40:32 +08:00

acrobot_dualdqn.py

Gymnasium Integration (#789 )

2023-02-03 11:57:27 -08:00

bipedal_bdq.py

Gymnasium Integration (#789 )

2023-02-03 11:57:27 -08:00

bipedal_hardcore_sac.py

Add warnings for duplicate usage of action-bounded actor and action scaling method (#850 )

2023-04-23 16:03:31 -07:00

lunarlander_dqn.py

Gymnasium Integration (#789 )

2023-02-03 11:57:27 -08:00

mcc_sac.py

Add warnings for duplicate usage of action-bounded actor and action scaling method (#850 )

2023-04-23 16:03:31 -07:00

README.md

Add BranchingDQN for large discrete action spaces (#618 )

2022-05-15 21:40:32 +08:00

README.md

Bipedal-Hardcore-SAC

Our default choice: remove the done flag penalty, will soon converge to ~280 reward within 100 epochs (10M env steps, 3~4 hours, see the image below)
If the done penalty is not removed, it converges much slower than before, about 200 epochs (20M env steps) to reach the same performance (~200 reward)

BipedalWalker-BDQ

To demonstrate the cpabilities of the BDQ to scale up to big discrete action spaces, we run it on a discretized version of the BipedalWalker-v3 environment, where the number of possible actions in each dimension is 25, for a total of 25^4 = 390 625 possible actions. A usaual DQN architecture would use 25^4 output neurons for the Q-network, thus scaling exponentially with the number of action space dimensions, while the Branching architecture scales linearly and uses only 25*4 output neurons.