NatureのDQNのCNN層
NatureのDQN:ttps://www.nature.com/articles/nature14236
これのCNN層はこんな構成
Stable baselines の CnnPolicy と多分一緒.
https://elix-tech.github.io/ja/2016/06/29/dqn-ja.html
seen from United States

seen from China

seen from United States
seen from Singapore

seen from Canada
seen from Costa Rica

seen from Finland
seen from Russia
seen from United States
seen from China
seen from Türkiye
seen from Germany
seen from United States
seen from Sweden
seen from United Kingdom

seen from Italy
seen from China

seen from United States
seen from United States
seen from Japan
NatureのDQNのCNN層
NatureのDQN:ttps://www.nature.com/articles/nature14236
これのCNN層はこんな構成
Stable baselines の CnnPolicy と多分一緒.
https://elix-tech.github.io/ja/2016/06/29/dqn-ja.html
Name Refactored [1] Recurrent Box Discrete Multi Processing A2C ✔️ ✔️ ✔️ ✔️ ✔️ ACER ✔️ ✔️ ❌ [4] ✔️ ✔️ ACKTR ✔️ ✔️ ✔️ ✔️ ✔️ DDPG ✔️ ❌ ✔️ ❌ ✔️ [3] DQN ✔️ ❌ ❌ ✔️ ❌ HER ✔️ ❌ ✔️ ✔️ ❌ GAIL [2] ✔️ ✔️ ✔️ ✔️ ✔️ [3] PPO1 ✔️ ❌ ✔️ ✔️ ✔️ [3] PPO2 ✔️ ✔️ ✔️ ✔️ ✔️ SAC ✔️ ❌ ✔️ ❌ ❌ TD3 ✔️ ❌ ✔️ ❌ ❌ TRPO ✔️ ❌ ✔️ ✔ ✔️ [3]
Name Name unomitted Year URL Detail 派生元 A3C Asynchronous Advantage Actor-Critic 2016 A2C Advantage Actor-Critic 2016 A3C ACER Actor-Critic with Experience Replay 2016 A3Cを方策オフ型に書き換えて, 経験再生を利用するアルゴリズムです. A3C ACKTR Actor Critic using Kronecker-Factored Trust Region 2017 PPO?; Actor-CritC,TRPO,Kronecker factorization DDPG Deep DPG 2015 DPG DQN 2013 https://arxiv.org/abs/1312.5602 HER Hindsight Experience Replay 2017 DDPG? GAIL Generative Adversarial Imitation Learning 2016 TRPO? PPO Proximal Policy Optimization 2017 https://openai.com/blog/openai-baselines-ppo/ TRPO SAC Soft Actor-Critic 2018 https://arxiv.org/abs/1801.01290 TD3 TD3 Twin Delayed DDPG 2018 https://arxiv.org/abs/1802.09477 DDPG TRPO 2015 https://arxiv.org/abs/1502.05477 DQN?; VPG?
https://qiita.com/namakemono/items/cfd4db78e2fec738a65c
https://qiita.com/shionhonda/items/ec05aade07b5bea78081