reinforcement learning - Is the reward related to previous state or next state?

In the reinforcement learning framework, I am a little bit confused about the reward and how it is related to states. For example, in Q-learning, we have the following formula for updating the Q table:

that means that the reward is obtained from the environment at the time t+1. I mean that after applying the action a_t, the environment gives s_t+1 and r_t+1.

It is often true that the reward is associated with the previous time step, that is using r_t in the above formula. See, for example the Wikipedia page for Q-learning (https://en.wikipedia.org/wiki/Q-learning). Why is this?

Accidentally, some Wikipedia pages about the same topic but in different languages, use r_t+1 (or unexpectedly R_t+1). See, for example, the Italian and Japanese pages:

https://it.wikipedia.org/wiki/Q-learning
https://ja.wikipedia.org/wiki/Q%E5%AD%A6%E7%BF%92

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

Categories

reinforcement learning - Is the reward related to previous state or next state?

reinforcement learning - Is the reward related to previous state or next state?

Please log in or register to add a comment.

Please log in or register to answer this question.

1 Answer

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags