PPO is commonly referred to as an on-policy algorithm. We argue that this is confusing, and show that truly on-policy PPO reduces to vanilla policy gradient REINFORCE with baseline.