Learn Deep Reinforcement Learning from the ground up. With a special case study on RLHF & RLVR for LLM tuning