Resources › ai › reinforcement_learning › llm Bookmarks 2026-08-31 bookmark magazine.sebastianraschka.com · 29 min From DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL Updates attention mechanisms 2026-07-16 bookmark lilianweng.github.io · 10 min Why We Think inference 2025-07-22 video youtube.com · 23:16 DeepSeek's GRPO (Group Relative Policy Optimization) | Reinforcement Learning for LLMs policy methods 2025-04-22 bookmark natolambert.substack.com · 3 min "Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?" evaluation 2025-02-20 bookmark arxiv.org · 1h 1m Execution-based Code Generation using Deep Reinforcement Learning 2024-01-20 bookmark arxiv.org · 1 min Self-Rewarding Language Models training 2024-01-10 bookmark huggingface.co · 1 min Paper page - Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models training 2024-01-03 bookmark whatdhack.medium.com · 29 min Some Core Principles of Large Language Model (LLM) Tuning training Subcategories