rl

45 projects share this GitHub topic

rl — Llama-Chinese ★14.7krldopamine — ★10.9kART — ★10.5kAReaL — ★5.6kEasyR1 — ★5.1kAlphaZero_Gomoku — ★3.6krl — ★3.5krl-baselines3-zoo — ★2.9kmuzero-general — ★2.8kPapers-in-100-Lines-of-Code — ★2.8kAwesome-RL-for-LRMs — ★2.5kall-rl-algorithms — ★1.9kPRIME — ★1.9kjudgeval — ★1kSDPO — ★1ktensorlake — ★975OpenDerisk — ★958mushroom-rl — ★939zeroth-bot — ★801rl-tutorial-jnrr19 — ★747mobilegym — ★741stable-baselines3-contrib — ★729awesome-monte-carlo-tree-search-papers — ★713Matterport3DSimulator — ★707irl-imitation — ★678RosettaStone — ★670BenchMARL — ★647Awesome-Long-Chain-of-Thought-Reasoning — ★647pytorch-DRL — ★617awesome-on-policy-distillation — ★572meta-agents-research-environments — ★533morl-baselines — ★533DeepThinkVLA — ★527gymfc — ★442Agentic-RAG-R1 — ★428drq — ★423rad — ★415torchtrade — ★411Pytorch-PCGrad — ★403learning-to-communicate-pytorch — ★357RL-Theory-book — ★350dopamine★ 10.9kART★ 10.5kAReaL★ 5.6kEasyR1★ 5.1kAlphaZero_Gomoku★ 3.6krl★ 3.5krl-baselines3-zoo★ 2.9kmuzero-general★ 2.8kPapers-in-100-Lines-of-C…★ 2.8kAwesome-RL-for-LRMs★ 2.5kall-rl-algorithms★ 1.9kPRIME★ 1.9kjudgeval★ 1kSDPO★ 1ktensorlake★ 975OpenDerisk★ 958mushroom-rl★ 939zeroth-bot★ 801rl-tutorial-jnrr19★ 747mobilegym★ 741stable-baselines3-contri…★ 729awesome-monte-carlo-tree…★ 713Matterport3DSimulator★ 707irl-imitation★ 678RosettaStone★ 670BenchMARL★ 647Awesome-Long-Chain-of-Th…★ 647pytorch-DRL★ 617awesome-on-policy-distil…★ 572meta-agents-research-env…★ 533morl-baselines★ 533DeepThinkVLA★ 527gymfc★ 442Agentic-RAG-R1★ 428drq★ 423rad★ 415torchtrade★ 411Pytorch-PCGrad★ 403learning-to-communicate-…★ 357RL-Theory-book★ 350

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
Llama-Chinese
★ 14.7k
dopamine
Dopamine is a research framework for fast prototyping of reinforcement learning algorithms.
★ 10.9k
ART
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents…
★ 10.5k
AReaL
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
★ 5.6k
EasyR1
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1k
AlphaZero_Gomoku
An implementation of the AlphaZero algorithm for Gomoku (also called Gobang or Five in a Row)
★ 3.6k
rl
A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.
★ 3.5k
rl-baselines3-zoo
A training framework for Stable Baselines3 reinforcement learning agents, with hyperparameter optimization…
★ 2.9k
muzero-general
MuZero
★ 2.8k
Papers-in-100-Lines-of-Code
Implementation of papers in 100 lines of code.
★ 2.8k
Awesome-RL-for-LRMs
A Survey of Reinforcement Learning for Large Reasoning Models
★ 2.5k
all-rl-algorithms
Implementation of all RL algorithms in a simpler way
★ 1.9k
PRIME
Scalable RL solution for advanced reasoning of language models
★ 1.9k
judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and…
★ 1k
SDPO
Reinforcement Learning via Self-Distillation (SDPO)
★ 1k
tensorlake
Tensorlake is a serverless runtime for sandboxes and deploying background agentic applications
★ 975
OpenDerisk
AI-Native Risk Intelligence Systems, OpenDeRisk——Your application system risk intelligent manager…
★ 958
mushroom-rl
Python library for Reinforcement Learning.
★ 939
zeroth-bot
3D-printed open-source humanoid robot platform for sim-to-real and RL
★ 801
rl-tutorial-jnrr19
Stable-Baselines tutorial for Journées Nationales de la Recherche en Robotique 2019
★ 747
mobilegym
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research ·…
★ 741
stable-baselines3-contrib
Contrib package for Stable-Baselines3 - Experimental reinforcement learning (RL) code
★ 729
awesome-monte-carlo-tree-search-papers
A curated list of Monte Carlo tree search papers with implementations.
★ 713
Matterport3DSimulator
AI Research Platform for Reinforcement Learning from Real Panoramic Images.
★ 707
irl-imitation
Implementation of Inverse Reinforcement Learning (IRL) algorithms in Python/Tensorflow. Deep MaxEnt, MaxEnt,…
★ 678
RosettaStone
Hearthstone simulator using C++ with some reinforcement learning
★ 670
BenchMARL
BenchMARL is a library for benchmarking Multi-Agent Reinforcement Learning (MARL). BenchMARL allows to…
★ 647
Awesome-Long-Chain-of-Thought-Reasoning
Latest Advances on Long Chain-of-Thought Reasoning
★ 647
pytorch-DRL
PyTorch implementations of various Deep Reinforcement Learning (DRL) algorithms for both single agent and…
★ 617
awesome-on-policy-distillation
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of…
★ 572
meta-agents-research-environments
Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic,…
★ 533
morl-baselines
Multi-Objective Reinforcement Learning algorithms implementations.
★ 533
DeepThinkVLA
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
★ 527
gymfc
A universal flight control tuning framework
★ 442
Agentic-RAG-R1
Agentic RAG R1 Framework via Reinforcement Learning
★ 428
drq
DrQ: Data regularized Q
★ 423
rad
RAD: Reinforcement Learning with Augmented Data
★ 415
torchtrade
Modular reinforcement learning framework for algorithmic trading
★ 411
Pytorch-PCGrad
Pytorch reimplementation for "Gradient Surgery for Multi-Task Learning"
★ 403
learning-to-communicate-pytorch
Learning to Communicate with Deep Multi-Agent Reinforcement Learning in PyTorch
★ 357
RL-Theory-book
Reinforcement learning theory book about foundations of deep RL algorithms with proofs.
★ 350
ReinFlow
[NeurIPS 2025] Flow x RL. "ReinFlow: Fine-tuning Flow Policy with Online Reinforcement Learning". Support…
★ 349
model-based-diffusion
Official implementation for the paper "Model-based Diffusion for Trajectory Optimization". Model-based…
★ 339
VADER
Video Diffusion Alignment via Reward Gradients. We improve a variety of video diffusion models such as…
★ 315
stable-baselines
Mirror of Stable-Baselines: a fork of OpenAI Baselines, implementations of reinforcement learning algorithms
★ 307
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.