PPO LunarLander-v3

This repository contains a Stable-Baselines3 PPO agent trained to solve the Gymnasium LunarLander environment.

Training setup

  • Algorithm: PPO
  • Environment: LunarLander-v3
  • Policy: MlpPolicy
  • Training timesteps: 1,000,000

Evaluation

The agent was evaluated on the LunarLander environment with deterministic rollout settings.

Notes

This model is intended for experimentation and educational purposes.

Downloads last month
1
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support