HD-PPO Agent playing Pendulum-v1

This is a trained HD-PPO (Hyperdimensional Proximal Policy Optimization) agent playing Pendulum-v1 using gradient-adaptive Fractional Power Encoding (FPE) with a prune-and-fine-tune pipeline.

Published by LTU-AI.

Pipeline

  1. Train a teacher at D=512 with gradient-adaptive single-beta FPE.
  2. Prune by actor-weight importance through D=512 → 128 → 32.
  3. Fine-tune each pruned checkpoint with PPO.

Published checkpoint: seed 2024, compact D=32 model (held-out eval mean reward -125.19 ± 88.07).

Usage

Install dependencies:

pip install -r requirements.txt

Evaluate the local checkpoint:

python enjoy.py --weights hdppo-Pendulum-v1/weights.npz --episodes 10

Render episodes:

python enjoy.py --weights hdppo-Pendulum-v1/weights.npz --render --episodes 3

Record a replay video:

python record_video.py --weights hdppo-Pendulum-v1/weights.npz --output replay.mp4

Load from Hugging Face Hub:

python enjoy.py --weights LTU-AI/hdppo-Pendulum-v1 --episodes 10

Training pipeline

Reproduce the teacher → prune → fine-tune workflow:

python run_prune_finetune_5seed.py

Hyperparameters

{
    "env": "Pendulum-v1",
    "algo": "HD-PPO (gradient-adaptive FPE, continuous)",
    "teacher_D": 512,
    "pruned_D": 32,
    "beta_base": 2.5,
    "timesteps_per_stage": 1000000,
    "seed": 2024
}

Environment Arguments

{
    "render_mode": "rgb_array"
}

Model files

File Description
hdppo-Pendulum-v1/weights.npz Published actor (+ critic if HD) and FPE encoder (D=32)
hdppo-Pendulum-v1/weights_D512_teacher.npz Teacher checkpoint (D=512)
replay.mp4 Sample rollout video from the published min-D checkpoint
results.json Evaluation summary for the published checkpoint
results_D512_teacher.json Evaluation summary for the teacher
config.yml Training hyperparameters
train_hdppo.py / training modules Self-contained training code

Citation

If you use this model, please cite the HD-PPO / Hybrid-HD-PPO work.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results