Training with Curriculum Learning

This example script demonstrates how to use curriculum learning (CL) and domain randomization (DR) during training with RLlib.

Three different examples of curriculums are shown, as well as the DR case. This example is part of the paper Improving Robustenss of Autonomous Spacecraft Scheduling Using Curriculum Learning and of a future publication. In CL, a sequence of different tasks with increasing difficulty are presented to the agent during training. Each task is seen as a different Markov decision process (MDP). For this problem, each task is characterized by a satellite with different battery capacity and exposed to different external torques, which would lead to different transition probabilities in the MDP.

Load Modules

[1]:
import time
from collections.abc import Callable
from typing import Any, ClassVar, TypeVar

import numpy as np
import ray
from Basilisk.architecture import bskLogging
from ray import tune
from ray.rllib.algorithms.ppo import PPOConfig
from ray.tune.registry import register_env

from bsk_rl import act, data, obs, sats, scene
from bsk_rl.gym import SatelliteTasking
from bsk_rl.sats import Satellite
from bsk_rl.sim import dyn, fsw
from bsk_rl.utils.rllib.callbacks import EpisodeDataWrapper, WrappedEpisodeDataCallbacks

bskLogging.setDefaultLogLevel(bskLogging.BSK_WARNING)

SatObs = TypeVar("SatObs")
MultiSatObs = tuple[SatObs, ...]
SatArgRandomizer = Callable[[list[Satellite]], dict[Satellite, dict[str, Any]]]

Creating an Environment with CL

In this example, the SatelliteTasking environment is modified to allow changes to the spacecraft parameters during training. Two extra method set_task and get_task are introduced to set the difficulty and get the difficulty of the environment. Additionally, update_sat_params is used to change specific spacecraft arguments as a function of the difficulty and is called before each environment reset.

[2]:
class SatelliteTaskingCL(SatelliteTasking):
    def __init__(
        self,
        satellite: Satellite,
        *args,
        difficulty=0.0,
        CL_params=None,
        **kwargs,
    ):
        super().__init__(
            satellite,
            *args,
            **kwargs,
        )

        self.difficulty = difficulty
        self.CL_params = CL_params or {}

    def reset(
        self,
        seed: int | None = None,
        options=None,
    ) -> tuple[MultiSatObs, dict[str, Any]]:
        self.update_sat_params()  # Update satellite parameters based on difficulty before resetting
        obs, info = super().reset(seed=seed, options=options)
        return obs, info

    def update_sat_params(self):
        """
        Update the satellite parameters based on the difficulty level.
        """
        if self.CL_params is not None:
            for satellite in self.satellites:
                for key, value in self.CL_params.items():
                    if key in satellite.sat_args_generator:
                        satellite.sat_args_generator[key] = value(self.difficulty)
                    else:
                        setattr(self, key, round(value(self.get_task())))

    def set_task(self, task):
        """
        Set the difficulty level.
        """
        self.difficulty = task

    def get_task(self):
        """
        Get the current difficulty level.
        """
        return self.difficulty

Registering the Custom Environment

Since a custom environment was created, it needs to be registered and made it compatible with RLlib.

[3]:
def _satellite_tasking_env_creator(env_config):
    """
    Create an environment compatible with RLlib.
    """

    if "episode_data_callback" in env_config:
        episode_data_callback = env_config.pop("episode_data_callback")
    else:
        episode_data_callback = None
    if "satellite_data_callback" in env_config:
        satellite_data_callback = env_config.pop("satellite_data_callback")
    else:
        satellite_data_callback = None

    return EpisodeDataWrapper(
        SatelliteTaskingCL(**env_config),
        episode_data_callback=episode_data_callback,
        satellite_data_callback=satellite_data_callback,
    )


register_env("SatelliteTaskingCL-RLlib", _satellite_tasking_env_creator)

Creating the Scanning Satellite

A nadir scanning satellite is created with personalized observation space properties, including the angle between the solar panels and the sun, and the angle between the instrument and nadir. A custom dynamics module is introduced to combine “GroundStationDynModel” and “ContinuousImagingDynModel”, allowing for scanning and downlink actions.

[4]:
def attitude_error_norm(sat) -> float:
    # Calculate this using the instrument unit vector and the spacecraft position
    # in inertial frame (-r_BN_P) and c_hat_P (get angle between then)
    r_BN_P_unit = sat.dynamics.r_BN_P / np.linalg.norm(sat.dynamics.r_BN_P)
    c_hat_P = sat.dynamics.satellite.fsw.c_hat_P  # Instrument unit vector in ECEF frame
    error_angle = np.arccos(np.dot(-r_BN_P_unit, c_hat_P))

    return error_angle / np.pi


def solar_angle_norm(sat) -> float:
    a = (
        sat.dynamics.world.gravFactory.spiceObject.planetStateOutMsgs[
            sat.dynamics.world.sun_index
        ]
        .read()
        .PositionVector
    )
    a_hat = a / np.linalg.norm(a)
    b = np.array([0, 0, -1])  # Solar panel opposite to instrument
    mat = np.transpose(sat.dynamics.BN)
    b_N = np.matmul(mat, b)
    error_angle = np.arccos(np.dot(b_N, a_hat))

    return error_angle / np.pi


class CustomDynamics(dyn.GroundStationDynModel, dyn.ContinuousImagingDynModel):
    pass


class ScanningSatellite(sats.AccessSatellite):
    observation_spec: ClassVar[list[obs.Observation]] = [
        obs.SatProperties(
            dict(prop="wheel_speeds_fraction"),
            dict(prop="battery_charge_fraction"),
            dict(prop="storage_level_fraction"),
            dict(prop="attitude_error_norm", fn=attitude_error_norm),
            dict(prop="solar_angle_norm", fn=solar_angle_norm),
        ),
        obs.Eclipse(norm=5700.0),
        obs.OpportunityProperties(
            dict(prop="opportunity_open", norm=5700.0),
            dict(prop="opportunity_close", norm=5700.0),
            type="ground_station",
            n_ahead_observe=1,
        ),
    ]
    action_spec: ClassVar[list[act.Action]] = [
        act.Scan(duration=180.0),  # Scan for 3 minute
        act.Charge(duration=180.0),  # Charge for 3 minutes
        act.Downlink(duration=180.0),  # Downlink for 3 minute
        act.Desat(duration=180.0),  # Desaturate for 3 minute
    ]
    dyn_type = CustomDynamics
    fsw_type = fsw.ContinuousImagingFSWModel

Defining Curriculum Function

The following functions are used to define how the satellite parameters vary as a function of the difficulty during training. For these cases, the difficulty is assumed to be between 0 and 1. Direct, inverse, and constant curriculums can be defined based on the initial and final levels.

[5]:
def capacity_fn(time_seed, init_val, final_val, difficulty):
    """
    Function to calculate the capacity of the a given satellite property (e.g. battery, storage, etc) based on the difficulty level.

    Args:
        time_seed (float, optional): Seed for random number generation. If None, CL will be used. Otherwise, DR will be used.
        init_val (float): Initial value of the capacity.
        final_val (float): Final value of the capacity.
        difficulty (float): Difficulty level.

    Returns:
        float: Capacity of the satellite.
    """

    if time_seed is not None:
        random_generator = np.random.default_rng(
            seed=int(time_seed * 100) * int(difficulty * 10000)
        )
        return random_generator.uniform(init_val, final_val)
    else:
        return init_val - (init_val - final_val) * difficulty


def capacity_init_fn(time_seed, init_val, final_val, difficulty, max_init, min_init):
    """
    Function to calculate the initial capacity of the a given satellite property (e.g. battery, storage, etc) based on the difficulty level.
    This function is necessary since the capacity is not constant and can change based on the difficulty level.

    Args:
        time_seed (float, optional): Seed for random number generation. If None, CL will be used. Otherwise, DR will be used.
        init_val (float): Initial value of the capacity.
        final_val (float): Final value of the capacity.
        difficulty (float): Difficulty level.
        max_init (float): Maximum initial value of the capacity.
        min_init (float): Minimum initial value of the capacity.
    Returns:
        float: Initial level of the given satellite property.
    """

    if time_seed is not None:
        random_generator = np.random.default_rng(
            seed=int(time_seed * 100) * int(difficulty * 10000)
        )
        capacity = random_generator.uniform(init_val, final_val)
        return np.random.uniform(min_init, max_init) * capacity
    else:
        capacity = init_val - (init_val - final_val) * difficulty
        return np.random.uniform(min_init, max_init) * capacity


def random_disturbance_vector(magnitude_disturbance, seed=None):
    """
    Function to generate a random disturbance vector with a given magnitude.

    Args:
        magnitude_disturbance (float): Magnitude of the disturbance vector.
        seed (int, optional): Seed for random number generation. Defaults to None.
    Returns:
        np.ndarray: Random disturbance vector with the given magnitude.
    """

    disturbance_rand_vector = np.random.normal(size=3)
    disturbance_rand_unit_vector = disturbance_rand_vector / np.linalg.norm(
        disturbance_rand_vector
    )
    disturbance_vector = disturbance_rand_unit_vector * magnitude_disturbance
    return disturbance_vector


def external_disturbance_fn(time_seed, init_val, final_val, difficulty):
    """
    Function to calculate the external disturbance vector based on the difficulty level.

    Args:
        time_seed (float, optional): Seed for random number generation. If None, CL will be used. Otherwise, DR will be used.
        init_val (float): Initial value of the disturbance vector.
        final_val (float): Final value of the disturbance vector.
        difficulty (float): Difficulty level.
    Returns:
        np.ndarray: External disturbance vector.
    """

    if time_seed is not None:
        random_generator = np.random.default_rng(
            seed=int(time_seed * 100) * int(difficulty * 10000)
        )
        disturbance_mag = random_generator.uniform(init_val, final_val)
        return random_disturbance_vector(disturbance_mag)
    else:
        disturbance_mag = init_val - (init_val - final_val) * difficulty
        return random_disturbance_vector(disturbance_mag)

Custom Callback to Enable CL

A custom Callback function is required to enable CL. The CLCallbacks reads the number of trained steps from the environment and determines the task (difficulty). Here, different functions could be used to implement more complex curriculums instead of a linear function, such as spring mass dynamics.

A custom episode_data_callback is also defined to collect information about the agent and the curriculum during training.

[6]:
class CLCallbacks(WrappedEpisodeDataCallbacks):
    def on_episode_start(
        self,
        *,
        episode,
        worker=None,
        env_runner=None,
        metrics_logger=None,
        base_env=None,
        env=None,
        policies=None,
        rl_module=None,
        env_index,
        **kwargs,
    ) -> None:
        try:
            n_steps = metrics_logger.peek("num_env_steps_sampled_lifetime")
            if n_steps is None:
                task = 0.0
            else:
                task = n_steps / 5_000_000  # 5M steps = 1.0 difficulty
        except KeyError:
            task = 0.0

        env.envs[env_index].unwrapped.set_task(task)


def episode_data_callback(env):
    reward = env.rewarder.cum_reward
    reward = sum(reward.values()) / len(reward)
    orbits = env.simulator.sim_time / (95 * 60)

    data_log = dict(
        reward=reward,
        # Are satellites dying, and how and when?
        alive=float(env.satellites[0].is_alive()),
        rw_status_valid=float(env.satellites[0].dynamics.rw_speeds_valid()),
        battery_status_valid=float(env.satellites[0].dynamics.battery_valid()),
        orbits_complete=orbits,
        # Is CL working? How is it varying during training?
        difficulty=env.get_task(),
        battery_capacity=env.satellites[0].dynamics.powerMonitor.storageCapacity,
        external_torque=np.linalg.norm(
            env.satellites[0].dynamics.extForceTorqueObject.extTorquePntB_B
        ),
    )
    if orbits > 0:
        data_log["reward_per_orbit"] = reward / orbits
    if not env.satellites[0].is_alive():
        data_log["orbits_complete_partial_only"] = orbits

    return data_log

Defining Satellite, Environment, and CL Options

Two different environment configurations are defined, the standard_90 and degraded_90, which can be used for training and testing. Additionally, different initialization ranges can be defined for the parameters during reset. Here, nominal corresponds to parameters being initialized in a range near their nominal operation values. In wide, parameters can vary from 0% to 100%.

Different CL and DR levels are also defined to be chosen from. Each case can include several different parameters from the spacecraft, each with different CL levels.

[7]:
sat_config = dict(
    standard_90=dict(
        # Nominal env parameters
        intervals=90,
        batteryStorageCapacity=400 * 3600,  # in Ws
        disturbance_vector_mag=0.0002,
        panelEfficiency=0.2,
    ),
    degraded_90=dict(
        # Degraded env parameters
        intervals=90,
        batteryStorageCapacity=400 * 3600 * 0.5,  # in Ws
        disturbance_vector_mag=0.0002 * 3,
        panelEfficiency=0.2 * 0.75,
    ),
    # Other sat parameters common to all
    sat_params=dict(
        imageAttErrorRequirement=0.1,  # norm of MRP ~ 20 degree
        imageRateErrorRequirement=0.1,  # norm of angular velocity (rad/s)
        dataStorageCapacity=5000 * 8e6,  # in bits
        instrumentPowerDraw=-30.0,  # in Watts
        instrumentBaudRate=0.5e6,  # bits per second
        transmitterPowerDraw=-25.0,  # in Watts
        transmitterBaudRate=-112.0e6,  # bits per second #size it to downlink in one downlink opportunity
        rwMechToElecEfficiency=0.0,
        rwElecToMechEfficiency=0.5,
        thrusterPowerDraw=-80.0,
        rwBasePower=10.0,
        maxWheelSpeed=6000,  # RPM
        desatAttitude="nadir",
        K=3.5,  # Derivative control gain (attitude)
        Ki=-1,  # Integral gain (turned off)
        P=17.5,  # Proportional gain (attitude))
    ),
)

init_range_options = dict(
    nominal=dict(
        battery_init_range=[0.375, 0.625],
        data_storage_init_range=[0, 1],
        reaction_wheel_init_range=[-4000, 4000],  # RPM
    ),
    wide=dict(
        battery_init_range=[0, 1],
        data_storage_init_range=[0, 1],
        reaction_wheel_init_range=[-6000, 6000],  # RPM
    ),
)

CL_options = dict(
    constant_BT_high=dict(
        battery={
            "name": "batteryStorageCapacity",
            "init_val": 0.40,
            "final_val": 0.40,
            "init_range_config": "battery_init_range",
            "name_init": "storedCharge_Init",
            "domain_randomization": False,
        },
        torque={
            "name": "disturbance_vector_mag",
            "var_name": "disturbance_vector",
            "init_val": 8.0,
            "final_val": 8.0,
            "domain_randomization": False,
        },
    ),
    direct_BT_high=dict(
        battery={
            "name": "batteryStorageCapacity",
            "init_val": 1.00,
            "final_val": 0.40,
            "init_range_config": "battery_init_range",
            "name_init": "storedCharge_Init",
            "domain_randomization": False,
        },
        torque={
            "name": "disturbance_vector_mag",
            "var_name": "disturbance_vector",
            "init_val": 1.0,
            "final_val": 8.0,
            "domain_randomization": False,
        },
    ),
    inverse_BT_high=dict(
        battery={
            "name": "batteryStorageCapacity",
            "init_val": 0.40,
            "final_val": 1.0,
            "init_range_config": "battery_init_range",
            "name_init": "storedCharge_Init",
            "domain_randomization": False,
        },
        torque={
            "name": "disturbance_vector_mag",
            "var_name": "disturbance_vector",
            "init_val": 8.0,
            "final_val": 1.0,
            "domain_randomization": False,
        },
    ),
    DR_BT_high=dict(
        battery={
            "name": "batteryStorageCapacity",
            "init_val": 0.40,
            "final_val": 1.00,
            "init_range_config": "battery_init_range",
            "name_init": "storedCharge_Init",
            "domain_randomization": True,
        },
        torque={
            "name": "disturbance_vector_mag",
            "var_name": "disturbance_vector",
            "init_val": 1.0,
            "final_val": 8.0,
            "domain_randomization": True,
        },
    ),
)

Choosing Curriculum for Training

Here, the direct_BT_high is selected with a nominal initialization range and standard environment with each episode lasting at most 90 steps.

[8]:
CL_params = {}
CL_enabled = True
CL_case = "direct_BT_high"
initialization_range = "nominal"
environment_mode = "standard_90"

sat = ScanningSatellite(
    "Scanner_1",
    sat_args=dict(
        **sat_config["sat_params"],
        batteryStorageCapacity=sat_config[environment_mode]["batteryStorageCapacity"],
        disturbance_vector=lambda: random_disturbance_vector(
            sat_config[environment_mode]["disturbance_vector_mag"]
        ),
        panelEfficiency=sat_config[environment_mode]["panelEfficiency"],
    ),
)

duration = (
    sat_config[environment_mode]["intervals"] * 180
)  # intervals of 180 seconds (3 minutes)

Assigning Curriculum Functions

After selecting the curriculum, the code below will populate the CL_params dictionary with functions, specifying how each of the parameters will vary during training.

[9]:
if CL_enabled:
    for key in CL_options[CL_case]:
        current_time = time.time()

        if CL_options[CL_case][key]["domain_randomization"] is False:
            current_time = None
        else:
            current_time = time.time()

        if key == "torque":
            capacity = sat_config[environment_mode][CL_options[CL_case][key]["name"]]
            init_val = CL_options[CL_case][key]["init_val"]
            final_val = CL_options[CL_case][key]["final_val"]
            CL_params[CL_options[CL_case][key]["var_name"]] = (
                lambda difficulty,
                capacity=capacity,
                init_val=init_val,
                final_val=final_val,
                time_seed=current_time: external_disturbance_fn(
                    time_seed,
                    capacity * init_val,
                    capacity * final_val,
                    difficulty,
                )
            )

        else:
            capacity = sat_config[environment_mode][CL_options[CL_case][key]["name"]]
            init_val = CL_options[CL_case][key]["init_val"]
            final_val = CL_options[CL_case][key]["final_val"]
            if "var_name" in CL_options[CL_case][key]:
                temp_name = CL_options[CL_case][key]["var_name"]
            else:
                temp_name = CL_options[CL_case][key]["name"]
            CL_params[temp_name] = (
                lambda difficulty,
                capacity=capacity,
                init_val=init_val,
                final_val=final_val,
                time_seed=current_time: capacity_fn(
                    time_seed,
                    capacity * init_val,
                    capacity * final_val,
                    difficulty,
                )
            )
            if "name_init" in CL_options[CL_case][key]:
                init_range = init_range_options[initialization_range][
                    CL_options[CL_case][key]["init_range_config"]
                ]
                init_val = CL_options[CL_case][key]["init_val"]
                final_val = CL_options[CL_case][key]["final_val"]
                CL_params[CL_options[CL_case][key]["name_init"]] = (
                    lambda difficulty,
                    capacity=capacity,
                    init_val=init_val,
                    final_val=final_val,
                    init_range=init_range,
                    time_seed=current_time: capacity_init_fn(
                        time_seed,
                        capacity * init_val,
                        capacity * final_val,
                        difficulty,
                        init_range[1],
                        init_range[0],
                    )
                )

Training

Training is performed using ray tune. Usually, the num_env_steps_sampled_lifetime should be set similar to the number of training steps in CLCallbacks. Originally, the paper Improving Robustenss of Autonomous Spacecraft Scheduling Using Curriculum Learning used the APPO algorithm with generalized advantage estimation instead of PPO.

[10]:
N_CPUS = 3

env_args = dict(
    satellite=sat,
    scenario=scene.UniformNadirScanning(value_per_second=1 / duration),
    rewarder=data.ScanningTimeReward(),
    time_limit=duration,
    failure_penalty=-1.0,
    difficulty=0.0,
    CL_params=CL_params,
)

training_args = dict(
    lr=0.00003,
    gamma=0.999,
    train_batch_size=250,  # originally 10,000
    num_sgd_iter=50,
    model=dict(fcnet_hiddens=[512, 512], vf_share_layers=False),
    lambda_=0.95,
    use_kl_loss=False,
    entropy_coeff=0.0,
    clip_param=0.2,
    grad_clip=0.5,
)

config = (
    PPOConfig()
    .training(**training_args)
    .env_runners(num_env_runners=N_CPUS - 1, sample_timeout_s=1000.0)
    .environment(
        env="SatelliteTaskingCL-RLlib",
        env_config=dict(**env_args, episode_data_callback=episode_data_callback),
    )
    .reporting(
        metrics_num_episodes_for_smoothing=1,
        metrics_episode_collection_timeout_s=180,
    )
    .checkpointing(export_native_model_files=True)
    .framework(framework="torch")
    .api_stack(
        enable_rl_module_and_learner=True,
        enable_env_runner_and_connector_v2=True,
    )
    .callbacks(CLCallbacks)
    # .evaluation(evaluation_interval=10, evaluation_duration=1, evaluation_parallel_to_training=True, evaluation_config={"env": unpack_config(env_class), "env_config": nominal_env_args, "explore":False}, evaluation_num_workers=1, always_attach_evaluation_results=True) #An evaluation environment can be configured with parameters different from the training environment by specifying the `nominal_env_args` argument. This is useful for evaluating the performance of the agent in a different environment than the one it was trained in.
)

ray.init(
    ignore_reinit_error=True,
    num_cpus=N_CPUS,
    object_store_memory=2_000_000_000,  # 2 GB
)

# Run the training
results = tune.run(
    "PPO",
    config=config.to_dict(),
    stop={
        "num_env_steps_sampled_lifetime": 750
    },  # Total number of steps to train the model. Originally 5M
    checkpoint_freq=10,
    checkpoint_at_end=True,
)

# Shutdown Ray
ray.shutdown()
2026-10-09 20:35:07,190 INFO worker.py:1783 -- Started a local Ray instance.
2026-10-09 20:35:09,748 INFO tune.py:616 -- [output] This uses the legacy output and progress reporter, as Jupyter notebooks are not supported by the new engine, yet. For more information, please see https://github.com/ray-project/ray/issues/36949
/opt/hostedtoolcache/Python/3.11.17/x64/lib/python3.11/site-packages/gymnasium/spaces/box.py:130: UserWarning: WARN: Box bound precision lowered by casting to float32
  gym.logger.warn(f"Box bound precision lowered by casting to {self.dtype}")
/opt/hostedtoolcache/Python/3.11.17/x64/lib/python3.11/site-packages/gymnasium/utils/passive_env_checker.py:164: UserWarning: WARN: The obs returned by the `reset()` method was expecting numpy array dtype to be float32, actual type: float64
  logger.warn(
/opt/hostedtoolcache/Python/3.11.17/x64/lib/python3.11/site-packages/gymnasium/utils/passive_env_checker.py:188: UserWarning: WARN: The obs returned by the `reset()` method is not within the observation space.
  logger.warn(f"{pre} is not within the observation space.")

Tune Status

Current time:2026-10-09 20:35:25
Running for: 00:00:16.17
Memory: 4.8/15.6 GiB

System Info

Using FIFO scheduling algorithm.
Logical resource usage: 3.0/3 CPUs, 0/0 GPUs

Trial Status

Trial name status loc iter total time (s) num_env_steps_sample d_lifetime num_episodes_lifetim e num_env_steps_traine d_lifetime
PPO_SatelliteTaskingCL-RLlib_ebb51_00000TERMINATED10.1.0.4:3934 3 3.747437508750
(PPO pid=3934) Install gputil for GPU system monitoring.

Trial Progress

Trial name env_runners fault_tolerance learners num_agent_steps_sampled_lifetime num_env_steps_sampled_lifetime num_env_steps_trained_lifetime num_episodes_lifetimeperf timers
PPO_SatelliteTaskingCL-RLlib_ebb51_00000{'num_module_steps_sampled': {'default_policy': 250}, 'num_agent_steps_sampled_lifetime': {'default_agent': 1500}, 'sample': np.float64(0.8950494590202573), 'episode_len_mean': 90.0, 'episode_return_max': 0.280679012345679, 'episode_len_max': 90, 'num_episodes': 4, 'num_module_steps_sampled_lifetime': {'default_policy': 1500}, 'module_episode_returns_mean': {'default_policy': 0.2732716049382716}, 'episode_len_min': 90, 'num_agent_steps_sampled': {'default_agent': 250}, 'episode_return_mean': 0.2732716049382716, 'agent_episode_returns_mean': {'default_agent': 0.2732716049382716}, 'num_env_steps_sampled_lifetime': 2250, 'episode_duration_sec_mean': 0.6905680585000198, 'num_env_steps_sampled': 250, 'episode_return_min': 0.26586419753086415, 'time_between_sampling': np.float64(0.3281797255881036), 'battery_capacity': np.float64(1439999.568), 'reward': np.float64(0.24778827160493824), 'reward_per_orbit': np.float64(0.08718476223136716), 'orbits_complete': np.float64(2.842105263157895), 'alive': np.float64(1.0), 'rw_status_valid': np.float64(1.0), 'difficulty': np.float64(5.05e-05), 'battery_status_valid': np.float64(1.0), 'external_torque': np.float64(0.0002000007)}{'num_healthy_workers': 2, 'num_in_flight_async_reqs': 0, 'num_remote_worker_restarts': 0}{'default_policy': {'gradients_default_optimizer_global_norm': 0.13789102435112, 'num_trainable_parameters': 139013.0, 'vf_explained_var': 0.16948705911636353, 'default_optimizer_learning_rate': 3e-05, 'total_loss': 0.05914388224482536, 'num_module_steps_trained': 250, 'policy_loss': 0.058846425265073776, 'mean_kl_loss': 0.0, 'entropy': 1.3194327354431152, 'vf_loss_unclipped': 0.0002974536910187453, 'num_non_trainable_parameters': 0.0, 'curr_entropy_coeff': 0.0, 'vf_loss': 0.0002974536910187453}, '__all_modules__': {'num_non_trainable_parameters': 0.0, 'total_loss': 0.05914388224482536, 'num_trainable_parameters': 139013.0, 'num_module_steps_trained': 250, 'num_env_steps_trained': 250}}{'default_agent': 750} 750 750 8{'cpu_util_percent': np.float64(44.25), 'ram_util_percent': np.float64(30.6)}{'env_runner_sampling_timer': 0.9016591849307023, 'learner_update_timer': 0.2980271735021473, 'synch_weights': 0.004593398862737405, 'synch_env_connectors': 0.005024259785006258}
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,749 utils.orbital                  WARNING    <0.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,756 utils.orbital                  WARNING    <180.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,763 utils.orbital                  WARNING    <360.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,770 utils.orbital                  WARNING    <540.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,778 utils.orbital                  WARNING    <720.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,786 utils.orbital                  WARNING    <900.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,793 utils.orbital                  WARNING    <1080.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,801 utils.orbital                  WARNING    <1260.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,809 utils.orbital                  WARNING    <1440.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,816 utils.orbital                  WARNING    <1620.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,823 utils.orbital                  WARNING    <1800.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,831 utils.orbital                  WARNING    <1980.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,838 utils.orbital                  WARNING    <2160.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,846 utils.orbital                  WARNING    <2340.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,854 utils.orbital                  WARNING    <2520.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,861 utils.orbital                  WARNING    <2700.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,869 utils.orbital                  WARNING    <2880.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,877 utils.orbital                  WARNING    <3060.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,884 utils.orbital                  WARNING    <3240.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,891 utils.orbital                  WARNING    <3420.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,897 utils.orbital                  WARNING    <3600.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,905 utils.orbital                  WARNING    <3780.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,913 utils.orbital                  WARNING    <3960.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,921 utils.orbital                  WARNING    <4140.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,929 utils.orbital                  WARNING    <4320.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,937 utils.orbital                  WARNING    <4500.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,944 utils.orbital                  WARNING    <4680.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,952 utils.orbital                  WARNING    <4860.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,959 utils.orbital                  WARNING    <5040.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,966 utils.orbital                  WARNING    <5220.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,973 utils.orbital                  WARNING    <5400.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,980 utils.orbital                  WARNING    <5580.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,987 utils.orbital                  WARNING    <5760.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:23,994 utils.orbital                  WARNING    <5940.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,000 utils.orbital                  WARNING    <6120.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,007 utils.orbital                  WARNING    <6300.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,014 utils.orbital                  WARNING    <6480.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,020 utils.orbital                  WARNING    <6660.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,027 utils.orbital                  WARNING    <6840.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,034 utils.orbital                  WARNING    <7020.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,040 utils.orbital                  WARNING    <7200.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,047 utils.orbital                  WARNING    <7380.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,053 utils.orbital                  WARNING    <7560.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,060 utils.orbital                  WARNING    <7740.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,066 utils.orbital                  WARNING    <7920.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,073 utils.orbital                  WARNING    <8100.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,081 utils.orbital                  WARNING    <8280.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,088 utils.orbital                  WARNING    <8460.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,094 utils.orbital                  WARNING    <8640.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,101 utils.orbital                  WARNING    <8820.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,107 utils.orbital                  WARNING    <9000.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,114 utils.orbital                  WARNING    <9180.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,121 utils.orbital                  WARNING    <9360.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,127 utils.orbital                  WARNING    <9540.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,133 utils.orbital                  WARNING    <9720.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,140 utils.orbital                  WARNING    <9900.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,147 utils.orbital                  WARNING    <10080.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,153 utils.orbital                  WARNING    <10260.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,160 utils.orbital                  WARNING    <10440.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,167 utils.orbital                  WARNING    <10620.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,173 utils.orbital                  WARNING    <10800.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,181 utils.orbital                  WARNING    <10980.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,189 utils.orbital                  WARNING    <11160.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,195 utils.orbital                  WARNING    <11340.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,202 utils.orbital                  WARNING    <11520.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,210 utils.orbital                  WARNING    <11700.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,217 utils.orbital                  WARNING    <11880.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,224 utils.orbital                  WARNING    <12060.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,230 utils.orbital                  WARNING    <12240.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,236 utils.orbital                  WARNING    <12420.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,243 utils.orbital                  WARNING    <12600.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,564 utils.orbital                  WARNING    <12780.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,571 utils.orbital                  WARNING    <12960.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,580 utils.orbital                  WARNING    <13140.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,588 utils.orbital                  WARNING    <13320.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,596 utils.orbital                  WARNING    <13500.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,610 utils.orbital                  WARNING    <13680.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,618 utils.orbital                  WARNING    <13860.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,626 utils.orbital                  WARNING    <14040.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,634 utils.orbital                  WARNING    <14220.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,641 utils.orbital                  WARNING    <14400.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,649 utils.orbital                  WARNING    <14580.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,656 utils.orbital                  WARNING    <14760.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,664 utils.orbital                  WARNING    <14940.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,672 utils.orbital                  WARNING    <15120.00> Could not find eclipse transitions in next 12000.0 seconds
(SingleAgentEnvRunner pid=3988) 2026-10-09 20:35:24,680 utils.orbital                  WARNING    <15300.00> Could not find eclipse transitions in next 12000.0 seconds
2026-10-09 20:35:25,956 INFO tune.py:1009 -- Wrote the latest version of all result files and experiment state to '/home/runner/ray_results/PPO_2026-10-09_20-35-09' in 0.0167s.
(PPO pid=3934) Checkpoint successfully created at: Checkpoint(filesystem=local, path=/home/runner/ray_results/PPO_2026-10-09_20-35-09/PPO_SatelliteTaskingCL-RLlib_ebb51_00000_0_2026-10-09_20-35-09/checkpoint_000000)
2026-10-09 20:35:26,320 INFO tune.py:1041 -- Total run time: 16.57 seconds (16.16 seconds for the tuning loop).
(SingleAgentEnvRunner pid=3987) 2026-10-09 20:35:25,463 utils.orbital                  WARNING    <16200.00> Could not find eclipse transitions in next 12000.0 seconds [repeated 96x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)

Checking Difficulty Over Training

After a few training steps, the difficulty started to increase

[11]:
results.results[next(iter(results.results.keys()))]["env_runners"]["difficulty"]
[11]:
np.float64(5.05e-05)