space_to_space_rso_imaging_evaluation
Evaluate an RSO imaging policy and save a replayable mission and plots.
Run python examples/space_to_space_rso_imaging_evaluation.py --help.
The default priority baseline needs no checkpoint or learning framework.
Matplotlib is optional with --no-plots. The
space-to-space imaging notebook explains
the environment, reward callbacks, replay, and policy adapters.
The priority, nearest-target, and random baselines first manage battery charge, wheel speed, and stored images. They then choose an accessible, illuminated candidate. A saved catalog fixes target states, while a mission manifest also fixes the imager, world, and action/reward settings. Output CSV files distinguish attempts, physically stored acquisitions, and fully completed deliveries.
The RLlib adapter loads one PyTorch RLModule for inference. Custom policies use
module:function factories returning policy(observation, context).
The context includes stable candidate IDs, physical access, resources, and action
indices. Empty slots remain empty. See the notebook for runnable commands.
- checkpoint_fingerprint(path)[source]
Hash an exact checkpoint file or directory without loading its contents.
- decision_context(env, imager)[source]
Expose IDs and physical diagnostics without changing observation columns.
- baseline_policy(kind, rng)[source]
Use resource guards, downlink stored products, then select an image slot.
- load_policy(kind, env, imager, rng, checkpoint=None, factory=None, stochastic=False)[source]
Load a stateless discrete policy, without creating workers or an algorithm.
- validate_action(action, context)[source]
Reject malformed, out-of-range and padded actions without substitution.
- write_csv(path, rows, fields)[source]
Keep empty event tables readable with the same explicit schema.
- plot_results(directory, states, decisions)[source]
Plot physical products, credited rewards, resources, and action intervals.
- evaluate(output, *, policy='priority', seed=0, horizon=None, max_steps=200, catalog_path=None, manifest_path=None, profile=None, checkpoint=None, factory=None, stochastic=False, plots=True)[source]
Run one bounded episode; never overwrite an existing output directory.
Times/resources are sampled at decision boundaries. Capture timestamps come from the physical gate; delivery timestamps are boundary recognition times. Catalog replay alone fixes targets; mission replay fixes scanner/world too.