space_to_space_rso_imaging_evaluation

Evaluate an RSO imaging policy and save a replayable mission and plots.

Run python examples/space_to_space_rso_imaging_evaluation.py --help. The default priority baseline needs no checkpoint or learning framework. Matplotlib is optional with --no-plots. The space-to-space imaging notebook explains the environment, reward callbacks, replay, and policy adapters.

The priority, nearest-target, and random baselines first manage battery charge, wheel speed, and stored images. They then choose an accessible, illuminated candidate. A saved catalog fixes target states, while a mission manifest also fixes the imager, world, and action/reward settings. Output CSV files distinguish attempts, physically stored acquisitions, and fully completed deliveries.

The RLlib adapter loads one PyTorch RLModule for inference. Custom policies use module:function factories returning policy(observation, context). The context includes stable candidate IDs, physical access, resources, and action indices. Empty slots remain empty. See the notebook for runnable commands.

checkpoint_fingerprint(path)[source]

Hash an exact checkpoint file or directory without loading its contents.

decision_context(env, imager)[source]

Expose IDs and physical diagnostics without changing observation columns.

baseline_policy(kind, rng)[source]

Use resource guards, downlink stored products, then select an image slot.

load_policy(kind, env, imager, rng, checkpoint=None, factory=None, stochastic=False)[source]

Load a stateless discrete policy, without creating workers or an algorithm.

validate_action(action, context)[source]

Reject malformed, out-of-range and padded actions without substitution.

write_csv(path, rows, fields)[source]

Keep empty event tables readable with the same explicit schema.

plot_results(directory, states, decisions)[source]

Plot physical products, credited rewards, resources, and action intervals.

evaluate(output, *, policy='priority', seed=0, horizon=None, max_steps=200, catalog_path=None, manifest_path=None, profile=None, checkpoint=None, factory=None, stochastic=False, plots=True)[source]

Run one bounded episode; never overwrite an existing output directory.

Times/resources are sampled at decision boundaries. Capture timestamps come from the physical gate; delivery timestamps are boundary recognition times. Catalog replay alone fixes targets; mission replay fixes scanner/world too.

main()[source]

Command-line entry point for a single reproducible evaluation.