Resume training with full state

Resume training with full state#

Use the same model, optimizer, scheduler, callback, and dataset configuration. Set the new epoch limit above the completed epoch count. From the quickstart:

python -m stable_pretraining.quickstart --epochs 1 --cache-dir ./first-run

Locate the checkpoint under the printed run directory, then resume:

python -m stable_pretraining.quickstart --epochs 2 --cache-dir ./resumed-run --resume /absolute/path/to/checkpoint.ckpt

Replace the checkpoint path with the actual file from your run. The quickstart passes its resolved path to spt.Manager(..., ckpt_path=..., weights_only=False). It delegates checkpoint restoration to Lightning. Use only a trusted checkpoint for full-state loading. A new invocation creates a new run directory; SLURM requeue has its own Manager-managed continuation behavior.

The regression suite checks deterministic CPU restart at an epoch boundary, including optimizer, scheduler, EMA teacher, probe, queue, and callback state. This is not a promise of bitwise equivalence for shuffled data resumed mid-epoch or across different devices/precision settings.

For Jet, preserve scale_parameterization and scale_eps. Its state dict records those settings and rejects incompatible transforms on load. A bounded checkpoint cannot be silently reinterpreted as an unbounded-scale experiment.