/goal for AI model training runs is so good - it really feels like the future. Very little babysitting now.
Mine: Launch a full training run on 4 nodes.
Continuously record everything in the experiment document (if it exists). Log hyperparameters, configs, periodic evals, performance insights, training stability analysis, and all important changes for future analysis and reproducibility.
Fix any major bugs you encounter while monitoring the training, but do not change the fundamental nature of the experiment without asking. If it crashes, resume from the latest reliable checkpoint and keep going.
Reach <num> steps.
