⚡Under $0.50/hr🧠VRAM Estimator⚖Compare GPUs🎁Free LLM APIs🎯Model Index
Fault Tolerance2026-09-089 min read

Zero-Loss Spot GPU Training: Automated S3 Checkpointing and SIGTERM Handlers

Slashing compute burn rates by 60% without losing training progress: Implementing robust SIGTERM interceptors, asynchronous S3 state flushes, and cluster resumption.

Table of Contents

Spot Preemption LifecyclePyTorch SIGTERM TrappingAsynchronous S3 StreamingResumption Automation

Spot vs On-Demand Pricing

Pricing data used in spot checkpointing cost calculations.

ProviderGPU & VRAMInterconnectSpot RateOn-DemandMonthlyStatusAction
Community
PCIe 4.0 (64 GB/s)$0.34 / hr$0.85 / hr$208 / moInstant
Deploy →
Bare Metal
PCIe 4.0 (64 GB/s)$0.69 / hr$1.73 / hr$422 / moInstant
Deploy →
Community
PCIe 4.0 (64 GB/s)$0.69 / hr$1.73 / hr$422 / moInstant
Deploy →
Cloud
PCIe 4.0 (64 GB/s)$0.74 / hr$1.85 / hr$453 / moInstant
Deploy →
Dedicated
PCIe 4.0 (64 GB/s)$0.89 / hr$2.23 / hr$545 / moInstant
Deploy →
Cloud
PCIe 4.0 (64 GB/s)$1.09 / hr$2.73 / hr$667 / moInstant
Deploy →
Bare Metal
PCIe 4.0 (64 GB/s)$1.19 / hr$2.97 / hr$728 / moInstant
Deploy →
Dedicated
PCIe 4.0 (64 GB/s)$1.49 / hr$3.73 / hr$912 / moInstant
Deploy →
Dedicated
N/A$1.59 / hr$3.98 / hr$973 / moInstant
Deploy →
Community
NVLink 4.0 (900 GB/s)$1.89 / hr$4.72 / hr$1,157 / moInstant
Deploy →
Bare Metal
NVLink 4.0 (900 GB/s)$2.29 / hr$5.73 / hr$1,401 / moInstant
Deploy →
Community
NVLink 4.0 (900 GB/s)$2.79 / hr$6.98 / hr$1,707 / moInstant
Deploy →
Dedicated
NVLink 4.0 (900 GB/s)$2.99 / hr$7.48 / hr$1,830 / moInstant
Deploy →
Bare Metal
NVLink 4.0 (900 GB/s)$3.19 / hr$7.98 / hr$1,952 / moInstant
Deploy →
Cloud
NVLink 4.0 (900 GB/s)$3.49 / hr$8.73 / hr$2,136 / moInstant
Deploy →
Community
NVLink 5.0 (1.8 TB/s)$3.99 / hr$9.98 / hr$2,442 / moInstant
Deploy →
Dedicated
NVLink 4.0 (900 GB/s)$3.99 / hr$9.98 / hr$2,442 / moInstant
Deploy →
Cloud
NVLink 4.0 (900 GB/s)$4.31 / hr$10.77 / hr$2,638 / moInstant
Deploy →
Bare Metal
NVLink 5.0 (1.8 TB/s)$4.49 / hr$11.23 / hr$2,748 / moInstant
Deploy →
Dedicated
NVLink 5.0 (1.8 TB/s)$5.49 / hr$13.73 / hr$3,360 / moInstant
Deploy →
Cloud
NVLink 5.0 (1.8 TB/s)$5.99 / hr$14.98 / hr$3,666 / moInstant
Deploy →
Data Freshness: Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x, PagedAttention v2, FlashAttention-3

Spot Preemption Lifecycle

Spot GPU preemption follows a predictable lifecycle: (1) Provider signals preemption (30-120 seconds notice depending on provider). (2) SIGTERM sent to running process. (3) GPU compute halted. (4) Instance terminated. The critical window is between SIGTERM and termination — you have 30 seconds to save state. For training jobs, this means: checkpoint every 30 minutes minimum, and implement SIGTERM trapping to trigger emergency checkpoint. Vast.ai gives 30s, RunPod gives 30s, Lambda gives 30s. All providers treat SIGTERM as 'save and exit'.

PyTorch SIGTERM Trapping

Standard PyTorch training loops do not handle SIGTERM. You must register a signal handler: import signal, torch, os def emergency_checkpoint(signum, frame): print(f"Received SIGTERM, saving emergency checkpoint...") torch.save({ 'epoch': epoch, 'model_state_dict': model.state_dict(), 'optimizer_state_dict': optimizer.state_dict(), 'loss': loss, }, f"/checkpoint/emergency_{os.getpid()}.pt") # Upload to S3 asynchronously os.system(f"aws s3 cp /checkpoint/emergency_{os.getpid()}.pt s3://bucket/checkpoints/") exit(0) signal.signal(signal.SIGTERM, emergency_checkpoint) The handler must be non-blocking — upload to S3 in a subprocess or background thread to avoid exceeding the 30s window.

Asynchronous S3 Streaming

Synchronous S3 uploads block the training loop — a 2GB checkpoint takes ~10 seconds to upload at 200 MB/s network speed. During those 10 seconds, you cannot compute gradients. Solution: asynchronous S3 streaming: import boto3, threading def async_upload(local_path, s3_path): s3 = boto3.client('s3') threading.Thread( target=s3.upload_file, args=(local_path, 'bucket', s3_path) ).start() # In training loop: torch.save(model.state_dict(), '/tmp/checkpoint.pt') async_upload('/tmp/checkpoint.pt', f'checkpoints/step_{step}.pt') This uploads in the background while training continues. The 30s SIGTERM window is now sufficient for both checkpoint and upload.

Resumption Automation

After spot preemption, the new instance must automatically resume training. Implementation: (1) On startup, check S3 for latest checkpoint. (2) Download checkpoint to local storage. (3) Load model weights, optimizer state, and scheduler state. (4) Resume from the saved step. (5) Log resumption event for monitoring. The full resumption script should be in the instance's user-data or Docker entrypoint. Target: <5 minutes from new instance creation to training resumption. For a 70B model, checkpoint download from S3 takes ~60 seconds (2GB at 200 MB/s), model loading takes ~90 seconds, and warmup takes ~30 seconds. Total: ~3 minutes.
OR

OpenGPU Radar Systems Engineering Team

Independent compute telemetry and infrastructure analysis. Not affiliated with NVIDIA, cloud providers, or hardware vendors.

Related Guides