Fault Tolerance2026-09-089 min read
Zero-Loss Spot GPU Training: Automated S3 Checkpointing and SIGTERM Handlers
Slashing compute burn rates by 60% without losing training progress: Implementing robust SIGTERM interceptors, asynchronous S3 state flushes, and cluster resumption.
Table of Contents
- 1. Spot Preemption Lifecycle
- 2. PyTorch SIGTERM Trapping
- 3. Asynchronous S3 Streaming
- 4. Resumption Automation
Spot Preemption LifecyclePyTorch SIGTERM TrappingAsynchronous S3 StreamingResumption Automation
Spot vs On-Demand Pricing
Pricing data used in spot checkpointing cost calculations.
| Provider | GPU & VRAM | Interconnect | Spot Rate | On-Demand | Monthly | Status | Action | |
|---|---|---|---|---|---|---|---|---|
Community | PCIe 4.0 (64 GB/s) | $0.34 / hr | $0.85 / hr | $208 / mo | Instant | |||
Bare Metal | PCIe 4.0 (64 GB/s) | $0.69 / hr | $1.73 / hr | $422 / mo | Instant | |||
Community | PCIe 4.0 (64 GB/s) | $0.69 / hr | $1.73 / hr | $422 / mo | Instant | |||
Cloud | PCIe 4.0 (64 GB/s) | $0.74 / hr | $1.85 / hr | $453 / mo | Instant | |||
Dedicated | PCIe 4.0 (64 GB/s) | $0.89 / hr | $2.23 / hr | $545 / mo | Instant | |||
Cloud | PCIe 4.0 (64 GB/s) | $1.09 / hr | $2.73 / hr | $667 / mo | Instant | |||
Bare Metal | PCIe 4.0 (64 GB/s) | $1.19 / hr | $2.97 / hr | $728 / mo | Instant | |||
Dedicated | PCIe 4.0 (64 GB/s) | $1.49 / hr | $3.73 / hr | $912 / mo | Instant | |||
Dedicated | N/A | $1.59 / hr | $3.98 / hr | $973 / mo | Instant | |||
Community | NVLink 4.0 (900 GB/s) | $1.89 / hr | $4.72 / hr | $1,157 / mo | Instant | |||
Bare Metal | NVLink 4.0 (900 GB/s) | $2.29 / hr | $5.73 / hr | $1,401 / mo | Instant | |||
Community | NVLink 4.0 (900 GB/s) | $2.79 / hr | $6.98 / hr | $1,707 / mo | Instant | |||
Dedicated | NVLink 4.0 (900 GB/s) | $2.99 / hr | $7.48 / hr | $1,830 / mo | Instant | |||
Bare Metal | NVLink 4.0 (900 GB/s) | $3.19 / hr | $7.98 / hr | $1,952 / mo | Instant | |||
Cloud | NVLink 4.0 (900 GB/s) | $3.49 / hr | $8.73 / hr | $2,136 / mo | Instant | |||
Community | NVLink 5.0 (1.8 TB/s) | $3.99 / hr | $9.98 / hr | $2,442 / mo | Instant | |||
Dedicated | NVLink 4.0 (900 GB/s) | $3.99 / hr | $9.98 / hr | $2,442 / mo | Instant | |||
Cloud | NVLink 4.0 (900 GB/s) | $4.31 / hr | $10.77 / hr | $2,638 / mo | Instant | |||
Bare Metal | NVLink 5.0 (1.8 TB/s) | $4.49 / hr | $11.23 / hr | $2,748 / mo | Instant | |||
Dedicated | NVLink 5.0 (1.8 TB/s) | $5.49 / hr | $13.73 / hr | $3,360 / mo | Instant | |||
Cloud | NVLink 5.0 (1.8 TB/s) | $5.99 / hr | $14.98 / hr | $3,666 / mo | Instant |
Data Freshness: Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)|Benchmark Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x, PagedAttention v2, FlashAttention-3
Spot Preemption Lifecycle
Spot GPU preemption follows a predictable lifecycle: (1) Provider signals preemption (30-120 seconds notice depending on provider). (2) SIGTERM sent to running process. (3) GPU compute halted. (4) Instance terminated. The critical window is between SIGTERM and termination — you have 30 seconds to save state. For training jobs, this means: checkpoint every 30 minutes minimum, and implement SIGTERM trapping to trigger emergency checkpoint. Vast.ai gives 30s, RunPod gives 30s, Lambda gives 30s. All providers treat SIGTERM as 'save and exit'.
PyTorch SIGTERM Trapping
Standard PyTorch training loops do not handle SIGTERM. You must register a signal handler:
import signal, torch, os
def emergency_checkpoint(signum, frame):
print(f"Received SIGTERM, saving emergency checkpoint...")
torch.save({
'epoch': epoch,
'model_state_dict': model.state_dict(),
'optimizer_state_dict': optimizer.state_dict(),
'loss': loss,
}, f"/checkpoint/emergency_{os.getpid()}.pt")
# Upload to S3 asynchronously
os.system(f"aws s3 cp /checkpoint/emergency_{os.getpid()}.pt s3://bucket/checkpoints/")
exit(0)
signal.signal(signal.SIGTERM, emergency_checkpoint)
The handler must be non-blocking — upload to S3 in a subprocess or background thread to avoid exceeding the 30s window.
Asynchronous S3 Streaming
Synchronous S3 uploads block the training loop — a 2GB checkpoint takes ~10 seconds to upload at 200 MB/s network speed. During those 10 seconds, you cannot compute gradients. Solution: asynchronous S3 streaming:
import boto3, threading
def async_upload(local_path, s3_path):
s3 = boto3.client('s3')
threading.Thread(
target=s3.upload_file,
args=(local_path, 'bucket', s3_path)
).start()
# In training loop:
torch.save(model.state_dict(), '/tmp/checkpoint.pt')
async_upload('/tmp/checkpoint.pt', f'checkpoints/step_{step}.pt')
This uploads in the background while training continues. The 30s SIGTERM window is now sufficient for both checkpoint and upload.
Resumption Automation
After spot preemption, the new instance must automatically resume training. Implementation: (1) On startup, check S3 for latest checkpoint. (2) Download checkpoint to local storage. (3) Load model weights, optimizer state, and scheduler state. (4) Resume from the saved step. (5) Log resumption event for monitoring. The full resumption script should be in the instance's user-data or Docker entrypoint. Target: <5 minutes from new instance creation to training resumption. For a 70B model, checkpoint download from S3 takes ~60 seconds (2GB at 200 MB/s), model loading takes ~90 seconds, and warmup takes ~30 seconds. Total: ~3 minutes.
OR
OpenGPU Radar Systems Engineering Team
Independent compute telemetry and infrastructure analysis. Not affiliated with NVIDIA, cloud providers, or hardware vendors.