Editorial

Sreedhar

AI Architect & Founder, OpenGPU Radar

12+ years in GPU-accelerated ML infrastructure ยท LLM inference optimization

Biography

SR

I'm Sreedhar, an AI Architect with over 12 years of experience building production infrastructure for machine learning systems. I specialize in GPU-accelerated inference pipelines, LLM serving architecture, and the cost optimization problems that every ML team faces when moving from prototype to production.

My career spans the full lifecycle of ML infrastructure โ€” from designing multi-node GPU clusters with InfiniBand interconnects to tuning vLLM PagedAttention kernels for maximum throughput. I've deployed inference systems serving millions of tokens per day across providers like NVIDIA H100, H200, and AMD MI300X clusters.

OpenGPU Radar started as an internal tool โ€” a way to track cloud GPU spot prices and calculate VRAM requirements without relying on vendor-provided estimates. When I realized how fragmented and opaque the cloud GPU market was, I decided to make the data public. Every price point, every benchmark, and every VRAM calculation on this site is independently sourced and reproducible.

My technical focus areas include LLM inference optimization (PagedAttention, speculative decoding, quantization-aware serving), multi-node GPU cluster architecture with NVLink and InfiniBand, and deterministic cost modeling for long-context workloads. I built the VRAM calculator and the entity graph that powers the model, GPU, and provider pages on this platform.

Expertise

LLM Inference

PagedAttention, FlashAttention-3, speculative decoding, continuous batching, KV-cache optimization, quantization (FP8, INT4, AWQ, GPTQ).

GPU Cluster Architecture

Multi-node training with InfiniBand NDR, NVLink mesh topologies, tensor parallelism, pipeline parallelism, NCCL AllReduce optimization.

Cost Modeling

Deterministic VRAM calculations, spot vs reserved pricing analysis, cost-per-token benchmarking, cloud GPU market telemetry.

Production ML Systems

vLLM, TensorRT-LLM, Ollama, Docker orchestration, Prometheus/Grafana monitoring, SIGTERM checkpointing for spot instances.

What I Believe

Transparency Over Hype

The AI infrastructure market is drowning in benchmarks that nobody can reproduce and pricing pages that change without notice. I believe in publishing the methodology behind every number โ€” not just the number itself.

Determinism Over Estimates

VRAM calculations should be deterministic, not rule-of-thumb. The entity graph on this site uses exact formulas โ€” weights bytes per parameter, GQA KV-cache head counts, layer counts โ€” not "10-20% overhead" placeholders.

Free Access to Compute Intelligence

Every tool on OpenGPU Radar โ€” the VRAM calculator, the GPU pricing table, the free LLM API directory โ€” is free and requires no account. Compute intelligence should not be gated behind enterprise sales calls.

Connect

About OpenGPU Radar

OpenGPU Radar is an independent AI compute intelligence platform. No affiliate partnerships. No sponsored placements. Every data point is independently sourced and reproducible. The platform provides deterministic VRAM calculations, verified cloud GPU spot pricing, and a curated directory of free LLM API endpoints.