Sreedhar
AI Architect & Founder, OpenGPU Radar
12+ years in GPU-accelerated ML infrastructure ยท LLM inference optimization
Biography
I'm Sreedhar, an AI Architect with over 12 years of experience building production infrastructure for machine learning systems. I specialize in GPU-accelerated inference pipelines, LLM serving architecture, and the cost optimization problems that every ML team faces when moving from prototype to production.
My career spans the full lifecycle of ML infrastructure โ from designing multi-node GPU clusters with InfiniBand interconnects to tuning vLLM PagedAttention kernels for maximum throughput. I've deployed inference systems serving millions of tokens per day across providers like NVIDIA H100, H200, and AMD MI300X clusters.
OpenGPU Radar started as an internal tool โ a way to track cloud GPU spot prices and calculate VRAM requirements without relying on vendor-provided estimates. When I realized how fragmented and opaque the cloud GPU market was, I decided to make the data public. Every price point, every benchmark, and every VRAM calculation on this site is independently sourced and reproducible.
My technical focus areas include LLM inference optimization (PagedAttention, speculative decoding, quantization-aware serving), multi-node GPU cluster architecture with NVLink and InfiniBand, and deterministic cost modeling for long-context workloads. I built the VRAM calculator and the entity graph that powers the model, GPU, and provider pages on this platform.
Expertise
PagedAttention, FlashAttention-3, speculative decoding, continuous batching, KV-cache optimization, quantization (FP8, INT4, AWQ, GPTQ).
Multi-node training with InfiniBand NDR, NVLink mesh topologies, tensor parallelism, pipeline parallelism, NCCL AllReduce optimization.
Deterministic VRAM calculations, spot vs reserved pricing analysis, cost-per-token benchmarking, cloud GPU market telemetry.
vLLM, TensorRT-LLM, Ollama, Docker orchestration, Prometheus/Grafana monitoring, SIGTERM checkpointing for spot instances.
What I Believe
The AI infrastructure market is drowning in benchmarks that nobody can reproduce and pricing pages that change without notice. I believe in publishing the methodology behind every number โ not just the number itself.
VRAM calculations should be deterministic, not rule-of-thumb. The entity graph on this site uses exact formulas โ weights bytes per parameter, GQA KV-cache head counts, layer counts โ not "10-20% overhead" placeholders.
Every tool on OpenGPU Radar โ the VRAM calculator, the GPU pricing table, the free LLM API directory โ is free and requires no account. Compute intelligence should not be gated behind enterprise sales calls.
Connect
About OpenGPU Radar
OpenGPU Radar is an independent AI compute intelligence platform. No affiliate partnerships. No sponsored placements. Every data point is independently sourced and reproducible. The platform provides deterministic VRAM calculations, verified cloud GPU spot pricing, and a curated directory of free LLM API endpoints.