Local AI2026-10-11โ€ขBy Sreeโ€ข5 min read

Can I Run AI Locally? A Practical Guide to Your PC, RAM, GPU and Windows Setup

Most PCs can run local AI. Your limits are system RAM, GPU VRAM and model size. Step-by-step Ollama setup on Windows, model shortlist, and upgrade triggers.

Direct answer

Yes โ€” most modern computers can run local AI models. Your system RAM decides the largest model that will load. Your GPU VRAM decides whether it runs fast enough to be usable. Model size and quantization decide how much of either you actually need.

Start here โ€” three steps:

  1. Install Ollama for Windows (~500 MB installer).
  2. Run ollama run qwen2.5:3b in PowerShell โ€” a 1.9 GB model that fits comfortably on 8 GB of RAM.
  3. If that was too small, run ollama run qwen2.5:7b (4.7 GB) or ollama run phi4 (9.1 GB).

Model sizes below are the published download sizes from the Ollama library, verified 2026-10-11. Throughput figures are estimates, not benchmarks โ€” your CPU, GPU and context length will move them.

Your hardwareRealistic starting modelEstimated speedNext step if you need more
8 GB RAM, no GPUqwen2.5:3b (1.9 GB)Slow โ€” roughly 5โ€“15 tok/sUpgrade RAM to 16 GB
16 GB RAM, no GPUqwen2.5:7b (4.7 GB), phi4 (9.1 GB)Slow โ€” roughly 3โ€“10 tok/sAdd a GPU with 8 GB+ VRAM
32 GB RAM, no GPUqwen2.5:14b (9.0 GB), qwen2.5:32b (20 GB)Slow โ€” roughly 2โ€“8 tok/sAdd a 24 GB GPU
8 GB VRAM GPUqwen2.5:7b, gemma2:9b (5.4 GB)Moderate โ€” roughly 15โ€“40 tok/sA solid entry point for local AI
24 GB VRAM GPUqwen2.5:32b (20 GB)Fast โ€” roughly 30โ€“80 tok/sThe practical sweet spot for most users
48 GB+ VRAM GPUllama3.3:70b (43 GB)Fast โ€” roughly 40โ€“120 tok/sResearch, local fine-tuning

These are starting points, not guarantees. Numbers assume a 4-bit (Q4) model and short context. Long context adds memory that these estimates do not include.


A. What can my computer run?

The single most important distinction for a beginner is system RAM vs. GPU VRAM vs. shared GPU memory. Mixing these up is the usual cause of "my model crashed."

ComponentWhat it doesTypical sizes
System RAMHolds the model file, the runtime, and your other apps. The model download must fit here.8 / 16 / 32 / 64 GB
GPU VRAMHolds the model when it runs on the GPU. If the model plus context fits here, inference is fast. If it does not, the runtime spills to system RAM and becomes slow.0 / 8 / 16 / 24 / 40 GB+ / 80 GB+
Shared GPU memoryIntegrated graphics (Intel Iris Xe, AMD Radeon) borrow from system RAM. Faster than pure CPU, slower than dedicated VRAM.Shared with system RAM

Hardware decision table

ConfigurationRAMDedicated GPU VRAMWhat realistically runsKnown limitations
Budget PC (2020+)8 GBNoneqwen2.5:3b (1.9 GB); qwen2.5:7b (4.7 GB) is tight7B models compete with your browser for RAM
16 GB laptop, integrated GPU16 GB0 GB dedicatedqwen2.5:7b (4.7 GB), mistral:7b (4.4 GB), phi4 (9.1 GB)No GPU acceleration; 32K context is the practical ceiling
32 GB workstation, no GPU32 GB0 GBqwen2.5:14b (9.0 GB), qwen2.5:32b (20 GB)Works, but CPU-only speed
Gaming PC with RTX 3060/406016 GB8 GBqwen2.5:7b, gemma2:9b (5.4 GB)7B fits with room for context; 9B is tight
RTX 4090 / RTX 6000 Ada32 GB24 GBqwen2.5:32b (20 GB), phi432B fits with short-to-medium context
Dual RTX 409064 GB2 ร— 24 GBllama3.3:70b (43 GB)Multi-GPU sharding; PCIe, not NVLink, is the limit
A100 / H200, 80โ€“141 GB VRAM256 GB+80โ€“141 GBllama3.3:70b, 70B-class models at long contextDatacenter tier; see DeepSeek R1 sizing

Quantized models (the default in the Ollama library) are what make these numbers work. An unquantized 14B model needs about 28 GB of VRAM; the 9.1 GB Q4 build is what actually fits an ordinary laptop. For the underlying VRAM math, use the LLM VRAM calculator.


B. Which local AI model should I choose?

Shortlist by task. Every model below is an open-weight model distributed through the Ollama library.

General chat and writing

  • qwen2.5:7b (4.7 GB) โ€” the best general-purpose starter. 32K context, multilingual.
  • phi4 (9.1 GB, 16K context) โ€” Microsoft's 14B model, stronger at reasoning and logic.

Coding assistance

  • qwen2.5-coder:7b (4.7 GB) โ€” code-specialized, same footprint as the general 7B.
  • qwen2.5-coder:32b (20 GB) โ€” 32B flagship; needs a 24 GB GPU or 32 GB of RAM.

Summarization and document questions

  • phi4 (9.1 GB) โ€” good at structured answers and long-ish passages within its 16K window.
  • qwen2.5:14b (9.0 GB) โ€” better fluency than 7B at the same class of cost.

Image generation

  • FLUX.1 [dev] โ€” 24 GB VRAM at FP16, ~12 GB at INT4. Image models are not run through Ollama โ€” use ComfyUI or a Stable Diffusion UI, which load FLUX from a local file.

Why a model that loads can still be unusable

Loading is not the same as running well. Four things catch people out:

  1. VRAM overflow. If weights plus context exceed VRAM, the runtime falls back to system RAM โ€” typically 10โ€“50ร— slower. A 14B model (~9 GB) on an 8 GB GPU will load, then crawl.
  2. Context length. Context grows memory usage on top of the weights. A model rated for 32K context does not mean you get 32K context for free at 8 GB VRAM.
  3. Quantization. Q4 is the Ollama default and the right call for most people. Higher quantizations (Q5, Q6, Q8) improve quality but grow the download.
  4. No GPU at all. CPU-only inference is fine for occasional questions. It is not fine for a chatbot you use all day.

C. How do I install AI on Windows?

Step 1: Install Ollama

Ollama is the simplest established runtime for local models on Windows. It handles downloads, quantization, and GPU detection automatically.

  1. Download the Windows installer from ollama.com/download/windows.
  2. Run it (~500 MB on disk after install).
  3. Open a new PowerShell window โ€” the installer updates your PATH, and existing windows will not pick that up.

Step 2: Run your first model

# Small, safe first model (1.9 GB download)
ollama run qwen2.5:3b

# Step up when that works (4.7 GB)
ollama run qwen2.5:7b

# Larger, stronger (9.1 GB) โ€” needs 16 GB of RAM
ollama run phi4

Type your question at the prompt. Type /bye or press Ctrl+C to exit. The first run downloads the model; later runs start instantly.

Step 3: Confirm your GPU is actually being used

Ollama uses NVIDIA CUDA automatically when it detects a compatible GPU and drivers.

# See which models are loaded and on which processor
ollama ps

In the PROCESSOR column, 100% GPU means the model is fully on the GPU. A CPU or partial-GPU value means it is spilling to system RAM.

If you see CPU when you expected GPU:

  • Install or update NVIDIA drivers (535 or newer).
  • Verify the card reports at least 6 GB of VRAM (nvidia-smi in the same window).
  • Confirm the model actually fits โ€” ollama ps shows the size it loaded.

Step 4: Pick a smaller model if memory is short

If a run fails or is unusably slow, step down rather than fighting the hardware:

# 8 GB RAM: stay at 3B
ollama run qwen2.5:3b        # 1.9 GB

# 16 GB RAM: 7B is the sweet spot
ollama run qwen2.5:7b        # 4.7 GB
ollama run mistral:7b        # 4.4 GB

# 32 GB RAM or 24 GB VRAM
ollama run qwen2.5:32b       # 20 GB

Common errors and fixes

ErrorCauseFix
out of memoryModel larger than available RAMRun ollama list to see what you have, then ollama rm <model> and start a smaller one
CUDA out of memoryModel plus context exceeds VRAMUse a smaller model, or close other GPU apps
ollama : The term 'ollama' is not recognizedPowerShell window predates the installClose PowerShell and open a fresh window
model '...' not foundTypo or wrong tagCheck the exact tag at ollama.com/library
Download stops / disk fullInsufficient disk spaceollama list shows disk usage; ollama rm <model> frees it

D. macOS and Linux

macOS

brew install ollama
ollama serve          # start the local server

# In a second terminal
ollama run qwen2.5:7b

On Apple Silicon (M1โ€“M4), Ollama uses Metal acceleration and performance is comparable to a mid-range NVIDIA GPU. On Intel Macs there is no GPU acceleration โ€” expect CPU-only speeds. Models are stored in ~/.ollama/models.

Linux

curl -fsSL https://ollama.com/install | sh
ollama serve

# In a second terminal
ollama run qwen2.5:7b

For NVIDIA GPU acceleration, install the driver (535+) and CUDA 12.x first โ€” Ollama detects them on first run. Models are stored in ~/.ollama/models.


E. How much storage and memory do I need?

Download size (Ollama library, verified 2026-10-11)

ModelDownloadContextNotes
qwen2.5:3b1.9 GB32KFits 8 GB RAM
qwen2.5:7b4.7 GB32KThe default recommendation
mistral:7b4.4 GB32KApache-licensed, function calling
gemma2:9b5.4 GB8KShort context window
phi49.1 GB16KStrong reasoning at 14B
qwen2.5:14b9.0 GB32K
qwen2.5-coder:32b20 GB32KNeeds 24 GB VRAM or 32 GB RAM
llama3.3:70b43 GB128KWeights alone fit 48 GB class GPUs at short context; long context needs multi-GPU

Downloads are cached, so each model is fetched once. Models are stored in ~/.ollama/models (C:\Users\<you>\.ollama\models on Windows) and can be deleted individually with ollama rm <model>.

RAM rules of thumb

A practical floor: the model download must fit in RAM and leave room for context and your other applications.

System RAMComfortable class
8 GB1โ€“3B models
16 GB7B models; 14B only if nothing else is running
32 GBUp to 14B comfortably; 32B (20 GB) with short context
64 GB32B comfortably; 70B at INT4 is still multi-GPU territory

VRAM

If the model plus context fits in VRAM, you get GPU speed. If it does not, you get CPU speed with a GPU sitting idle. The exact numbers depend on quantization, context length and KV-cache โ€” OpenGPU Radar's LLM VRAM calculator computes them, and the 8/16/24/48/80 GB tier table shows which models fit where.


F. When should you upgrade โ€” or use the cloud?

Upgrade RAM first if you are on 8 GB and running CPU-only. Going to 16 GB is the cheapest way to unlock 7B models, which is the point where local AI becomes genuinely useful.

Add a dedicated GPU if you do not have one. Even an entry-level card in the 8 GB VRAM class gives a large multiple over CPU inference โ€” for the models that fit. See current GPU pricing.

Go to 24 GB VRAM when you regularly want 14Bโ€“32B models with real context. This is the practical enthusiast tier.

Use the cloud or an API when the model you want is larger than anything you will own:

SituationRealistic choice
Occasional questions, prototyping, privacy-sensitive dataLocal, on what you already own
Regular use of 32B-class models24 GB local GPU
70B+ models, long context, high concurrencyCloud GPU or an API โ€” see DeepSeek R1 for the upper end
Bursty image or video generationCloud โ€” local GPU cost is hard to justify for bursty demand

The honest upgrade path: start free on existing hardware, add RAM, add a GPU, and only then consider renting. Most people stop at step two.


Troubleshooting cheat sheet

ProblemFix
Model will not loadStep down a size: ollama run qwen2.5:3b instead of qwen2.5:7b
Very slow repliesollama ps โ€” if PROCESSOR is not 100% GPU, you are on CPU
CUDA out of memorySmaller model, or shorter context, or close other GPU apps
ollama not recognizedOpen a new PowerShell / terminal window
Disk fullollama list to check usage, ollama rm <model> to free space
GPU not detectedUpdate NVIDIA drivers; confirm nvidia-smi works

Next steps

  1. LLM VRAM calculator โ€” exact memory for your model, precision and context
  2. VRAM tier breakdown โ€” which models fit which GPUs
  3. Model directory โ€” verified specs for every model we track
  4. Cheapest GPUs โ€” what the hardware costs when you are ready to upgrade

Methodology

  • Model download sizes and context windows are the published values on the Ollama library for each tag, verified 2026-10-11.
  • Installation commands are taken from Ollama's official Windows, macOS and Linux install instructions, verified 2026-10-11.
  • VRAM requirements for the larger models use OpenGPU Radar's canonical engine via the VRAM calculator; see the methodology page.
  • Throughput figures are estimates for a single interactive stream at short context. They are not benchmark results, and real-world numbers vary by CPU, GPU, context length and runtime version.
  • Pricing references link to OpenGPU Radar's tracked GPU pricing. No prices are hardcoded in this article.