Can I Run AI Locally? A Practical Guide to Your PC, RAM, GPU and Windows Setup
Most PCs can run local AI. Your limits are system RAM, GPU VRAM and model size. Step-by-step Ollama setup on Windows, model shortlist, and upgrade triggers.
Direct answer
Yes โ most modern computers can run local AI models. Your system RAM decides the largest model that will load. Your GPU VRAM decides whether it runs fast enough to be usable. Model size and quantization decide how much of either you actually need.
Start here โ three steps:
- Install Ollama for Windows (~500 MB installer).
- Run
ollama run qwen2.5:3bin PowerShell โ a 1.9 GB model that fits comfortably on 8 GB of RAM. - If that was too small, run
ollama run qwen2.5:7b(4.7 GB) orollama run phi4(9.1 GB).
Model sizes below are the published download sizes from the Ollama library, verified 2026-10-11. Throughput figures are estimates, not benchmarks โ your CPU, GPU and context length will move them.
| Your hardware | Realistic starting model | Estimated speed | Next step if you need more |
|---|---|---|---|
| 8 GB RAM, no GPU | qwen2.5:3b (1.9 GB) | Slow โ roughly 5โ15 tok/s | Upgrade RAM to 16 GB |
| 16 GB RAM, no GPU | qwen2.5:7b (4.7 GB), phi4 (9.1 GB) | Slow โ roughly 3โ10 tok/s | Add a GPU with 8 GB+ VRAM |
| 32 GB RAM, no GPU | qwen2.5:14b (9.0 GB), qwen2.5:32b (20 GB) | Slow โ roughly 2โ8 tok/s | Add a 24 GB GPU |
| 8 GB VRAM GPU | qwen2.5:7b, gemma2:9b (5.4 GB) | Moderate โ roughly 15โ40 tok/s | A solid entry point for local AI |
| 24 GB VRAM GPU | qwen2.5:32b (20 GB) | Fast โ roughly 30โ80 tok/s | The practical sweet spot for most users |
| 48 GB+ VRAM GPU | llama3.3:70b (43 GB) | Fast โ roughly 40โ120 tok/s | Research, local fine-tuning |
These are starting points, not guarantees. Numbers assume a 4-bit (Q4) model and short context. Long context adds memory that these estimates do not include.
A. What can my computer run?
The single most important distinction for a beginner is system RAM vs. GPU VRAM vs. shared GPU memory. Mixing these up is the usual cause of "my model crashed."
| Component | What it does | Typical sizes |
|---|---|---|
| System RAM | Holds the model file, the runtime, and your other apps. The model download must fit here. | 8 / 16 / 32 / 64 GB |
| GPU VRAM | Holds the model when it runs on the GPU. If the model plus context fits here, inference is fast. If it does not, the runtime spills to system RAM and becomes slow. | 0 / 8 / 16 / 24 / 40 GB+ / 80 GB+ |
| Shared GPU memory | Integrated graphics (Intel Iris Xe, AMD Radeon) borrow from system RAM. Faster than pure CPU, slower than dedicated VRAM. | Shared with system RAM |
Hardware decision table
| Configuration | RAM | Dedicated GPU VRAM | What realistically runs | Known limitations |
|---|---|---|---|---|
| Budget PC (2020+) | 8 GB | None | qwen2.5:3b (1.9 GB); qwen2.5:7b (4.7 GB) is tight | 7B models compete with your browser for RAM |
| 16 GB laptop, integrated GPU | 16 GB | 0 GB dedicated | qwen2.5:7b (4.7 GB), mistral:7b (4.4 GB), phi4 (9.1 GB) | No GPU acceleration; 32K context is the practical ceiling |
| 32 GB workstation, no GPU | 32 GB | 0 GB | qwen2.5:14b (9.0 GB), qwen2.5:32b (20 GB) | Works, but CPU-only speed |
| Gaming PC with RTX 3060/4060 | 16 GB | 8 GB | qwen2.5:7b, gemma2:9b (5.4 GB) | 7B fits with room for context; 9B is tight |
| RTX 4090 / RTX 6000 Ada | 32 GB | 24 GB | qwen2.5:32b (20 GB), phi4 | 32B fits with short-to-medium context |
| Dual RTX 4090 | 64 GB | 2 ร 24 GB | llama3.3:70b (43 GB) | Multi-GPU sharding; PCIe, not NVLink, is the limit |
| A100 / H200, 80โ141 GB VRAM | 256 GB+ | 80โ141 GB | llama3.3:70b, 70B-class models at long context | Datacenter tier; see DeepSeek R1 sizing |
Quantized models (the default in the Ollama library) are what make these numbers work. An unquantized 14B model needs about 28 GB of VRAM; the 9.1 GB Q4 build is what actually fits an ordinary laptop. For the underlying VRAM math, use the LLM VRAM calculator.
B. Which local AI model should I choose?
Shortlist by task. Every model below is an open-weight model distributed through the Ollama library.
General chat and writing
qwen2.5:7b(4.7 GB) โ the best general-purpose starter. 32K context, multilingual.phi4(9.1 GB, 16K context) โ Microsoft's 14B model, stronger at reasoning and logic.
Coding assistance
qwen2.5-coder:7b(4.7 GB) โ code-specialized, same footprint as the general 7B.qwen2.5-coder:32b(20 GB) โ 32B flagship; needs a 24 GB GPU or 32 GB of RAM.
Summarization and document questions
phi4(9.1 GB) โ good at structured answers and long-ish passages within its 16K window.qwen2.5:14b(9.0 GB) โ better fluency than 7B at the same class of cost.
Image generation
FLUX.1 [dev]โ 24 GB VRAM at FP16, ~12 GB at INT4. Image models are not run through Ollama โ use ComfyUI or a Stable Diffusion UI, which load FLUX from a local file.
Why a model that loads can still be unusable
Loading is not the same as running well. Four things catch people out:
- VRAM overflow. If weights plus context exceed VRAM, the runtime falls back to system RAM โ typically 10โ50ร slower. A 14B model (~9 GB) on an 8 GB GPU will load, then crawl.
- Context length. Context grows memory usage on top of the weights. A model rated for 32K context does not mean you get 32K context for free at 8 GB VRAM.
- Quantization. Q4 is the Ollama default and the right call for most people. Higher quantizations (Q5, Q6, Q8) improve quality but grow the download.
- No GPU at all. CPU-only inference is fine for occasional questions. It is not fine for a chatbot you use all day.
C. How do I install AI on Windows?
Step 1: Install Ollama
Ollama is the simplest established runtime for local models on Windows. It handles downloads, quantization, and GPU detection automatically.
- Download the Windows installer from ollama.com/download/windows.
- Run it (~500 MB on disk after install).
- Open a new PowerShell window โ the installer updates your
PATH, and existing windows will not pick that up.
Step 2: Run your first model
# Small, safe first model (1.9 GB download)
ollama run qwen2.5:3b
# Step up when that works (4.7 GB)
ollama run qwen2.5:7b
# Larger, stronger (9.1 GB) โ needs 16 GB of RAM
ollama run phi4
Type your question at the prompt. Type /bye or press Ctrl+C to exit. The first run downloads the model; later runs start instantly.
Step 3: Confirm your GPU is actually being used
Ollama uses NVIDIA CUDA automatically when it detects a compatible GPU and drivers.
# See which models are loaded and on which processor
ollama ps
In the PROCESSOR column, 100% GPU means the model is fully on the GPU. A CPU or partial-GPU value means it is spilling to system RAM.
If you see CPU when you expected GPU:
- Install or update NVIDIA drivers (535 or newer).
- Verify the card reports at least 6 GB of VRAM (
nvidia-smiin the same window). - Confirm the model actually fits โ
ollama psshows the size it loaded.
Step 4: Pick a smaller model if memory is short
If a run fails or is unusably slow, step down rather than fighting the hardware:
# 8 GB RAM: stay at 3B
ollama run qwen2.5:3b # 1.9 GB
# 16 GB RAM: 7B is the sweet spot
ollama run qwen2.5:7b # 4.7 GB
ollama run mistral:7b # 4.4 GB
# 32 GB RAM or 24 GB VRAM
ollama run qwen2.5:32b # 20 GB
Common errors and fixes
| Error | Cause | Fix |
|---|---|---|
out of memory | Model larger than available RAM | Run ollama list to see what you have, then ollama rm <model> and start a smaller one |
CUDA out of memory | Model plus context exceeds VRAM | Use a smaller model, or close other GPU apps |
ollama : The term 'ollama' is not recognized | PowerShell window predates the install | Close PowerShell and open a fresh window |
model '...' not found | Typo or wrong tag | Check the exact tag at ollama.com/library |
| Download stops / disk full | Insufficient disk space | ollama list shows disk usage; ollama rm <model> frees it |
D. macOS and Linux
macOS
brew install ollama
ollama serve # start the local server
# In a second terminal
ollama run qwen2.5:7b
On Apple Silicon (M1โM4), Ollama uses Metal acceleration and performance is comparable to a mid-range NVIDIA GPU. On Intel Macs there is no GPU acceleration โ expect CPU-only speeds. Models are stored in ~/.ollama/models.
Linux
curl -fsSL https://ollama.com/install | sh
ollama serve
# In a second terminal
ollama run qwen2.5:7b
For NVIDIA GPU acceleration, install the driver (535+) and CUDA 12.x first โ Ollama detects them on first run. Models are stored in ~/.ollama/models.
E. How much storage and memory do I need?
Download size (Ollama library, verified 2026-10-11)
| Model | Download | Context | Notes |
|---|---|---|---|
qwen2.5:3b | 1.9 GB | 32K | Fits 8 GB RAM |
qwen2.5:7b | 4.7 GB | 32K | The default recommendation |
mistral:7b | 4.4 GB | 32K | Apache-licensed, function calling |
gemma2:9b | 5.4 GB | 8K | Short context window |
phi4 | 9.1 GB | 16K | Strong reasoning at 14B |
qwen2.5:14b | 9.0 GB | 32K | |
qwen2.5-coder:32b | 20 GB | 32K | Needs 24 GB VRAM or 32 GB RAM |
llama3.3:70b | 43 GB | 128K | Weights alone fit 48 GB class GPUs at short context; long context needs multi-GPU |
Downloads are cached, so each model is fetched once. Models are stored in ~/.ollama/models (C:\Users\<you>\.ollama\models on Windows) and can be deleted individually with ollama rm <model>.
RAM rules of thumb
A practical floor: the model download must fit in RAM and leave room for context and your other applications.
| System RAM | Comfortable class |
|---|---|
| 8 GB | 1โ3B models |
| 16 GB | 7B models; 14B only if nothing else is running |
| 32 GB | Up to 14B comfortably; 32B (20 GB) with short context |
| 64 GB | 32B comfortably; 70B at INT4 is still multi-GPU territory |
VRAM
If the model plus context fits in VRAM, you get GPU speed. If it does not, you get CPU speed with a GPU sitting idle. The exact numbers depend on quantization, context length and KV-cache โ OpenGPU Radar's LLM VRAM calculator computes them, and the 8/16/24/48/80 GB tier table shows which models fit where.
F. When should you upgrade โ or use the cloud?
Upgrade RAM first if you are on 8 GB and running CPU-only. Going to 16 GB is the cheapest way to unlock 7B models, which is the point where local AI becomes genuinely useful.
Add a dedicated GPU if you do not have one. Even an entry-level card in the 8 GB VRAM class gives a large multiple over CPU inference โ for the models that fit. See current GPU pricing.
Go to 24 GB VRAM when you regularly want 14Bโ32B models with real context. This is the practical enthusiast tier.
Use the cloud or an API when the model you want is larger than anything you will own:
| Situation | Realistic choice |
|---|---|
| Occasional questions, prototyping, privacy-sensitive data | Local, on what you already own |
| Regular use of 32B-class models | 24 GB local GPU |
| 70B+ models, long context, high concurrency | Cloud GPU or an API โ see DeepSeek R1 for the upper end |
| Bursty image or video generation | Cloud โ local GPU cost is hard to justify for bursty demand |
The honest upgrade path: start free on existing hardware, add RAM, add a GPU, and only then consider renting. Most people stop at step two.
Troubleshooting cheat sheet
| Problem | Fix |
|---|---|
| Model will not load | Step down a size: ollama run qwen2.5:3b instead of qwen2.5:7b |
| Very slow replies | ollama ps โ if PROCESSOR is not 100% GPU, you are on CPU |
CUDA out of memory | Smaller model, or shorter context, or close other GPU apps |
ollama not recognized | Open a new PowerShell / terminal window |
| Disk full | ollama list to check usage, ollama rm <model> to free space |
| GPU not detected | Update NVIDIA drivers; confirm nvidia-smi works |
Next steps
- LLM VRAM calculator โ exact memory for your model, precision and context
- VRAM tier breakdown โ which models fit which GPUs
- Model directory โ verified specs for every model we track
- Cheapest GPUs โ what the hardware costs when you are ready to upgrade
Methodology
- Model download sizes and context windows are the published values on the Ollama library for each tag, verified 2026-10-11.
- Installation commands are taken from Ollama's official Windows, macOS and Linux install instructions, verified 2026-10-11.
- VRAM requirements for the larger models use OpenGPU Radar's canonical engine via the VRAM calculator; see the methodology page.
- Throughput figures are estimates for a single interactive stream at short context. They are not benchmark results, and real-world numbers vary by CPU, GPU, context length and runtime version.
- Pricing references link to OpenGPU Radar's tracked GPU pricing. No prices are hardcoded in this article.