Developer Tools2026-09-2411 min read
Free AI Models in Cursor, Windsurf & Local Editors: The Complete Setup Guide
Configure Cursor, Windsurf, and local AI editors with free LLM backends. Step-by-step instructions for Groq, SambaNova, and Google AI Studio integration.
Table of Contents
- 1. Cursor BYOK Configuration
- 2. Windsurf Free Model Integration
- 3. Local Model Integration with Ollama
- 4. Free Provider Comparison
Cursor BYOK SetupWindsurf ConfigurationLocal Model IntegrationFree Provider Comparison
Cursor BYOK Configuration
Cursor supports custom OpenAI-compatible API endpoints through its Settings โ Models panel. Unlike VS Code Continue (which uses a JSON config file), Cursor provides a GUI for adding custom models. To configure free models:
1. Open Cursor Settings (Ctrl+,)
2. Navigate to Models โ Add Custom Model
3. Set Model Name: 'Llama 3.3 70B (Groq)'
4. Set Override OpenAI Base URL: https://api.groq.com/openai/v1
5. Set API Key: gsk_YOUR_GROQ_KEY
6. Click 'Verify' โ should show green checkmark
7. Set as Default Model for Chat
Cursor also supports environment variable configuration: set OPENAI_BASE_URL and OPENAI_API_KEY in your shell profile, and Cursor will automatically detect them. This is useful for switching between providers without reconfiguring Cursor's settings panel.
Windsurf Free Model Integration
Windsurf (formerly Codeium) supports custom API endpoints through its Cascade configuration. To add free models:
1. Open Windsurf Settings โ AI Configuration
2. Add Custom Provider โ OpenAI Compatible
3. Enter Base URL: https://api.sambanova.ai/v1
4. API Key: sb_YOUR_SAMBANOVA_KEY
5. Model: qwen/qwen2.5-72b
Windsurf's Cascade feature (multi-file editing) works best with SambaNova's Qwen 2.5 72B, which handles complex code generation with good instruction following. For autocomplete, Windsurf's built-in Codeium model is free and works well for simple completions.
Local Model Integration with Ollama
For fully offline, zero-API-key coding assistance, use Ollama to run models locally. Ollama supports Llama 3.3, Qwen 2.5, DeepSeek, and hundreds of other models. Configuration for each tool:
Ollama + Continue: Set baseUrl to http://localhost:11434/v1, use any pulled model.
Ollama + Cursor: Add custom model with baseUrl http://localhost:11434/v1.
Ollama + VS Code Terminal: Use ollama run llama3.3 directly in terminal.
Local models are slower (10-50 tok/s vs 300+ tok/s on Groq) but unlimited โ no rate limits, no API keys, no internet required. Best for: learning, prototyping, working with sensitive code that cannot leave your machine.
Free Provider Comparison
Provider Selection Matrix for Coding:
Groq (Recommended for Speed): 330+ tok/s on Llama 3.3 70B, 30 RPM, 14,400 req/day. Best for: real-time autocomplete, fast iteration. Limitation: 30 RPM can be hit during intensive coding sessions.
SambaNova (Recommended for Quality): ~20 RPM, daily quota. Best for: complex code generation, multi-file refactoring, reasoning tasks. Qwen 2.5 Coder 32B is specifically trained for code.
Google AI Studio (Recommended for Context): 1M token context, 15 RPM. Best for: understanding large codebases, working with long files, analyzing entire repos. Gemini 2.0 Flash is fast and handles massive context windows.
OpenRouter Free (Recommended for Fallback): 3 RPM for :free models. Best as: secondary provider when primary hits rate limits. Routes to multiple free model endpoints automatically.
Optimal Strategy: Primary = Groq (speed), Secondary = SambaNova (quality), Tertiary = Google AI Studio (context). Configure your tool to cascade through providers when rate limits are hit.
OR
OpenGPU Radar Systems Engineering Team
Independent compute telemetry and infrastructure analysis. Not affiliated with NVIDIA, cloud providers, or hardware vendors.