The best open-weight, self-hostable alternatives to GPT-3.5 Turbo Instruct in 2026 — compared on price, context window and capabilities. Run them locally and cut API costs.
Refreshed from live data · olud.ai
GPT-3.5 Turbo Instruct is a proprietary, API-only model. These open-weight models can be self-hosted, run offline and used at a fraction of the cost — here's how the top ones stack up.
Within 0.5 pts of GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 10.0× cheaper per million output tokens
Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling. Featuring a...
Reka Flash 3 vs GPT-3.5 Turbo Instruct →Within 1.3 pts of GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 14.3× cheaper per million output tokens
[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...
Phi 4 vs GPT-3.5 Turbo Instruct →Within 2.8 pts of GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 4.0× cheaper per million output tokens
Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruction-following scenarios. Its capacity enables strong performance on demanding evaluation tasks and...
Olmo 3 32B Think vs GPT-3.5 Turbo Instruct →Live ranking of open-weight models with pricing, context windows and capabilities.
Open the leaderboard →