The closest open-weight model to the proprietary frontier. A large Mixture-of-Experts model with a 1M-token context, excelling at reasoning, coding and agentic tasks — at a tiny fraction of the cost of closed APIs.
View DeepSeek models →Gemini brings strong multimodal skills and very large context windows, tightly woven into Google's ecosystem — but it is proprietary and cloud-only. If you want multimodal, long-context AI you can self-host and control, open-weight models now deliver. Here are the best open-source alternatives to Gemini, what each is best at, and how to run one yourself.
Updated regularly · curated by OpenSourceAI.tech
Run models on your own machine or servers so your prompts and data never leave your control — no third party sees them.
Run locally for free, or use a hosted option that is often far cheaper per token, with no monthly subscription.
Fine-tune on your own data, change behaviour, and integrate the tool deeply into your own products and workflows.
The software is yours to keep. No surprise deprecations, no forced upgrades, no sudden price hikes pulling the rug out.
These are open-weight models you can download, self-host, and use commercially (check each license). They run from most capable to most lightweight — pick based on your hardware and needs.
The closest open-weight model to the proprietary frontier. A large Mixture-of-Experts model with a 1M-token context, excelling at reasoning, coding and agentic tasks — at a tiny fraction of the cost of closed APIs.
View DeepSeek models →The most widely-adopted open LLM family, with by far the largest ecosystem of tools, fine-tunes and guides. A reliable general-purpose assistant that runs well locally in its smaller sizes. If unsure where to start, start here.
View Llama models →A top-tier family with outstanding multilingual ability, strong coding, and excellent quality across every size. Frequent releases keep it cutting-edge, and permissive licensing on most variants makes it easy to build on.
View Qwen models →Efficient, European-built models that consistently punch above their weight. A great balance of speed, quality and openness with strong multilingual support — appealing if you want to keep your stack inside the EU.
View Mistral models →A reasoning-focused family that shines at long-horizon, project-level coding and autonomous agent workflows — able to work continuously on a task rather than just answering single questions.
View GLM models →Built for very long context and end-to-end coding, with multimodal input. Handles large codebases and long documents in a single pass, making it well suited to agentic, multi-step work over big inputs.
View Kimi models →Google's open models offer some of the best quality-for-size available, with native multimodal input — and they are among the easiest frontier-adjacent models to run on a single GPU or a Mac.
View Gemma models →OpenAI's own open-weight models — a familiar option if you like ChatGPT's style but want something self-hostable and extremely cheap to run. The smaller variant runs on consumer hardware.
View gpt-oss models →Open-source doesn't always mean you run it yourself — many of these models are also available through low-cost hosted APIs. Here is how today's most-used open models compare, pulled live from our leaderboard.
Running a model on your own machine means total privacy and zero per-token cost. These tools make it straightforward — no machine-learning expertise required.
The easiest way to start. Install it, then pull and run a model with a single command on macOS, Windows or Linux.
A friendly desktop app with a graphical model browser and chat interface — ideal if you would rather avoid the command line.
Run quantized models efficiently on modest hardware, including laptops without a dedicated GPU.
For production serving — high-throughput inference engines used to host open models at scale behind an API.
Hardware in brief: small models (≈7–12B parameters) run on a modern laptop or a consumer GPU. Mid-size models want a 16–24GB GPU. The largest Mixture-of-Experts models need a workstation — for those, a cheap hosted API is often the practical choice.
Yes. Open models such as Gemma (from Google itself), Qwen and Llama are free to run locally, with no subscription.
Gemma, Qwen's vision models and Kimi handle image (and in some cases video) input, making them strong multimodal alternatives.
Yes — context windows of 1M tokens are now available on several open models.
Yes. Smaller models run on a laptop or consumer GPU; Gemma in particular is easy to run locally.
Self-hosted, yes — your data stays entirely on your own hardware.
Compare 150+ open-weight models by price, context and popularity — updated daily, with rankings that track how the field shifts over time.
Open the leaderboard →