Meta

MLlama 4 ScoutOPEN

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B.

Context window
1.3M
tokens
Input price
$0.1
per M tokens
Output price
$0.3
per M tokens
Provider
Meta

Prices update automatically — checked hourly against provider list prices.

See model comparisons → Compare all model prices

Benchmarks & performance

Independent benchmark scores for Llama 4 Scout, measured by Artificial Analysis. Higher is better.

5.5×
More intelligence per dollar than Claude Opus 5 (Fast).
For the same budget, Llama 4 Scout delivers 5.5 times more capability. Open weights also mean you can self-host it and pay nothing per token.
Intelligence index10
Coding index8.2
Math index14
GPQA58.7%
MMLU-Pro75.2%
Humanity's Last Exam4.3%
Long Context Reasoning25.8%
LiveCodeBench29.9%
SciCode17%
MATH-50084.4%
AIME28.3%
AIME 202514%
IFBench39.5%
τ²-Bench15.5%
τ-Bench Banking3.3%
Terminal-Bench3.7%
Terminal-Bench Hard1.5%
⚡ Speed83.3 tokens/sec
⏱ Latency0.61s to first token
💰 Blended price$0.3 / 1M tokens
📈 Value33.3 intelligence points per $
vs. models measured heretop 81%
Scores higher than 19% of the 239 models measured by Artificial Analysis and tracked here.
Models at this level cost $0.2 per 1M tokens (median of 21) — this one costs $0.15.
Cheaper and better on this index: Hy3 preview · Gemma 4 26B A4B · gpt-oss-120b and 10 more
Benchmark data by Artificial Analysis

About this model

Llama 4 Scout is an open-weight AI model by Meta. You can download and self-host it for free; the prices below are hosted-API list prices, tracked hourly, for when you prefer convenience over self-hosting.

Frequently asked questions

What is Llama 4 Scout?

Llama 4 Scout is an AI language model from Meta. It is open-weight: you can download it and run it on your own hardware, for free. It scores 10 on the Artificial Analysis intelligence index.

Is Llama 4 Scout free?

The weights are free and open — you can self-host Llama 4 Scout and pay nothing per token. If you prefer a hosted API, list prices are $0.1 per million input tokens and $0.3 per million output tokens.

What is Llama 4 Scout good at?

Independent benchmarks from Artificial Analysis give it GPQA 58.7%, MMLU-Pro 75.2%, Humanity's Last Exam 4.3%, Long Context Reasoning 25.8%, LiveCodeBench 29.9%, SciCode 17%, MATH-500 84.4%, AIME 28.3%, AIME 2025 14%, IFBench 39.5%, τ²-Bench 15.5%, τ-Bench Banking 3.3%, Terminal-Bench 3.7%, Terminal-Bench Hard 1.5%. It is particularly used for code generation.

How fast is Llama 4 Scout?

It generates about 83.3 tokens per second, with a median 0.61s delay before the first token. Measured independently by Artificial Analysis.

Can I self-host Llama 4 Scout?

Yes. Llama 4 Scout has open weights, so you can download it and run it on your own GPU or server with tools like Ollama, vLLM or llama.cpp — with no per-token cost.

Related models

Muse Spark 1.1MetaLlama 3.1 8B InstructMetaLlama 4 MaverickMetaLlama 3.3 70B InstructMetaLlama 3.1 70B InstructMetaLlama Guard 4 12BMeta