Llama 4 Scout

Metatextimage

Open-weights MoE (17B active, 16 experts) with an industry-leading 10M token context window. Host-dependent pricing. Open weights — API prices vary by hosting provider.

Llama 4 Scout strengths

  • Industry-leading 10M context
  • Open weights / self-hostable
  • Very low cost
  • Natively multimodal

Pricing & context

Context window10M tokens
Input price /1M≈$0.08
Output price /1M≈$0.30
Modalitiestext, image

Cost guide: a typical call of about 10K input + 2K output tokens costs roughly $0.001 at list prices. Worth modelling against cheaper tiers before committing high-volume traffic.

When to choose Llama 4 Scout

Llama 4 Scout is best for extreme long-context tasks (whole codebases, large document sets) on an open, cheap model. If your workload is more cost-sensitive, weigh it against gpt-oss-120b (≈$0.03 input /1M) first.

Llama 4 Scout FAQ

How much does Llama 4 Scout cost?

Llama 4 Scout is priced at ≈$0.08 per 1M input tokens and ≈$0.30 per 1M output tokens (public API list price), with a 10M tokens context window. A typical call of about 10K input and 2K output tokens costs roughly $0.001.

What is Llama 4 Scout best for?

Llama 4 Scout by Meta is best for extreme long-context tasks (whole codebases, large document sets) on an open, cheap model.

How does Llama 4 Scout pricing compare to Gemini 2.5 Flash-Lite?

Llama 4 Scout input costs ≈$0.08 per 1M tokens versus $0.10 for Gemini 2.5 Flash-Lite, roughly 1.3x less expensive on input. Output is ≈$0.30 vs $0.40.

Is Llama 4 Scout multimodal?

Llama 4 Scout supports text, image.

Tools that use Llama 4 Scout

Other models

All models →
01Claude Fable 5Anthropic$10.00
02Claude Opus 5Anthropic$5.00
03GPT-5.6 SolOpenAI$5.00
04GPT-5.5OpenAI$5.00