HomeTemplates › Llama 4 Scout
🧠 Self-hosted LLM

Llama 4 Scout with long context

Meta Llama 4 Scout — MoE with ultra-long context

From
loading…
cheapest machine that meets this template
Deploy in one click →
Video memory
200 GB
Setup
~15 min
Access
port 8000
Billing
hourly, in BRL

Scout is Meta's variant built for very long context: whole documents, repositories, complete histories. It needs access approved on the model repository before the first download.

Meta's Llama 4 Scout served via vLLM with an OpenAI-compatible API — MoE with a huge context window (up to ~10M tokens). IMPORTANT: GATED on Hugging Face → set HF_TOKEN with approved access. Large model: prefer multiple H100/H200 GPUs; long context uses a lot of VRAM.

What it is for

How to deploy

  1. Create your account and add balance (card or Pix, no subscription).
  2. In the console, pick the Llama 4 Scout template and a machine — the console hides the ones that do not meet the requirement.
  3. In about 15 minutes the setup finishes and the access address shows up in the panel, on port 8000.

Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.

FAQ

Do I need to request access?

Yes, it is a gated model. Accept the terms on your repository account and provide the token in the template.

How much video memory?

Around 200 GB, meaning a block of large cards.

Can I use maximum context right away?

Very long context costs extra memory. Start smaller and raise it as the card allows.

Run Llama 4 Scout today
You only pay for the hours the machine is running.
Deploy in one click →

Related templates

vLLM
OpenAI-compatible API for Llama, Qwen, Mistral and more
TGI (HuggingFace)
HuggingFace Text Generation Inference — vLLM alternative
LiteLLM Proxy
1 OpenAI endpoint routing to 100+ providers (cloud + local)
Ollama
Run DeepSeek, Qwen3, Llama and Mistral with one command