Run Llama 4 Locally: The 67GB Bill Behind the 17B Label

Screenprint illustration of hands holding a llama figurine whose belly window reveals dozens of tiny llamas packed inside, with a laptop and server racks behind

10 min read · 2,617 words

This article was written with AI. It was drafted from the sources it cites and checked against the full text of those sources before publishing. How we make articles

Ollama’s page lists the default Llama 4 Scout tag, digest bf31604e25c2, at 67GB (Ollama), not the lightweight pull that Meta’s “17 billion active parameter” description might suggest (Meta). Anyone setting out to run Llama 4 locally with those two numbers in hand is about to make a storage-plan mistake. Everything below closes the gap between spec and disk: the correct tag for your hardware, a served model verified on the GPU, a reproducible tokens-per-second benchmark, an OpenAI-compatible endpoint with vision support, and a residency-flow ledger that prices local serving against hosted API billing using numbers you measured yourself.

What You’ll Build

By the end you will have Llama 4 Scout (or Maverick, if the machine can hold it) answering requests on http://localhost:11434, residency confirmed in ollama ps, a Python standard-library benchmark logging tokens per second from the model’s own eval metrics, a /v1/chat/completions endpoint any OpenAI SDK can call, and a ledger calculation that says when owning the weights stops paying. Everything runs from a terminal. No cloud account, no GPU cluster login, no mystery.

Prerequisites

  • Ollama, installed in Step 1 from the official installers for macOS, Windows, Linux, or Docker (GitHub).
  • Hardware matched to a tag, not to a spec. Scout’s default Q4_K_M build is a 67GB file (Ollama), and Meta’s stated target for it is a single NVIDIA H100 with Int4 quantization (Meta).
  • Python 3 with standard library only, so there is nothing to pip-install and nothing to version-pin.
  • curl for endpoint tests.
  • At least 67GB of free disk for Scout’s default build plus a few GB of headroom, or 245GB plus headroom if Maverick is the goal, per the registry’s listed sizes (Ollama).

One clause of background explains every size in this tutorial. So why does a model labeled 17B weigh 67GB? Scout is a mixture-of-experts (MoE) model that Meta’s announcement describes as “a 17 billion active parameter model with 16 experts” (Meta). Active means consulted per token. Resident means sitting in memory. All 16 experts stay resident whether or not they fire, and the download bills you for residency.

Step 1: Install Ollama and Confirm the Server

Official installers exist for macOS, Linux, Windows, and Docker (GitHub).

# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh

# Linux, pinned to the current release tag for reproducible setups
export OLLAMA_VERSION="$(curl -fsSL https://api.github.com/repos/ollama/ollama/releases/latest | sed -n 's/.*"tag_name": *"v\{0,1\}\([^"]*\)".*/\1/p')"
curl -fsSL https://ollama.com/install.sh | sh
# Windows (PowerShell)
irm https://ollama.com/install.ps1 | iex
# Docker, with NVIDIA GPU passthrough
docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

Pinning matters more than it looks. Version drift between machines is the silent killer of reproducible benchmarks, so pin it: Ollama’s Linux install script accepts an OLLAMA_VERSION environment variable to install a specific version (Ollama Docs). Strip the leading v from the GitHub tag, as the command above does. In our test, passing the raw tag (v0.34.4) instead made the download fail with a 404 after the script had already removed the previous install’s libraries, leaving an Ollama that could only use the CPU. This excerpt is from that failed run; the same install with 0.34.4 succeeded:

(excerpt, progress bars removed)
>>> Cleaning up old version at /usr/local/lib/ollama
>>> Installing ollama to /usr/local
>>> Downloading ollama-linux-amd64.tgz
curl: (22) The requested URL returned error: 404

For the Docker route, install the NVIDIA Container Toolkit first (Ollama Docs). In our test, docker run --gpus=all on a machine without the toolkit failed with this error, and the same command worked once the toolkit was installed (our rented machine also needed one reboot, for a driver update its image had applied on its own):

docker: Error response from daemon: could not select device driver "" with capabilities: [[gpu]]

AMD GPU owners also download and extract the additional ROCm package (Ollama Docs).

Verify it works:

ollama --version
curl http://localhost:11434/

Success looks like a printed version string, the check Ollama’s Linux docs use to verify that Ollama is running (Ollama Docs), with the API answering on localhost:11434 (GitHub).

Step 2: Pick the Tag Your Machine Can Actually Hold

Every size below is what the Ollama registry lists per tag, not an estimate (Ollama).

Tag Quantization Download size Advertised context (up to)
llama4:scout (= llama4:latest) q4_K_M 67GB 10M tokens, text+image
Scout q8_0 q8_0 117GB 10M tokens
Scout fp16 fp16 217GB 10M tokens
llama4:maverick q4_K_M 245GB 1M tokens
Maverick q8_0 q8_0 428GB 1M tokens
Maverick fp16 fp16 803GB 1M tokens

Read that table as a hardware bill, and start with Meta’s own targets. Meta’s announcement states Scout “fits on a single H100 GPU (with Int4 quantization)” and Maverick “fits on a single H100 host” (Meta), and the registry lists Scout’s default build at 67GB and Maverick’s at 245GB (Ollama). Those are targets for Meta’s own builds, not the only way to run Scout: InsiderLLM reports that Unsloth’s 1.78-bit quantization fits in 24GB of VRAM at about 20 tokens per second (InsiderLLM), at the cost of much heavier quantization. The default Q4_K_M tag this tutorial uses is a 67GB download (Ollama), so check any guide’s smaller figure against the tag it actually runs.

Verdict on tag choice: default q4_K_M unless your own testing says otherwise. Whether any fidelity gain from q8_0 justifies its 117GB download against q4_K_M’s 67GB (Ollama) is a question to answer on your machine, not in a spec sheet, so run the Step 4 harness against both tags on identical prompts and read the diffs yourself. If you cannot point to what improved, take the smaller tag.

Step 3: Serve Llama 4 on Your Own Machine

ollama pull llama4:scout
ollama run llama4:scout "Explain mixture-of-experts routing in two sentences."
ollama ps
ollama show llama4:scout

Success looks like this: the pull completes, and the run command prints an answer to the prompt. Then ollama ps lists the loaded model with its size and a Processor column, where 100% GPU means the model was loaded entirely into the GPU, 100% CPU entirely into system memory, and a split such as 48%/52% CPU/GPU partly into each (Ollama Docs). Scout’s registry entry accepts images as well as text, which becomes useful in Step 5 (Ollama).

Three verification moves turn a working serve into a documented one.

Match the digest. The default Scout build carries digest bf31604e25c2 and default Maverick carries b4832b93e292 (Ollama). Check that your pulled model’s ID matches before comparing results across machines.

Check the context. ollama show reports the model’s maximum, and ollama ps reports the context the loaded model actually runs with. In our run they differed by a factor of 40 (10,485,760 ÷ 262,144):

$ ollama show llama4:scout        (excerpt)
    context length      10485760
$ ollama ps
NAME            ID              SIZE     PROCESSOR    CONTEXT    UNTIL
llama4:scout    bf31604e25c2    66 GB    100% GPU     262144     4 minutes from now

Size prompts to the ollama ps context. ollama ps lists the models currently loaded into memory (Ollama Docs), so its SIZE column, in the output above, describes the loaded model, not the 67GB download the registry lists.

Build the unload habit. ollama stop llama4:scout releases the weights. Leaving the model resident keeps its memory in use.

Step 4: Benchmark Tokens per Second Without External Tools

What does the model actually deliver on your silicon? Quick check first: as SitePoint notes, Ollama’s --verbose flag outputs token-per-second statistics after each response (SitePoint).

ollama run llama4:scout --verbose "Write a 150-word explanation of KV caching."

For a loggable, repeatable number, call the native generate endpoint. Its response includes eval_count (tokens generated) and eval_duration (generation time, in nanoseconds) (Ollama API), which the script below turns into tokens per second.

#!/usr/bin/env python3
"""bench.py — tokens/sec from Ollama's own eval metrics. Stdlib only."""
import json
import time
import urllib.request

HOST = "http://localhost:11434/api/generate"
PROMPT = "Write a 200-word summary of how mixture-of-experts routing works."
RUNS = 3

def bench(model: str) -> None:
    payload = json.dumps({"model": model, "prompt": PROMPT, "stream": False}).encode()
    for i in range(1, RUNS + 1):
        req = urllib.request.Request(
            HOST, data=payload, headers={"Content-Type": "application/json"})
        start = time.perf_counter()
        with urllib.request.urlopen(req, timeout=1800) as resp:
            out = json.load(resp)
        wall = time.perf_counter() - start
        toks = out.get("eval_count", 0)
        gen = out.get("eval_duration", 1) / 1e9
        print(f"run {i}: {toks} tok | {toks/gen:8.1f} tok/s (eval) | wall {wall:6.1f}s")

if __name__ == "__main__":
    bench("llama4:scout")

Success looks like three printed lines, each with a token count, an eval-based rate, and a wall time. Interpret before celebrating: when wall time dwarfs eval time, the model loaded mid-request or spilled to CPU. Warm the model with one throwaway request, then benchmark, because first-request numbers are load-time measurements wearing a throughput costume.

Throughput numbers attached to unnamed hardware say little about yours; numbers attached to your digest on your machine are data. Log every candidate like this:

Host Tag ID (ollama ps) tok/s (eval) Resident (ollama ps)
your machine llama4:scout bf31604e25c2 from bench.py from ollama ps

Run the identical prompt on every tag and machine. Same prompt, same RUNS, comparable rows. Report the eval-based rate, not wall-clock, and keep the output; Step 6 needs it.

Step 5: Expose an OpenAI-Compatible Chat and Vision Endpoint

Ollama’s compatibility docs say it “supports a subset of the OpenAI API”, including /v1/chat/completions, so existing OpenAI SDK code moves over with a base-URL change (Ollama Docs).

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama4:scout",
    "messages": [{"role": "user", "content": "Name three mixture-of-experts models."}]
  }'

Success is a JSON body whose choices array carries a message. Scout’s registry entry covers image input, and Llama 4 has been tested for image understanding up to 5 input images (Ollama). The vision.py script below sends an image to the same /v1/chat/completions endpoint:

#!/usr/bin/env python3
"""vision.py — send an image through the OpenAI-compatible endpoint."""
import base64, json, urllib.request

with open("chart.png", "rb") as f:
    b64 = base64.b64encode(f.read()).decode()

payload = json.dumps({
    "model": "llama4:scout",
    "messages": [{"role": "user", "content": [
        {"type": "text", "text": "Describe this image in one sentence."},
        {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}},
    ]},
]}).encode()

req = urllib.request.Request(
    "http://localhost:11434/v1/chat/completions",
    data=payload, headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=1800) as resp:
    print(json.load(resp)["choices"][0]["message"]["content"])

In our run, with a 64×32 test image whose left half was red and right half blue, the script printed:

The image is a simple graphic consisting of two solid-colored vertical rectangles, one red on the left and one blue on the right, with a clear and sharp border between them.

Success is a sentence that could only come from seeing your specific image. If the description is generic, check the file path and the base64 payload first. Apple Silicon users can run Scout through Ollama’s MLX backend, which SitePoint’s guide walks through (SitePoint).

Step 6: Run the Residency-Flow Ledger

This article treats cost as a duty-cycle question: hardware bills for residency (hours the weights sit in memory, and for a mixture-of-experts model all of them sit there even though only a subset is active per token (Meta AI)) while APIs bill for flow (tokens actually consumed). Compute your own ledger for your usage shape, line by line:

measured_tok_s = Step 4 output, eval-based
monthly_tokens = your real usage, from app logs
generation_hours = monthly_tokens / measured_tok_s / 3600
duty_cycle = generation_hours / 730
power_cost_month = wall_watts * 730 * price_per_kWh / 1000
hardware_month = purchase_price / planned_months_of_service
residency_bill = hardware_month + power_cost_month
flow_bill = monthly_tokens / 1_000_000 * price_per_million
break_even_tokens = residency_bill * 1_000_000 / price_per_million

For the API side, take price_per_million from your hosted Llama 4 provider’s current pricing page. Look the rate up when you run the ledger rather than relying on a figure printed here. Work through it in order, starting with hours.

Hours. Suppose Step 4 reports 40 tok/s and application logs show 10M tokens per month. Then generation_hours = 10,000,000 / 40 / 3600 ≈ 69.4.

Duty cycle. Those 69.4 hours are 69.4 / 730 ≈ 9.5% of the month; for the rest of it, the weights sit resident and generate nothing.

Residency against flow. residency_bill is what owning costs each month whether or not a token is generated; flow_bill is what the same tokens would cost from an API at your looked-up rate. break_even_tokens is the monthly volume at which the two are equal: above it, owning is cheaper; below it, renting is.

Whether that threshold sits above or below your usage is now a division rather than an argument, and Maverick changes only the inputs: on this article’s reading, the larger machine it will likely need means a higher wall_watts and purchase_price, which the ledger absorbs.

Common Pitfalls

  1. The llama4:latest ambush. It carries the same digest, bf31604e25c2, and the same 67GB size as the default Scout build (Ollama), not a lightweight one.
  2. Silent CPU offload. If ollama ps shows any CPU percentage in the Processor column, part of the model was loaded into system memory rather than the GPU (Ollama Docs). Re-check the Step 2 table and pick the tag that fits.
  3. Docker without GPU passthrough. Omitting --gpus=all on NVIDIA, or skipping the ROCm package on AMD, leaves the container without GPU access. Ollama’s Linux docs cover the CUDA and ROCm install paths (Ollama Docs).
  4. Benchmarking a cold model. ollama run --verbose reports load duration separately from the eval rate; send one warm-up request before timing, so the log does not record load time as throughput.
  5. Trusting the advertised context window. The registry lists a 10M-token window for Scout (Ollama), and ollama show repeats it; in our run, ollama ps showed the loaded model’s CONTEXT as 262,144, a fortieth of that. Size prompts to the ollama ps figure.

When This Approach Is Wrong

When does renting beat owning outright? Steelman the hosted side honestly: zero ops, zero 67GB residency, elasticity on demand, and no electricity line charging while you sleep. At a looked-up price of P dollars per million tokens, spiky usage of T tokens costs P × T / 1M per month with nothing resident at 3 a.m. Hardware frames the fork: InsiderLLM’s guide says that on a single GPU under 24GB you should skip Llama 4 entirely (InsiderLLM).

Response, in one sentence: the Step 6 ledger decides it, and when generation hours per month land in single digits, the honest answer is to rent. Owning weights to generate two hours a month is a hobby with a power bill attached, not infrastructure. Local is right when the arithmetic says so, and the arithmetic takes inputs you already have.

What’s Next

  • Repeat Step 4 against llama4:maverick, whose default tag is a 245GB download (Ollama), on hardware that can hold it, log a second row in the residency-flow ledger, and compare both models on your actual prompts rather than public leaderboards.
  • Point an existing OpenAI SDK application at the localhost base URL from Step 5; Ollama supports a subset of the OpenAI API, so test the endpoints your application uses (Ollama Docs).
  • Keep bench.py in the repository. Log a before-and-after row for every hardware or tag change.

References

Leave a Comment

Scroll to Top