HomeModelsLlama 3.1 405B
MetaPower Tier

Llama 3.1 405B

Meta · Released July 2024

Input Price

$2.70

per 1M tokens

Output Price

$2.70

per 1M tokens

Context Window

128K

tokens

Speed Score

65

out of 100

Benchmarks

Avg 87/100
Coding88/100
Reasoning88/100
Extraction87/100
Creative86/100
VisionN/A

Pricing

Input

Prompt / context tokens

$2.70/ 1M

Output

Generated tokens

$2.70/ 1M

Best For

reasoningcodinganalysis

Llama 3.1 405B in context

On a typical 70% input / 30% output mix, Llama 3.1 405B blends to $2.70 per million tokens. That puts it above 47% of the 32 models tracked here, and 1 other models in the same flagship tier undercut it. Tier is a capability label, not a price band — the spread inside one tier is often wider than the gap between tiers, which is why picking by tier alone tends to overspend.

Input runs $2.70 per million and output $2.70 — a 1.0x ratio. That ratio matters more than either number on its own: a summarisation or classification workload reads far more than it writes and will track the input price, while a code-generation or long-form writing workload inverts that and will track the output price. Compare models on the ratio your own traffic actually has, not on the input price alone.

The benchmark profile is unusually flat — coding leads at 88/100 but only 2 points separate the strongest category from the weakest. That makes Llama 3.1 405B a safe default for mixed workloads where you cannot predict in advance which capability a request will lean on, and a poor choice if you need a specialist.

The context window is 128K tokens and the throughput score is 65/100. Context only earns its keep if you fill it, and filling it is also what makes a request expensive — a large window is an option, not a discount. For a flagship-tier model, judge the speed score against what you are doing: it decides everything for interactive chat and almost nothing for overnight batch work.

Worth checking before you commit: DeepSeek V3 scores at least as well on combined coding and reasoning and blends to $0.32 per million against Llama 3.1 405B's $2.70. That does not make Llama 3.1 405B the wrong choice — context window, latency, vendor lock-in, and specific capabilities all sit outside a benchmark score — but it does mean the price difference needs a reason behind it.

The closest comparison is Claude 3 Opus from Anthropic, at $33.00 per million against $2.70. 6 models in this tier score higher on combined coding and reasoning, which is the useful framing: the question is rarely whether a model is good, it is whether it is the cheapest thing that is good enough for the specific work you are sending it.

What it costs in production

At 10M + 2M tokens a month — a realistic mid-size production workload — Llama 3.1 405B runs about $32.40. There is no batch discount on this model, so that figure is the floor — the usual lever of shifting background work to a cheaper asynchronous tier is not available here.

Price History

No price changes since launch

Strengths

  • Largest open-source model — competitive with GPT-4
  • Fully self-hostable — no per-token cost on own infra
  • Active community fine-tuning ecosystem
  • Strong coding and reasoning benchmarks

Avoid For

  • Vision tasks (text-only)
  • Teams without substantial GPU infrastructure
  • Low-latency real-time requirements