Llama 3.3 70B Instruct FP8 Fast is Meta's language model with a 24K context window, starting at $0.293 / 1M input and $2.25 / 1M output. Llama 3.3 70B instruction-tuned model quantized to FP8 precision and further optimized for throughput-focused fast inference deployments.
Specifications
Canonical IDmeta-llama-3-3-70b-instruct-fp8-fast
TypeLanguage
StatusActive
CreatorMetaMeta
Providers
Context Window24K tokens
Input ModalitiesText
Output ModalitiesText

Capabilities

Input1/5
Text✓
Image·
Audio·
Video·
PDF·
Output1/5
Text✓
Image·
Audio·
Video·
Embedding·
Capabilities1/13
Reasoning·
Adaptive Reasoning·
Function Calling✓
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Cloudflare Workers AI logo
Cloudflare Workers AI
@cf/meta/llama-3.3-70b-instruct-fp8-fast
$0.293$2.25

Cost Calculator

US Dollar ($)
Preset:

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Llama 3.3 70B Instruct131K$0.100$0.200Deprecated
Llama 3.2 3B Instruct131K$0.015$0.020Deprecated
Llama 3.2 1B Instruct131K$0.020$0.020Deprecated
Llama 3.1 405B Instruct131K$0.120$0.300Deprecated
Llama 3.1 8B Instruct200K$0.020$0.030Deprecated
Llama 3.1 70B Instruct131K$0.100$0.100Deprecated
Llama 3.1 70B128K$0.360$0.360Available
Llama 3.1 8B131K$0.030$0.050Available
Llama 3 70B Instruct131K$0.120$0.300Deprecated
Llama 3 8B Instruct32K$0.030$0.040Deprecated
Llama 3.3 70B Instruct FP8 Fast—24K$0.293$2.25Current

Model IDs

@cf/meta/llama-3.3-70b-instruct-fp8-fast
cloudflare/@cf/meta/llama-3.3-70b-instruct-fp8-fast
meta-llama-3-3-70b-instruct-fp8-fast