Llama 2 70B Chat is Meta's language model with a 4K context window and up to 4K output tokens, available from 6 providers, starting at $0.5 / 1M input and $0.9 / 1M output. A 70B Llama 2 model fine-tuned with RLHF for dialogue, providing high-quality conversational responses at the largest Llama 2 scale.
Specifications
Canonical IDmeta-llama-2-70b-chat
TypeLanguage
StatusActive
CreatorMetaMeta
Providers
Context Window4K tokens
Max Output4K tokens
Input ModalitiesText
Output ModalitiesText
Parameters70B
Benchmarks
Intelligence Index
#498
MMLU-Pro
#313
GPQA
#472
HLE
#345
LiveCodeBench
#310
AIME
#177
Time to First Token
#188
MATH-500
#181
Output TPS
#423

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Cache Write 5m
$ / 1M
Amazon Bedrock logo
Amazon Bedrock
meta.llama2-70b-chat-v1
$1.95$2.56N/AN/A
Anyscale logo
Anyscale
anyscale/meta-llama/Llama-2-70b-chat-hf
$1.00$1.00N/AN/A
Databricks logo
Databricks
databricks/databricks-llama-2-70b-chat
$0.5$1.50$0.5$0.5
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/llama-v2-70b-chat
$0.9$0.9N/AN/A
Perplexity logo
Perplexity
perplexity/llama-2-70b-chat
$0.7$2.80N/AN/A
Replicate logo
Replicate
replicate/meta/llama-2-70b-chat
$0.65$2.75N/AN/A

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 2 70B Chat, ranked by cheapest on-demand price. The model needs about 168 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
Standard_NP40sAzure4× AMD Alveo U250 FPGA (64GB)256 GB$6.60/hr
g2-standard-96GCP8× nvidia-l4192 GB$7.98/hr
g7e.12xlargeAWS2× RTX PRO Server 6000192 GB$8.29/hr
7 more instances can run Llama 2 70B Chat
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Llama 3.3 70B Instruct131K$0.120$0.200Deprecated
Llama 3.2 3B Instruct131K$0.015$0.020Deprecated
Llama 3.2 1B Instruct131K$0.020$0.020Deprecated
Llama 3.2 11B128K$0.160$0.160Available
Llama 3.1 405B Instruct131K$0.120$0.300Deprecated
Llama 3.1 8B Instruct200K$0.020$0.030Deprecated
Llama 3.1 70B Instruct131K$0.100$0.100Deprecated
Llama 3.1 70B128K$0.360$0.360Available
Llama 3.1 8B131K$0.030$0.050Available
Llama 3 70B Instruct131K$0.120$0.300Deprecated
Llama 2 70B Chat4K$0.500$0.900Current

Model IDs

anyscale/meta-llama/Llama-2-70b-chat-hf
databricks/databricks-llama-2-70b-chat
fireworks_ai/accounts/fireworks/models/llama-v2-70b-chat
llama-2-chat-70b
meta-llama-2-70b-chat
meta.llama2-70b-chat-v1
perplexity/llama-2-70b-chat
replicate/meta/llama-2-70b-chat
snowflake/llama2-70b-chat