Llama 2 7B Chat is Meta's language model with a 4K context window and up to 4K output tokens, available from 3 providers, starting at $0.05 / 1M input and $0.15 / 1M output. A 7B Llama 2 model fine-tuned with RLHF for dialogue use cases, offering an efficient and accessible conversational LLM.
Specifications
Canonical IDmeta-llama-2-7b-chat
TypeLanguage
StatusActive
CreatorMetaMeta
Providers
Context Window4K tokens
Max Output4K tokens
Input ModalitiesText
Output ModalitiesText
Parameters7B
Benchmarks
Intelligence Index
4.3
#425
MMLU-Pro
0.2
#320
GPQA
0.2
#471
HLE
0.1
#256
LiveCodeBench
0.0
#323
AIME
0.0
#168
Time to First Token
20.41s
#485
SciCode
0.0
#472
MATH-500
0.1
#192
Output TPS
96.4
#130

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Anyscale logo
Anyscale
anyscale/meta-llama/Llama-2-7b-chat-hf
$0.15$0.15
Fireworks AI logo
Fireworks AI
fireworks_ai/accounts/fireworks/models/llama-v2-7b-chat
$0.2$0.2
Replicate logo
Replicate
replicate/meta/llama-2-7b-chat
$0.05$0.25

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Llama 2 7B Chat, ranked by cheapest on-demand price. The model needs about 17 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
g2-standard-4GCPnvidia-l424 GB$0.705/hrus-east4
g6.xlargeAWSL422 GB$0.805/hrus-east-1
g2-standard-8GCPnvidia-l424 GB$0.851/hrus-east4
7 more instances can run Llama 2 7B Chat
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Llama 3.3 70B Instruct131K$0.120$0.200Deprecated
Llama 3.2 3B Instruct131K$0.015$0.020Deprecated
Llama 3.2 1B Instruct128K$0.027$0.080Deprecated
Llama 3.2 11B128K$0.160$0.160Available
Llama 3.1 405B Instruct131K$0.120$0.300Deprecated
Llama 3.1 8B Instruct200K$0.020$0.030Deprecated
Llama 3.1 70B Instruct131K$0.120$0.300Available
Llama 3.1 70B128K$0.360$0.360Available
Llama 3.1 8B131K$0.030$0.050Available
Llama 3 70B Instruct131K$0.120$0.300Deprecated
Llama 2 7B Chat4K$0.050$0.150Current

Model IDs

@cf/meta/llama-2-7b-chat-fp16
accounts/fireworks/models/llama-v2-7b-chat
anyscale/meta-llama/Llama-2-7b-chat-hf
cloudflare/@cf/meta/llama-2-7b-chat-fp16
cloudflare/@cf/meta/llama-2-7b-chat-int8
fireworks_ai/accounts/fireworks/models/llama-v2-7b-chat
llama-2-chat-7b
meta-llama-2-7b-chat
replicate/meta/llama-2-7b-chat