Llama 3.2 90B Instruct Vision is an AI model from Meta. Meta's 90B instruction-tuned vision-language model from Llama 3.2, combining large-scale multimodal understanding with instruction-following capabilities.
Specifications
Canonical IDmeta-llama-3-2-90b-instruct-vision
StatusActive
CreatorMetaMeta
Benchmarks
Intelligence Index
6.2
#373
MMLU-Pro
0.7
#224
GPQA
0.4
#354
HLE
0.0
#303
LiveCodeBench
0.2
#246
AIME
0.1
#137
Time to First Token
0.58s
#267
SciCode
0.2
#311
MATH-500
0.6
#141
Output TPS
46.4
#233

Capabilities

Input0/5
Text·
Image·
Audio·
Video·
PDF·
Output0/5
Text·
Image·
Audio·
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Llama 3.3 70B Instruct131K$0.100$0.200Available
Llama 3.2 3B Instruct131K$0.015$0.020Deprecated
Llama 3.2 1B Instruct131K$0.027$0.080Deprecated
Llama 3.2 11B128K$0.160$0.160Available
Llama 3.1 405B Instruct131K$0.120$0.300Deprecating
Llama 3.1 70B Instruct131K$0.120$0.300Available
Llama 3.1 8B Instruct200K$0.020$0.030Available
Llama 3.1 70B128K$0.360$0.360Available
Llama 3.1 8B131K$0.030$0.050Available
Llama 3 70B Instruct131K$0.120$0.300Deprecated
Llama 3.2 90B Instruct VisionCurrent

Model IDs

llama-3-2-instruct-90b-vision
meta-llama-3-2-90b-instruct-vision