Inkling Small is Thinkingmachines's language model with a 1.0M context window and up to 262K output tokens, available from 3 providers, starting at $0.45 / 1M input and $1.20 / 1M output. A lightweight 12B-parameter reasoning LLM with vision and tool-use capabilities, offering lower cost and latency than the full Inkling model.
Specifications
Canonical IDthinkingmachines-inkling-small
TypeLanguage
StatusActive
CreatorThinkingmachines
Providers
Context Window1.0M tokens
Max Output262K tokens
Input ModalitiesAudioImagePDFText
Output ModalitiesText
Reasoning Effortsdefault
Parameters266B
HuggingFace Likes201
HuggingFace Downloads (30d)2,971
HuggingFace Downloads (all-time)2,971
Release Date · 13 days ago
Benchmarks
Intelligence Index
41.2
#39
Coding Index
52.9
#41
GPQA
0.9
#33
HLE
0.3
#44
Time to First Token
1.47s
#476
SciCode
0.5
#38
LCR
0.7
#80
Output TPS

Capabilities

Input4/5
Text
Image
Audio
Video·
PDF
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities5/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Hugging Face logo
Hugging Face
together_ai:thinkingmachines/Inkling-Small
$0.5$1.20N/A
OpenRouter logo
OpenRouter
thinkingmachines/inkling-small
$0.45$1.20$0.1
Vercel AI Gateway logo
Vercel AI Gateway
thinkingmachines/inkling-small
$0.5$1.20$0.1

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Inkling Small, ranked by cheapest on-demand price. The model needs about 638 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
p4de.24xlargeAWS8× A100640 GB$27.45/hrus-east-1
Standard_ND96amsr_A100_v4Azure8× NVIDIA A100 (80GB)640 GB$32.77/hreastus
g7e.48xlargeAWS8× RTX PRO Server 6000768 GB$33.14/hrus-east-1
7 more instances can run Inkling Small
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Model IDs

accounts/fireworks/models/inkling-small
inkling-small
thinkingmachines-inkling-small
thinkingmachines/inkling-small
thinkingmachines/Inkling-Small