Inkling Small is Thinkingmachines's language model with a 1.0M context window and up to 262K output tokens, available from 4 providers, starting at $0.45 / 1M input and $1.20 / 1M output. A lightweight 12B-parameter reasoning LLM with vision and tool-use capabilities, offering lower cost and latency than the full Inkling model.
Specifications
Canonical IDthinkingmachines-inkling-small
TypeLanguage
StatusDeprecated
CreatorThinkingmachines
Providers
Context Window1.0M tokens
Max Output262K tokens
Input ModalitiesAudioImagePDFText
Output ModalitiesText
Reasoning Effortsdefault
Parameters266B
HuggingFace Likes201
HuggingFace Downloads (30d)2,971
HuggingFace Downloads (all-time)2,971
Release Date · 2 months ago
Deprecation Date
Benchmarks
Intelligence Index
#88
Coding Index
#61
GPQA
#52
HLE
#69
Time to First Token
#453
SciCode
#57
LCR
#78
Output TPS
#38

Capabilities

Input4/5
Text
Image
Audio
Video·
PDF
Output1/5
Text
Image·
Audio·
Video·
Embedding·
Capabilities5/13
Reasoning
Adaptive Reasoning·
Function Calling
Parallel Function Calling·
Structured Outputs
Native JSON Schema
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandardFree
Input
$ / 1M
Output
$ / 1M
Cache Read
$ / 1M
Input
$ / 1M
Output
$ / 1M
DeepInfra logo
DeepInfra
deepinfra/thinkingmachines/Inkling-Small
$0.45$1.20$0.1
OpenRouter logo
OpenRouter
thinkingmachines/inkling-small
$0.45$1.20$0.1N/AN/A
OpenRouter logo
OpenRouter
thinkingmachines/inkling-small:free
$N/A$N/A
Together AI logo
Together AI
together_ai/thinkingmachines/Inkling-Small
$0.5$1.20$0.1
Vercel AI Gateway logo
Vercel AI Gateway
thinkingmachines/inkling-small
$0.5$1.20$0.1

Cost Calculator

US Dollar ($)
Preset:

Cheapest Instances to Run It

Cloud GPU instances that can host Inkling Small, ranked by cheapest on-demand price. The model needs about 638 GB of GPU memory at FP16 precision (estimated from its parameter count), so treat the fit as guidance rather than a guarantee.

All clouds
FP16 (full precision)
US Dollar ($)
Instance
Cloud
GPU
VRAM
Price
Cheapest region
p4de.24xlargeAWS8× A100640 GB$27.45/hr
Standard_ND96amsr_A100_v4Azure8× NVIDIA A100 (80GB)640 GB$32.77/hr
g7e.48xlargeAWS8× RTX PRO Server 6000768 GB$33.14/hr
7 more instances can run Inkling Small
Unlock the full ranked list and FP8 / INT4 quantization with a CloudPrice subscription.

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Inkling Small1.0M$0.450$1.20Current
Inkling1.0M$1.00$4.05Available

Model IDs

accounts/fireworks/models/inkling-small
deepinfra/thinkingmachines/Inkling-Small
huggingface-llm-inkling-small
inkling-small
openrouter/thinkingmachines/inkling-small
thinkingmachines-inkling-small
thinkingmachines/inkling-small
thinkingmachines/Inkling-Small
together_ai/thinkingmachines/Inkling-Small