Thinking Machines LabInkling

External Link

Inkling is Thinking Machines Lab's open-weight multimodal mixture-of-experts model with 975B total parameters, 41B active parameters, native text, image, audio, and video input, and a 1M-token context window.

Context
1,048,576 tokens
Inputs
Text, Image, Video, Audio
Outputs
Text
Released
Jul 15, 2026

Providers & Pricing

Estimate workload cost →
ProviderInput / 1MOutput / 1MCache read / 1MCache write / 1MObservedEvidence
$0.95$4.05$0.16Sep 3, 2026Provider-reported
$1.00$4.05$0.17Sep 3, 2026Provider-reported
ARC-AGI-1verified-v1-b2e2a31b3d5379.50percentInklingunknownOriginal sourcewinner eligibleThird-party benchmark
ARC-AGI-2verified-v2-326661568d5f36.53percentInklingunknownOriginal sourcewinner eligibleThird-party benchmark
LMArena Agent Arenaagent-2026-08-31-011508720696-7.02score39,537Inkling; 95% CI [-8.08412500, -5.94789916]; sessions 39537; observations 1486696; rank 47unknownOriginal sourcewinner eligibleThird-party benchmark
LMArena Text Arenatext-2026-09-01-0115087206961,438.25rating21,973inkling; 95% CI [1433.11517161, 1443.38005000]; votes 21973; rank 67unknownOriginal sourcewinner eligibleThird-party benchmark
ToneBench2026-08-28-10-task-cd9819ab6e4d79.41points5,78550Inklingaverage per caseOriginal sourcewinner eligibleThird-party benchmark
ToneBench2026-08-28-10-task-cd9819ab6e4d78.74points5,28250Inkling (high)average per caseOriginal sourcewinner eligibleThird-party benchmark

At a Glance

Model Facts

Identity

Status
active
Developer ID
thinkingmachines/Inkling
Released
Jul 15, 2026
Version
Inkling
License
apache-2.0

Capacity

Context
1,048,576 tokens
Maximum output
Unknown
Total parameters
975,000,000,000
Active parameters
41,000,000,000
Knowledge cutoff
Not reported

Interface and Access

Inputs
Text, Image, Video, Audio
Outputs
Text
API available
Yes
Open weights
Yes
Self-hostable
Yes
Reasoning
Yes
Vision
Yes
Tool calling
Yes
Capabilities
agents, chat, coding, generation, reasoning, tools, vision

This page represents one immutable developer model ID. Serving endpoints, pricing, regions, quantizations, and service tiers attach separately and never create duplicate model records.

Version Lineage

PredecessorUnknown
SuccessorsUnknown
Aliasesinkling, thinkingmachines/Inkling, Thinking Machines Inkling

Explicit Unknowns

  • Maximum output
  • Knowledge cutoff

Primary Evidence

Sources and Observation Date

Comparable Peers

Related Model Comparisons

All comparisons →
APairBContext
vscross-developer peerstext
vscross-developer peersimage, text
vscross-developer peerstext
vscross-developer peersaudio, image, text, video
vscross-developer peersaudio, image, text, video
vscross-developer peersimage, text, video
vscross-developer peersimage, text
vscross-developer peersaudio, text
vscross-developer peerstext
vscross-developer peersaudio
vscross-developer peersaudio
vscross-developer peersvideo

Questions

Inkling FAQs

What is Inkling?+

Inkling is Thinking Machines Lab's open-weight multimodal mixture-of-experts model with 975B total parameters, 41B active parameters, native text, image, audio, and video input, and a 1M-token context window. It is developed by Thinking Machines Lab and its current sourced lifecycle status is active.

What inputs and outputs does Inkling support?+

Inkling's current record lists Text, Image, Video, and Audio as input and Text as output.

How much context does Inkling support?+

The current record lists a 1,048,576-token context window and does not report a maximum output length.

Is Inkling available through an API?+

Inkling is marked as API-available. The directly sourced provider records currently include Deepinfra, Fireworks Ai, Openrouter, and Together Ai.

How large is Inkling?+

The current record lists 975,000,000,000 total parameters and 41,000,000,000 active parameters. Its architecture is InklingForConditionalGeneration.

What capabilities and tasks does Inkling support?+

Inkling's recorded capabilities are agents, chat, coding, generation, reasoning, tools, and vision. Its supported tasks are audio, coding, image, reasoning, text, and video.

How much does Inkling cost through an API?+

The lowest directly sourced prices currently attached to Inkling are $0.95 per million input tokens through Deepinfra and $4.05 per million output tokens through Deepinfra. Prices are provider-specific and should be checked against each cited observation date.

Can Inkling be self-hosted?+

Inkling is an open-weight model. Self-hosting is supported, and the recorded license is apache-2.0.

Does Inkling have related versions or aliases?+

No predecessor is recorded. No successor is recorded. Recorded aliases are inkling, thinkingmachines/Inkling, and Thinking Machines Inkling.

Which models can Inkling be compared with?+

The current compatible peer set includes NVIDIA Nemotron 3.5 Lightning 30B-A3B, Kimi-K3, NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16, Seed 2.0 Lite, and Muse Spark 1.1. Compatibility requires shared sourced modalities and tasks; it does not imply that one model is better. Browse model comparisons

Send Feedback