Inkling icon
Inkling
#33 in LLM models
4.4/5
« Discover Thinking Machines’s first LLM: an open, multimodal model with a context size of up to 1 million tokens. It was designed primarily for code, tool manipulation, and adaptive reasoning (with fine-tuning capabilities via Tinker) »
Free 99425

Inkling opens up a near-trillion-parameter multimodal model, Apache license included

Inkling is the first open-weights AI model from Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati. The weights download free of charge under an Apache 2.0 license; only inference, on your own hardware or through partner APIs, carries a cost. Under the hood sits a mixture-of-experts design with 975 billion parameters, 41 billion of them active per query, plus a context window of one million tokens. It reads text, images and audio, then answers in text, code included.

Pros
  • Weights free to download and modify under Apache 2.0
  • Handles text, images and audio natively
  • One-million-token context window
  • Adjustable thinking effort to keep costs in check
  • Inkling-Small variant for lighter hardware
Cons
  • Outputs are text only, no image generation
  • Full model calls for a serious GPU cluster
  • Raw scores sit below the very top models

Three input senses and a dial for thinking effort

Inkling takes in text, images and audio, and writes back text, from code to structured data. The architecture is a mixture of experts: of its 975 billion parameters, only 41 billion wake up for any given query, which keeps inference bills down. Training chewed through 45 trillion tokens of text, images, audio and video.

Its handiest trick day to day is the thinking-effort dial. You turn it down, answers land faster and cheaper; you turn it up, the model digs deeper. On one coding benchmark, Thinking Machines reports matching NVIDIA's Nemotron 3 Ultra with a third of the tokens. And rather than bluffing, Inkling was trained to flag what it is unsure about.

Benchmarks, architecture and hands-on tests of Inkling
First look at Inkling, from coding to multimodal tasks

Inkling, raw material to shape rather than a trophy shelf

Thinking Machines states plainly that Inkling is not the strongest model on the market, open or closed. The wager sits elsewhere: ship a balanced foundation that each organization then fine-tunes on its own data through Tinker, the company's customization service, with 64K or 256K context options. At release it was the largest American open-weights model, a direct answer to DeepSeek V4, GLM 5.2 and Kimi K2.6, the names crowding the top of the LLM model rankings.

A few distinguishing marks noted at launch:

  • 77.6% on SWE-bench Verified, ahead of Nemotron 3 on software engineering
  • among the strongest open models for audio understanding (VoiceBench, MMAU)
  • able to write its own fine-tuning scripts through a coding agent
  • trained to answer directly on censorship-prone topics
  • a smaller sibling, Inkling-Small, 276 billion parameters with 12 billion active

Where to try and download Inkling

The quickest route is the Inkling Playground, a free browser chat put online at launch. For real work, several paths coexist (your desktop, for its part, will politely sit the full model out).

Access routeWhat you getBest for
Inkling Playgroundbrowser-based chat, nothing to installa first taste
Hugging Facefull BF16 weights, NVFP4 and GGUF variantsteams running a GPU cluster
Tinkerfine-tuning on your data, 64K or 256K contextdomain customization
Partner APIs (Together AI, Fireworks, Baseten, Modal, Databricks)hosted inference billed per useplugging it into your apps

Frequently asked questions

Is Inkling free?

Yes, Inkling's weights download free from Hugging Face under an Apache 2.0 license. The cost shifts to running it: either your own hardware or pay-per-use APIs at Together AI, Fireworks, Baseten, Modal or Databricks. Fine-tuning through Tinker is a paid service as well.

Is Inkling really open source?

Strictly speaking, Inkling is open weights rather than fully open source: Apache 2.0 covers commercial use, modification and redistribution royalty-free, but the training data and training code stay private. That is the same standard set by Llama and DeepSeek releases.

What hardware does Inkling take to run locally?

Heavy iron: the BF16 checkpoint asks for 2 TB of pooled VRAM, the quantized NVFP4 version gets by on roughly 600 GB, and compressed GGUF builds circulate for llama.cpp. On modest machines, the realistic option is Inkling-Small with its 12 billion active parameters.

Inkling or DeepSeek, which one to pick?

Both play in the big open-model league, DeepSeek V4 on the Chinese side, Inkling on the American one. Early comparisons hand some rivals the edge on terminal coding, yet Inkling remains the only one of the pack to hear audio and read images natively, with a one-million-token context.

Verdict: A leaderboard king, no; a base to build on, yes: if your team is after a multimodal model to fine-tune on its own data without leaning on a closed API, Inkling ranks among the most ambitious open foundations west of the Pacific.

★ Featured AI Tools ★
AI Alternatives for Inkling
Paid
Freemium
Freemium
Paid
Free
Paid
Free
Paid