QUICK TIPS
1
Feed audio as 16kHz WAV under about 20 minutes, which is the range it was tuned for
2
Keep images between roughly 40px and 4096px on the long edge
3
Use the controllable thinking effort setting rather than paying for maximum reasoning on every call
4
Treat the vendor benchmark numbers as an upper bound, they were run at effort 0.99 and temperature 1.0
5
Try a hosted inference provider before committing to 975B of self-hosted infrastructure
❓
FREQUENTLY ASKED QUESTIONS
Q
Is Inkling free?
A
Yes, Inkling is completely free to use.
Q
Does Inkling have an API?
A
Yes, Inkling offers an API for programmatic integration.
Q
What is Inkling best for?
A
Inkling is best for Best Open-Weight Multimodal Base.
Q
What platforms does Inkling support?
A
Inkling supports api, local.
Q
Is Inkling open source?
A
Yes, Inkling is open source.
Q
What can I do with Inkling?
A
Inkling is designed for Fine-tuning an open multimodal base, Audio and image input without a separate encoder, Commercial self-hosting under Apache 2.0. Inkling is the first open-weights model from Thinking Machines Lab, released 15 July 2026 under Apache 2. Key strengths include Apache 2.0, the least restrictive licence of any model at this scale and Native text, image and audio input in a single open checkpoint.
Q
How do I use Inkling?
A
Inkling is a large language model for text generation, analysis, and conversation. Use the API for programmatic access. Enter prompts or questions to get responses. It excels at apache 2.0, the least restrictive licence of any model at this scale.
Q
How do I get started with Inkling?
A
Pull the weights from huggingface.co/thinkingmachines/Inkling, or skip the hardware and call it through Together AI, Fireworks, Modal, Databricks or Baseten. For local work, the NVFP4 checkpoint is the one to use on Blackwell cards; everyone else sho...
Work on Inkling? You're hand-reviewed in our directory. Add this badge to your site — it links back to this profile.