Inkling now comes in smaller size 🤗 Inkling-Small is out by Thinking Machines Lab. We have updated this post with performance , and deployment configurations for the Inkling-Small and the Inkling-Small-NVFP4 variants. Here’s the collection with all the Inkling models. We made it easier for you to deploy Inkling-Small with one-click on Inference Endpoints (getting up to 160 TPS). We also ship a real-time voice and image demo where you can interact with the model. Inkling is a large (1T params!) open model to natively accept image, text, and audio inputs. TLDR; Inkling by Thinking Machines is out on Hugging Face. Inkling is a huge multimodal LLM that understands all modalities (image, audio, text), has agentic capabilities, and supports 1M context. It comes in full BF16 and a well-calibrated NVFP4 variant, and includes speculative MTP layers for faster inference. …