AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Thinking Machines has introduced Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. Its open availability contrasts with high hardware requirements and limited independent evaluation. The development signals a step forward in large-scale AI but raises questions about accessibility and performance, as detailed in Welcome Inkling By Thinking Machines.

Thinking Machines has released Inkling on Hugging Face, a 975-billion-parameter multimodal model designed to process text, images, and audio within a one-million-token context window. This release makes Inkling available to developers and researchers, although its substantial hardware requirements limit widespread use. The release is notable because it pairs an unprecedented scale with open access, raising both opportunities and challenges for AI development.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data types. It employs a sparse architecture, activating only a subset of parameters per input, which aims to optimize inference efficiency given its size. The model architecture features 256 experts, global and sliding-window attention, and hierarchical image patching, with audio converted into mel-spectrograms for multimodal reasoning.

Hugging Face reports that running the full model requires approximately 2 TB of VRAM for BF16 precision and about 600 GB for NVFP4, making it accessible primarily through hosted inference services or on specialized hardware. The release includes support in popular frameworks like Transformers and llama.cpp, with some speculative multi-token prediction layers aimed at speeding up generation. However, independent benchmarks, safety evaluations, and licensing details remain unavailable, and the model’s real-world performance and safety are still unverified.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines has released Inkling, a massive multimodal model, on Hugging Face, marking a significant development in open AI models with high hardware demands.

Implications of Inkling’s Scale and Accessibility

The release of Inkling demonstrates the ongoing development of large-scale multimodal models and their potential applications. Its availability could support research and practical applications across various fields that require reasoning across text, images, and audio. However, the high hardware requirements may limit its use to organizations with substantial computational infrastructure. This highlights ongoing challenges in balancing model size, accessibility, and safety considerations in AI research.

Amazon

high VRAM GPU for AI inference

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large Multimodal AI Models

Recent years have seen rapid growth in the size and complexity of AI models, particularly in natural language processing and multimodal understanding. While models like GPT-4 and PaLM have demonstrated advanced capabilities, they often remain proprietary or have limited accessibility. Open models such as Meta’s Llama and open-source initiatives on Hugging Face aim to broaden access, though hardware constraints continue to pose challenges. Inkling’s release continues this trend by offering a large, open multimodal model with significant technical requirements, which may restrict its immediate use to organizations with substantial infrastructure.

“This model is large in scale.”

— Hugging Face’s Inkling release article

Unverified Aspects and Performance Unknowns

Independent benchmark results, safety evaluations, and real-world performance data for Inkling are not yet publicly available. Details regarding licensing, usage restrictions, and fine-tuning options are also not specified. Additionally, the model’s performance on video processing tasks or its adaptability for specific domains has not been evaluated or disclosed.

Upcoming Evaluations and Accessibility Developments

In the coming months, developers and organizations are expected to test Inkling through supported inference frameworks, which will provide insights into its latency, accuracy, and resource requirements. Independent researchers are likely to conduct safety assessments and benchmark comparisons. Future updates may clarify licensing terms, demonstrate domain-specific fine-tuning, and explore methods to reduce hardware demands through quantization or optimized deployment techniques. These developments will help assess Inkling’s practical utility and safety profile.

Key Questions

What is Inkling?

Inkling is a large-scale, multimodal AI model developed by Thinking Machines, capable of processing text, images, and audio, with 975 billion parameters and a one-million-token context window.

Can I run Inkling on my personal computer?

It is unlikely that the full model can be run on a standard personal computer due to its high hardware requirements, which include approximately 2 TB of VRAM for BF16 precision. Access is primarily expected through cloud services or specialized hardware setups.

What are the potential uses of Inkling?

Potential applications include scientific research, media analysis, enterprise data processing, and other fields that benefit from integrated multimodal reasoning capabilities.

Has Inkling been independently tested?

No, independent benchmarking, safety assessments, and detailed performance evaluations have not yet been published.

What are the licensing terms for Inkling?

The release notes that Inkling is an open model but do not specify licensing details, restrictions, or whether training code is available for further development or fine-tuning.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Battery Health Drops Faster for Some Users Than Others

Theories suggest that frequent charging habits, heat exposure, and intense app use cause faster battery health decline, but the full story might surprise you.

Microsoft Patches Windows And Excel – Breaks Audio, Remote Access, And Paste

Recent Microsoft patches for Windows and Excel have led to widespread problems including audio failures, remote access disruptions, and paste functionality loss, with details still emerging.

Anchor. The Schwarz Group model.

Schwarz Group commits €11B to Europe’s largest AI data center, exemplifying a new industrial-anchor investment model at scale in Europe.

The $725 Billion Question: Hyperscaler Capex Q1 2026 and What the Earnings Don’t Answer

The Big Four hyperscalers announced a combined $725 billion AI infrastructure investment in Q1 2026, raising questions about future revenue growth and market impact.