🔍 Read the full analysis: Transformers Now Runs Llama.cpp Quants on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has added support for loading GGUF quantized checkpoints through Transformers’ `from_pretrained` API, initially for Qwen3.5 models on Apple Silicon. The feature is on the library’s main branch; wider hardware and architecture support, stable release timing and full benchmark details have not been specified.
Users can select a GGUF checkpoint hosted on the Hugging Face Hub and pass its file through the `gguf_file` argument to `from_pretrained`, then generate text using Transformers. Hugging Face says the integration reuses llama.cpp’s ggml kernels; the announcement also compares performance with llama.cpp across three checkpoints: a small dense model, a larger dense model and a mixture-of-experts model.
For supported Apple Silicon setups, Transformers can load compatible ggml/Metal layer kernels when weights remain packed on Metal. The feature uses `ggml-org/ggml-attn` for attention when available. If that kernel cannot be fetched, it falls back to standard SDPA attention with a warning; users can also select SDPA explicitly. Without a compatible quantization kernel, the loader dequantizes the model, which uses more memory.
The stated requirements include an Apple Silicon Mac, a PyTorch version supported by the published kernel builds—generally one of the two latest releases—and current Transformers plus a compatible kernels library. The same checkpoints can also be served using `transformers serve`, which provides a local OpenAI-compatible API for clients configured to connect to that endpoint.
GGUF Joins the Transformers Workflow
The change gives developers already using Transformers and PyTorch a direct route to GGUF checkpoints, a format commonly used by local inference tools including Ollama, LM Studio and Jan. Previously, users generally relied on llama.cpp-derived software to run these files. The new path may let developers use a broader range of Hub checkpoints without changing their model-loading workflow, subject to the current platform and architecture limits.
Quantization can reduce the memory needed to run a model. Hugging Face lists Unsloth’s Qwen3.5-4B at 8.42 GB in BF16 and 2.74 GB in Q4_K_M. That smaller footprint can make local inference feasible on machines with less memory, though reduced precision can affect output quality. The announcement advises users to evaluate quantization on their own models and tasks; it does not establish that one setting will work equally well across workloads.
Performance is another part of the rationale: the integration is designed to use ggml kernels rather than simply unpacking weights into a conventional format. Hugging Face says it benchmarked against llama.cpp, but speed depends on the hardware, model and configuration. The announcement’s benchmark references alone do not establish a universal performance result.
A New Route for Quantized Checkpoints
GGUF, developed for the llama.cpp ecosystem, packages model weights and metadata in a single file. Its quantization variants trade some numerical precision for a smaller memory footprint. A label such as Q4_K_M indicates a mixed-precision format, with most weights stored at four bits and some tensors kept at higher precision.
Hugging Face’s example sizes for Qwen3.5-4B are 3.53 GB for Q6_K, 3.14 GB for Q5_K_M and 2.74 GB for Q4_K_M, compared with 8.42 GB for BF16. The company suggests starting with Q4_K_M and trying higher-precision variants when more memory is available, while warning that quality effects depend on the model and task.
The feature arrives amid growing interest in running models locally. Hugging Face co-founder Julien Chaumond recently described a demonstration of Qwen3.6 27B running in the Pi coding agent through llama.cpp on a MacBook Pro. That was his assessment of a particular demonstration, not a benchmark establishing comparable performance across tasks or systems.
““We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.””
— Hugging Face announcement
Platform and Release Scope Remain Open
The announced support is focused on Apple Silicon and Qwen3.5. Hugging Face has not specified a timeline for CUDA, Linux or Windows support, or said which model architectures will be added next. The feature is on the Transformers main branch, and no date has been announced for a stable release.
The announcement refers to comparisons with llama.cpp across three checkpoints, but the results depend on the selected models and hardware. The information provided here does not establish a performance advantage across devices. It also remains unclear how quickly additional kernel builds and architecture support will be made available.
Hugging Face cautions that quality changes from quantization depend on the model and task. The listed file sizes show memory differences, but they do not by themselves measure output quality or suitability for a particular workload.
Stable Release and Broader Support
The next practical milestone is a stable Transformers release that includes GGUF loading; Hugging Face has not announced when that will happen. Until then, users can access the capability from the main branch, provided their hardware, PyTorch version and kernels library meet the stated requirements.
Users following the rollout can check Hugging Face’s GGUF documentation and kernels library for changes to supported formats and compatibility. The announcement leaves expansion to other hardware backends and model architectures as open areas; it does not provide dates or a confirmed rollout sequence for them.
Key Questions
How do users load a GGUF checkpoint in Transformers?
On a compatible setup, users can select a Hub checkpoint and pass its file with the `gguf_file` argument to `from_pretrained`. The feature is currently on the Transformers main branch.
Which hardware and model family are supported initially?
The initial rollout targets Apple Silicon Macs and Qwen3.5. Hugging Face has not announced a timeline for CUDA, Linux or Windows support.
What happens if the quantization kernel is unavailable?
The loader falls back to dequantizing the model, which uses more memory. If the ggml attention kernel cannot be fetched, the system falls back to SDPA attention with a warning; users can select SDPA directly.
Is the feature in a stable Transformers release?
No. It is available on the main branch. Hugging Face has not announced a stable release date.
Does a smaller GGUF file guarantee the same output quality?
No. Quantization reduces file size, but Hugging Face says its effect on quality depends on the model and task. Users should evaluate the model on their own workload.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
