🔍 Read the full analysis: How Holo4 Powers Generalist Computer-Use Agents on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
H Company has released Holo4, a pair of open-weight models designed to use graphical interfaces, code, MCP tools and APIs within one system. The company reports a 61.7% OSWorld 2.0 score for the 27B model, but the results have not been independently verified and comparisons use differing evaluation setups.
H Company has released Holo4, a series of open-weight models built to operate software through graphical user interfaces, code, MCP tools and APIs. The company reports that its 27B dense model scored 61.7% on OSWorld 2.0; that result could make the release relevant to developers seeking lower-cost software agents, though independent evaluation is still pending.
The series includes a 27B dense model and a 35B-A3B Mixture of Experts model. Both are available through the H Models API and for download from Hugging Face in FP16, FP8 and GGUF formats, according to the announcement. H Company says the models can choose among screen interaction, writing and running code, and calling MCP or API tools as a task requires.
H Company reports OSWorld 2.0 scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. It compares the first result with 81.8% for Opus 5.5, which the company identifies as the strongest closed model in its comparison. The announcement says Holo4 uses far fewer parameters and costs less, but the comparisons involve different releases, harnesses and task subsets.
The company says Holo4 was trained with supervised and reinforcement learning across environments and tasks, including tasks generated by its Agentic Task Factory. It also reports improvements over its Qwen base models, citing side-by-side examples in FreeCAD and Godot. H Company has published trajectories behind its public benchmark scores on its trajectory site and says they can also be downloaded from Hugging Face.
One Model Across Software Interfaces
Many workplace tasks cross between a screen, code and a business service. An agent that can move among those interfaces could handle more of a workflow without requiring separate models or integrations for each step. H Company presents Holo4 as a way to make that kind of cross-interface automation available in open weights.
If outside evaluations reproduce the reported results, organizations may have another option for running software agents with more control over deployment and costs. The published trajectories also give researchers and developers material to inspect when checking how the models produced their benchmark results. The release itself does not establish that Holo4 is reliable on broad, live business workloads.
From Holo1 to Holo4
Holo4 follows H Company’s earlier agentic model work and arrives alongside an updated Holotron 3 version called Holotron4 Nano, according to the source material. The company’s benchmark notes identify Qwen3.8 27B as the dense model’s base and Qwen3.6 35B-A3B as the MoE model’s base.
H Company says its cost comparisons use H Models API rates for Holo4, Alibaba Cloud list prices for Qwen, and OpenAI launch data for GPT and Opus effort sweeps. It notes that model releases, harnesses and task subsets differ. On AutomationBench, the company measured Holo4 in its internal harness version 1.0.6; other models’ public-set scores and private-set cost figures do not provide a like-for-like comparison.
“Real work is not siloed that way, and a single business task can require combining these different approaches.”
— H Company
Benchmark Results Await Replication
Holo4’s headline benchmark figures are company-reported, and no independent reproduction is included in the source material. H Company acknowledges differences in releases, evaluation harnesses and task subsets among the models it compares. The AutomationBench comparison also mixes public-set scores with private-set cost data, and Holo4 has not yet been evaluated on that private set.
The announcement does not explain why the 35B-A3B model scored substantially below the 27B model on OSWorld 2.0. It also does not establish how either model performs across varied business workflows outside benchmark tasks and selected examples. Those questions remain open pending further evaluation.
Private Set and Independent Tests
H Company says it will report Holo4’s AutomationBench private-set results once that evaluation is complete. Independent submissions and reproductions can test the public benchmark claims, with the released weights and trajectories giving evaluators material to examine. Developers can access the models through the H Models API or download them from Hugging Face.
Key Questions
What is Holo4?
Holo4 is H Company’s open-weight agent model series, designed to operate software through graphical interfaces, code, MCP tools and APIs.
Which Holo4 models are available?
The release includes a 27B dense model and a 35B-A3B Mixture of Experts model. H Company offers them through its API and on Hugging Face in several formats.
How did Holo4 score on OSWorld 2.0?
H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. The results have not been independently verified in the supplied material.
Have independent evaluators confirmed the results?
No independent confirmation is described in the announcement. H Company has published benchmark trajectories, and third-party evaluations remain a next step.
Where can developers get Holo4?
Developers can access the models through the H Models API or download them from Hugging Face.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
