H Company released Holo4 on September 28, with a 27B dense model and a 35B-A3B mixture-of-experts model. The release targets tasks that combine graphical interfaces, code execution, MCP and API calls. That is relevant to workflows where one step has a reliable API and the next exists only in a desktop or browser interface. H Company’s announcement
The two downloadable checkpoints also have different licenses. That distinction should come before any benchmark comparison.
Select the checkpoint by its actual terms
| Checkpoint | Architecture | Published weight license |
|---|---|---|
| Holo4-27B | Dense, 27B | CC BY-NC 4.0 |
| Holo4-35B-A3B | MoE, 35B total and about 3B active | Apache 2.0 |
The first-party 27B license carries a non-commercial restriction. The 35B-A3B license is Apache 2.0. Downloadable weights therefore do not imply the same commercial-use permissions across the family. Hosted service terms are a separate check.
For deployment planning, use the complete checkpoint and runtime requirements. The active parameter count describes computation per token; it does not mean only that many parameters need storage. Our Open Models guide explains that distinction and the machine guide covers memory headroom.
The harness is part of the system
The 35B-A3B model card describes a harness that passes screenshots and tool results to the model, executes requested actions and returns their results. It also links published evaluation trajectories. A checkpoint alone does not supply your application’s permissions, execution environment or result verification.
Consider a workflow that copies a value from a desktop application into an internal record. An evaluation should check the source value, the destination record and the final saved state. A screenshot showing the intended field filled in is insufficient if the save failed. This is an illustrative test design, not a result measured with Holo4.
Keep the test environment repeatable. Record application versions, starting state, enabled tools and action limits. Preserve failed trajectories alongside successful ones. If one model receives an API tool and another must navigate the interface, the comparison includes a difference in tooling as well as model capability.
Read the failures in the trajectories
H Company publishes a trajectory viewer and evaluation dataset. Those artifacts can help an evaluator inspect how an agent reaches a result. They do not establish reliability in a different application or under different permissions.
When reviewing a trace, look for actions based on stale screen state, repeated attempts after an error, and completion claims that lack a corresponding saved result. For your own tests, include a recoverable interruption, such as a validation error, and check whether the agent repairs it without modifying unrelated records.
S5 Labs has not reproduced H Company’s benchmark results or measured local runtime performance. Holo4 is worth evaluating when a task spans interfaces. Start with the checkpoint whose terms fit the use, then judge the complete agent by verified task outcomes and the work required to recover from its failures.
