Halo brings distributed training to stock Hugging Face models

White Circle has open-sourced Halo, the training framework it uses for its released models. Halo targets teams whose models have outgrown standard Hugging Face TRL workflows but do not justify the engineering cost of porting architectures into Megatron-LM or NeMo. It adds distributed training and reinforcement learning while preserving Hugging Face model classes and checkpoints.

Halo accepts an unmodified Hugging Face model and produces a checkpoint loadable through the standard from_pretrained interface. Teams can keep existing architectures, avoid conversion scripts, and add a model family with roughly 100 lines of wrapper code, according to White Circle.

A controlled throughput comparison

White Circle reports 2.3 to 2.8 times the throughput of stock TRL when training OpenAI’s gpt-oss-20b, along with lower peak memory use. Both systems use FlashAttention-4, Liger kernels with fused linear cross-entropy, and grouped-GEMM expert kernels. The comparison therefore isolates differences in the trainer, optimizer, and parameter-sharding strategy.

The project’s benchmark report provides these results:

These are project-reported benchmarks, and distributed-training results depend heavily on GPU topology, software versions, sequence length, and batch shape. The repository and report provide the configurations needed to inspect or reproduce the comparisons.

Four ways to divide the work

Halo combines four parallelism strategies so teams can distribute model weights, experts, and long sequences according to the workload: