Bolmo: Byteifying the Next Generation of Language Models | Summary
17 Dec 2025 | Paper Review Byte-Level Language Models Knowledge Distillation Adaptive ComputationContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Byteified Olmo
- 4 Experiment Setup
- 5 Main Results
- 6 Ablations
- 7 Conclusion
- 8 Future Directions
- Appendix
- Brief Thoughts
This article explains the key points of Bolmo: Byteifying the Next Generation of Language Models.
- 2025-12-17 (arXiv)
- Minixhofer, Benjamin, Murray, Tyler, Limisiewicz, Tomasz, Korhonen, Anna, Zettlemoyer, Luke, Smith, Noah A., Ponti, Edoardo M., Soldaini, Luca, Hofmann, Valentin.
- Allen Institute for AI, University of Cambridge, University of Washington, University of Edinburgh
- Paper
- Reviewed PDF: arXiv v2, revised 2026-02-09. The selection week is 2025-W52.
Summary
- Bolmo converts pretrained subword LLMs into byte-level Latent Tokenizer Language Models (LTLMs). Its main contributions are patch boundary prediction with one byte of future context during prefill and a two-stage procedure that first distills the source model with its Transformer frozen, then trains the complete model.
- Bolmo 7B achieves a character-understanding aggregate of 75.1 versus 56.0 for Olmo 3 7B and a code aggregate of 40.7 versus 39.5, while scoring below its source model in math and several QA categories. Stage 1 processes 9.8B tokens and Stage 2 processes 39.3B tokens, reusing the source model’s existing pretraining investment.
- Higher patch compression enables Bolmo 7B to surpass its source model’s inference speed at approximately 6.6 bytes per patch in batchsize=1 measurements on H100 GPUs. Task Arithmetic transfers instruction-following capabilities without additional Bolmo training in an Olmo 3 RL-Zero case study, but success depends on compatibility with the base embedding spaces.
1 Introduction
Subword tokenization obscures character structure, introduces prompt-dependent tokenization effects, constrains the vocabulary, and allocates compute at token granularity. The authors hypothesize that byte-level models have struggled to keep pace with leading subword models partly because development has concentrated on training from scratch, requiring separate investment in data, architecture, and post-training.
- Bolmo 7B starts from Olmo 3 7B, while Bolmo 1B starts from OLMo 2 1B. Byteification seeks to preserve their learned capabilities while replacing subword input and output processing with byte-level components.
- The paper describes 39.3B tokens as less than 1% of a typical pretraining budget. This is the Stage 2 budget; Stage 1 separately processes 9.8B tokens.
- The released resources are https://huggingface.co/allenai/Bolmo-7B, https://huggingface.co/allenai/Bolmo-1B, https://huggingface.co/datasets/allenai/bolmo_mix, and https://github.com/allenai/bolmo-core.
2 Related Work
The related work distinguishes direct byte processing from hierarchical architectures that use lightweight local models to pool bytes into patches, process patches with a deep global model, and reconstruct byte representations. Dynamic patching allows patch granularity and global-model compute allocation to vary with the input.
- DTP, BLT, and H-Net provide precedents for LTLMs and byteification. Bolmo develops an architecture and training procedure specifically suited to transferring an existing subword model.
- Tokenizer transfer and retrofitting often use embedding initialization or cross-tokenizer distillation. Bolmo combines supervised boundary matching, representation alignment, and patch-likelihood distillation before full-model training.
- Byte-level processing still uses a tokenizer: UTF-8 maps text to 256 possible byte values. Its Latin-centric encoding means that removing a finite subword vocabulary does not establish equal efficiency across scripts.
3 Byteified Olmo
Byteified Olmo retains the pretrained Transformer as a global model operating on latent byte patches. Local components translate between bytes and the representation spaces expected by that Transformer; their architecture and training objectives are designed to recover the source model’s behavior before full-model adaptation.
3.1 Architecture
At each byte position, the architecture adds a source subword embedding selected by the longest common suffix with the byte sequence up to that position. A one-layer mLSTM encoder contextualizes these embeddings, and the final byte representation of each patch is pooled for the retained Transformer. Depooling adds the latest available Transformer patch representation to a linear projection of local byte representations, followed by a four-layer mLSTM decoder and a byte prediction head.
- Local and global representation dimensions are equal, avoiding an up-projection from a smaller local representation that could constrain embedding rank. The dimensions are 4096 for Bolmo 7B and 2048 for Bolmo 1B.
- Retaining subword suffix embeddings adds sparsely activated capacity and permits a shallow encoder. The authors report that a larger encoder can achieve similar performance without this retention, with a less favorable performance–efficiency tradeoff.
- Removing the source output embedding matrix offsets much of the added local-model capacity: Bolmo 7B has 7.63B parameters versus 7.30B for Olmo 3 7B, and Bolmo 1B has 1.47B versus 1.48B for OLMo 2 1B.
The Transformer processes pooled patches, while local mLSTM components process every byte. The residual path from local byte representations to the decoder preserves a route for fine-grained information that bypasses patch pooling.
3.1.1 Non-Causal Patch Boundary Prediction addresses a mismatch: subword tokenizers can inspect future bytes when determining token endings, whereas causal LTLM boundary predictors cannot. Bolmo’s prefill predictor uses half the cosine distance between projected representations of the current byte and the next byte. During decoding, the language modeling head predicts the next byte together with whether it ends a patch.
- For example, the boundary after r can differ between _Hello_Wor! and _Hello_World! despite the identical prefix _Hello_Wor. One future byte makes this distinction available to the prefill boundary predictor, although subword tokenizers can inspect more distant context.
- Boundary Symbol Fusion expands the output byte vocabulary from 256 to 512, pairing each byte with a version followed by a boundary. This avoids an additional decoder step for a separate <b> symbol.
- The boundary predictor receives external supervision rather than gradients through continuous boundary scores. The authors identify a potential failure mode in which non-causal bfloat16 scores could transmit 16 bits about a future byte to the decoder.
Bolmo separates future-aware segmentation during prefill from joint byte-and-boundary prediction during generation. The illustrated causal patch-start approach shifts patch content by one byte, whereas Bolmo’s non-causal patch-end prediction preserves alignment with the source tokenization.
3.2 Byteifying Procedure
3.2.1 Stage 1: Subword-to-Byte Distillation freezes the global Transformer while training the local encoder, local decoder, boundary predictor, and language modeling head. Its objectives supervise subword patch endings, align encoder representations after four frozen Transformer layers, and match source-token likelihoods to products of the corresponding byte-and-boundary likelihoods. The likelihood comparison assumes matching patch boundaries and trains the decoder using the source model’s contextualized representations.
- Boundary supervision uses binary cross-entropy and rapidly reaches >99% accuracy. True subword boundaries are used for pooling during representation alignment to preserve sequence correspondence.
- Matching representations after n=4 Transformer layers improves performance over direct embedding matching at n=0 while limiting backward computation through the global model.
- Patch-likelihood matching uses temperature-modulated binary cross-entropy with τ=5 and is computed in log-space. The combined objective also includes next-byte cross-entropy; its weights are λ_B=4, λ_E=1, λ_D,Distill=1, and λ_D,CE=1.
3.2.2 Stage 2: End-to-End Training unfreezes the global model and optimizes all parameters using boundary supervision and next-byte cross-entropy. The next-byte objective now uses Bolmo’s own contextualized patch representations, allowing adaptation to imperfect learned boundaries and local representations. This stage can also increase the target number of bytes per patch.
- Stage 1 requires a full global-model forward pass but backward computation only through its first four layers, together with the local components. This enables faster architectural experiments than repeatedly training the complete model.
- Boundary decisions remain externally supervised in Stage 2; gradients are not propagated through discrete patch selection.
4 Experiment Setup
Bolmo Mix contains approximately 172B tokens from the Dolma 3 pretraining mix and 75M tokens of CUTE-style character tasks, with token counts measured using the Dolma2 Tokenizer. Stage 1 processes 9.8B tokens, approximately 43B bytes, and Stage 2 processes 39.3B tokens, approximately 173B bytes. Training covers less than one epoch of the mix.
- Bolmo 7B starts from the Olmo 3 7B checkpoint after mid-training and long-context extension. Architecture development primarily uses OLMo 2 models, and the resulting procedure is applied to Olmo 3 7B without adjustments.
- A continued-training Olmo 3 baseline receives the same 39.3B-token global-model update budget on the same data, helping separate conversion effects from continued-training effects.
- The stated 7B evaluation setup adapts OlmoBaseEval, excludes GSM Symbolic and BigCodeBench, and adds CUTE and multilingual EXECUTE. The 1B suite adapts Base Easy Suite and adds CUTE.
- Code sampling uses temperature=0.6 and top_p=0.6 for both source and byteified models. The 1B core tasks used in sweeps are ARC, MMLU, CSQA, HellaSwag, WinoGrande, SocialIQA, PIQA, Basic Skills, and CUTE.
5 Main Results
Bolmo 7B leads the evaluated byte-level baselines in the character-understanding, code, math, and both multiple-choice QA aggregates. Its scores are 75.1 for Char, 40.7 for Code, 48.9 for Math, 65.5 for MC STEM, and 75.8 for MC Non-STEM. TFree-Hat remains slightly higher on GenQA, at 71.3 versus 70.9.
- Against Olmo 3 7B, CUTE improves from 56.9 to 78.6 and EXECUTE from 55.1 to 71.6. Math declines from 55.3 to 48.9, MC STEM from 66.3 to 65.5, MC Non-STEM from 77.7 to 75.8, and GenQA from 72.4 to 70.9.
- The code aggregate improves from 39.5 to 40.7 through stronger pass@16 despite generally weaker pass@1. HumanEval changes from 49.0/71.1 to 40.6/74.7 for pass@1/pass@16.
- The authors interpret the pass@16 pattern as greater continuation diversity under the tested sampling settings. They leave broader claims about byte-level quality–diversity tradeoffs unresolved.
- The continued-training control also loses performance on many tasks and improves character understanding. Bolmo retains a character advantage over that control, whose Char score is 71.0, while scoring below it in Math, both multiple-choice QA aggregates, and GenQA.
Bolmo’s largest advantage over its source is character understanding, at 75.1 versus 56.0. Its code aggregate improves to 40.7, while Math remains at 48.9 versus 55.3; the category results therefore show substantial recovery of source capabilities with task-dependent gains and losses.
Bolmo 1B averages 58.2 on its full suite, close to OLMo 2 1B at 58.3 and BLT 1B at 58.5, and above the one-stage and two-stage H-Net XL variants at 52.5 and 53.2. These are not matched total parameter budgets: BLT 1B contains 4.53B parameters including embeddings, versus 1.47B for Bolmo.
- CUTE improves from 27.5 for OLMo 2 1B to 60.0 for Bolmo 1B. Lambada improves from 60.1 to 65.2, while MMLU falls from 40.4 to 37.2.
- The paper’s prose reports a +3.3% CoQA improvement, but Table 2 displays 81.7 versus 77.4, a difference of 4.3 percentage points.
Bolmo 1B’s full-suite average is close to its source and BLT, while CUTE improves much more substantially. Counting embedding parameters shows that BLT 1B has 4.53B total parameters, so the nominal model names do not represent matched parameter budgets.
5.1 Training at Higher Compression Factors
Higher compression is trained by removing source subword boundaries until the supervision reaches a target average patch size. The tested strategies merge frequent pairs using BPE per example, pairs with the lowest summed next-token entropy, or pairs with the lowest summed next-token cross-entropy.
- Entropy and cross-entropy supervision use a 370M-parameter auxiliary model trained on 74.3B tokens. It is required only during training and is absent from inference.
- The target compression t differs from the attained compression c because boundary supervision competes with next-byte prediction. The boundary-loss weight remains λ_B=4.
- The comparison transfers OLMo 2 to SuperBPE vocabularies of 200k and 400k using a 10GB text sample and FOCUS embedding initialization. At these larger vocabulary sizes, output softmax cost limits further improvements in the subword performance–compute frontier.
- BPE merging gives the strongest tested Bolmo tradeoff: target t=16 attains c=7.2, whereas entropy and cross-entropy merging attain c=8.5 and c=8.3 with lower core-task performance. The authors’ explanation involving proximity to the pretrained model remains a hypothesis.
The plotted tradeoff measures core-task performance against GFLOPs per byte rather than wallclock speed. BPE-supervised Bolmo compression preserves more performance than the tested entropy-based alternatives, while increasing the SuperBPE vocabulary from 200k to 400k does not further improve the subword frontier.
5.2 Post-Training Byteified Models via Task Arithmetic
Task Arithmetic transfers a post-training update by adding the difference between the post-trained and base Olmo 3 Transformer weights to the corresponding Bolmo Transformer weights. In the RL-Zero instruction-following case study, Bolmo’s IFEval score rises from 31.1% to 67.4%, compared with 35.4% for base Olmo 3 and 66.9% for its post-trained checkpoint.
- The operation requires no additional Bolmo training, but uses a previously post-trained source checkpoint.
- Only the global Transformer has matching parameters across the architectures. Bolmo’s local encoder and decoder remain in their base-model representation spaces.
- Success requires embedding resettability: the post-trained Transformer must remain compatible with the base input and output embedding spaces. The appendix shows that this condition varies across models and post-training procedures.
Adding the source instruction-following weight update raises Bolmo’s IFEval score from 31.1% to 67.4%, close to the source post-trained checkpoint’s 66.9%. This demonstrates successful reuse of the tested RL-Zero update without further Bolmo training.
6 Ablations
The ablations test non-causal boundaries, Stage 1 initialization, and local architectures. They compare boundary accuracy and representation alignment, training loss under approximate compute matching, and measured inference latency and throughput.
6.1 Impact of Non-Causal Patch Boundaries
Causal prediction of patch starts achieves accurate boundary labels but offsets patch content by one byte; causal prediction of patch ends preserves content alignment but is harder for a shallow encoder. Non-causal patch-end prediction combines accurate boundaries with aligned representations and improves downstream performance over both causal settings.
- After Stage 1, causal patch-end prediction reaches 96.0% boundary accuracy and a core-task average of 42.5. Fused non-causal patch-end prediction reaches 99.2% accuracy and an average of 55.6.
- Using oracle boundaries raises the fused non-causal setting to 57.9, supporting the authors’ attribution of much of the remaining Stage 1 gap to boundary errors. This comparison does not directly establish the cause of every final Stage 2 benchmark gap.
- Separate non-causal boundary symbols score 55.7, slightly above the fused setting, but require 9.8 local-model invocations per global invocation versus 8.8 for fusion. Fusion reduces this overhead with nearly identical task performance.
6.2 Is Stage 1 Training Necessary?
Stage 1 improves bits-per-byte under an approximately compute-matched comparison, with a larger benefit for the 1B model than the 7B model. Omitting it does not cause catastrophic degradation within the tested budget. Stage 1 therefore provides a useful initialization and a shorter architectural experimentation cycle without being a strict prerequisite.
- The comparison estimates Stage 1 global-model compute at approximately two forward-pass equivalents and Stage 2 at three. Stage 2-only runs therefore receive an additional 9.8B×2/3=6.5B tokens, increasing their Stage 2 length by 17%.
- Memory usage is not matched, and the authors note that the approximation may favor Stage 2-only training. The bits-per-byte advantage narrows over training but remains positive within the tested budget.
- Extrapolation to larger budgets is uncertain because the trajectories may depend on the learning-rate schedule and data mix.
6.3 Selecting the Right Local Model Architecture for Fast Inference
Local architectures are selected using measured decoding throughput and prefill latency because FLOPs per byte do not fully predict execution speed. With batchsize=1 on H100 GPUs, the paper reports approximately 125 bytes/s for Bolmo versus 150 bytes/s for its source at equal compression, and approximately 1 s versus 0.8 s to prefill 72K bytes.
- Increasing Bolmo’s compression to approximately 6.6 bytes per patch overtakes the source model in the reported inference measurements. Further compression improves speed while reducing task performance.
- mLSTM implemented with Tiled Flash Linear Attention achieves higher decoding throughput than the tested Mamba2 and Gated DeltaNet alternatives at comparable FLOPs per byte. The reported FLOP–speed fits have R² approximately 0.63–0.66.
- FLOP accounting can distort comparisons when embedding lookups are counted as dense one-hot matrix multiplications. The speed measurements apply to the tested hardware and implementations; efficient batched serving remains future work.
The left panels show Bolmo overtaking its source at approximately 6.6 bytes per patch in the reported inference measurements. The right panels show different throughput and latency at similar FLOPs per byte across local architectures, supporting selection by measured execution speed under the H100 batchsize=1 conditions.
7 Conclusion
The conclusion presents byteification as a complement to training byte-level models from scratch. The experiments recover much of the source models’ benchmark performance, improve character understanding, demonstrate adjustable patch compression, and transfer one compatible post-training update. Performance relative to the source model still varies by task.
8 Future Directions
The eight Future Directions cover training the architecture from scratch, learning non-causal boundaries end-to-end, jointly scaling patch size and local capacity, multi-byte prediction, less destructive continued training, specialized LTLM sampling, alternative atomic input units, and batched inference optimization.
- The current boundary predictor uses external labels, and the decoder predicts only the immediate next byte. Benefits from training the architecture from scratch and from multi-byte prediction have not been tested here.
- The paper proposes PEFT methods to reduce forgetting and patch-position-dependent sampling to investigate generation behavior.
- UTF-8 encoding bias remains unresolved. Efficient batching must also handle the variable correspondence between bytes and latent patches; the paper points to https://github.com/Aleph-Alpha/vllm as related infrastructure.
Appendix
- A Additional Ablations compares oracle and learned boundaries and contrasts byteification with standard continued training. Oracle boundaries substantially reduce the Stage 1 gap, while continued training on Bolmo Mix also degrades several source-model capabilities. Bolmo improves character understanding over the continued-training control, but additional gaps remain in other categories.
- B Benchmark Details specifies OLMES evaluation formats, few-shot conditions, metrics, and sampling parameters for the 7B and 1B suites. The main experimental setup says BigCodeBench is excluded, whereas the appendix evaluation specification lists it; the main result table reports no BigCodeBench score.
- C CUTE-Style Training Data describes approximately 75M synthetic tokens, approximately 0.04% of the data mix, generated using https://github.com/Leukas/CUTE from 150000 words with zero overlap with CUTE test words. English-only tasks include spelling, reversal, character swapping, deletion, and substitution, while improvements also occur on multilingual EXECUTE. The authors report that byte-level models otherwise fail to acquire this character knowledge during their short training schedule.
- D Does Post-Training Byteified Models via Task Arithmetic Always Work? examines embedding resetting using cross-entropy on Tulu 3 examples across Olmo, Qwen, Gemma, and Llama models. Compatibility varies substantially, with Olmo 3 RL-Zero models preserving performance especially well. The weak relationship with model size does not establish a general transfer guarantee.
- E Embedding Rank Analysis finds substantial high-rank structure in most tested input and output embedding matrices, with Qwen3-4B-Base input embeddings as a notable exception. A lower-dimensional local representation followed by linear up-projection imposes a hard rank limit, motivating Bolmo’s use of equal local and global dimensions.
- F Full Hyperparameters specifies one encoder layer and four decoder layers, each combining mLSTM and FFN, with SwiGLU and RMSNorm. Both model sizes use 75K Stage 1 steps with batch size 32 and 150K Stage 2 steps with batch size 64, AdamW, weight decay 0.1, β_1=0.9, β_2=0.95, maximum gradient norm 0.5, and warmup followed by linear learning-rate schedules. Tables 7 and 8 also provide model-specific dimensions, learning rates, and per-accelerator throughput estimates from H100 training runs.
Brief Thoughts
The boundary ablations provide the clearest evidence for Bolmo’s architectural choices: accurate labels alone are insufficient when patch representations do not align with the source model. The continued-training control helps distinguish conversion effects from adaptation effects and shows that Bolmo’s character advantage persists when the source model receives the same targeted training data. Byteification reuses an existing model’s pretraining investment. The experiments do not compare total development costs against training a byte-level model from scratch. Character gains depend partly on synthetic training tasks, and stronger code pass@16 coexists with generally weaker pass@1. These findings do not establish a universal advantage for byte-level modeling. The speed results apply to the tested H100 batchsize=1 setting. Efficient production batching, transfer across a wider range of source models, and reliable reuse of arbitrary post-training updates remain open questions.