Gorio Tech Blog search

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel | Summary

|

Contents

This article explains the key points of Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel.

Read this article in Korean


Summary

  • Taliesin persists a verified solution as KV state and grafts it into a fresh inference context; Galahad uses this mechanism to solve, verify, deposit, and retrieve knowledge without changing model weights. The paper evaluates both the answer quality of the frozen model with stored knowledge and the cost of avoiding repeated prefill or reasoning.
  • Under a pinned deterministic configuration, own-position graft produces SHA-256-equal logit bytes to fresh computation in the reported trials. On AIME 2025, a confidence-gated Gemma-4-12B system moves from 80.0% to 93.3% with a verified library; when eight stored problems recur, it returns all eight answers in 61 decode tokens, versus 401,026 tokens and no solutions for a base best-of-five run.
  • The background distinguishes this proposal from persistent KV-cache infrastructure such as LMCache and CacheGen, prefix reuse such as RadixAttention, approximate non-prefix reuse such as CacheBlend, and knowledge added through retrieval tokens or weight updates. The paper’s specific claim combines persistent state reuse with tested byte equality at the block’s captured position.
  • Recurrence reads verified answers and does not demonstrate new reasoning; held-out transfer instead tests adaptation to new instances. Byte equality is established for tested single-block restorations at captured positions under matched configurations, not for arbitrary relocation, monolithic positional composition, or transfer of identical logit bytes across GPU architectures.

3 Methodology

The proprietary state-capture and restore engine is described by its input-output behavior rather than its implementation. Local tests use Gemma-4-12B on an RTX 5090, with CPU and CUDA fixture checks; scale-up uses Gemma-4-31B on H100 and B200 GPUs. Board power is sampled at 0.5 to 2 Hz, and published Gemma 4 and Qwen AIME 2026 model-card scores serve as anchors, not same-benchmark baselines for the paper’s AIME 2025 runs.

3.2 The exactness guarantee, defined

A graft is byte-exact when SHA-256(bytes(ℓgraft)) = SHA-256(bytes(ℓfresh)) at the corresponding position; identical logits imply zero distributional KL divergence and the same argmax. The tested block is restored as a prefix at [0, N), with a new query computed at [N, N + q), under GGML_DETERMINISTIC=1 and CUBLAS_WORKSPACE_CONFIG=:4096:8. This is a within-configuration and within-architecture guarantee, not a claim that GPU inference is generally deterministic.

The diagram separates the deposit path from later use: verification precedes storage, and a new query is routed to one block before grafting. Its two exits correspond to the separately evaluated recurrence and transfer regimes.

Galahad’s solve–verify–deposit–route–graft workflow, with separate recurrence and transfer paths.
Galahad’s solve–verify–deposit–route–graft workflow, with separate recurrence and transfer paths.

3.3 The flywheel protocol (Galahad)

Galahad spends additional inference effort on a problem the frozen model cannot reliably solve, externally verifies the result, and deposits its KV state. For AIME, the check executes generated code and confirms its printed answer against the known answer; for LiveBench, it checks ground truth. Recurrence retrieves the stored answer by exact lookup, whereas transfer classifies a new instance, grafts one relevant block, and asks the model to adapt its method; selecting one block avoids the interference observed with a flat prefix.

4 Empirical Results

Sequentially prefilling block B after restoring A yields a median KL of 0.012 against a monolithic fresh concatenation, with 3 of 50 argmax flips; stitching independently captured final states yields KL 13.85 and 50 of 50 flips because B never attended to A. Sequential composition therefore substantially reduces divergence but is not byte-equal to the monolithic reference.

Eight solution blocks copied between H100 pods of the same architecture retained matching file SHA-256 digests and functioned 8 of 8 without re-deposit; matching file hashes alone do not establish cross-machine logit-byte equality. The reported negatives include no benefit on an already-solved 216-of-216 task family, harm from an inferior cached procedure, and LiveBench recurrence of only 5 of 20 for multi-exemplar blocks.

4.1 The graft is bit-exact at its own position (KL = 0, SHA-equal)

On a two-layer fixture, 50 graft samples had median and p99 KL of 0, no graft failures, and no argmax disagreements; logit-byte hashes matched fresh computation in 5 of 5 trials. On Gemma-4-12B, the 50-sample grafted KL matched the fresh self-comparison floor on the order of 10^-27, with zero argmax disagreements, and the reported SHA equality test passed. Restoration into a separately created inference context also reproduced the expected logits under the tested conditions.

The plotted KL values place own-position restoration at the measured self-comparison floor, relocation and sequential composition near 0.01, and independent stitching far higher. Only the tested own-position case carries the reported byte-equality result.

Log-scale KL comparison of own-position graft, positional relocation, sequential composition, and naive stitching.
Log-scale KL comparison of own-position graft, positional relocation, sequential composition, and naive stitching.

4.2 The numerical regime of positional graft: own-position is the unique exact point

Moving a captured block to a different absolute offset gives graft-versus-fresh KL near 0.015 across tested offsets {8, 128, 1024, 4096}. A graft-free control finds about 0.014 KL between fresh computations of identical tokens at positions 0 and 128, while a fresh self-comparison is zero. The authors attribute the residual to floating-point position sensitivity under rotary encoding; their demonstrated byte-exact operating point is restoration at the captured position.

4.3 The cost engine: an 85.6x prefill subsidy

For an 11,994-token prompt on Gemma-4-12B, cold prefill takes 1,547.3 ms; restoring cached state and advancing one token takes 18.1 ms, an 85.6x speedup. A separate 5,569-token continuation falls from 1,087 ms to 13 ms after enabling full-size KV for windowed layers and pinning the request to its cache-bearing slot; the former fix raises the stated 64k-context memory use from 19.7 to 25.5 GB.

4.4 Widening the usable context 87x at zero extra memory

A disk store holds 2,854,766 tokens in 88 blocks, or 87.1x the configured 32,768-token serving window, using 40.6 GB on disk without extra accelerator memory relative to that serving setup. Seven of seven planted needles are recovered at sampled depths from 0 to 2.82M tokens; restoring a selected block takes about 0.29 s plus prefill of roughly 35 new tokens. Querying the last needle while the window remains on the first block returns a wrong number, showing that the whole store is not resident in live attention; a smaller test restores a 464-token skill manual and routes six of six tasks.

Capability, recurrence, and transfer

The capability tests distinguish a full-benchmark system score from exact-answer retrieval on recurring problems and adaptation to independently verified or temporally held-out instances. Only the latter tests directly assess application of cached methods to new problems.

On the local AIME 2025 evaluation, adding the verified library raises lean routing accuracy at a modest increase in decode tokens, whereas the sampled configuration uses substantially more tokens. The dashed model-card anchors refer to AIME 2026 and are not same-benchmark comparisons.

AIME 2025 accuracy against decode tokens per problem for frozen Gemma-4-12B configurations, with published AIME 2026 model-card anchors.
AIME 2025 accuracy against decode tokens per problem for frozen Gemma-4-12B configurations, with published AIME 2026 model-card anchors.

4.5 The flywheel makes a 12B smarter: 80% to 93.3% on post-cutoff AIME

On 30 AIME 2025 problems, the frozen 12B’s best reported reasoning-plus-verification configuration scores 24 of 30 (80.0%). A confidence gate marks eight problems unsure; a grafted library recovers six, yielding 22 confidently correct direct answers plus six library answers, or 28 of 30 (93.3%). In a separate lean code-routing comparison, a grafted library scores 90.0% at 4,360 decode tokens per problem, versus 76.7% at roughly 4,100 tokens without cached knowledge or 76.7% at roughly 25,000 tokens with sampling-and-voting; these are different configurations from the 93.3% run.

The log-scale bars contrast an unsuccessful 401,026-token base sampling run with 61 tokens used to return eight verified stored answers. The comparison establishes a recurrence cost reduction, not transfer to unseen problems.

Decode-token comparison for base best-of-five and verified-block recurrence on eight AIME problems.
Decode-token comparison for base best-of-five and verified-block recurrence on eight AIME problems.

4.6 The flywheel makes a 12B cheaper: 6,574x fewer tokens on recurrence

When the same eight AIME problems recur, exact routing to individual verified blocks returns 8 of 8 answers in 61 total decode tokens, or 7.6 per problem. A base best-of-five run solves 0 of 8 despite using 401,026 tokens, giving the reported 6,574x token reduction; this comparison measures stored-answer retrieval, not unseen-problem solving. The 2.8 s run has an integrated energy estimate of 0.053 Wh, but limited power-sampling resolution motivates a conservative 0.145 Wh bound and a reported energy-reduction range of roughly 3,000x to 8,700x.

4.7 Transfer: the same store generalizes to new data, within a stated boundary

For seven independently verified AIME variants with swapped constants, the 12B routes 7 of 7 to the intended block but autonomously solves 5 of 7. It times out on a hard-coded radical in #11 and over-adapts a structural constant in #28; mechanical substitution of the intended literals solves all seven. No clean same-structure variant was available for #10, so this test covers seven rather than eight problems.

Scale and operational behavior

The scale-up reports functional outcomes for a frozen 31B on Hopper and a separate byte-level replay on Blackwell. Portability, routing, and paging tests address whether stored blocks work on another same-architecture server and whether restoring them retains a cost advantage over fresh prefill.

4.9 Cross-architecture scale-up: a frozen 31B on an H100

On a rented H100, frozen Gemma-4-31B with Taliesin scores 8 of 8 on recurrence, 7 of 7 on held-out AIME variants, and 30 of 30 for a full-benchmark system that includes cached answers for the eight hard problems; no bare-model H100 re-run is reported. Transfer uses 25,361 decode tokens, 72.15 Wh, and 8.4 min; the full-30 run uses 56,189 tokens, 194.45 Wh, and 21.9 min. On a temporally held-out LiveBench split with no question-id overlap, single-block routing scores 43 of 60 (71.7%), versus 56.7% with merged blocks.

4.10 Cross-architecture byte-exactness: a pre-registered B200 replay

A pre-registered B200 replay with frozen Gemma-4-31B reports zero graft failures, zero argmax disagreements in 50 trials, grafted KL at the fresh self-comparison floor, and SHA-equal logit bytes in 10 of 10 trials. Graft-free prefills at different positions still diverge: the measured median near 7 × 10^-4 is about nineteen times smaller than the 12B result. Byte exactness was measured on the RTX 5090 and B200, but not on the functionally tested H100; the authors report that their prediction of a nonzero positional effect held, whereas its predicted magnitude did not.

4.12 Systems behavior: router misrouting and disk paging, measured

In-library one-shot transfer routing is correct on 15 of 15 local cases, with a reported Wilson 95% interval of about 79.6% to 100%; four of four off-distribution prompts abstain. A confident wrong graft remains possible because there is no confidence gate between graft and answer. Across 748 to 5,977 tokens, re-prefill grows from 96.1 to 748.3 ms while first-touch disk restore grows from 57.7 to 85.8 ms, increasing the measured restore subsidy from 1.7x to 8.7x; the page cache was not force-dropped, and the two-slot test makes no high-concurrency claim.

5 Discussion

The cost argument depends on reusing an externally verified deposit rather than repeating sampling for every query. The exactness argument applies to matched own-position single-block restoration: relocation is not byte-exact, and independent stitching fails; sequential composition is functional but diverges from monolithic prefill. The paper reports pre-run input hashes and post-run output hashes with retained generations and solver code for scoring checks, while generation reproduction requires the closed Merlin/Taliesin suite.

5.3 Limitations

Transfer can fail when cached code embeds problem-specific constants, and multi-exemplar blocks show weak recurrence; caching an already-known procedure can also harm performance. The router lacks a post-graft confidence gate, byte equality requires matched architecture and configuration, and paging evidence is single-node rather than a concurrency study. The conclusion’s 93.3% and 100% AIME figures are system scores, distinct from the respective 5-of-7 and 7-of-7 held-out transfer results; the reproducibility statement specifies what can be checked without the closed engine and credits Sietse Schelpe as sole author and inventor.

Appendix

  • The supplied PDF has no separate experimental appendix. Its reproducibility statement records hashes for datasets and runner scripts before runs and for results and raw generations afterward; retained solver code permits score checks without access to the proprietary engine, but reproducing generation requires its closed benchmark suite.

Brief Thoughts

The strongest numerical evidence pairs matched fresh-versus-graft logit-byte tests with a graft-free positional control; the systems evidence separately measures prefill and restore costs. The main interpretive distinction is between cheap retrieval of stored answers, mixed live-and-cached benchmark scores, and transfer to new problems, which depends in part on whether the cached method can be adapted without re-deriving problem-specific constants.