Shaping capabilities with token-level data filtering | Summary
29 Jan 2026 | Paper Review Token-Level Data Filtering LLM Safety LLM Pretraining Weak SupervisionContents
- Summary
- 1. Introduction
- 2. Motivation and related work
- 3. Setting and approach
- 4. Token-level data filtering works and scales
- 5. How to train your classifier
- 6. How bad are bad labels?
- 7. Wrapping up
- Appendix
- Brief Thoughts
This article explains the key points of Shaping capabilities with token-level data filtering.
- 2026-01-29 (arXiv), Preprint
- Rathi, Neil, Radford, Alec.
- Anthropic, Stanford, Independent
- Paper
- Github
Summary
- Shaping capabilities with token-level data filtering studies whether pretraining interventions can suppress unwanted capabilities while preserving useful ones. Using medicine as a proxy forget domain and biology plus general non-medical tasks as retain domains, the authors compare document filtering, token loss masking, and token removal.
- Across Transformers from 61M to 1.8B parameters, token filtering achieves a better forget–retain tradeoff than document filtering in the tested loss comparisons. At 1.8B parameters, token removal produces over a 7000× loss-matched effective compute slowdown on medical text and requires 13× more adversarial finetuning tokens than RMU to recover baseline medical loss; these are different measurements, not interchangeable estimates of adversarial cost.
- The classifier pipeline uses sparse autoencoder features to generate noisy token labels and distills them into inexpensive probes on bidirectional language-model representations. Token removal also improves held-out medical refusal rates, but the experiments do not establish universal capability removal or deployment safety beyond the medical proxy, tested scales, and attack setup.
1. Introduction
Capability shaping aims to reduce unwanted abilities without substantially damaging useful ones. The paper motivates filtering during pretraining by arguing that capabilities left in a pretrained model can remain accessible despite post hoc safeguards; Figure 1 summarizes the increasing medical suppression observed with scale.
- Short subsequences within otherwise benign documents can teach a capability, motivating token-level rather than exclusively document-level selection.
- The experiments separately examine suppression, relearning robustness, refusal training, classifier cost, and imperfect supervision.
- The source provides the repository identifier neilrathi/token-filtering.
Medical loss-matched slowdown rises much faster for token removal than for document filtering as pretraining compute increases. The biology panel also shows a retention cost, so the medical result is a domain-specific effective-compute comparison rather than suppression without collateral damage.
2. Motivation and related work
The related work distinguishes post hoc behavioral safeguards, machine unlearning, and interventions during pretraining. The motivation for finer-grained filtering is that removing acquired capabilities is difficult, while document filtering can discard benign content or miss informative spans inside mostly benign documents.
- Prior document-filtering studies report suppression of CBRN-related capabilities, but do not establish that document granularity is optimal.
- Token-level data-attribution results show that tokens within one document can differ substantially in their influence on downstream generations.
- Gradient routing and distillation from unlearned models are discussed as alternatives that exploit model representations rather than relying solely on external content classifiers.
3. Setting and approach
The objective is near-baseline performance on a retain set and maximally reduced performance on a forget set. Medicine serves as a proxy because the available training budget cannot produce models sufficiently capable for informative evaluation of actual dangerous capabilities.
- The workflow labels a pretraining corpus, filters forget-related content, trains models at several compute budgets, and evaluates text loss, multiple-choice accuracy, and free-response quality.
- Biology is a closely related retain domain, exposing collateral damage near the medical classification boundary.
3.1. Data and data filtering
Models are pretrained on FineWeb-Edu because early experiments with default FineWeb yielded weak biomedical benchmark performance even at 1.8B parameters. Unless otherwise specified, classifier thresholds maximize F1 on a held-out FineWeb-Edu subset.
- Loss masking excludes losses for forget targets but leaves their tokens visible in the forward-pass context.
- Removal replaces classifier-identified forget tokens with <|hidden|> and masks their losses, hiding their identities while disrupting contextual coherence.
- Figure 2 distinguishes suppressed target losses from removed input content; neither operation guarantees that all capability-relevant information has been identified.
Both variants suppress training loss at classifier-identified medical target positions. Only removal hides their identities from subsequent input context, explaining why masking and removal can produce different representations and suppression effects.
3.2. Model training
The authors train compute-optimal, modified GPT-2-style Transformers from 61M to 1.8B parameters using AdamW and µP learning-rate transfer. Models through 521M parameters use 2×NVIDIA H200s, while the 1B and 1.8B models use 8×NVIDIA H200s.
- The main multiple-choice and chat evaluations use 1.8B models because smaller unfiltered baselines were too weak to provide useful suppression signal.
- A custom multiple-choice instruction mix elicits accuracy under a limited compute budget; chat training uses smol-smoltalk.
- These evaluations supplement text loss, which is closely tied to the training objective directly modified by filtering.
3.3. Evaluation
Text-loss evaluations use medical PubMed articles, biology bioRxiv articles, and general non-medical arXiv and PhilPapers articles. Claude Sonnet 4 removes documents assigned to the wrong domain, and the token-labeling pipeline excludes unrelated tokens.
- Medical multiple-choice tests are MedMCQA, MedQA-USMLE, and MMLU Medicine; retain tests use MMLU biology, non-biomedical STEM, and non-STEM subsets.
- Chat-trained models answer HealthSearchQA questions, with Claude Sonnet 4 judging correctness, relevance, and coherence under the definitions reproduced in Appendix E.
- The free-response control is medically filtered Alpaca, not AlpacaEval or its evaluation harness.
4. Token-level data filtering works and scales
The results assess the forget–retain frontier, suppression across pretraining scales, recovery through adversarial finetuning, and control through refusal training. These tests distinguish weaker domain performance from resistance to relearning and from selective refusal behavior.
4.1. Token filtering Pareto dominates document filtering
For 521M-parameter models, the authors sweep classifier thresholds to filter 3%, 6%, 12%, 25%, or 50% of pretraining tokens. Figure 3 shows that token filtering achieves lower biology loss at matched medical loss than document filtering, supporting a Pareto improvement in this experiment.
- Higher medical loss represents stronger suppression; lower biology loss represents better retention.
- The comparison measures downstream loss tradeoffs, not classifier precision and recall directly.
Threshold sweeps for 521M-parameter models give token filtering a more favorable medical–biology loss tradeoff than document filtering. Marker sizes encode the filtered fraction; equal medical loss can be achieved with lower biology loss, rather than merely removing fewer documents.
4.2. Filtering works, and filtering scales
Figure 4 traces forget–retain loss frontiers across model sizes, with token filtering generally achieving higher medical loss at a given retain loss. Estimates from the unfiltered compute-to-loss relationship give over a 7000× medical slowdown for the largest token-removal model, compared with around 30× for document filtering.
- This is a loss-matched effective-compute estimate, not a measured wall-clock slowdown or a guarantee that an adversary needs 7000× more compute.
- Suppression strengthens relative to the baseline over the tested scale range, although token removal also incurs a substantial biology text-loss penalty.
- The paper describes log-log linear interpolation; Figure 19 extends the baseline relationship below the measured baseline compute range for the largest token-removal comparison.
Each point represents a model size, showing how the forget–retain tradeoff changes across scale. Token removal gives the highest medical loss, but its biology loss also makes collateral damage visible; the non-medical and biology panels should therefore be read separately.
Instruction-tuned filtered models perform near chance on MedMCQA and MedQA-USMLE, while retain-domain multiple-choice scores remain broadly similar. Token filtering also reduces the quality of HealthSearchQA responses, which often become repetitive, irrelevant, or fluent but factually incorrect.
- Figure 5 shows reductions in biology and non-STEM accuracy, particularly for token removal, so retention is not literally unchanged.
- The main text reports 4× lower coherence and relevance and 10× lower correctness; Figure 6’s caption reports up to 20× lower correctness and 3× lower relevance and coherence. These source statements should not be collapsed into one numerical estimate.
- The Alpaca control shows generally comparable free-response performance with slight degradation, but its appendix results use a single random seed.
Filtered models approach the 25% chance line on MedMCQA and MedQA-USMLE, while MMLU Medicine retains more residual performance. Retain-domain scores are broadly similar, but the bars show biology and non-STEM reductions, particularly for token removal.
Loss masking produces especially few responses judged factually correct, while removal preserves more fluency without restoring medical correctness. The criteria separate factual accuracy from attempted relevance and linguistic fluency; the caption’s fold-change wording differs from the main-text summary.
4.3. Filtering is more robust than unlearning
The robustness experiment grants an adversary open-weight access and finetunes on Common Pile PubMed text until medical loss matches the unfiltered baseline. At 1.8B parameters, token removal requires 13× as many recovery tokens as RMU; the corresponding ratios are 3× for token loss masking and 1.5× for document filtering.
- RMU uses PubMed as its forget set and Project Gutenberg as its retain set, preserving retain representations while steering forget representations toward a random vector.
- Finetuning hyperparameters are selected to recover baseline loss in the fewest steps, rather than imposing one shared learning rate.
- Figure 7 shows generally declining recovery-token requirements relative to pretraining tokens as scale increases, with a steeper decline for RMU. Filtering is more robust in this comparison, not immune to relearning or proven superior to every unlearning method.
Recovery tokens are normalized by pretraining tokens, and the requirement generally decreases as models grow. Token removal has the largest requirement at the largest scale, while RMU becomes easier to reverse relative to pretraining size; the curves do not imply immunity to recovery.
4.4. Token-level filtering makes alignment easier
Linear probes trained on a 2.05M-token subset test whether filtered models still distinguish medical from non-medical tokens. Figure 8 shows that the classification gap relative to unfiltered models narrows with scale; the separately trained pretraining filter provides an additional reference.
- Domain recognition does not imply preserved medical question-answering competence.
- A separate medRxiv experiment finds weaker discrimination between neurology and infectious disease under token filtering, despite strong discrimination between either subdomain and non-medical content.
- The authors interpret this pattern as consistent with distinguishing trained from untrained content rather than preserving detailed distinctions inside the forget domain.
Filtered representations remain useful for medical-versus-non-medical classification, and their gap relative to the unfiltered model series narrows with scale. The horizontal reference is the pretraining filter trained on 4× as many classification tokens, not another point in the unfiltered model series.
Already chat-trained 1.8B models are finetuned to give single-sentence refusals on HealthSearchQA and normal completions on Alpaca, then evaluated on held-out examples across three random seeds. Token removal roughly doubles medical refusal rates relative to the baseline without a notable increase on Alpaca, whereas document filtering also refuses benign questions.
- Loss masking generates slightly fewer medical refusals than the baseline while maintaining low Alpaca refusal rates.
- Figure 9 supports improved selective refusal specifically for token removal, not every token-filtering variant.
- Refusal rates remain far below complete coverage, and the experiment is not a broad jailbreak-resistance evaluation.
Token removal increases held-out medical refusals while keeping Alpaca refusals low. Document filtering produces a less selective policy, and loss masking does not improve the medical refusal rate over the baseline, limiting the alignment advantage to the removal variant in this test.
5. How to train your classifier
The classifier predicts whether a token contains knowledge relevant to the forget domain, using relevance as a proxy for whether the token teaches the unwanted capability. The paper explicitly distinguishes relevance from causal influence: capability-improving data can escape the classifier, and identified content need not improve the capability.
5.1. Sourcing ground-truth labels
Token annotations are generated using the Gemma 2 9B layer-31 SAE with width 16k. Claude 3.5 Haiku explains latents through the Neuronpedia API, Claude Sonnet 4 identifies medical explanations, and embedding scores below 0.9 are discarded, leaving 600 medical latents.
- A token is labeled medical when activation is at least 4SD above the mean on at least two medical latents; adjacent tokens with positive activation on at least one medical latent are then added iteratively until convergence.
- Propagation aims to label spans such as “insert the catheter,” not only medical keywords. Figure 10 illustrates the resulting annotations and their remaining noise.
- The 128k-document annotation corpus is roughly 75% academic and 25% web material. Classifiers use 8.2M balanced forget–retain training tokens, 1.64M validation tokens, and 0.82M FineWeb-Edu test tokens.
- SAE inference labels a subset for distillation, not the full pretraining corpus. Retokenization labels a whole new token as forget when it overlaps a Gemma forget token.
Highlighted spans distinguish medical information from neighboring text, including disease-related material inside a hydroponics passage. Fragmented selections and uncertain boundaries illustrate why the SAE annotations are treated as noisy supervision rather than independently verified ground truth.
5.2. A good representation is hard to find
The classifier is an L-BFGS linear probe on bidirectional contextual representations, with layers swept to select the best probe. The biLM concatenates representations from separately trained left-to-right and right-to-left autoregressive models; the probe uses frozen representations rather than fully finetuning on noisy labels.
- ModernBERT-large achieves validation/test F1 of 0.794/0.812; a 65M RoBERTa-like model achieves 0.808/0.834.
- Two 61M autoregressive models, 122M parameters altogether, each trained for 4.8B tokens at 4× Chinchilla, yield 0.830/0.880. A 50% PubMed and 50% FineWeb-Edu pretraining mix yields 0.834/0.877.
- Scaling the recipe gives 0.844/0.885 at the 113M stage and 0.856/0.894 at the final 224M stage. Table 1 shows that the improvements are not strictly monotonic on every metric.
- Figure 11 illustrates contextual predictions and ambiguous biological terminology. It is a qualitative example, not an independent quantitative audit.
The final recipe improves validation/test F1 from ModernBERT-large’s 0.794/0.812 to 0.856/0.894. Intermediate rows change architecture, objective, corpus composition, and scale; PubMed upsampling improves validation F1 but slightly reduces test F1, so not every change improves both metrics.
Context-dependent probabilities select spans within mixed passages rather than uniformly labeling whole documents. Predictions in the cheese passage expose ambiguity around biological terminology, making the examples useful for inspecting boundaries but insufficient for establishing classification accuracy.
5.3. Document-level classification
The document classifier uses a probe on the 224M biLM and the same 128k-document corpus, with Claude Sonnet 4 supplying document labels. It achieves validation F1 of 0.922 and test F1 of 0.941; these document-level scores are not directly comparable to token-level F1, and the resulting filter does not match the token-filtering loss frontier.
6. How bad are bad labels?
The label-quality experiments examine noise in filtering decisions, changes to the precision–recall threshold, and weaker or coarser classifier supervision. Filtering is sensitive to label errors, but aggressive thresholds and token-level generalization from weak labels help under the tested conditions.
6.1. They’re pretty bad…
The authors randomly flip outputs of the final 224M biLM classifier and train filtered models through 521M parameters. Figure 12 shows a steep reduction in loss-matched slowdown when errors increase from a low starting level, followed by diminishing additional degradation at high error rates.
- For classifier accuracy a = 0.89 and flip rate r, the plotted error rate is 1 − a(1 − r) − r(1 − a).
- The stated 11% initial error is relative to SAE-generated labels, not independently established human ground truth.
- Random flips test one noise model and do not represent all systematic errors in real capability classifiers.
The steep initial decline shows that filtering effectiveness is sensitive to added label errors. At high error rates, slowdown approaches the unfiltered level across the tested sizes; the starting error rate is measured against SAE-generated reference labels.
6.2. …but good things come to those who scale
A threshold sweep using token loss masking and models through 521M parameters filters 3%, 6%, 12%, 25%, or 50% of tokens. Figure 13 shows stronger medical suppression with aggressive filtering, but also greater retain loss at a fixed scale.
- The proposed mechanism is that high-recall filters remove proportionally more forget content while leaving enough retain data for larger, more sample-efficient models.
- The claim about indefinitely scaling compute is an extrapolation; this experiment stops at 521M parameters.
- The preferred plotted direction is high medical loss and low retain loss, despite the main text’s inconsistent reference to the bottom right.
Filtering larger token fractions increases medical loss but also worsens retain loss at fixed scale. Larger models partially compensate for lost retain examples, making the benefit conditional on scale rather than eliminating the precision–recall tradeoff.
6.3. Token-level classifiers generalize from weak labels
Token probes on the 61M biLM are trained using sentence or document labels from Claude Sonnet 4, assigning each token its enclosing unit’s label. When these probes filter corpora for models through 521M parameters, Figure 14 shows somewhat weaker suppression and scaling than token-supervised probes, but still effective filtering.
- This is token-level prediction trained with coarse supervision, not document-level filtering.
- The result supports using coarser annotations when token labels are unavailable while retaining a token-level filtering interface.
Sentence- and document-supervised token classifiers still suppress medicine, although token supervision generally gives a more favorable frontier, particularly against non-medical loss. The comparison concerns supervision granularity while preserving token-level filtering.
For weak-to-strong generalization, a 13M biLM is pretrained at 4× Chinchilla and probed using varying amounts of labeled data from up to half the classifier training set. Its predictions label the other half for a probe on the 224M biLM; Figure 15 shows that the stronger token probe improves over its weak teacher, while the stronger document probe consistently underperforms its teacher.
- Token and document experiments use their corresponding reference labels, making this a within-task teacher–student comparison.
- The result applies to this biLM and linear-probe recipe, not to all forms of document-level weak supervision.
The stronger token probe improves on its weak teacher, whereas the stronger document probe falls below its teacher. This demonstrates task-dependent weak-to-strong generalization within the tested biLM and linear-probe setup, not a universal failure of document-level weak supervision.
7. Wrapping up
The conclusion presents token filtering as an effective pretraining intervention while emphasizing that content relevance is an imperfect approximation to capability-improving data. The authors propose influence-based selection, representation-aware pretraining, gradient routing, and distillation as directions for improving capability shaping.
- The supervision pipeline uses annotators more capable than the trained models; maintaining that capability gap becomes harder at frontier scale.
- The authors raise possible U-shaped scaling if larger models learn from rare leaked examples or externally supplied in-context information, and call for evaluations better suited to capability shaping.
- They recommend defense in depth rather than replacing posttraining or deployment safeguards, particularly for retrieval attacks and dual-use information.
- Applications to character priors, misalignment-related data, and scheming capabilities remain hypotheses rather than demonstrated results.
Appendix
- A. Implementation Details, A.1. Architecture, and A.2. Optimization and Hyperparameters specify RoPE, ReLU², pre-RMSNorm, width-to-depth ratio 64, and the cl100k_base tokenizer. Pretraining models use block size 2048, classifier models use 1024, and effective batch size is 327,680; AdamW uses β1 = 0.9, β2 = 0.95, 10% linear warmup, and cosine decay to 0.1× maximum learning rate.
- A.3. Instruction Tuning reports one pass through 122k multiple-choice examples at learning rate 10⁻⁴, with held-out accuracy increasing from 0.23 to 0.66. The text identifies the MMLU auxiliary training set as the core and lists ARC Easy, ARC Challenge, BIG-Bench Abstract Narrative Understanding, BoolQ, MCTest, OpenBookQA, PIQA, and RACE Middle and High; chat training uses one pass through 460k smol-smoltalk examples at learning rate 10⁻⁵, while the full 1.1M-example smoltalk mix degraded coherence.
- B. Evaluation Details and B.1. Estimating loss-matched baseline compute estimate baseline compute from the log-log relationship Lb ∝ Cb^(−α). The appendix writes relative slowdown as Cb*/Cf*, whereas the reported greater-than-one slowdowns and Figure 19’s horizontal comparisons use the reciprocal interpretation; this is a source-level ratio-direction inconsistency. B.2. Multiple choice evaluations evaluates base models by selecting the answer string with the lowest conditional loss and finds broadly similar suppression trends.
- B.3. Robustness specifies RMU learning rate 1 × 10⁻⁴, weight decay 0.01, batch size 8192, α = 100.0, c = 20.0, and 1,000 optimization steps, with loss computed at the middle layer and updates restricted to MLPs in that layer and the two preceding layers. Adversarial finetuning sweeps learning rates from 1 × 10⁻⁵ to 1 × 10⁻³ and weight decay in {0.01, 0.1}, using effective batch size 40,960. B.4. Training to generate refusal tokens reproduces the advantage of token removal with <|refusal|>, while B.5. Training dynamics shows that delaying filtering until 40% of training makes it around an order of magnitude less effective.
- C. Classifier Details and C.1. Defining the forget and retain sets define medicine around clinically useful information, the medical and pharmaceutical industries, devices and procedures, human physiology, virology, immunology, pathology, neurology, and medical genetics. Exclusions include non-medical biochemistry or genetics, healthcare policy or education, psychiatry, mental illness and psychology, public health and epidemiology, pregnancy and childcare, wellness and meditation, cosmetic surgery, animal behavior and cognition, and colloquial anatomy references. These boundaries are task-specific rather than a general definition of biomedical content.
- C.2. How much text is filtered? reports that around 23% of documents contain no classifier-identified medical tokens and 37% contain more than 10%. The document classifier marks 18% of documents as medical, but the SAE pipeline labels only 50% of their tokens medical. C.3. Are better classifiers actually better filters? fixes the positive fraction at exactly 20% of tokens and generally associates stronger classifiers with better downstream loss frontiers, though the normalized-AUC comparisons are not uniformly monotonic across domains.
- D. Example responses to free-response medical questions provides five randomly selected HealthSearchQA questions, with responses truncated at 128 tokens or <|im_end|>. Filtered models often produce repetition, irrelevant interpretations, or incorrect medical statements. These examples illustrate model failures and are not medical guidance.
- E. Prompts includes Identifying medical SAE features (claude-sonnet-4-20250514), Identifying medical documents (claude-sonnet-4-20250514), and Scoring HealthSearchQA responses (claude-sonnet-4-20250514). Correctness assesses factual accuracy in isolation without requiring an answer to the question, relevance accepts any attempted relevance, and coherence measures fluent English without requiring logical consistency. Several classification definitions and examples are explicitly omitted for brevity, limiting reproduction from the displayed prompts alone.
Brief Thoughts
The most informative contribution is the comparison of filtering granularity across threshold sweeps, scaling, question answering, relearning, and refusal training. Together, these experiments support token-level filtering as a complement to posttraining safeguards when wanted and unwanted information coexist within documents.
Transfer beyond the medical proxy and tested scales remains unresolved. Modest baseline accuracy on some medical exams, model-generated reference labels, one unlearning comparator, and narrow refusal tests limit conclusions about dangerous capabilities; the text-loss slowdown is neither guaranteed adversarial cost nor evidence of complete information removal.