Trust The Typical | Summary
04 Feb 2026 | Paper Review Out-of-Distribution Detection LLM Safety Unsupervised LearningContents
- Summary
- 1 Introduction
- 2 Related Works
- 3 Methodology
- 4 Results
- Evaluation Setup
- RQ1: How does T3 compare against specialized safety models and traditional OOD methods on diverse harm detection benchmarks?
- RQ2: Can T3, trained only on safe data, generalize to detect novel, unseen adversarial and jailbreaking attacks?
- RQ3: How effectively does T3 mitigate the common problem of overrefusal on benign-but-challenging prompts? How is the cold-start performance with limited ID data?
- RQ4: Does T3’s performance generalize across different languages and specialized domains without retraining?
- RQ5: Can T3 be integrated into a high-performance inference engine for practical, real-time deployment with minimal latency?
- 5 Discussion
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Trust The Typical.
- 2026-02-04 (arXiv), ICLR 2026
- Ganguly, Debargha, Sankar, Sreehari, Zhang, Biyao, Singh, Vikash, Gupta, Kanan, Kavuru, Harshini, Luo, Alan, Chen, Weicong, Morningstar, Warren, Machiraju, Raghu, et al.
- Case Western Reserve University, University of Pittsburgh, The Ohio State University, Google Research
- Paper
Summary
- Trust The Typical introduces T3, a guardrailing framework that flags deviations from a curated distribution of safe text. It adapts Forte to text by combining neighborhood statistics from three embedding models and fitting a GMM or One-Class SVM using safe examples only.
- T3+OCSVM reaches AUROC 0.9940 with FPR@95 0.0201 on OffensEval and AUROC 0.9982 with FPR@95 0.0039 on PolyGuard HR. Its advantage is not universal: Perspective API performs better on CivilComments, specialized models outperform it on several adversarial benchmarks, and LlamaGuard3-1B has lower FPR@95 on OR-Bench.
- An NVIDIA H200 deployment with vLLM and OPT-125M reports 1.5% generation overhead for 500 prompts and 6.0% for 5,000 prompts, excluding approximately 10 seconds of T3 initialization. T3 performs near chance on an HH-RLHF chosen-versus-rejected setup, illustrating that typicality cannot reliably distinguish safety when the relevant representations overlap.
1 Introduction
The introduction motivates T3 through the dependence of supervised safety filters on observed harmful content and attack patterns. T3 instead asks whether an input is compatible with a curated distribution of acceptable usage, avoiding harmful examples during detector fitting.
- The authors connect this approach to the information-theoretic typical set and the concentration of benign interactions in semantic representation space.
- The assertion that adversarial inputs must depart from typical language is a motivating hypothesis, not a guarantee that harmful and benign intent are geometrically separable.
2 Related Works
The paper situates T3 among likelihood-based OOD detection, representation-based detection, and synthetic outlier exposure. T3 retains pretrained embedding representations, avoids dual-model likelihood evaluation, and does not require generating examples of anticipated attacks.
- The related-work discussion emphasizes that high likelihood need not imply typicality in high-dimensional spaces and that task-specific fine-tuning can disrupt geometry useful for OOD detection.
- Figure 1 examines 10,000 Alpaca embeddings with d=1024: the caption reports a 90% typical annulus and a Q-Q fit of R² > 0.99.
- The distance concentration motivates studying typicality; the PCA projections do not establish a safe-versus-harmful decision boundary.
The distance histogram and Q-Q plot motivate studying a concentrated typical region rather than equating safety with proximity to the mean. The PCA panels are illustrative projections; neither they nor the dimensionality trend establish separation from harmful content.
3 Methodology
T3 describes each text embedding through its relationship to safe reference neighborhoods, then fits a detector to these compact descriptions. It combines multiple semantic views with per-point Precision, Recall, Density, and Coverage rather than estimating density directly in the original embedding spaces.
Problem Formulation:
Given a safe reference corpus X and a test set Y, T3 identifies test points that are incompatible with the reference distribution. Qwen3-Embedding-0.6B, BGE-M3, and E5-Large-v2 each produce an L2-normalized embedding, reducing encoder-specific magnitude effects.
- Figure 2 compares 10,000 safe and 2,000 toxic embeddings. Mahalanobis distance achieves ROC AUC 0.944, compared with 0.733 for Euclidean distance from the safe mean.
- The formulation groups adversarial prompts, toxic content, and off-topic queries as OOD; detecting deviation from intended use is therefore broader than detecting harmfulness.
Accounting for covariance improves separation in this example: Mahalanobis ROC AUC is 0.944 versus 0.733 for Euclidean distance. This motivates richer geometric descriptors but is not an evaluation of the complete T3 detector.
Precision indicates whether a test point lies inside any reference k-nearest-neighbor ball; Density counts the reference balls containing it, normalized by km. Recall measures the fraction of reference points inside the test point’s neighborhood, while Coverage indicates whether that neighborhood contains any reference point.
- Recall and Coverage use neighborhoods constructed within Y, so their values depend on the test set, not solely on an isolated incoming prompt.
- In Table 6, CivilComments FPR@95 falls from 0.3343 for Precision+Density to 0.0366 for Recall+Coverage and 0.0157 for all four metrics.
- The full combination is not best in every cell: Recall+Coverage gives XSafety FPR@95 0.0072, compared with 0.0108 for all four metrics.
Recall and Coverage sharply reduce false positives relative to Precision and Density alone. The full representation usually performs best, but Recall+Coverage has lower XSafety FPR@95, and these ablation values differ materially from the main benchmark results.
Theorem 3.1 states the null expectations E[Recall]=k/n and E[Density]=1/m, the bound E[Coverage] ≤ 1−(1−k/n)^m, and an asymptotic Precision limit of 1. The appendix’s symmetry argument assumes continuous distributions, while its Precision-limit argument additionally assumes compact support, a density bounded and bounded away from zero on the support, and a measure-zero boundary.
- Under the stated support regularity conditions, the paper reports an asymptotic expected Precision of 1−α when the harmful distribution assigns probability mass α outside the safe support.
- For common-support density shifts, the reported Coverage limit depends on r(y)=p_safe(y)/p_harmful(y), with continuous positive bounded densities and a finite positive limiting sample-size ratio λ.
- These expectation-level arguments do not provide operational thresholds or establish reliable classification of every harmful prompt. The local-perturbation gap also contains a sign error, discussed in Brief Thoughts.
The four features from each encoder are concatenated into T(y), a representation in R^(4K), giving 12 features for three encoders. T3 fits a GMM with component selection by Bayesian Information Criterion or an RBF-kernel OCSVM, and the paper describes sigmoid-normalized anomaly scores.
- The source calls the score negative log-likelihood for both approaches, but an OCSVM provides a decision score rather than a fitted probability density.
- Neither GMM nor OCSVM consistently dominates across benchmarks.
- The comparison with classical nearest-neighbor two-sample tests emphasizes asymmetric reuse of reference structure while acknowledging that within-set PRDC neighborhoods differ from pooled-neighbor tests.
Implementation Details.
Reference examples are randomly divided into two equal halves to reduce self-similarity bias, and embeddings are cached in PyTorch format. The implementation uses L2 distance on normalized vectors, exploiting ∥x−y∥₂²=2(1−cos(x,y)).
- The GMM component grid is {1, 2, 4, 8, 16, 32, 64}, constrained by sample size.
- The OCSVM ν grid is {0.01, 0.05, 0.1, 0.2, 0.5}; the methodology describes selection using validation accuracy without detailing that validation protocol here.
4 Results
The evaluation addresses harm detection, unseen attacks, overrefusal and sample efficiency, domain and language transfer, and serving latency. Toxicity and domain-separation results are generally stronger than the results on adversarial and benign-but-challenging prompts.
Evaluation Setup
The main text describes approximately 40K safe examples drawn equally from Alpaca, Dolly, and OpenAssistant, with no harmful examples used to fit T3. Evaluation covers toxicity and hate speech, adversarial benchmarks, PolyGuard domains, multilingual datasets, and OR-Bench.
- Safety baselines include LlamaGuard, WildGuard, MD-Judge, DuoGuard, and PolyGuard; commercial comparisons include Perspective API and OpenAI Omni Moderation.
- Representation-based baselines are adapted to text embeddings. Additional experiments examine larger embedding backbones and text-native Energy, kNN, and Mahalanobis detectors.
- The appendix specifies task-dependent ID corpora: toxicity and domain experiments include Anthropic/hh-rlhf in a four-source mixture, multilingual experiments use 30,000 OpenAssistant/oasst2 prompts, and OR-Bench uses its own safe training split.
Reported metrics are AUROC, AUPRC, F1, and FPR@95TPR. FPR@95 measures safe examples flagged at the ROC operating point that detects 95% of harmful examples; outside OR-Bench, safe evaluation examples generally come from a held-out reference corpus rather than benign examples matched to each harmful benchmark.
- The paper states that OOD labels and test data are not used to fit T3’s density estimators.
- Although the source says thresholds are not tuned per benchmark, its FPR@95 protocol selects an operating point from each evaluation ROC curve. This statistic does not demonstrate one fixed deployment threshold achieving 95% harmful recall across benchmarks.
- Separating benchmark-specific harmful text from a general-purpose safe corpus can reflect source or domain differences as well as harmfulness.
RQ1: How does T3 compare against specialized safety models and traditional OOD methods on diverse harm detection benchmarks?
Table 1 favors T3+OCSVM on five of the six toxicity benchmarks. It achieves AUROC/FPR@95 of 0.9913/0.0350 on Davidson, 0.9895/0.0408 on HatEval, and 0.9940/0.0201 on OffensEval.
- Perspective API performs better on CivilComments, with AUROC 0.9711 and FPR@95 0.1413 versus T3+OCSVM’s 0.9678 and 0.1722.
- The narrative’s approximately 37× OffensEval improvement compares against DuoGuard’s FPR@95 0.7516. Perspective API has 0.5171 and PolyGuard has 0.5315, so DuoGuard is not the best listed baseline for that metric.
- LlamaGuard3-1B’s continuous logit variant improves over discrete scoring but still trails T3 in Table 1, illustrating the importance of score extraction for ROC-based comparisons.
T3+OCSVM leads the listed AUROC and FPR@95 results on five datasets, including OffensEval FPR@95 0.0201. Perspective API leads on CivilComments, and its OffensEval result also contradicts the narrative's identification of DuoGuard as the best baseline.
RQ2: Can T3, trained only on safe data, generalize to detect novel, unseen adversarial and jailbreaking attacks?
Table 2 provides mixed evidence for unseen-attack generalization. T3+GMM leads the listed methods on AdvBench with AUROC 0.9675 and FPR@95 0.1577, while T3+OCSVM leads on JailbreakBench with 0.8622 and 0.4539.
- On HarmBench, LlamaGuard3-1B-LOGITS reaches AUROC 0.8887 and FPR@95 0.3600, outperforming T3+OCSVM’s 0.8102 and 0.5850.
- On BeaverTails, DuoGuard’s AUROC 0.8525 exceeds both T3 variants, and PolyGuard’s FPR@95 0.7269 is lower than either variant.
- XSTest remains difficult for T3, with AUROC 0.6794 for GMM and 0.5800 for OCSVM. The results do not support uniformly superior attack detection.
AdvBench and JailbreakBench favor T3, whereas HarmBench, BeaverTails, and XSTest contain substantial exceptions. Several T3 false-positive rates remain high, limiting claims of generally low-false-alarm adversarial defense.
RQ3: How effectively does T3 mitigate the common problem of overrefusal on benign-but-challenging prompts? How is the cold-start performance with limited ID data?
On OR-Bench, T3+OCSVM has the highest listed AUROC, 0.9342, and T3+GMM achieves FPR@95 0.2217. LlamaGuard3-1B has lower FPR@95, 0.1483, so the table does not establish T3 as best under every overrefusal-related metric.
- The appendix specifies 3,500 safe training prompts, 1,500 safe test prompts, and 600 toxic prompts from OR-Bench subsets.
- Prepending GPT-OSS-20B safety analysis degrades both variants: GMM AUROC drops from 0.9108 to 0.8594, and OCSVM AUROC drops from 0.9342 to 0.8581.
- The experiment measures detector false alarms on safe prompts, not changes in an integrated assistant’s refusal behavior. LlamaGuard’s favorable result alone does not establish the dataset-specific overfitting suggested in the source.
T3+OCSVM leads AUROC at 0.9342, while LlamaGuard3-1B has the lowest FPR@95 at 0.1483. Both augmented variants worsen relative to their base versions, showing that prepended safety analysis does not necessarily improve detection.
Figure 3 shows increasing AUROC and generally decreasing FPR@95 as safe training data grows toward 2,000 examples. Its caption highlights approximately 90% AUROC with 1,000 samples, but fluctuations and benchmark-dependent false-positive trajectories support sample efficiency rather than the absence of cold-start limitations.
The curves show gains from small safe training sets and further improvement as more examples are added. Fluctuations and differing false-positive trajectories support sample efficiency, not complete elimination of cold-start limitations.
RQ4: Does T3’s performance generalize across different languages and specialized domains without retraining?
In Table 4, T3+GMM reaches AUROC 0.9959 and FPR@95 0.0089 on Code, while T3+OCSVM reaches 0.9982 and 0.0039 on HR. Social Media is harder, with FPR@95 0.1208 for GMM and 0.1485 for OCSVM.
- The target domains supply harmful examples for evaluation, not detector fitting; the safe side remains the general held-out reference corpus.
- In Table 5, unaugmented T3+OCSVM AUROC ranges from 0.9768 to 0.9821 across 14 RTP_LX languages and from 0.9762 to 0.9815 across nine XSafety languages.
- These scores demonstrate a narrow observed range across the evaluated languages, not universal superiority: DuoGuard exceeds T3 on several RTP_LX subsets.
- The multilingual experiment uses 30,000 OpenAssistant/oasst2 ID prompts and at most 800 OOD examples per language. This differs from the toxicity experiment’s ID mixture, qualifying the claim that one fitted model covers all reported experiments.
Code and HR are well separated from the safe reference corpus, with T3 FPR@95 below 1%; Social Media is harder, with rates above 12%. The comparison is against general-purpose safe data rather than matched benign examples within each domain.
LLM-Enhanced Variant (Augment): GPT-OSS-20B generates structured safety analysis that is prepended before embedding. Augmentation improves T3+GMM on many non-English RTP_LX subsets but reduces English AUROC from 0.9554 to 0.9071 and usually hurts XSafety; it also degrades OCSVM across the listed languages.
- The authors attribute some degradation to safety analysis being produced in the prompt’s native language, reducing consistency of the embedding inputs.
- The reported results do not establish a consistent gain that justifies the additional computation.
Unaugmented OCSVM scores occupy a narrow range across the evaluated languages, although specialized baselines exceed them on some RTP_LX subsets. Augmentation improves GMM on many non-English RTP_LX subsets but degrades English and most XSafety results; OCSVM augmentation is consistently worse.
RQ5: Can T3 be integrated into a high-performance inference engine for practical, real-time deployment with minimal latency?
Streaming Results: T3 intercepts outputs in vLLM’s Main Process while Worker Processes continue generation, batches safety checks, and marks unsafe requests for termination. On an NVIDIA H200 with vLLM v0.10.2 and OPT-125M, generation time rises from 6.342 to 6.439 seconds for 500 prompts and from 38.011 to 40.292 seconds for 5,000 prompts.
- The experiment uses three embedding models, a detector trained on 1,000 safe Alpaca instructions, and a prediction batch size of 32.
- The prose describes checks every 20 tokens, whereas Table 7 and the scheduling code use words; the evaluation interval is inconsistent in the source.
- Table 7 reports 1.5% and 6.0% generation overhead, with initialization costs reported separately as 10.5 and 9.8 seconds.
- Figure 4 illustrates overlapping guardrail and inference kernels. The measured overhead applies to the tested configuration, not necessarily larger models or more heavily utilized GPUs; the experiment does not report end-to-end safety effectiveness during streaming.
The traces and schematic illustrate Main Process guardrail kernels overlapping Worker Process inference and reducing idle GPU periods. They support the proposed latency-hiding mechanism in the tested workload, not equivalent overhead under every serving configuration.
Generation overhead is 1.5% for 500 prompts and 6.0% for 5,000 prompts; initialization adds 10.5 and 9.8 seconds separately. The caption specifies a 20-word interval, unlike the 20-token description in the prose.
Post-Generation Results: With OPT-125M generating up to 256 tokens on OR-Bench, T3 post-processing takes 59.84–155.81 ms across batch sizes 8–64. DuoGuard is faster, while PolyGuard, MD-Judge, and the listed LlamaGuard measurements are slower.
- Runtime is averaged over 20 runs following five warm-up iterations on an NVIDIA H200.
- At batch size 64, T3+GMM takes 155.81 ms, T3+OCSVM 146.49 ms, DuoGuard 108.04 ms, PolyGuard 373.74 ms, and MD-Judge 1524.49 ms.
- The appendix prose says LlamaGuard does not support batch sizes ≥32, but Table 8 reports 2675.74 ms at 32 and N/A only at 64.
DuoGuard is fastest, while T3 takes substantially less time than PolyGuard and MD-Judge across the listed batches. A LlamaGuard measurement is present at batch size 32 despite the surrounding statement that such batches are unsupported; only batch size 64 is N/A.
5 Discussion
The discussion identifies a boundary for typicality-based safety: with HH-RLHF chosen responses as ID and rejected responses as OOD, all tested approaches perform near chance. The paper reports cosine similarity above 0.95 between paired responses and attributes the difficulty to toxic vocabulary in the chosen responses and subtle differences between pairs.
- Regardless of the proposed explanation, the result illustrates that overlapping semantic representations can defeat the detector; appropriate reference-data curation is consequential.
- The HILL evaluation uses 1,250 safe Dolly training prompts, 250 held-out safe prompts, and 46 learning-style jailbreaks. Table 13 reports AUROC 0.9803 for GMM and 0.9783 for OCSVM, with FPR@95 0.0435 for both.
- The authors propose combining typicality screening with reasoning-based evaluation for near-boundary cases, but do not evaluate that hybrid architecture.
- For comparison, the discussion describes mathematical problems and solutions from MATH or GSM8K as ID and unrelated topics as OOD. It reports that even traditional detectors perform well in this clearly separated domain-adherence setting, which is distinct from detecting harmful intent.
6 Conclusion
The conclusion presents safe-distribution modeling as an alternative to enumerating harmful patterns and emphasizes transfer and serving efficiency. The evidence supports a useful guardrail for several distributionally separable tasks, with limits from mixed adversarial results, reference-data curation, and the narrow serving configuration.
Appendix
- A THEORETICAL ANALYSIS derives general expected PRDC values and, for continuous identical distributions, null expectations through symmetry and a Coverage bound through Jensen’s inequality. A.2 CONSISTENCY presents asymptotic separation arguments for support mismatch and common-support density shifts under additional regularity conditions; these arguments concern metric expectations and should not be read as guarantees of individual-prompt safety classification. A.3 CONNECTION WITH TWO-SAMPLE TESTS distinguishes within-set PRDC neighborhoods from Schilling’s pooled-neighbor statistic, explains reference-computation reuse, and introduces a related B statistic whose consistency is explicitly conjectured rather than proved.
- B EXPERIMENT TECHNIQUE DETAILS adapts vision-oriented detectors to 1024-dimensional Qwen3 embeddings. An auxiliary logistic classifier trained on ID text versus synthetic background data supplies logits, weights, or pseudo-gradients where needed; other adaptations include direct cosine-neighbor retrieval, embedding clipping for ReAct, and pseudo-binary statistics for RMD. B.2 converts API outputs to 1.0−max_toxicity_score and describes caching and error handling. B.3 maps categorical LlamaGuard outputs to fixed scores, MD-Judge severity to scaled scores, and DuoGuard probabilities through max aggregation; PolyGuard receives the dummy response “I cannot and will not provide that information” and uses a hierarchical output-to-score mapping. These adaptations, particularly discrete scoring and the dummy response, qualify comparisons with the models’ native use.
- C EXPERIMENT PARAMETERS specifies cuda:0, seed 42, and the T3 encoders Qwen/Qwen3-Embedding-0.6B, BAAI/bge-m3, and intfloat/e5-large-v2. C.1 uses 40,000 ID samples equally drawn from tatsu-lab/alpaca, databricks/databricks-dolly-15k, Anthropic/hh-rlhf, and OpenAssistant/oasst2, at most 10,000 OOD samples per benchmark, and batch size 32. Its baseline settings include CIDER K=5, NNGuide K=100 and α=1.0, an eight-cluster tied-covariance GMM without dimensionality reduction, OpenMax tail size 20 and α=3 with Euclidean distance, ReAct clipping at the 90th percentile, VIM d=512, and AdaScale percentile range (90.0, 99.0), k1=k2=50.0, λ=1.0, and o=0.1. DuoGuard/DuoGuard-0.5B uses threshold 0.5 and maximum sequence length 512; OpenSafetyLab/MD-Judge-v0_2-internlm2_7b uses temperature 0.1, 128 maximum new tokens, and GPU memory utilization 0.7. C.2 uses 30,000 OpenAssistant/oasst2 ID prompts, at most 800 OOD examples per language, batch size 24, and BAAI/bge-m3 for representation baselines. The paragraph headed Why you should consider using Representation Typicality Estimation describes Forte’s architecture-independent applicability, but its claim of no detector fitting conflicts with T3’s GMM/OCSVM fitting. C.3 uses 5,000 safe OR-Bench prompts split into 3,500 training and 1,500 test examples, 600 toxic examples, batch size 16, and C.1 baseline settings. C.4 specifies the four-source ID mixture with 40,000 training and 15,000 test examples, batch size 32, and the same baseline settings.
- C.5 T3/FORTE ALGORITHMIC ABLATIONS compares Precision+Density, Recall+Coverage, and all four metrics. Table 6 generally favors the full representation, although Recall+Coverage has lower XSafety FPR@95; its scores differ materially from the main benchmark tables, so the configurations should not be assumed interchangeable. Local-covariance ellipsoids were abandoned because high-dimensional estimates were unstable and costly. The appendix also reports heuristic robustness to the choice of sentence transformer.
- D INTEGRATING T3 WITH VLLM FOR ONLINE GENERATION GUARDRAILING describes vLLM’s PagedAttention, continuous batching, and three-tier Main Process, Engine Core, and Worker Process architecture. T3 patches OutputProcessor.process_outputs to synchronize request state, evaluate selected batches, and construct outputs with unsafe requests marked ABORT. Listing 1 and the surrounding description explain self.reqs metadata, word-count scheduling, near-threshold and completion candidates, fallback batching, and concurrency through CUDA MPS or temporal slicing. Runtime evaluation uses NVTX annotations and Nsight Systems on an H200 with vLLM v0.10.2, OPT-125M, three encoders, and 1,000 safe Alpaca training examples; D.1 supplies the separate OR-Bench post-generation timing protocol.
- E ADDITIONAL EXPERIMENTS covers WildGuardMix, 4B/8B embedding ablations, text-native OOD baselines, and HILL attacks. E.1 treats prompt_harm_label=”harmful” OR adversarial=True as OOD and contrasts the human-annotated Test split with the GPT-4-labeled Train split against a held-out safe mixture; T3+OCSVM achieves Test AUROC 0.8882 and FPR@95 0.3663. E.2 retains a T3 false-positive-rate advantage over the listed representation baselines with larger backbones, but T3’s scores are materially weaker than Table 1, and not every AUROC comparison favors it. E.3 reports high FPR@95 for Energy, kNN, and Mahalanobis; this does not establish an inherent limitation under every possible implementation, and the accompanying claim of T3 FPR@95 in the 1–5% range conflicts with many main-table results. E.4 evaluates 46 HILL prompts against a Dolly-based safe split and reports FPR@95 0.0435 for both variants; only GMM exceeds AUROC 0.98.
Brief Thoughts
The clearest contribution is the compact, multi-view neighborhood representation. The component ablation supports including test-set neighborhood information, and the serving experiment provides a concrete integration with measured runtime rather than only offline detection results.
The main evaluation uncertainty is whether T3 detects harmfulness or broader differences between curated safe data and each benchmark source. Matched benign examples from target domains, fixed-threshold evaluation, and tests of sensitivity to incoming batch composition would help separate these effects, especially because Recall and Coverage depend on Y.
Claims of broad superiority and production readiness require qualification given the benchmark exceptions, inconsistent corpus and interval descriptions, and OPT-125M-only serving experiment. The theory also warrants scrutiny: the printed local-perturbation gap δ(e^(−λk)−e^(−λk(1−η))) is negative for positive λ and k with η∈(0,1), despite being described as positive; differences in metric expectations alone do not establish consistent individual-sample decisions.