On Data Engineering for Scaling LLM Terminal Capabilities | Summary
24 Feb 2026 | Paper Review Terminal Agents Synthetic Data Supervised Fine-TuningContents
- Summary
- 1. Introduction
- 2. Related Work
- 3. Background
- 4. Synthetic Data Generation
- 5. Experiments
- 6. Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of On Data Engineering for Scaling LLM Terminal Capabilities.
- 2026-02-24 (arXiv)
- Pi, Renjie, Lam, Grace, Shoeybi, Mohammad, Jannaty, Pooya, Catanzaro, Bryan, Ping, Wei.
- NVIDIA
- Paper
- Project Page
Summary
- On Data Engineering for Scaling LLM Terminal Capabilities studies how training-data design affects command-line agents. Terminal-Task-Gen combines adapters for existing math, code, and software engineering prompts with seed-based and skill-based task generation, producing Terminal-Corpus for supervised fine-tuning.
- Fine-tuning Qwen3 models yields Nemotron-Terminal-8B, -14B, and -32B, which score 13.0 ± 2.2, 20.2 ± 2.7, and 27.4 ± 2.4 on Terminal-Bench 2.0, respectively. Their corresponding baselines score 2.47 ± 0.5, 4.04 ± 1.3, and 3.37 ± 1.6 under the Terminus 2 agent framework.
- The tested recipe favors mixed single-stage training, retaining the full synthetic trajectory pool, and the standard context configuration. Filtering also changes dataset size, and evaluation centers on one 89-task benchmark, so these results do not isolate the benefit of failed traces or establish general terminal competence.
1. Introduction
Terminal-agent training faces two data bottlenecks: scarce task prompts, dependency files, and prepared environments, and expensive collection of multi-turn trajectories in freshly instantiated environments. The paper addresses these bottlenecks with complementary strategies: adapters supply broad coverage efficiently, while synthetic generation targets terminal-specific skills, difficulty levels, and task compositions.
- The contributions are Terminal-Task-Gen, a study of filtering, data mixing, context length, and scaling, and the Nemotron-Terminal model family.
- Model checkpoints and most synthetic datasets are released at https://huggingface.co/collections/nvidia/nemotron-terminal.
2. Related Work
The related work distinguishes improvements to agent scaffolding from improvements to the underlying model. It also connects terminal-task generation to earlier synthetic instruction-data methods, emphasizing generation efficiency and downstream SFT evaluation.
Agent Design.
Prior terminal agents improve performance through scaffolding, but such designs can be model-specific and require substantial engineering. This study keeps the Terminus 2 scaffold fixed and targets model capabilities through supervised fine-tuning; it does not compare alternative scaffold designs experimentally.
Dataset Adapters.
Dataset adapters reuse existing prompts by rolling them out in terminal environments, enabling data collection across domains such as competitive coding and mathematics. The paper studies adapter sources and filtering choices, while noting that source datasets were not originally designed for sequential environment interaction.
Synthetic Task Generation.
The paper draws on Evol-Instruct, Code Evol-Instruct, AgentInstruct, LAB, and MAGPIE as precedents for synthetic instruction generation. Compared with terminal-task pipelines that use multiple agents to brainstorm tasks, construct environments, and validate artifacts, Terminal-Task-Gen simplifies coordination and uses shared environments.
3. Background
The experimental interface combines executable, containerized tasks with a model-agnostic terminal agent. Task specification and verification are separate from the language model that selects commands.
3.1. Terminal-Bench
Terminal-Bench 2.0 contains 89 hand-crafted, human-verified tasks spanning scientific computing, software engineering, machine learning, security, system administration, and data science. These tasks require end-to-end workflows such as compilation, model training, system configuration, and environment debugging rather than isolated function generation.
- A benchmark task includes a natural-language instruction, a Docker environment, verification tests, and an oracle solution.
- The directory structure contains instruction.md, task.toml, environment/, solution/, and tests/; the study uses Terminus 2 for consistent evaluation across checkpoints.
The directory layout separates the agent-visible instruction and environment from solution and verification artifacts. This defines the executable task interface reused by the pipeline, although generated tasks omit oracle solutions and adapters omit tests.
3.2. Terminus 2 Agent Framework
Terminus 2 exposes an interactive tmux session inside a sandboxed Docker container rather than a collection of specialized tools. At each step, the model receives terminal output and emits structured JSON containing analysis, plan, commands, and an optional task_complete flag.
- Command objects specify keystrokes and an optional duration in seconds, allowing the agent to invoke available command-line tools and wait for execution.
The response example represents terminal actions through command keystrokes and durations, alongside analysis, planning, and completion status. The same action representation supports demonstration collection and trained-model evaluation.
4. Synthetic Data Generation
The data pipeline separates broad prompt adaptation from targeted executable-task construction. Both streams feed DeepSeek-V3.2 trajectory generation through Terminus 2, followed by decontamination and quality filtering to form the SFT dataset.
- The coarse-to-fine design describes complementary data-generation strategies; the adopted training procedure mixes their trajectories in one stage.
Two data streams converge on the same trajectory collector: prompt adaptation supplies existing problems, while seed and skill synthesis supply task files and tests. Shared domain images support the synthetic stream, and post-processing produces Terminal-Corpus.
4.1. Dataset Adapters
Dataset adaptation converts existing prompts into Terminal-Bench-style instructions and environments without an LLM in the conversion loop. The prompt datasets are Nemotron-Cascade Stage-2 math and code prompts and its SWE code-repair prompts; because these sources provide prompts only, the adapted training tasks have no associated verification tests.
- Math Prompts. The source contains 163K unique OpenMathReasoning prompts, excluding questions whose DeepSeek-R1 response is shorter than 2K tokens.
- Code Prompts. The source contains 79K OpenCodeReasoning prompts, filtered and deduplicated to a 35K prompt subset.
- SWE Prompts. The source contains 127K instances from SWE-Bench-Train, SWE-reBench, SWE-Smith, and SWE-Fixer-Train, filtered and deduplicated to 32K unique prompts; supplied buggy code files are instantiated in the environment.
- Adapter Format. The original prompt fills {instruction} with a domain-specific suffix: math answers go to /app/solution.txt, Python solutions to /app/solution.py, and SWE SEARCH/REPLACE edits to /app/solution.patch.
4.2. Synthetic Task Generation
4.2.1. Generation from Seed Data uses existing problems as inspiration for new terminal tasks rather than merely wrapping their prompts. A seed contains a problem description, an optional domain label, and an optional reference solution; the generator adds concrete package, input-path, implementation, and output-path requirements.
- Seed Data Structure. Reference solutions, when available, provide ground truth for test expectations but are not exposed to the solving agent.
- Task Adaptation. The generator creates input files and pytest checks for file existence, formats, numerical accuracy with suitable floating-point tolerances, and edge cases.
- Conversion Guidelines. Complex problems can be decomposed into verifiable units, with explicit input sizes, precision requirements, and unambiguous output formats.
4.2.2. Generation from Primitive Skills synthesizes tasks by combining a curated taxonomy of terminal-operation skills. The main text defines 9 domains—data processing, data querying, data science, debugging, dependency management, file operations, scientific computing, security, and software engineering—and typically combines 3–5 primitives per task.
- Domain-Specific Generation. Dedicated prompts guide tasks within each domain, such as statistical transformations or cryptographic and access-control checks.
- Skill Taxonomy. Primitive families include algorithmic, systems, data processing, mathematical, testing, and web/security skills.
- Compositional Task Synthesis. Prompts request novel integrated scenarios rather than isolated skill exercises or reproductions of standard tutorials.
- Table 10 and the appendix modules use System Administration where the main-text domain list names dependency management; the document does not explain this naming difference.
The taxonomy combines operational skills with algorithms, mathematics, and testing rather than treating terminal use as shell syntax alone. Its System Administration domain differs from the dependency management label in the main-text list.
4.2.3. Task Format and Execution Environment standardizes generated tasks as instructions, weighted pytest tests, supplementary input files, and domain-specific Docker environments. Unlike Terminal-Bench’s human-verified tasks, these synthetic tasks have no generated oracle solutions; synthesized tests assess agent outputs instead.
- Solution Isolation. Generation prompts prohibit exposing algorithms, implementations, or solution code in the agent-visible task; seed solutions are used to derive test expectations.
- Pre-Built Docker Images. The pipeline uses 9 shared domain-specific base images instead of generating a Dockerfile per task, avoiding repeated Dockerfile repair and reducing image storage.
- Stable base images decouple environment preparation from task generation, while agents can still install runtime dependencies.
4.3. Teacher Model
DeepSeek-V3.2 generates both tasks and trajectories, selected partly for its 38.2 ± 2.9 Terminal-Bench 2.0 score. Under Terminus 2 on adapted benchmarks, it achieves pass@1 scores of 93.33 on the combined AIME 2024, AIME 2025 entry, 67.20 on LiveCodeBench v6, and 52.40 on SWE-bench Verified.
- These checks support the teacher’s ability to solve adapted tasks; they do not independently validate every generated task or test suite.
DeepSeek-V3.2 solves substantial portions of standard math, coding, and SWE benchmarks when they are routed through Terminus 2. These scores support teacher selection for adapter rollouts but do not certify synthetic-task validity.
4.4. Data Filtering
The pipeline removes training prompts with a 14-gram overlap with Terminal-Bench 2.0 test samples, then applies quality filters that remove identity leaks and responses containing Chinese characters. Additional experiments discard incomplete trajectories or, when verification tests exist, retain only trajectories that pass them.
- The experimental label No filter refers to the additional completion or success filters, not to the absence of decontamination and basic quality controls.
5. Experiments
Experiments measure overall and category-level Terminal-Bench 2.0 performance and ablate data sources, trajectory filtering, context length, curriculum, and data scale. Qwen3-8B is the primary ablation model, with Qwen3-14B and Qwen3-32B used to examine results at larger model sizes.
5.1. Experimental Setup
Unless otherwise specified, SFT uses a learning rate of 2e-5, weight decay of 1e-4, 2 epochs, a maximum sequence length of 32,768 tokens, global batch size 128, and micro-batch size 1 per GPU. Optimization uses AdamW with β = 0.9, 0.95, a cosine learning-rate schedule with 10% warmup, and gradient clipping at 1.0.
- The 8B and 14B models use 4 nodes with 8 GPUs each, totaling 32 GPUs, and sequence parallelism of 2; the 32B model uses 16 nodes, totaling 128 GPUs. All experiments use CPU offloading.
- Harbor orchestrates trajectory generation, with a Singularity extension for HPC deployment; the authors accept rare fakeroot-overlay failures during synthetic generation.
- Daytona provides isolated cloud sandboxes for evaluation, and veRL is used for SFT.
- Scores are accompanied by uncertainty values, but the experimental setup does not specify evaluation repetition counts or explain how all tabular uncertainty values are computed.
5.2. Main Results
Nemotron-Terminal improves over each corresponding Qwen3 baseline: 8B rises from 2.47 ± 0.5 to 13.0 ± 2.2, 14B from 4.04 ± 1.3 to 20.2 ± 2.7, and 32B from 3.37 ± 1.6 to 27.4 ± 2.4. The 14B model’s mean exceeds GPT-OSS (high) 120B at 18.7 ± 2.7, and the 32B model’s mean exceeds Qwen3-Coder 480B at 23.9 ± 2.8, though the reported uncertainty ranges overlap.
- These comparisons show competitiveness with selected larger systems, not parity with the strongest evaluated models: DeepSeek-V3.2 scores 38.2 ± 2.9 and Claude Opus 4.5 scores 57.8 ± 2.5. The paper does not report significance tests for the larger-model comparisons.
All three Nemotron-Terminal models improve over their Qwen3 counterparts. Selected larger models have lower mean scores, but the strongest evaluated systems remain ahead, and some larger-model comparison ranges overlap.
Category results show large but uneven gains. For 32B, Data Processing improves from 5.00 to 50.0, Security from 2.50 to 27.5, Software Engineering from 5.00 to 31.7, and Model Training from 0.00 to 50.0.
- Data Querying reaches 60.0, but that category contains only 1 task; Debugging contains 3 tasks and reaches 33.3.
- Mathematics, Games, and Video Processing remain at 0.00 for all trained model sizes. Scientific Computing remains weak, with scores of 0.00, 2.90, and 0.00 for 8B, 14B, and 32B.
- Small category counts and persistent failures limit conclusions about broad terminal competence.
The breakdown reveals gains in data processing, security, software engineering, and model training alongside persistent weaknesses in mathematics and scientific computing. Parenthesized task counts show why results in one-task categories cannot establish broad capability.
5.3. Ablation on Dataset Components
Combining adapter sources raises performance: Math alone scores 5.39 ± 1.65, Code 6.29 ± 1.65, and SWE 7.02 ± 2.13, while all 226,313 adapter samples yield 9.66 ± 2.11. Skill-based synthetic data scores 12.4 ± 2.38, compared with 6.18 ± 1.91 for seed-based data.
- The adapter splits contain 162,692 Math, 31,960 Code, and 31,661 SWE samples.
- Synthetic splits contain 124,366 seed-based and 139,841 skill-based samples; combining them yields 264,207 samples and 12.4 ± 2.29.
- Adding seed-based data does not improve the reported mean over skill-based data alone. The slightly smaller uncertainty value does not by itself demonstrate the increased robustness claimed in the paper.
- Source combinations also increase sample count, so this ablation does not isolate domain diversity from data-volume effects.
Combining adapter domains improves the mean beyond any individual adapter source. Skill-based tasks yield the strongest individual synthetic-source result; adding seed-based tasks leaves the mean unchanged and slightly reduces the reported uncertainty value, without independently establishing robustness.
5.4. Filtering Strategies
Trajectory Filtering favors retaining the full pool when training on all adapters or all synthetic tasks. For adapters, No filter scores 9.66 ± 2.11 with 226,313 samples versus 8.09 ± 1.84 with 196,940 complete-only samples; synthetic data scores 12.4 ± 2.29 without additional filtering, 6.74 ± 2.20 with complete-only filtering, and 5.06 ± 2.11 with success-only filtering.
- The synthetic complete-only and success-only sets contain 104,603 and 83,448 samples, respectively, compared with 264,207 without additional filtering.
- Adapter results vary by source: complete-only Math has a higher mean than unfiltered Math, Code changes little, and SWE favors no filtering.
- The authors suggest that unsuccessful traces expose useful error states and recovery patterns. Strict filtering also removes over half the synthetic data, so the experiment does not isolate failure-trace content from retained data quantity or task difficulty.
Completion filtering has inconsistent effects across adapter sources: Math has a higher complete-only mean, SWE favors retaining all traces, and Code changes little. The combined dataset favors no additional filtering, which also retains more samples.
Keeping 264,207 synthetic trajectories yields a higher mean than either restricted subset. Complete-only and success-only filtering reduce the pool to 104,603 and 83,448 samples, so the comparison changes both trajectory content and training-data scale.
5.5. Long Context Training and Evaluation
Most trajectories fit within 32,768 tokens, motivating a comparison with 65,536-token SFT and YaRN2 context extension for the longer tail. The default configuration—32,768-token SFT and 40,960-token evaluation—has the highest reported score, 13.0 ± 2.2, while the tested extended-context settings range from 10.3 ± 2.0 to 11.9 ± 2.1.
- Applying YaRN2 only at evaluation with a 65,536-token window yields 11.9 ± 2.0; 65,536-token SFT and evaluation without YaRN2 yields 10.3 ± 2.0; using YaRN2 in both yields 11.9 ± 2.1.
- Figure 5 reports adapter token counts with mean 17,507 and median 13,861, and synthetic-task counts with mean 17,363 and median 17,836. The adapter distribution has a heavier long-token tail.
- Figure 6 reports mean turn counts of 17.5 for adapters and 16.3 for synthetic tasks, with median 16.0 for both.
- The authors attribute the lack of benefit to noisy, less informative long trajectories, but do not directly measure quality by length. The results show no benefit from the tested extensions, not that longer context is generally harmful.
The standard training and evaluation context configuration has the highest mean score. Longer windows and YaRN2 do not improve the tested setup; the reported uncertainty ranges overlap, and no significance test is supplied.
Both trajectory groups have median 16.0 turns, with means of 17.5 for adapters and 16.3 for synthetic tasks. The histograms document multi-turn interactions and a broader long-turn tail for adapter traces.
5.6. Curriculum Learning
The curriculum comparison trains adapters first and synthetic tasks second, versus training on all sources concurrently. Mixed single-stage SFT scores 13.03 ± 2.16, compared with 10.39 ± 1.71 for the two-stage curriculum, and is adopted for the other experiments.
- The comparison finds no advantage for the particular adapter-then-synthetic schedule tested; it does not evaluate other difficulty-based or adaptive curricula.
Mixed single-stage training reaches 13.03 ± 2.16 compared with 10.39 ± 1.71 for adapter-then-synthetic training. This favors the simpler recipe over the tested schedule, rather than ruling out curriculum learning generally.
5.7. Scaling Experiments
Scaling experiments fine-tune Qwen3-8B and Qwen3-14B using 0%, 1%, 2%, 5%, 10%, and 100% of the synthetic training data. Figure 4 shows an overall increase in Terminal-Bench 2.0 performance for both sizes, with 14B performing better at every tested fraction and gaining more from the smallest to the largest data setting.
- The plotted 14B mean dips slightly between 1% and 2%, so the trend is not strictly monotonic.
- The figure includes 95% confidence intervals. The tested fractions establish an empirical data-scaling trend, not a fitted scaling law or evidence about saturation beyond the available dataset.
Both models improve overall across the tested data fractions, and 14B remains above 8B, although its mean dips slightly from 1% to 2%. Shaded 95% confidence intervals communicate uncertainty; the figure neither fits a scaling law nor establishes where further-data gains would saturate.
6. Conclusion
The paper concludes that dataset adaptation and targeted task synthesis improve the tested Qwen3 terminal agents. It releases models and most synthetic datasets, including adapter and skill-based subsets, and proposes reinforcement learning with verifiable execution feedback as future work rather than an evaluated component.
- The results support the pipeline under the tested conditions, but do not establish the conclusion’s broader claim that trajectory quality is more important than parameter scale.
Appendix
- A.1. Synthetic Trajectory Analysis provides token-count and turn-count histograms for 226,313 adapter trajectories and 264,207 synthetic-task trajectories. These distributions document the length variation behind the long-context experiment and show a longer tail for adapter trajectories.
- A.2. Details for Trajectory Generation supplies the Terminus 2 template with {instruction} and {terminal_state} placeholders, structured JSON actions, and command timing conventions. Figures 8–10 document the adapter suffixes for /app/solution.txt, /app/solution.py, and /app/solution.patch, including the SWE requirement to localize the bug and generate SEARCH/REPLACE edits.
- A.3. Details for Synthetic Task Generation covers Curation of Primitive Skills and Prompts for Synthetic Task Generation. Table 10 gives domain examples ranging from graph traversal for dependency resolution and structured-file parsing to statistical transformations and service configuration. Figures 11–20 provide the shared generation template and nine domain modules; generated artifacts use <prompt>, <tests>, <weights>, <info>, <files>, and <test_requirements>, with requirements for self-contained specifications, programmatic verification, and no solution leakage.
Brief Thoughts
The main contribution is a reproducible data recipe with downstream ablations: shared environments, compositional skill tasks, and mixed SFT improve performance without changing the agent scaffold. The synthetic filtering result argues against automatically discarding failed traces, although a sample-count-controlled comparison is needed to explain why retaining them helps.
The evidence is narrower than the paper’s broadest conclusions. Generated tasks lack independently verified oracle solutions, several benchmark categories contain very few tasks, and some capabilities remain at zero; additional terminal benchmarks and controlled comparisons of data quantity, difficulty, and trajectory quality would strengthen the findings.