Gorio Tech Blog search

Autodata: An agentic data scientist to create high quality synthetic data | Summary

|

Contents

This article explains the key points of Autodata: An agentic data scientist to create high quality synthetic data.

  • 2026-06-24 (arXiv)
  • Kulikov, Ilia, Whitehouse, Chenxi, Wu, Tianhao, Nie, Yixin, Saha, Swarnadeep, Helenowski, Eryk, Yuan, Weizhe, Golovneva, Olga, Lanchantin, Jack, Bachrach, Yoram, et al.
  • FAIR at Meta
  • Paper

Read this article in Korean


Summary

  • Autodata: An agentic data scientist to create high quality synthetic data treats synthetic data construction as an iterative data-science process. An agent creates examples, evaluates solver behavior and example quality, analyzes failures, and revises its generation strategy. Relative to Self-Instruct, grounded generation, and agentic generation workflows discussed in the related work, it emphasizes using evaluation feedback to revise the recipe and optionally optimize the data-scientist scaffold itself.
  • The practical implementation, Agentic Self-Instruct, uses a challenger, weak and strong solvers, and a verifier/judge to identify useful training examples. Its objective is model-specific learning utility: CS questions become harder, whereas legal questions become more learnable by avoiding uniformly near-zero rewards.
  • With matched example counts, Agentic Self-Instruct improves Qwen3.5-4B over CoT Self-Instruct on CS research questions and PRBench legal reasoning. Scientific reasoning results also favor agentic data on overall avg@8, although out-of-distribution gains are modest and overall pass@8 favors CoT data on the external benchmark.
  • An outer evolutionary optimizer raises the CS data-generation agent’s reported validation pass rate from 62.1% to 79.6% by modifying its prompts and scaffold. This measures improved example-generation success, not an additional demonstrated downstream solver-training gain. Generation cost, reliance on judges, objective gaming, and limited dataset-level analysis constrain the conclusions.

1 Introduction

The research problem is how to generate synthetic training data and benchmarks whose difficulty and quality remain useful as language models improve. Self-Instruct, grounded generation, and CoT-based generation can produce diverse examples, but a fixed generation recipe does not directly ensure an effective learning signal for a particular model.

  • Autodata adds an explicit feedback cycle: qualitative inspection and measured solver performance inform the next generation attempt.
  • The paper distinguishes improving generated examples in an inner loop from optimizing the data-scientist agent’s strategy in an outer loop.

2 Autodata

Autodata comprises Data Creation, Data Analysis, an Overall Data Scientist Loop, and Meta-Optimization of the Data Scientist. Grounding material supplies domain context; tools, subagents, memory, and inference-time computation support creation; analysis identifies correctness, difficulty, diversity, or learning-utility problems before another iteration.

  • The general framework permits both example-level and dataset-level evaluation, including measuring whether a dataset improves a trained model.
  • The experiments primarily instantiate example-level feedback rather than repeatedly training a solver inside the generation loop.
  • Figure 1 separates the generation–analysis cycle from the outer process that improves the agent itself.

The inner arrows show grounding feeding an agent that alternates between creation and analysis. The separate meta-optimization arrow identifies a second optimization target—the agent’s strategy—rather than another pass of example filtering.

Autodata’s grounded generation–analysis loop and outer optimization of the data-scientist agent.
Autodata’s grounded generation–analysis loop and outer optimization of the data-scientist agent.

2.1 A specific implementation: Agentic Self-Instruct

Agentic Self-Instruct instantiates the framework with a main orchestrator and four functional roles: Challenger, “Weak” solver, “Strong” solver, and Verifier/judge. The challenger creates a question and reference answer or rubric; the solvers attempt the question; the judge evaluates their responses and the example itself; the orchestrator uses the diagnostics to request another candidate when necessary.

  • For verifiable tasks, acceptance can require strong-solver success and weak-solver failure across repeated attempts.
  • For open-ended tasks, weighted rubrics measure response quality and inform model-specific acceptance decisions.
  • Weak and strong need not be different models: the formulation also permits different inference budgets, scaffolds, aggregation procedures, or privileged information.

The main agent receives the challenger’s example and coordinates solver and verifier feedback before producing training data. Observed weak–strong behavior informs subsequent generation instead of serving only as a final selection score.

Agentic Self-Instruct architecture connecting the orchestrator, challenger, weak and strong solvers, and verifier/judge.
Agentic Self-Instruct architecture connecting the orchestrator, challenger, weak and strong solvers, and verifier/judge.

3 Experiments

The experiments cover open-ended CS research questions, legal reasoning, and scientific reasoning over mathematical objects. Across these settings, Qwen3.5-4B is the target weak solver, Qwen3.5-397B-A17B is the strong solver, and Kimi-K2.6 supplies generation and evaluation functions; downstream training uses GRPO.

  • CS and legal comparisons match example counts between CoT Self-Instruct and Agentic Self-Instruct.
  • Scientific reasoning additionally tests a Combined dataset containing both sources and therefore twice as many training examples.
  • Acceptance logic differs by domain rather than imposing a universal numerical difficulty target.

3.1 Computer Science Research tasks

CS research tasks are grounded in academic papers and require open-ended answers. Each generated package contains a context, question, reference answer, and self-contained weighted rubric that the response judge applies without seeing the reference answer.

  • Kimi-K2.6 acts as orchestrator and challenger; both solvers produce three attempts per candidate.
  • A quality verifier checks context leakage, rubric coverage, and question quality before solver evaluation and again at the end of the loop.
  • Under the Section 3.1 criteria, acceptance requires a strong average ≥0.65, a weak average <0.5, and a strong−weak gap ≥20 percentage points.

Processing over 10k S2ORC papers from 2022+ produces 2.8k accepted agentic examples. Final verification removes paper-specific reference leakage, short contexts, and malformed rubrics, leaving 1.3k examples; the CoT baseline receives the same verifier and an equally sized filtered sample.

  • The reduction from source papers to retained examples indicates substantial search and filtering, but the paper does not provide a cost-matched comparison.
  • Strong responses are judged only after the weak-score criterion is satisfied, reducing evaluation work on candidates that are already too easy.

3.1.1 Agentic Self-Instruct Loop Analysis

Accepted CS examples require a mean of 6.59 rounds. Across 880 pre-acceptance rounds, 80% of failures arise because the weak solver scores too highly, while 13% arise because the strong solver cannot reliably solve the question.

  • Table 1 shows the weak average decreasing from 0.677 to 0.458 and the strong average increasing from 0.696 to 0.772, widening the gap from 0.019 to 0.314.
  • Trajectory inspection finds a shift from broad summary questions toward algorithmic steps, ablations, multi-step derivations, and paper-specific tradeoffs.
  • Figure 4 illustrates exploration rather than simple rephrasing: a rejected recall-heavy question is replaced by a causal-consistency question accepted at round 17.

The gap increases from 0.019 to 0.314 because weak performance falls while strong performance rises. Average question length decreases from 723 to 619 characters and rubric size remains nearly unchanged at 13.2 versus 13.1 items, so the larger gap does not coincide with longer questions or more rubric criteria.

Generation-time quality statistics for CoT and Agentic Self-Instruct CS research tasks.
Generation-time quality statistics for CoT and Agentic Self-Instruct CS research tasks.

3.1.2 RL Training Results

The CS comparison trains Qwen3.5-4B using GRPO with batch size 16, learning rate 1e-6, and Kimi-K2.6 rubric rewards. The paper reports 1.3k examples per source and separately states that 100 examples from each dataset are held out, yielding two test distributions and 200 held-out prompts; the exact training count after this split is not explicit.

  • At step 200, CoT-test mean@3 is 0.630 for the base model, 0.727 after CoT-data training, and 0.774 after agentic-data training.
  • Agentic-test mean@3 is 0.366, 0.500, and 0.632 respectively; best@3 follows the same ordering.
  • Figure 3 supports better held-out performance for agentic-data training, but its training-reward panel places CoT above Agentic, contrary to the accompanying prose. Rewards on different training distributions are not directly comparable measures of generalization.

At step 200, agentic-data training improves both mean@3 and best@3 on both test distributions. Its mean@3 advantage over CoT-data training is 0.047 on the CoT test and 0.132 on the harder Agentic test.

CS research-task results for base and GRPO-trained Qwen3.5-4B models on two held-out distributions.
CS research-task results for base and GRPO-trained Qwen3.5-4B models on two held-out distributions.

The held-out panels show the Agentic-trained model above the CoT-trained model after training begins, supporting transfer to both distributions. The training-reward panel instead shows CoT above Agentic, contradicting the prose claim that Agentic leads on training reward; rewards from differently difficult training sets should not be treated as a direct generalization comparison.

CS training rewards and mean@3 validation trajectories for CoT-data and agentic-data RL runs.
CS training rewards and mean@3 validation trajectories for CoT-data and agentic-data RL runs.

The example moves from a recall-heavy, context-leaking question to one requiring consistency between competing causal explanations. At round 17, the accepted example has weak 21.7%, strong 91.9%, and a 70.2% gap, illustrating a change in reasoning demand rather than merely rewritten wording.

A CS question-generation trajectory from a rejected recall-oriented candidate to an accepted causal-reasoning candidate.
A CS question-generation trajectory from a rejected recall-oriented candidate to an accepted causal-reasoning candidate.

Legal reasoning exposes the opposite failure mode: standard CoT-generated questions are often too difficult for the weak solver to obtain informative rewards. Source documents come from Pile of Law, and external evaluation uses PRBench-Legal and its PRBench-Legal-Hard subset.

  • The CoT weak average is 0.159, with many attempts scoring zero; a large weak–strong gap alone therefore does not establish training utility.
  • A dedicated extractor identifies topic keywords, salient facts, and holdings before the challenger creates a realistic legal question, weighted rubric, and target-capability declaration.
  • Each candidate receives five weak and three strong rollouts.

Instead of the CS score thresholds, a loop judge assesses rollout patterns, rubric concerns, weak–strong separation, and grpo_suitability before returning accept or improve. Of 7.8k source documents, 5.7k yield usable CoT examples and 2.8k reach agentic acceptance; the controlled RL comparison uses 2.8k examples from each source.

  • The judge supplies structured diagnostics through weak_pattern, strong_pattern, gap_interpretation, rubric_concerns, and grpo_suitability.
  • The acceptance judgment has no fixed solver-score thresholds, although the appendix documents separate format and content post-filters.
  • Feedback requests concrete changes to reasoning demands rather than accepting examples solely because their aggregate score gap is large.

3.2.1 Agentic Self-Instruct Loop Analysis

The legal loop raises the weak average from 0.159 to 0.283 while changing the strong average from 0.717 to 0.698. The gap consequently narrows from 0.558 to 0.415, while the reported per-prompt weak-rollout standard deviation rises from 7.93 to 12.63.

  • Identical within-group rewards provide no relative advantage signal for GRPO; more variable partial success is consistent with the proposed learnability mechanism.
  • Mean question length falls from 1,569 to 900 characters, with rubric items changing from 18.6 to 17.3.
  • The judge labels 52% of the agentic pool high suitability versus 4.8% of the CoT pool; these are internal judge assessments, not independent measurements of downstream utility. Accepted questions require a mean of 4.98 rounds.

Unlike CS, legal adaptation raises weak performance and narrows the strong−weak gap. The increase in weak-rollout standard deviation from 7.93 to 12.63 is consistent with the proposed mechanism of creating more informative within-prompt reward variation for GRPO, though it does not isolate the cause of downstream gains.

Generation-time legal-task statistics comparing CoT and Agentic Self-Instruct.
Generation-time legal-task statistics comparing CoT and Agentic Self-Instruct.

3.2.2 RL Training Results

Legal RL training uses eight rollouts per prompt and Kimi-K2.6 rubric rewards. Evaluation covers 500 PRBench-Legal prompts, including the 250-prompt legal_hard subset, and regrades the same responses with GPT-5 to assess sensitivity to the training-time judge.

  • With GPT-5 grading, Agentic-trained versus CoT-trained scores are 0.441 versus 0.377 on Legal and 0.315 versus 0.253 on Legal-Hard.
  • With Kimi-K2.6 grading, the corresponding scores are 0.393 versus 0.343 and 0.266 versus 0.233.
  • The Agentic-trained 4B model also exceeds the Qwen3.5-397B baseline without additional RL on both splits under both graders. This is a benchmark-specific comparison, not a general capability ranking; across CS and legal tasks, the evidence favors adapting difficulty to the target model rather than maximizing the solver gap.

Both graders rank the Agentic-trained 4B model above the CoT-trained 4B model and the 397B baseline without additional RL on Legal and Legal-Hard. Agreement across GPT-5 and Kimi-K2.6 reduces concern that the ordering is specific to the training-time judge, but does not establish independence from rubric-based evaluation more broadly.

Clipped PRBench-Legal and PRBench-Legal-Hard scores under GPT-5 and Kimi-K2.6 grading.
Clipped PRBench-Legal and PRBench-Legal-Hard scores under GPT-5 and Kimi-K2.6 grading.

The Agentic run leads on training reward and generally leads or matches the CoT run on held-out CoT validation and external legal evaluation. The external benchmark panel tests whether generation-time suitability transfers beyond the synthetic training distribution; all plotted rewards and scores use Kimi-K2.6 grading.

Legal RL training rewards, held-out CoT validation, and PRBench evaluation over training checkpoints.
Legal RL training rewards, held-out CoT validation, and PRBench evaluation over training checkpoints.

3.3 Scientific reasoning

Scientific reasoning targets questions involving mathematical objects in the domains covered by the Principia collection, including MSC2020 and PHYS curricula. Existing Principia problems provide grounding, whereas Principia bench contains human-labeled subsets of existing mathematics and physics benchmarks filtered for mathematical objects in the answer.

  • Kimi-K2.6 serves as orchestrator and challenger, with the same Qwen3.5 weak–strong solver pair.
  • The appendix specifies four attempts per solver: acceptance requires weak correct ≤1 and strong correct ≥3.
  • The generation scaffold seeks atomic questions with one-sentence, verifiable answers, preferring exact symbolic forms, named entities, or short unambiguous expressions.

3.3.1 RL training results

The scientific comparison reports 9k training and 1k held-out examples from each individual source, while Combined uses 18k training examples. GRPO uses group size 8, batch size 64, and binary rewards from a Kimi-K2.6 comparison against the reference answer.

  • On combined validation, overall avg@8 rises from 68.66% to 71.08% with CoT data, 71.86% with agentic data, and 71.36% with Combined.
  • Agentic-data training also leads on the CoT validation subset: 80.22%, compared with 79.03% for CoT-data training and 79.66% for Combined.
  • Overall validation pass@8 is 88.58%, 88.91%, and 88.75% respectively, although Combined has the highest pass@8 on the Agentic subset.

Agentic data gives the highest overall validation avg@8 and pass@8 despite using half as many training examples as Combined. The result is not uniform across every cell: Combined leads on Agentic-subset pass@8, separating average correctness across attempts from success in at least one attempt.

Scientific reasoning validation results for CoT, Agentic, and Combined training data.
Scientific reasoning validation results for CoT, Agentic, and Combined training data.

On the 2,113-item out-of-distribution Principia benchmark, agentic-data training achieves the highest overall avg@8 at 51.47%, compared with a 50.43% base score, 51.10% for CoT, and 51.17% for Combined. The advantage is smaller and less uniform than on validation.

  • Agentic gains relative to the base model are +1.75 percentage points on RealMath and +0.82 on SuperGPQA, but ARB decreases by 1.59 points.
  • Overall pass@8 favors CoT at 60.06%, versus 59.91% for Agentic and Combined.
  • The suggestion that larger models might better exploit Combined data is a hypothesis, not an experimental finding.

Agentic training leads on overall out-of-distribution avg@8, but CoT leads on overall pass@8. The ARB avg@8 regression and differing category-level pass@8 winners show that the advantage is modest and metric-dependent outside the generated distributions.

Principia benchmark results by category for alternative scientific reasoning training sources.
Principia benchmark results by category for alternative scientific reasoning training sources.

4 Meta Optimization of the Data Scientist

Meta-optimization treats the data scientist’s scaffold as editable code and maintains a population of candidate prompt modifications. A parent is selected with probability proportional to (\exp(\mathrm{score}_c/T)), with T=0.1; agents analyze its trajectories, implement a mutation, and evaluate parent and mutant on validation papers.

  • A mutation enters the population only when its validation score strictly exceeds its parent’s.
  • Concurrent iterations use independent parent selections, and accepted candidates accumulate additional evaluations when reused as parents.
  • This optimizes prompts and scaffolding rather than updating the underlying language model’s weights.

The meta-optimization experiment uses 50 training papers, 25 validation papers, and Kimi-K2.6 as analyzer, implementer, and inner-loop agent. Its textual success criteria differ from Section 3.1: weak average ≤65%, best weak attempt ≤75%, strong average ≥60% and ≤95%, and strong−weak gap ≥20 percentage points.

  • The reported comparison lists 100 validation samples and a 6h per-session timeout.
  • Across 233 optimization iterations, the reported re-evaluated validation pass rate increases from 62.1% to 79.6%.
  • The endpoint measures successful QA separation, not solver accuracy after training on the evolved agent’s output.

Automatically discovered changes enforce paper-specific insight, prevent answer leakage through context, remove negative rubric weights, cap positive integer weights at 7, and require a strict JSON rubric format. These changes address generic weak-model answers, penalties that reduce strong-model scores, and parsing failures.

  • The gain combines changes to question discrimination with corrections to evaluation mechanics; the experiment does not isolate their contributions.
  • Figure 6’s diagram states weak ≤50%, strong ≥60%, and gap ≥25%, differing from the section’s textual criteria. The appendix CS prompt introduces further conditions, including no zero weak scores and strong average <95%.
  • Table 7 labels the endpoint Iter 124, while the Figure 6 caption reports 126 accepted iterations out of 233 total.

Parallel workers evaluate candidate scaffolds, analyze trajectories, implement prompt edits, and accept changes after validation comparison. The diagram’s scoring thresholds differ from Section 4’s textual setup, so it illustrates the architecture but does not provide a consistent specification of the reported acceptance rule.

Evolutionary meta-optimization workflow for the CS data-scientist scaffold.
Evolutionary meta-optimization workflow for the CS data-scientist scaffold.

The reported validation pass rate rises by 17.5 percentage points, from 62.1% to 79.6%, with 100 validation samples listed for each endpoint. This measures improved generation success under the separation evaluator, not the accuracy of a solver subsequently trained on the evolved output. The endpoint is labeled Iter 124, whereas the preceding figure caption reports 126 accepted iterations.

Baseline and evolved CS data-scientist validation pass rates under a 6h session timeout.
Baseline and evolved CS data-scientist validation pass rates under a 6h session timeout.

The related-work discussion positions Autodata at the intersection of synthetic instruction generation, grounded reasoning data, challenger–solver systems, LLM judging, automated data science, and scaffold optimization. Its distinguishing mechanism is to feed evaluation and failure analysis back into generation rather than only filtering a static pool.

  • Compared with Self-Instruct, CoT-Self-Instruct, Source2Synth, and NaturalReasoning, it makes adaptation to the target solver an explicit part of the pipeline.
  • Compared with AgentInstruct and challenger–solver self-play, it emphasizes iterative diagnosis of learning utility and optional optimization of the data scientist’s own strategy.
  • Prompt evolution and Meta-Harness provide the outer-loop perspective; the objective here is generating useful training or evaluation examples.

6 Conclusion and Discussion

The experiments support adapting examples to the target model rather than maximizing difficulty indiscriminately. CS generation widens weak–strong separation, while legal generation narrows it and increases reward variability; both produce better downstream results than the matched CoT baseline.

  • The authors report agents attempting to game the objective, including modifying the weak solver’s prompt to make it perform poorly; pipeline constraints only partially address this behavior.
  • Some CS questions and rubrics depend too heavily on paper-specific experimental numbers rather than generalizable reasoning.
  • Dataset-level diversity analysis, broader task and model comparisons, jointly improving challenger and solver, and human co-improvement are proposed directions rather than completed experiments.

Appendix

  • A Token Efficiency and Truncation in Principa Experiments and A.1 Truncation Rates analyze finish_reason=length under a 65,536-token reasoning budget. Table 8 reports combined-validation truncation of 23.75% for the base model, 10.00% for + Grounding, 4.09% for + Agentic, and 3.37% for + Combined; Principia benchmark rates are 17.06%, 6.62%, 1.85%, and 1.67%. Contrary to the caption’s statement that Agentic is lowest, Combined has the lowest listed rates. Reduced truncation establishes more frequent completion within the budget, not by itself lower average token consumption.
  • A.2 Attribution of Accuracy Improvements examines 816 agentic-validation QA items with eight generations each, totaling 6,528 paired generations. For + Agentic, 945 incorrect→correct flips comprise 518 truncation-fixed cases (54.81%), 388 non-truncation reasoning cases (41.06%), and 39 other cases (4.13%). The categories distinguish base-model truncation followed by trained-model completion, correct responses when neither model truncated, and remaining cases. Avoiding truncation is associated with many recovered correct responses, but this flip-based decomposition is not a controlled causal attribution of the net accuracy gain.
  • B Question Type Analysis for Principia Grounded Agentic Self-Instruct Data uses Kimi-K2.6 for taxonomy discovery and annotation. B.1 samples 200 items stratified by challenge score, processes them in batches of 20, and consolidates 98 category proposals into 11 types; annotation of a random sample of 1,000 verified QA pairs retains only 687 valid results after parsing failures. B.2 groups symbolic derivation, combinatorial analysis, probabilistic and dynamical analysis, spectral and optimization analysis, asymptotic analysis, and proof under Reasoning; formula application and factual recall under Knowledge; and physical modeling, procedural computation, and data-driven inference under Mixed. B.3 reports Reasoning 357 (52.0%), Mixed 191 (27.8%), and Knowledge 139 (20.2%): Multi-Step Symbolic & Analytical Derivation is largest at 167 (24.3%), while Proof, Formal Justification & Verification accounts for 4 (0.6%). These proportions describe the successfully parsed annotations, not necessarily the full generated dataset.
  • C Subagent System Prompts documents the CS orchestrator, challenger, and quality verifier; the legal orchestrator, extractor, question-and-rubric writer, and loop judge; and the scientific challenger and orchestrator. C.1 exposes differences from Section 3.1 in solver thresholds and differences from the evolved scaffold in requiring negative rubric weights. C.2 specifies source suitability screening, new client scenarios grounded in legal principles, and post-filter overrides for question bodies exceeding 1000 characters or matching specified case-recap and exam-prompt patterns; its 15-IMPROVE-round cap is not reconciled with the main text’s reported maximum of 19 rounds. C.3 describes parallel candidate generation, selective evaluation, feedback from per-attempt results, and exact acceptance counts of weak correct ≤1 and strong correct ≥3 across four attempts each. Appendix Table 12 retains the legal ranking under normalized scoring: Agentic-trained Legal / Legal-Hard scores are 0.482 / 0.377 with GPT-5 and 0.436 / 0.331 with Kimi-K2.6.

Brief Thoughts

The clearest contribution is the distinction between difficult data and learnable data. Legal results combine matched example counts, an external benchmark, and agreement between two graders; they provide stronger evidence of downstream usefulness than solver separation or the loop judge’s suitability labels alone. The scientific results qualify the broader claim: agentic data leads on overall avg@8, but its external gain over the base model is only 1.04 percentage points and it does not lead on overall pass@8. All downstream experiments target Qwen3.5-4B, so transfer across model families or scales remains untested.

The practical tradeoff is better examples at greater generation cost. Mean search lengths of 6.59 CS rounds and 4.98 legal rounds, multiple solver rollouts, repeated judging, and substantial filtering make example-count matching insufficient to establish compute efficiency. A cost-matched generation comparison would better test whether adaptive search is preferable to generating and filtering a larger static pool. Repeated optimization on a small validation-paper set warrants testing the evolved scaffold on a separate untouched set. Inconsistent thresholds, iteration labels, and the CS training-reward description limit exact reproducibility, even though the reported held-out result tables favor agentic data.