Agentic Rubrics as Contextual Verifiers for SWE Agents | Summary
07 Jan 2026 | Paper Review AI Coding Agents Rubric-Based Evaluation Test-Time Scaling Knowledge DistillationContents
- Summary
- 1. Introduction
- 2. Preliminaries
- 3. Experimental Design
- 4. Results
- 5. Ablations
- 6. Related Work
- 7. Conclusion
- 8. Limitations
- Appendix
- Brief Thoughts
This article explains the key points of Agentic Rubrics as Contextual Verifiers for SWE Agents.
- 2026-01-07 (arXiv)
- Raghavendra, Mohit, Gunjal, Anisha, Liu, Bing, He, Yunzhong.
- Scale AI
- Paper
- Project Page
Summary
- Agentic Rubrics separates repository-grounded verification into two stages: an expert agent explores the repository and writes explicit criteria, then an LLM judge scores candidate patches without executing them. The method addresses the repeated cost of execution-based verification and the limited repository context and interpretability of direct patch classifiers.
- On all 500 SWE-Bench Verified problems with 16 independently generated candidates per problem, Agentic Rubrics achieves BEST@16 resolution rates of 40.6% for Qwen3-32B and 54.2% for Qwen3-Coder-30B-A3B. These exceed the strongest compared baseline, Patch Classifier, by 3.5 and 4.0 percentage points, respectively.
- Rubric scores discriminate ground-truth test outcomes with ROC-AUC 0.886 and PR-AUC 0.722 and provide criterion-level diagnoses. Repository access improves selection, and supervised distillation improves an open-weight rubric generator; however, the model-based utility audit also identifies over-prescription and rubric–test mismatches, and reinforcement learning with rubric rewards remains untested.
1. Introduction
Verification must determine whether a software patch is correct, complete, safe, and consistent with intended behavior. Unit tests provide environment-aware evidence but require setup and can produce sparse or brittle signals; execution-free classifiers and similarity measures reduce repeated execution costs but can overlook repository contracts or rely on superficial cues.
- Agentic Rubrics uses repository exploration to formulate task-specific acceptance criteria, rather than asking a judge to infer all relevant constraints directly from an issue and patch.
- The demonstrated contributions concern best-of-K patch selection, alignment and utility analysis, and supervised distillation of rubric generation. Although the introduction also discusses rewards for post-training, the paper does not demonstrate reinforcement learning with rubric rewards.
- The paper identifies its research page as https://scale.com/research/agenticrubrics.
Repository exploration produces a rubric that can be reused to grade multiple candidate patches. Both the rubric agent and SWE agent interact with the environment, but the final grading step evaluates the submitted diff against explicit criteria without executing candidate code.
2. Preliminaries
The preliminaries define patch verification and rubric-based scoring. They frame repository grounding as a way to retain task-specific constraints while keeping repeated candidate grading lightweight.
2.1 Verification for SWE Agents
A verifier assigns a score to a candidate patch for a given issue. Execution-based methods run code, usually human-authored or generated tests; execution-free methods rank patches using classifiers, similarity metrics, or LLM judgments without running the repository.
- Execution-based verification can incur sandbox setup overhead and provide sparse or brittle signals, including limited distinguishability among candidates and test toxicity.
- Execution-free verification trades lower operational overhead for possible losses in grounding, reliability, and interpretability, and can be sensitive to stylistic or non-semantic artifacts.
2.2 Rubric-based Verification
A rubric decomposes correctness into explicit natural-language criteria, optional axis groupings, importance weights, and an aggregation rule. A judge evaluates each criterion and combines the judgments into a score for selection or learning; criteria derived only from the issue description may omit repository-specific interfaces, constraints, and conventions.
3. Experimental Design
The experiments hold candidate-generating policies fixed and vary the verifier used for parallel test-time selection. They compare artifact-free scoring with methods that first interact with the repository to generate tests, a proxy patch, or a rubric.
3.1 Agentic Rubrics
The rubric-generation agent uses the SWE-Agent scaffold for repository navigation, file inspection, editing, and shell commands, and submits rubrics.yaml as its artifact. Each criterion has a short description and an importance weight of 1, 2, or 3, corresponding to nice-to-have, important, or must-have; an unparseable artifact makes the generation attempt invalid.
- File Change contains 4–8 items assessing whether edits are minimal, local, and sufficient.
- Spec Alignment contains 3–6 items assessing issue requirements.
- Integrity contains 3–6 items covering test weakening, unnecessary refactors, mass renames, dependency churn, and related hygiene constraints.
- Runtime contains 3–6 items describing intended behavior and obvious execution-time risks, assessed from textual evidence rather than executed tests.
- The generation template requests atomic, self-contained, instance-specific criteria and avoidance of double-counting. Figure 1 labels the illustrative artifact rubric.yaml, whereas the implementation uses rubrics.yaml.
For grading, the LLM judge assigns each criterion a binary satisfaction score. The weighted aggregation is (S=\frac{\sum_i w_i s_i}{\sum_i w_i}), where (s_i\in{0,1}) and (w_i\in{1,2,3}), producing a verifier score in [0, 1].
- Repository interaction occurs during artifact construction; candidate grading uses the problem statement, rubric, and patch without executing candidate code.
3.2 Test-time Scaling with Agentic Rubrics
For each of the 500 SWE-Bench Verified problems, the authors sample 16 independent SWE-Agent rollouts at temperature 1.0. Qwen3-32B uses the Instruct version with a maximum of 30 turns, while Qwen3-Coder-30B-A3B receives a maximum of 50 turns.
BEST@K measures whether the highest-scoring selected patch passes both ground-truth Fail-To-Pass and Pass-to-Pass tests. For K < 16, candidate subsampling is repeated over 100 trials; ORACLE PASS@K selects using ground-truth tests as an upper bound, and RANDOM@K selects uniformly.
- This protocol evaluates selection over a fixed candidate pool, not improvements to the generating policy.
- Hidden evaluation tests determine benchmark resolution but are unavailable to the verifier agents.
3.3 Baselines
Non-agentic Verifiers (no artifact). Self-Consistency selects the diff with the highest average similarity to the other K−1 candidates, using difflib’s SequenceMatcher.ratio() on unified-diff strings. Patch Classifier asks an LLM judge to predict correctness with a continuous score in [0, 1].
Agentic Verifiers (artifact-based). Agentic Tests generates test_issue.py and executes its cases against candidate patches; Agentic Patch Similarity generates a repository-grounded proxy reference patch and uses an LLM judge to score similarity on a 1–5 scale. Agentic Rubrics instead generates explicit criteria and grades candidates against them.
The main comparison specifies Claude Sonnet-4.5 with a 30-turn artifact-generation budget and GPT-5 with low reasoning for LLM-based scoring. Repositories are reset to pre-PR snapshots: agents may inspect existing code and tests but cannot access hidden evaluation tests, reference patches, or git history.
- Missing or invalid artifacts receive a score of 0.
- LLM scoring excludes rollout trajectories and tool traces, using the issue, artifact when applicable, and final patch to reduce trace-related bias and verifier hacking.
- Figure 2 compares these methods at BEST@16 and across candidate counts.
The left panel reports 40.6% and 54.2% BEST@16 for Agentic Rubrics, exceeding Patch Classifier by 3.5 and 4.0 percentage points. The right panel shows a rubric advantage as candidate counts increase beyond one for Qwen3-32B, while the oracle curve reveals remaining selection headroom. The tabular panel belongs to this figure, not a separately numbered table.
4. Results
The results assess patch-selection accuracy, agreement with ground-truth tests and reference patches, and the diagnostic utility of criterion-level judgments. The utility analysis is a separate model-based audit rather than another execution-based correctness evaluation.
4.1 Test Time scaling with Agentic Rubrics
Agentic Rubrics leads the compared verifiers at BEST@16 for both generators. On Qwen3-32B, it reaches 40.6% versus 37.1% for Patch Classifier and 35.0% for Agentic Patch Similarity; on Qwen3-Coder-30B-A3B, it reaches 54.2% versus 50.2% and 49.6%, respectively.
- The gains over Patch Classifier are +3.5 and +4.0 percentage points; the gains over Agentic Patch Similarity are +4.6 points in both settings.
- Agentic Tests reaches 33.6% and 49.0%, while Self-Consistency reaches 33.2% and 47.6%.
- Oracle selection reaches 51.4% and 65.6%, leaving substantial selection headroom.
- The Qwen3-32B scaling curve shows a rubric advantage as the candidate pool grows beyond K = 1, rather than only at K = 16.
The authors attribute the advantage to specifying properties that should hold rather than constructing executable tests or matching one proxy implementation. Generated tests depend on successful execution and discrimination, while proxy-patch similarity can penalize semantically correct alternatives; these are proposed explanations, not separately isolated causal findings.
4.2 Analysis of Agentic Rubrics
4.2.1 Rubric Score Alignment Analysis. For Qwen3-32B candidates graded with Sonnet-4.5 rubrics, test-passing patches typically concentrate at scores of 0.85–1.0, whereas failing patches have lower, more dispersed scores, often around 0.4–0.5. Predicting ground-truth pass/fail from rubric scores yields ROC-AUC 0.886 and PR-AUC 0.722.
- Figure 3 shows strong separation with overlap: a high rubric score does not guarantee test success.
- Human-written reference patches generally receive high scores from frontier-generated rubrics. File Change is a recurring exception because exact edit-location constraints can be more prescriptive than the reference implementation.
Test-passing candidates concentrate near the upper end of the weighted score range, while failures span lower and intermediate scores. Overlap between the distributions shows that the verifier is informative but imperfect, consistent with the reported ROC-AUC 0.886 and PR-AUC 0.722.
The axis-level distributions distinguish different grading patterns. Test-failing patches score lower on File Change, Spec Alignment, and Runtime, but often retain high Integrity scores; test-passing patches nearly saturate Spec Alignment and Integrity while still receiving some edit-scope and runtime penalties.
- Figure 4 shows that hygiene compliance alone does not establish functional correctness and that axis-level judgments provide more detail than a single pass/fail result.
- Intermediate rubric scores demonstrate graded discrimination, but the analysis does not independently validate them as calibrated measures of partial functional correctness.
Integrity scores remain high for many incorrect patches, so hygiene compliance does not establish functional correctness. File Change, Spec Alignment, and Runtime show stronger separation, while scope penalties remain visible even among test-passing candidates.
4.2.2 Rubric Utility Analysis. On a subset of 100 instances, GPT-5 with medium reasoning classifies judgments using the issue, ground-truth tests, and reference patch as the specification. With a rubric acceptance threshold of 0.7, 78% of judgments in the reported agreement regime are labeled high-utility; when tests pass but the rubric score is below 0.7, 54% of rubric failures are labeled high-utility and 46% low-utility.
- Useful aligned judgments mainly concern Core Semantics, API/Compatibility, Structure/Scope, and Edge Coverage.
- Useful stricter-than-test judgments often flag Root Cause Missed or Missing Edges; low-utility cases include over-specified fixes, redundant signals, evaluation errors, and specification or test mismatches.
- These percentages describe an LLM-based audit, not independent human adjudication or execution-confirmed discovery of additional bugs.
- Figure 5(a)’s subcaption specifies test reward = 1 and rubric reward ≥ 0.7, while its panel heading and the main text describe agreement across both acceptance and rejection. The exact composition of the reported agreement regime is therefore unclear.
The model-based audit labels 78% of aligned judgments and 54% of rubric failures on test-passing patches as high-utility. Root Cause Missed and Missing Edges dominate useful disagreements, but the 46% low-utility share prevents interpreting stricter grading as automatic evidence of better correctness. The first panel’s subcaption describes accepted patches only, whereas its heading and the main text include aligned rejections.
5. Ablations
The ablations vary the rubric-generation model, repository access, and grading model. They test how generation capability, contextual grounding, and judge capability affect downstream patch selection.
5.1 Rubric-Agent Model Choice
With Qwen3-Coder-30B-A3B candidates and GPT-5 low-reasoning grading fixed, frontier rubric generators outperform the tested open-weight generators. Figure 6 also shows substantial differences in criterion counts: Sonnet-4.5 typically generates more than 20 items, whereas Qwen3-32B and Meta-CWM produce smaller sets.
- The rubric generators use their default reasoning effort when applicable.
- Criterion count is associated with selection performance but does not explain it alone: Gemini-3-Pro generates fewer criteria while retaining strong performance.
- The section prose summarizes frontier performance at 54%, open-weight coding models near 45%, and Qwen3-32B near 43%. Figure 6 shows different approximate endpoints for several models, so these prose summaries should not be read as exact per-model measurements.
- Generation failures also affect selection because malformed or missing artifacts are assigned zero scores.
With the candidate generator and judge fixed, rubric-generation models produce different selection curves and criterion-count distributions. Gemini-3-Pro performs strongly despite generating fewer criteria, showing that count alone is insufficient to explain the ordering. Several plotted endpoints differ from the section’s rounded numerical summaries, so the figure supports the comparative pattern more clearly than those exact summaries.
5.1.1 Training Open-Weight Rubric Agents. The authors fine-tune Qwen3-32B on 2,000 Sonnet-4.5 rubric-generation trajectories, without on-policy rubric samples. The classifier comparison uses 2,000 R2E-Gym prompts and 4,696 approximately balanced, test-labeled examples from Sonnet-4.5 and Qwen3-32B rollouts, with YES/NO token probabilities used for scoring.
- Training uses AdamW for 2 epochs, learning rate 1.0e−5 with cosine scheduling, batch size 32, and 4 nodes of 8xH100 GPUs.
- The datasets differ in supervision and trajectory content, so this compares practical training recipes rather than matched information budgets.
Figure 7 shows that the fine-tuned rubric generator improves selection over its base version and both base and fine-tuned patch classifiers. The fine-tuned classifier’s curve flattens and slightly declines at larger K, while the fine-tuned rubric curve rises and then largely plateaus; exact endpoint values are not tabulated.
- Appendix Figure 10 shows improved valid-artifact production and a criterion-count distribution closer to the teacher’s after fine-tuning.
- This demonstrates trainability of rubric generation, not reinforcement-learning improvement of a patch-generating agent.
5.2 Impact of Repository Grounding
Removing repository access from Sonnet-4.5 rubric generation reduces BEST@16 by 4.0 percentage points on Qwen3-32B candidates and 1.4 points on Qwen3-Coder-30B-A3B candidates. Table 1 illustrates how grounding replaces general descriptions with checks tied to concrete symbols such as visit_Compare in EvaluateFalseTransformer and KeyboardTransform in transforms.py.
- Search and file-inspection calls identify implementation entities that issue-only rubrics do not name.
- The ablation supports a benefit from context gathering, but concrete edit-location requirements can also introduce the over-prescription identified in the utility audit.
Repository searches and file inspection turn general criteria into checks naming visit_Compare in EvaluateFalseTransformer and KeyboardTransform in transforms.py. These examples illustrate the grounding mechanism; the repository-access ablation supplies the quantitative evidence for improved selection.
5.3 Sensitivity to Judge Model Choice
Table 2 reports BEST@16 of 52.6 ± 2.20 for GPT-5-mini, 54.2 ± 2.22 for GPT-5 Low Reasoning, 54.3 ± 2.25 for GPT-5 Medium Reasoning, and 55.0 ± 2.21 for GPT-5 High Reasoning. The modest differences support low reasoning as a practical grading choice in this experiment, but do not establish a statistically significant ordering.
- The section prose names Qwen3-32B candidates, whereas Table 2’s caption names Qwen3-Coder-30B-A3B. The numbers are attributed here to the table’s stated condition.
- The paper does not define the meaning of the ± values in this table.
- Appendix A.4 separately assesses repeated-grading determinism.
6. Related Work
The related work connects repository-based coding benchmarks and agent scaffolds with test-time scaling, learned patch verifiers, and generated tests. It also situates the method within rubric-based LLM evaluation and rubric rewards, distinguishing the present contribution as repository-aware criterion generation evaluated through SWE patch selection.
7. Conclusion
The conclusion is supported most directly for parallel test-time selection: repository-grounded rubrics outperform the tested verifiers and provide inspectable criterion-level judgments. Alignment analyses and examples suggest complementary diagnostic value, while supervised distillation improves an open-weight rubric generator; integrating rubric rewards into post-training remains future work.
8. Limitations
The paper does not integrate rubric signals into reinforcement-learning post-training. It leaves reward hacking, non-stationarity as policies improve, and credit assignment over multi-step agent behavior for future work.
- Automatically generated rubrics can be redundant, over-prescriptive, or inconsistent with tests and intended behavior. Human review, editing, template reuse, and targeted refinement are proposed but not evaluated.
- The empirical scope is SWE-Bench Verified with two fixed candidate generators; generalization to broader repositories or open-ended software tasks is not demonstrated.
- Execution-free grading does not eliminate the initial repository sandbox or expert-model cost needed to construct the artifact.
Appendix
- A.1 Analyzing agentic rubric scores against Ground-Truth patch. Table 3 reports mean weighted scores on human-written patches: Opus-4.5 0.8658, GPT-5 0.8413, Sonnet-4.5 0.8233, Gemini-3-Pro 0.8082, Meta-CWM 0.8015, Qwen3-Coder-30B-A3B 0.8037, and Qwen3-32B 0.6729. Figure 8 breaks scores down by axis for a representative subset, showing particularly uneven File Change alignment.
- A.2 Agentic abilities of rubric generation models. Figure 9 reports parseable-file rates of 97.6% for Opus-4.5, 99.4% for Sonnet-4.5, 96.8% for Gemini-3-Pro, 86.2% for GPT-5, 88.2% for Qwen3-Coder-30B-A3B, 74.6% for Qwen3-32B, and 69.4% for Meta-CWM. The accompanying prose instead states 97.8% for Sonnet-4.5; Figure 10 reports that fine-tuning Qwen3-32B raises parseability to 88.8% and reduces parse errors from 18.2% to 3.6%, while moving criterion counts closer to the teacher’s distribution.
- A.3 Cost analysis for agentic verification methods. Table 4 gives artifact cost, per-candidate grading cost, and average API calls as $0.640/$0.006/48.5 for Patch Similarity, $0.499/$0.001/29.2 for Tests, and $0.245/$0.003/22.9 for Rubrics. Reported total BEST@16 verification costs are $0.736, $0.515, and $0.293 per instance, respectively; test grading includes sandbox execution on Modal. This comparison gives proxy-patch generation 50 steps rather than 30 and reports 36.6% for its Qwen3-32B selection result, versus 35.0% in the main comparison, so its conditions differ from the main baseline setup.
- A.4 Rubric Flakiness Study. For each of Sonnet-4.5 and Qwen3-32B, the study samples 20 instances and 5 criteria per instance, then grades each of the 100 criteria 5 independent times with GPT-5 low reasoning. Any disagreement marks an item as flaky: rates are 2% for Sonnet-4.5 criteria and 9% for Qwen3-32B criteria, providing small-sample evidence of grading stability rather than correctness.
- A.5 Hybrid verifiers using rubrics v/s classifier. Figure 11 shows that combining Agentic Rubrics with Agentic Tests outperforms either alone and the classifier-plus-tests alternative in the plotted comparison. The appendix does not tabulate exact gains or specify the aggregation formula, limiting reproducibility of the hybrid result.
- A.6 Categories of rubric utility classification. Table 5 presents a manually constructed taxonomy based on inspected rollouts and failure modes. High-utility categories are Rule Break, Band-Aid Fix, Wrong Layer, Root Cause Missed, Missing Edges, Scope Creep, Perf Risk, API Break, and Security Risk. Low-utility categories are Test Rules Mismatch, Over-Specified Fix, API Over-Strict, Style Nit, Spec Clash, Ref Patch Conflict, Redundant Signal, Eval Bug, and Irrelevant Rule.
- A.7 SWE Agent Setup. All agentic methods use SWE-Agent with remaining-turn reminders every 5 turns. The authors modify response parsing to require a short per-turn planning summary alongside tool use, facilitating collection and distillation of rubric-generation trajectories.
- A.8 Rubric and their grading - Illustrative Examples. For matplotlib__matplotlib-26291, a candidate returns a zero-size bounding box when renderer=None or selected rendering errors occur, avoiding the reported crash and passing ground-truth tests. The paper argues that this workaround causes tight layout to ignore the inset, conflicting with R4’s positioning requirement, SA4’s absolute and relative size requirements, SA5’s bbox_inches=”tight” requirement, and R3’s backward-compatibility requirement.
- A.9 Rubric Examples and grading. The appendix prints diffs and criterion-level outcomes for three test-passing, low-rubric-scoring patches. For matplotlib__matplotlib-26291, the zero-size-box workaround receives both behavioral penalties and implementation-specific penalties for not setting self.figure. For matplotlib__matplotlib-25332, adding Grouper pickle methods receives penalties involving Grouper behavior, repeated alignment, and collected axes, but also prescriptive requirements to edit figure.py and add a test; for django__django-13417, repeated clear_ordering(True) calls receive penalties for preserved ordering behavior and state mutation, alongside requirements to modify QuerySet.ordered at a specified location.
- A.10 Prompts - Baselines and Agentic Rubrics. The appendix documents Agentic Patch Similarity (Rollout Generation), Agentic Patch Similarity (Judge), Agentic Rubrics, Agentic Tests, Patch Classifier, and Non-Agentic Rubrics. These specify minimal proxy fixes, 1–5 similarity grading, parseable rubrics.yaml criteria, standalone test_issue.py generation, YES/NO classification, and issue-only rubric generation, respectively. The rubric-generation documentation makes atomicity, self-containment, repository specificity, and avoidance of double-counting explicit design requirements.
- A.11 Rubric Judge Prompt. The grading interface supplies PR_DESCRIPTION, PATCH, and RUBRIC and requests one binary satisfaction score per criterion identifier. It confirms that the grader receives patch text rather than execution results or full agent trajectories.
- A.12 Rubric Utility Analysis Prompt. The utility classifier receives the issue, golden patch, candidate patch, golden tests, test reward, rubric descriptions, and selected criterion identifiers. It classifies judgments as Valid or Spurious with subcategories and evidence-based justifications; the printed template focuses on aligned positive and negative judgments, so it does not fully document the separate stricter-than-tests audit described in the main text.
Brief Thoughts
The main contribution is separating contextual evidence gathering from repeated patch grading. The repository-access ablation, selection gains across two candidate generators, and measured verification costs support this design for best-of-K selection; explicit criteria also expose why a patch was penalized, unlike a single classifier score.
Rubrics provide complementary evidence, not executable guarantees. The 46% low-utility share among stricter-than-test rejections, model-based utility labeling, and inconsistent reporting of several conditions warrant caution; the hybrid-verifier result supports combining rubric judgments with tests rather than treating them as a complete replacement.