Gorio Tech Blog search

FrontierCS: Evolving Challenges for Evolving Intelligence | Summary

|

Contents

This article explains the key points of FrontierCS: Evolving Challenges for Evolving Intelligence.

  • 2025-12-17 (arXiv)
  • Mang, Qiuyang, Chai, Wenhao, Li, Zhifei, Mao, Huanzhi, Zhou, Shang, Du, Alexander, Li, Hanchen, Liu, Shu, Chen, Edwin, Wang, Yichuan, et al.
  • UC Berkeley, Princeton University, UCSD, X-camp Academy, Independent, Georgia Tech, Stanford University, University of Washington, Nanyang Technological University, University of Toronto, UIUC, University of Michigan, New York University, MIT
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • FrontierCS introduces 156 expert-curated computer science problems: 107 algorithmic problems and 49 research problems. Models submit executable programs, and deterministic evaluators measure solution quality when global optima are unknown or practically unattainable.
  • In single-round evaluation without code execution or external tools, Gemini 3.0 Pro leads the algorithmic track with Score@1 29.37 and Score@5 52.06, compared with the human reference score of 95.41. Claude Opus 4.5 leads research Score@1 with 29.40, while GPT 5.1 Thinking leads research Score@5 with 47.21.
  • Additional sampling improves scores, but larger reasoning budgets do not consistently improve algorithmic performance. A Polyomino Packing case study reports substantial gains from guidance on internal representation, illustrating the difference between producing workable code and selecting effective optimization strategies.

1. Introduction

FrontierCS evaluates open-ended computer science problem solving beyond tasks with known answers and binary correctness. Its design requires an unknown or intractable optimum, deterministic validity checks and quantitative scoring, and a parametric problem generator supporting fresh instances of varying difficulty.

  • The evaluator runs a self-contained solver under task-specific time and memory limits; the model receives the specification and any required I/O or API stubs.
  • In the introductory Polyomino Packing example, both solutions are valid, but the human expert achieves 87% density and GPT-5 Thinking achieves 47%. Validity alone misses this quality difference.
  • The authors propose using the evaluators as reward signals for reinforcement learning or self-play, but the paper reports evaluation experiments rather than training results.

The human packing fills 87% of the grid, whereas GPT-5 Thinking fills 47%. Both satisfy validity requirements, demonstrating why a quantitative quality metric reveals differences that binary acceptance would conceal.

Polyomino Packing solutions illustrating valid outputs with substantially different packing densities.
Polyomino Packing solutions illustrating valid outputs with substantially different packing densities.

The paper compares FrontierCS with closed-form coding and reasoning benchmarks, optimization benchmarks, and open-ended research evaluations. HumanEval, MBPP, SWE-bench, BFCL, and LiveCodeBench emphasize correctness; ALE-Bench evaluates heuristics with objective scores; UQ uses human validation; MLR-Bench evaluates research with an automated judge; and KernelBench measures kernel correctness and runtime.

  • FrontierCS combines partial scoring with tasks across multiple computer science domains and deterministic evaluation that does not require an LLM judge.
  • Its distinction from ALE-Bench concerns problem and author diversity and the inclusion of research workflows; objective partial scoring already appears in prior work.
  • FrontierMath and LiveCodeBench Pro provide difficult expert-authored tasks, but the paper emphasizes their known reference answers or solutions. NP-Engine provides a benchmark and training framework for 10 classic NP-hard tasks.

3. Problem Collection

The collection contains 29 Optimization, 27 Constructive, and 51 Interactive algorithmic problems. Its 49 research problems comprise 8 Operating Systems, 19 High-Performance Computing, 6 Artificial Intelligence, 7 Databases, 5 Programming Languages, and 4 Security problems.

  • Constructive tasks synthesize valid structures under global constraints; Optimization tasks search a parameterized space; Interactive tasks adapt actions to previous responses.
  • Research tasks are assigned to the domain their proposed solution primarily addresses, including when the underlying topic spans multiple domains.
  • Both tracks use proposal, implementation, and review stages.

Interactive tasks account for 51 of 107 algorithmic problems, and High-Performance Computing accounts for 19 of 49 research problems. The benchmark spans multiple domains with an uneven distribution, which affects the composition of aggregate scores.

Distribution of 156 problems across three algorithmic categories and six research domains.
Distribution of 156 problems across three algorithmic categories and six research domains.

3.1. Algorithmic Problems

Experts with qualifications equivalent to ICPC World Finalists propose adaptations of programming-contest and classical computer science problems. Implementation introduces open-ended objectives and partial scoring, standardizes interfaces, and supplies deterministic verifiers, test generators, baselines, reference solutions, and evaluation scripts.

  • A second expert checks openness, scoring quality, evaluator correctness, test coverage, and whether the human reference significantly outperforms the best model.
  • Cost, density, and query count are typical quality metrics. Time and memory generally serve as feasibility constraints, with violations receiving no score; runtime contributes to scoring only when explicitly specified.
  • Unlike contest subtasks with discrete all-or-nothing credit, task-specific continuous or piecewise-continuous scores reward incremental improvements relative to a trivial baseline and a strong human reference.

3.2. Research Problems

CS PhD students propose research problems from unsolved questions and implement specialized evaluation environments. Accepted tasks provide a description, pinned dependencies or a Docker image, a deterministic evaluator with diagnostics, an expert reference, and trivial baselines; scoring must not use an LLM judge.

  • The README specifies VM or Docker requirements, data layout, and the solution’s input/output contract. The text names set_up_env.sh and evaluate.sh, while the pipeline diagram uses setup.sh, install_env.sh, and start_solve.py.
  • SkyPilot manages heterogeneous compute infrastructure. Review requires scripts to execute unattended in a fresh VM and an isolated, deterministic evaluation environment without external dependencies.
  • Research objectives may combine accuracy, latency, memory, and cost, with resource-limit violations receiving no score. Variants with different constraints, hardware, or objectives count as separate problems in the reported total of 49.

A proxy combines proposer-provided VM requirements and evaluation scripts with contestant-provided installation and solving scripts. It launches a testing VM, executes the setup and solution stages, and returns a score, placing task-specific environments within the evaluation pipeline.

Research evaluation pipeline connecting contestant inputs, proposer configuration, a proxy server, and a testing VM using SkyPilot.
Research evaluation pipeline connecting contestant inputs, proposer configuration, a proxy server, and a testing VM using SkyPilot.

3.3. Update Policy

The Update Policy supports adding new tasks, increasing existing task difficulty without rewriting the problem statement, and refining human references or evaluation thresholds. Changes can tighten resource limits, alter workloads or datasets, or raise optimization targets.

  • Separating specifications from instantiated environments lets the benchmark retain task continuity while changing difficulty.
  • References and scoring thresholds can change, so comparisons across releases require the associated benchmark version. The paper proposes an update policy but does not present longitudinal results demonstrating comparability.

4. Evaluation Results

Evaluation Results report the metric design and model performance on both tracks. The experiments distinguish one-attempt performance, average quality across five attempts, and the best result among five attempts.

4.1. Setup

The reported tables contain nine model rows per track, including GPT 5 Thinking, GPT 5.1 Thinking, Gemini 2.5 Pro, Gemini 3.0 Pro, Claude Opus 4.1, Claude Sonnet 4.5, Claude Opus 4.5, and DeepSeek 3.2. The remaining row is labeled Grok 4 in the algorithmic table and Grok 4 Fast in the research table; each LLM request has a 20-minute timeout.

  • GPT-5 / GPT-5.1 Thinking, Grok 4, and Claude Opus 4.5 use reasoning_effort = high. Claude Opus 4.1 / Sonnet 4.5 use max_tokens = 32,000 and reasoning_budget = 20,000; Gemini 2.5 Pro / Gemini 3.0 Pro use thinking_budget = -1.
  • Each attempt is a single-round, text-only generation. Models cannot execute code, inspect tests, revise using feedback, or access an editor, Python environment, or external tools.
  • The setup prose says nine models but lists eight names, and subsequent results prose refers to eight despite nine table rows. The tables provide the explicit record of the reported comparisons.

4.2. Metrics

Scores are calibrated against a trivial baseline and an expert reference, with nontrivial performance bounds used where available. The general scheme assigns zero below the baseline and full credit at the designated reference threshold or bound, but the examples specify task-dependent mappings, including complexity penalties and a human reference level of 95 for Square Packing.

  • Score@k is the maximum score among k trials, and Avg@k is their mean.
  • The metrics text defines Pass@1 and Pass@5 as the fraction of problems receiving a nonzero score, meaning a valid solution that exceeds the trivial baseline. Table 1’s caption instead refers to a proportion of runs, leaving a wording inconsistency.
  • Normalized scores measure progress relative to task-specific anchors; they do not establish proximity to an unknown global optimum.

4.3. Algorithmic Problems

Gemini 3.0 Pro achieves the strongest algorithmic results: Score@1 29.37, Avg@5 29.51, Score@5 52.06, Pass@1 65.42%, and Pass@5 83.18%. The human expert score is 95.41, so even the best-of-five model result remains substantially below the reference.

  • Other Score@1 values are Claude Opus 4.5 14.95, Grok 4 13.67, GPT 5 Thinking 10.87, GPT 5.1 Thinking 11.80, DeepSeek 3.2 12.65, Gemini 2.5 Pro 12.53, Claude Opus 4.1 6.91, and Claude Sonnet 4.5 5.84.
  • Score@5 exceeds Score@1 by 6.40–22.69 points across the reported models. Selecting the best of five attempts improves the measured score under this protocol; it does not demonstrate feedback-based refinement.

Gemini 3.0 Pro leads all five model metrics, but its Score@5 of 52.06 remains below the human reference score of 95.41. Selecting from five attempts improves scores without eliminating the reference gap.

Algorithmic results reporting Score@1, Avg@5, Score@5, Pass@1, and Pass@5.
Algorithmic results reporting Score@1, Avg@5, Score@5, Pass@1, and Pass@5.

4.4. Research Problems

Claude Opus 4.5 leads research Score@1 with 29.40 and Avg@5 with 32.31. GPT 5.1 Thinking leads Score@5 with 47.21 and Pass@5 with 83.67%, closely followed by Gemini 3.0 Pro’s Score@5 of 47.20.

  • Research Score@5 improves over Score@1 by 6.72–21.04 points across models.
  • Frequent nonzero scores alongside modest normalized quality scores support the authors’ interpretation that many submissions exceed trivial baselines while remaining under-optimized. Pass rates and quality scores measure different quantities.
  • Table 2 does not include a separate human aggregate row, so the explicit 95.41 human comparison belongs to the algorithmic track.

Claude Opus 4.5 has the highest Score@1 and Avg@5, while GPT 5.1 Thinking has the highest Score@5 and Pass@5. The leading model therefore depends on whether evaluation measures a single attempt, average quality, or selection from several attempts.

Research results reporting the same five evaluation metrics across nine model rows.
Research results reporting the same five evaluation metrics across nine model rows.

5. Discussion

The Discussion examines reasoning effort, internal representation, and the distinction between producing workable software and optimizing its quality. The evidence combines a reasoning-budget experiment with task-level submission analysis.

5.1. Improving Reasoning Effort Does Not Yield Further Gains

With three attempts per algorithmic problem, GPT 5 low effort averages 4,389 reasoning tokens and score 7.903; medium effort averages 11,554 tokens and 15.336; high effort averages 19,763 tokens and 12.626. GPT 5.1 high effort averages 20,402 tokens and 12.508.

  • The increase from low to medium improves the mean score, while the increase from medium to high lowers it despite greater token consumption.
  • The tested settings show non-monotonic performance. The figure reports no confidence intervals or significance tests and does not establish a universal upper limit on reasoning benefits.

Medium GPT 5 effort averages score 15.336 with 11,554 tokens, exceeding high effort’s 12.626 with 19,763 tokens. Many attempts receive zero across effort levels, and longer reasoning does not consistently produce stronger solutions in this experiment.

Per-attempt reasoning tokens and scores, with group means from three attempts per problem.
Per-attempt reasoning tokens and scores, with group means from three attempts per problem.

5.2. Misleading Micro-Optimization Trap

In Polyomino Packing, GPT 5-Thinking often uses the required output transformation list as its internal representation, making overlap detection and free-space search cumbersome. The authors report invalid code in about 30% of attempts, with scores of 20–70 in the remaining 70%.

  • The intervention adds: “Please use a 2D array to maintain the rectangle state, and convert to the required format only at the end”.
  • With that instruction, the zero-score rate drops to about 10%, and nearly 80% of cases achieve scores in the 80–85 range through an efficient search strategy.
  • The case study supports the value of representation guidance for this task. The section does not report the sample count or uncertainty estimates for these percentages.

5.3. The Research–Engineering Dilemma of Claude

The authors’ inspection finds that Claude Sonnet 4.5 often generates executable, resource-compliant programs that lack competitive optimization strategies and fall below the algorithmic trivial baseline. Research tasks also require integrating tools and configuring systems, allowing workable solutions to earn meaningful partial credit.

  • Claude Sonnet 4.5 scores 5.84 on algorithmic Score@1 and 24.65 on research Score@1; Claude Opus 4.5 scores 14.95 and 29.40, respectively.
  • The Symbolic Regression example requires invoking PySR, configuring expression search, tuning evolutionary parameters, and validating formulas. The authors report that workable but limited solutions can earn around 50 points.
  • The proposed connection to software-engineering tuning is an interpretation. The paper does not experimentally isolate training choices as the cause of the cross-track disparity.

6. Example Problems

The ten Example Problems cover compact constructions, constrained optimization, adaptive querying, symbolic regression, systems tradeoffs, vulnerability reproduction, kernel acceleration, and strategic decisions. Each example specifies validity requirements and a quantitative objective, and several compare human and model strategies directly.

Problem 1: World Map

Problem 1: World Map adapts an IOI 2025 task: construct a K × K grid whose orthogonal country adjacencies match a specified graph, while minimizing K. For valid outputs, R = K/N and score = 100 × clamp((6 − R)/(6 − R′), 0, 1), where R′ is the human reference ratio; invalid adjacencies receive zero.

  • Every required country pair must share at least one cell edge, and every edge between distinct country labels must correspond to a required pair. Diagonal contact does not count.
  • The paper reports a known construction with R′ = 1.5 across all possible inputs, while leaving the best attainable ratios unresolved.
  • For the illustrated case with N = 5 and M = 6, the human construction uses K = 7 and GPT 5 uses K = 15. Both are valid, but their compactness differs.

For the same five-country adjacency graph, the human construction uses a 7 × 7 grid and GPT 5 uses a 15 × 15 grid. Both realize valid adjacencies, so grid compactness supplies the discriminating objective.

World Map adjacency graph and valid human and GPT 5 grid constructions.
World Map adjacency graph and valid human and GPT 5 grid constructions.

Problem 2: Treasure Packing

Problem 2: Treasure Packing adapts the 2024 ICPC North America Championship NSA Challenge. It maximizes total value across C = 12 treasure categories under mass and volume capacities, with nonnegative integer counts bounded by available quantities.

  • The objective is max Σc v_c x_c, subject to Σc m_c x_c ≤ M, Σc ℓ_c x_c ≤ L, and 0 ≤ x_c ≤ q_c. Test ranges are 1 ≤ q_c ≤ 10⁴, 1 ≤ v_c ≤ 10⁶, 1 ≤ m_c ≤ 20 × 10⁶, and 1 ≤ ℓ_c ≤ 25 × 10⁶.
  • Valid solutions must finish within 1 second and 1024 MB. Their score is 100 × clamp((value − value_base)/(value_ref − value_base), 0, 1); invalid or over-budget outputs receive zero.
  • The human reference combines greedy and randomized improvement. GPT 5 combines a greedy pass with branch-and-bound, earning 74 points against the human reference’s 100.

Problem 3: Permutation Guess

Problem 3: Permutation Guess asks a program to identify a hidden permutation of n = 1000 using as few queries as possible. Each query submits a length-n integer sequence, which need not itself be a permutation, and the response counts positions matching the hidden permutation.

  • An incorrect final permutation receives zero. A correct result scores 100 × clamp((Q_base − Q)/(Q_base − Q_ref), 0, 1), using approximately 10,000 queries for a naive binary-search baseline and approximately 6,000 for the human divide-and-conquer reference.
  • For the toy instance n = 4 with hidden permutation 1432, the illustration lists 5 human steps and 12 GPT 5 steps, including final submission. These correspond to 4 and 11 query calls.
  • The displayed transcript contains responses inconsistent with the stated hidden permutation: for example, query 1411 matches 1432 in one position, although the figure reports 2. The transcripts therefore illustrate the authors’ efficiency comparison but do not provide a correct worked derivation.

The transcripts list 5 human steps and 12 GPT 5 steps, including final submission, and show repeated model queries. Some responses conflict with the stated hidden permutation 1432: query 1411 should return 1 rather than the displayed 2, and query 4222 should return 2 rather than 1. The listed step counts can be described, but the transcript is not a correct worked solution.

Permutation Guess strategy transcripts for n = 4 and hidden permutation 1432.
Permutation Guess strategy transcripts for n = 4 and hidden permutation 1432.

Problem 4: Square Packing

Problem 4: Square Packing places 1 ≤ n ≤ 10,000 unit squares inside an axis-aligned square container while minimizing its side length L. Individual squares may rotate arbitrarily, and touching is allowed, but interiors cannot overlap.

  • Invalid packings receive zero. The lower bound is L_B = √n and the naive upper bound is L_0 = ⌈√n⌉. Valid outputs receive 100 at L = L_B; 95 + 5 × min(1.0, 1.1 × (s − L)/(s − L_B)) for L_B < L ≤ s; 94 × min(1.0, 1.1 × (L_0 − L)/(L_0 − s)) + 1 for s < L < L_0; and zero for L ≥ L_0.
  • The reference s(n) uses best known human packings for n ≤ 100 and s(n) = 2 × s(⌈n/4⌉) for n > 100. The authors describe 95 as the human benchmark level, with the displayed formula allowing additional credit for sufficiently improved packings.
  • For n = 10, the human packing has L = 3.707, while Gemini 2.5 Pro uses the naive L = 4.

The human configuration uses a rotated square to achieve container side length 3.707, whereas Gemini 2.5 Pro uses side length 4. The comparison illustrates improvement beyond a valid naive grid arrangement.

Valid human and Gemini 2.5 Pro packings of ten unit squares.
Valid human and Gemini 2.5 Pro packings of ten unit squares.

Problem 5: Polyomino Packing

Problem 5: Polyomino Packing places every distinct polyomino of sizes 1 through n inside a W × H grid. Permitted transformations are integer translation, rotations of 0/90/180/270°, and optional reflection across the y-axis.

  • All pieces must lie inside the grid without overlapping. The objective minimizes W × H, equivalently maximizing density ρ = packed cells/(W × H).
  • Invalid placements receive zero. Scores scale linearly from zero at the supplied baseline density ρ_base to 100 at the reference density ρ_ref, with ρ_ref ≤ 1.
  • The illustrated valid packings achieve 87% density for the human expert and 47% for GPT 5-Thinking.

The shapes illustrate how differing piece geometry constrains a common packing arrangement. A solver must track each piece’s occupied cells under the permitted translations, rotations, and reflections.

Examples of polyomino shapes used to illustrate the packing task.
Examples of polyomino shapes used to illustrate the packing task.

The expert arrangement contains visibly less empty space than the model arrangement, with densities of 87% and 47%. The introductory comparison reappears alongside the full specification, connecting compactness to the task’s density objective.

Human and GPT 5-Thinking Polyomino Packing outputs for the same instance.
Human and GPT 5-Thinking Polyomino Packing outputs for the same instance.

Problem 6: Symbolic Regression

Problem 6: Symbolic Regression searches for a simple expression predicting a supervised target from features. The allowed grammar includes addition, subtraction, multiplication, division, exp, log, sin, cos, parentheses, and input variables; expressions must produce finite real values on every data row or receive zero.

  • Expression complexity is C = 2 × (# binary ops) + 1 × (# unary ops). Prediction quality uses MSE = (1/n) Σi (y_i − ŷ_i)².
  • The score is 100 × clamp((m_base − MSE)/(m_base − m_ref), 0, 1) × 0.99^max(C − C_ref, 0), where m_base is the best linear predictor’s MSE and m_ref is the reference MSE. If m_base = m_ref, the stated rule assigns 100 when MSE ≤ m_ref and zero otherwise.
  • For the McCormick function example, the human expression has complexity 12 and GPT 5’s expression has complexity 19. The visual comparison provides no numerical MSE values.

The data surface is compared with expressions of complexity 12 from the human expert and 19 from GPT 5. The figure permits a visual comparison of fitted shape and reports expression simplicity, but supplies no numerical prediction-error measurement.

McCormick function data and symbolic-regression surfaces discovered by a human expert and GPT 5.
McCormick function data and symbolic-regression surfaces discovered by a human expert and GPT 5.

Problem 7: Vector Database Design — Recall–Latency Tradeoff

Problem 7: Vector Database Design — Recall–Latency Tradeoff evaluates an approximate nearest-neighbor index on SIFT1M with 1M base vectors and 10K queries, both 128D, using the L2 metric and k = 1. It reports Recall@1 and average search latency in ms, with a 10-hour end-to-end build-and-search limit.

  • Validity requires the specified API, finite distances, valid integer indices, r ≥ r_min, and t ≤ t_thr. Under the latency threshold, Score = 100 × clamp((r − r_min)/(r_base − r_min), 0, 1).
  • Variants express different tradeoffs, including Recall80, which minimizes latency while maintaining at least 80% recall.
  • Human references tune established index parameters such as nprobe in IVF and efSearch in HNSW. The plotted human results offer a better recall–latency tradeoff than the GPT-5 Thinking results.

The plotted human index configurations provide higher recall at comparable latency or lower latency at comparable recall than the model configurations. The comparison demonstrates the importance of index configuration after implementing a functional ANN search system.

Recall@1 versus average query latency on SIFT1M for human and GPT-5 Thinking index variants at k = 1.
Recall@1 versus average query latency on SIFT1M for human and GPT-5 Thinking index variants at k = 1.

Problem 8: Minimal PoC Generation

Problem 8: Minimal PoC Generation asks a program to produce a short proof-of-concept input for a specified vulnerability in an open-source codebase. CyberGym verifies that the input triggers a sanitizer crash before the patch and produces no sanitizer crash after the patch.

  • A successful input scores 60 + 40 × 2^(−L/L_g), where L is its length and L_g is the ground-truth PoC length; failure to trigger the target vulnerability receives zero.
  • The PHP heap-use-after-free example compares a human-generated 79-byte PoC with a GPT 5-generated 577-byte PoC. Both are valid, so the difference concerns input minimization.
  • The task description permits a generated solution to incorporate an agentic workflow for processing large codebases, while the reported model-generation protocol remains single-round.

Both displayed inputs reproduce the target PHP heap-use-after-free vulnerability, but the human PoC is 79 bytes and GPT 5’s is 577 bytes. The example compares minimization quality after successful vulnerability reproduction.

Hexadecimal representations of valid human and GPT 5 proof-of-concept inputs.
Hexadecimal representations of valid human and GPT 5 proof-of-concept inputs.

Problem 9: Kernel Rewrite & Warp Specialization for GDPA

Problem 9: Kernel Rewrite & Warp Specialization for GDPA asks for a Triton implementation of gated dot-product attention and a further warp-specialized implementation using TLX. With α = 1/√d, the operation computes Q̃ = Q ⊙ σ(G_Q), K̃ = K ⊙ σ(G_K), S = αQ̃K̃ᵀ, P = softmax(S) over keys, and O = PV.

  • Inputs are float16, accumulation must use float32, and outputs default to float16. Correctness uses rtol=1e-3 and atol=5e-4.
  • A baseline _pt_gdpa implementation is supplied in PyTorch; the requested Triton routine is triton_gdpa. Query and output tensors have batch, head, query-length, and head-dimension axes, while key and value tensors use the key/value length.
  • Score = 100 × (1 − T_solution/T_baseline), measuring the fraction of runtime saved across benchmark shapes. This example supplies the task definition but no numerical model-versus-human result.

Problem 10: Poker Strategy Optimization

Problem 10: Poker Strategy Optimization evaluates heads-up Texas Hold’em against a fixed opponent. The opponent estimates Call EV with 100 Monte Carlo trials, assumes both players subsequently Check, and calls only when Call EV exceeds Fold EV, defined as current chips − 100.

  • The submitted routine returns Check, Fold, or Raise(x), where 1 ≤ x ≤ remaining chips. Scores depend on average profit per hand ω: zero for ω ≤ 8.0; 13.3(ω − 8) for 8.0 < ω ≤ 11.0; 40 + 14(ω − 11) for 11.0 < ω ≤ 14.0; 82 + 3(ω − 14) for 14.0 < ω ≤ 20.0; and 100 for ω ≥ 20.0.
  • GPT 5 scores 25 and Gemini 2.5 Pro scores 36. A simple human strategy using Monte Carlo estimates and All-in above 0.75 winning probability scores 54.
  • The accompanying diagram labels its decision threshold Win Rate ≥ 75%, while the prose states > 0.75; the source does not reconcile this boundary difference.

The diagram selects All-in when its displayed Win Rate ≥ 75% condition holds and Check otherwise. The accompanying text reports a score of 54 for the simple human strategy, exceeding GPT 5’s 25 and Gemini 2.5 Pro’s 36, although its prose uses a strictly greater-than threshold.

Human poker strategy selecting All-in or Check from an estimated winning probability.
Human poker strategy selecting All-in or Check from an estimated winning probability.

7. Conclusion

The Conclusion presents FrontierCS as a testbed for deterministically verifiable, partially scored computer science tasks with unknown or impractical global optima. The experiments identify weaknesses in single-round optimization and systems tradeoffs, while proposed versioned updates aim to retain discrimination as models improve.

Appendix

  • The supplied 28-page PDF contains no appendix. Project resources are listed as www.frontier-cs.org and https://github.com/FrontierCS/Frontier-CS; the paper is identified as arXiv:2512.15699v1.

Brief Thoughts

FrontierCS separates validity from optimization quality through executable evaluators and explicit reference anchors. The measured human gap applies to a restricted single-round protocol, and algorithmic curation explicitly requires a strong human advantage, so the aggregate comparison is conditional on that selection criterion. The reasoning-budget experiment shows non-monotonic returns in the tested settings. The representation intervention suggests a task-specific mechanism, but its sample count and uncertainty are not reported. Research variants count as separate problems, score mappings differ by task, and changing references can alter normalized scores. Benchmark versions, variant composition, and per-task metrics matter when comparing systems. The source contains inconsistencies in model counts, Grok labels, Pass@k wording, and the Permutation Guess transcript. Iterative, tool-assisted model-generation systems remain unevaluated in this paper.