sankalp's blog

Honey, I Looked at the Data of a Frontier Benchmark and Found some Issues: Porting Agents' Last Exam Linux CLI subset to Verifiers v1

Intro

2 months back, I started the Prime Residency with Florian Brand as my peer or shall I say my verifier. My first task was to port the Agents’ Last Exam Linux CLI subset to verifiers. Then one day Florian asked me to look into the data and this quest slowly evolved into finding issues in the benchmark and the making of what we call ALE-Gold.

It was more challenging than I expected. I was new to this kind of work, so there was some skill issue on my part. Along the way, I also hit benchmark defects, inference provider failures, and gaps in the freshly released Verifiers v1 (that have been fixed now so you don't need to worry). This post is a work log of my porting work, what the full runs revealed about the benchmark, how that led to ALE-Gold, followed by an analysis of how the models performed.

An ALE odyssey: a winding journey through porting, infrastructure issues, and benchmark defects to ALE-Gold

A small odyssey from porting ALE to ALE-Gold, with a few unexpected detours.

Table of contents

Porting ALE Linux CLI subset to verifiers

Agents’ Last Exam is considered a frontier benchmark made by Berkeley RDI, Dawn Song’s group at UC Berkeley. They worked with more than 300 experts across 100+ institutions to make long horizon economically valuable tasks spanning over 55 domains. Of the 1,500+ tasks in the corpus, 147 are public.

ALE was also featured in OpenAI’s GPT-6 Astra launch, where Astra scored 59.3% on Agents’ Last Exam.

My task was to port the Linux CLI subset, which consisted of 105 terminal-only tasks spanning 41 subdomains across 12 broad domain categories.

Why port to verifiers? ALE comes with its own evaluation implementation rather than being built on a standardized framework like Harbor or verifiers. Porting it to verifiers brings the tasks into a shared evaluation framework, with standardized rollout traces and a consistent way to configure models, harnesses, and runtimes. This lets others run their own evaluations using the same integration, and makes the resulting traces easier to inspect and compare.

What’s in ALE Linux CLI?

105 tasks across 12 broad domain categories

Health & medicine
22 · 21.0%
Computing & mathematics
19 · 18.1%
Life sciences
19 · 18.1%
Business & finance
14 · 13.3%
Physical sciences
11 · 10.5%
Engineering
8 · 7.6%
Education & information
3 · 2.9%
Transport & safety
3 · 2.9%
Legal
2 · 1.9%
Psychology & neuroscience
2 · 1.9%
Social sciences
1 · 1.0%
Other
1 · 1.0%

Share of the original 105-task CLI subset. Domains are grouped by task-name prefix throughout this post. Percentages are rounded.

One may wonder, “What does it mean to port a benchmark to verifiers?” Verifiers is a library by Prime Intellect that enables you to run “rollouts” in an easy and elegant manner. I mention the word rollout multiple times in this post so it's worth doing a small detour to understand it.

What is a Rollout?

A rollout in the context of evaluation and reinforcement learning is defined as a single run of a policy/model in an environment from start to finish, recording the states, actions, and rewards along the way.

For example when doing LLM-RL, you will be running multiple rollouts of a single task and these can have different scores/rewards. The policy parameters are updated using the spread of these rewards. If you use GRPO, differences in rewards within a group of rollouts provide a learning signal by showing which rollouts performed better relative to the others.

Parallel rollouts of the same task, each with a harness inside a sandbox and an illustrative reward

Same task, different attempts. Each rollout has its own sandbox and harness, takes a different path, and receives a score. Timelines and scores are illustrative.

If you look at benchmarks like FrontierSWE or TerminalBench (Harbor framework), they are usually divided into three parts - the task’s input in the form of files or a Docker image, a harness where the model will run like Claude Code and Codex and the runtime (where the rollout happens). The runtime can be your local machine, a Docker instance, or an AWS instance.

Verifiers provides this exact abstraction for you to port the benchmark and then run the benchmark on Prime Intellect’s infra including support for several sandbox providers.

Verifiers taskset, harness, and runtime

A taskset defines the work and scoring, a harness controls the agent’s interaction loop, and a runtime provides the execution environment.

What it took to port ALE

To give some context, here’s what the porting process generally looks like.

The porting work mainly involves porting the taskset/dataset to verifiers and then writing relevant adapter code between verifiers and the original public benchmark code (which I refer to as “upstream”) to support evaluation (as it’s defined by upstream). In short, port the input and hook up the scorer code. In addition to this, one may need to write some more compatibility related code depending on the benchmark. The “porting work” itself can vary depending on if the benchmark providers have provided sandbox images or not.

In ALE’s case, images were not provided so I had to make sandbox images for 105 tasks with the mentioned dependency specifications and input files. This was a code intensive and time consuming process and was only made possible due to coding agents.

One must minimise the rollout time

Once the images were made, I would run a couple of rollouts of a task to verify if the image was fine and the rollout happened successfully.

In my first attempt at building the images, I wasn’t baking the “input files” into them. Instead, I was downloading them at runtime, which meant fetching them from the Hugging Face archive during every rollout. Now, most of the input files were small, but for a few tasks, they were several GBs in size. Halfway through, I started hitting HF’s resolver limits…

This is when I realised my mistake and discussed it with Florian too. I was supposed to bake the inputs onto the image. Baking the inputs into the image adds storage cost, but saves time and network requests because the files no longer need to be downloaded during each rollout.

From an evaluation point of view, it’s a redundant operation and in general people make an effort to minimise the rollout time.

From an RL post-training point of view, this is an even bigger issue because you run multiple rollouts of the same task repeatedly across several training steps, so downloading the same input files for every rollout adds up quickly. RL researchers and engineers strive to reduce rollout time, often by increasing inference throughput, because rollout generation is frequently the biggest bottleneck. In a synchronous RL training setup, you need to wait for all rollouts to finish before running the gradient steps so the slowest rollout becomes the bottleneck. You can read a bit about it here (on how Prime Intellect engineers improve inference throughput).

Missing dependency specifications

Another challenge was that the task cards did not always list everything needed to run the supplied software. In physical_sciences/glm_lake_calibration, the card listed Python and pointed to a bundled lake-simulation binary, but declared no required system packages. The binary still depended on four libraries that were missing from the image: libgd3, libgfortran5, libgomp1, and libx11-6. My agent added those dependencies to the image and included a check for unresolved libraries.

Other failures came from images not matching what the task promised. In physical_sciences/exact_diag_heisenberg_j1j2, the task specified a Python runtime with NumPy and SciPy already installed, but the image was missing that setup. The agent ended up downloading and unpacking the libraries during its rollout. After I added the expected runtime, packages, and wrapper to the image, a rerun scored 1.0 without any installation commands.

As I mentioned earlier, we want to minimise rollout time, so we should preinstall known dependencies rather than have the agent install them during each rollout.

I also adjusted the memory or disk requirements for eight tasks and configured their sandbox resources accordingly.

Full eval runs

Once the images were ready, I started doing full evaluation runs with different models and harnesses to verify my implementation. This is when you start realising how important reliable infrastructure is and how fast the money burns when you use models via API (Not that I didn't have this realisation earlier but we all have been running the hedonic treadmill thanks to Tibo's resets).

A full eval run with a model like GPT-5.6 Sol could take 6-12 hours, including retries. Some individual tasks ran for hours, and sandbox or inference provider failures meant retrying affected rollouts, making the feedback loop quite slow.

Occasionally, I ran into more specific issues. During my Kimi runs, failures in the ACP integration caused many rollouts to end midway through a task. This was total suffering.

Running these full evaluations also surfaced infrastructure issues, which I reported to the Prime team. Verifiers v1 had just been released, and these reports helped the team identify and address issues in the new release. I also dabbled a bit in the sandbox infra code to see if I could use a “separate sandbox judging runtime” for my benchmark. Some ALE tasks produced output artifacts larger than 1 GB, but transferring artifacts from the rollout sandbox to the judge sandbox was limited to 32 MB at the time. I found out the bottleneck was the sandbox SDK’s lack of support for chunked file streaming.

These runs also helped uncover issues in the benchmark itself. Florian also encouraged me to "look at the data". A closer look revealed more issues. Note: look at the data at your own risk. You never know what crimes you will uncover.

The next section is all about my suffering and the issues we found in the ALE Linux CLI subset and the fixes. 9 tasks had to be excluded because a fix wasn't possible from our side for them. The resulting subset, ALE-Gold, contains 96 of the original 105 tasks, spanning 40 subdomains.

Look at the data, they said

Before we’re too hard on the creators, I should mention that frontier benchmarks are hard to build. These tasks span specialized domains, involve complicated dependencies, and need verifiers that can distinguish a wrong answer from a valid solution the author didn’t anticipate. Getting all of this right takes a lot of care, both when building the tasks and when validating them.

With that context, here are some of the issues I found while inspecting the tasks and rollout traces.

Underspecified task contracts and evaluator defects

I observed agents were failing 5-6 tasks despite producing correct solutions because the task prompt / input didn’t specify the expected output format correctly. To be more specific, essential schemas, enums, or semantics were absent or ambiguous in the visible prompt. Below are a few examples.

>"frontier benchmark"
>rollout has partial score
>verifier is deterministic scorer
>look inside
>verifier expects output format with field names agent can't infer from task prompt even by hallucinating pic.twitter.com/qGpDddm1Mn

sankalp (@dejavucoder) August 18, 2026

Five points for a different spelling

education_info/moodle_gradebook_closeout_reconciliation

The Moodle task asked the agent to reconcile a course gradebook against a grading policy. One decision here was how to handle the missing work. The task specification was clear about the behavior, but it did not specify the accepted values for empty_grade_behavior.

From the grading policy

“Missing work counts as zero. Excused work is excluded from the denominator.”

The agent wrote"count_as_zero"
The grader expected"zero"
0.95Four Sol rollouts. The same five-point mismatch.
See the grader check and the one-sentence fix
From the grader
if policy.get("empty_grade_behavior") == expected_policy["empty_grade_behavior"]:
    policy_points += 5

The expected value was "zero". The check required an exact match.

Added to the task requirements

In gradebook/policy.json, valid empty_grade_behavior values are exclude and zero; select the value that implements the published grading policy.

Across four rollouts, Sol chose "count_as_zero" and scored 0.95 despite otherwise producing the correct deliverables. I added one sentence listing the accepted values—"exclude" and "zero"—and a subsequent rollout scored 1.0.

Eight deliverables, no schemas

health_medicine/healthcare_sap_group_sequential_nsclc

The clinical-trial task had a larger version of the same problem. It asked for a statistical analysis plan, a reproducible R script, and six other output files. The instructions described the statistical work and named the files, but left out the JSON keys and CSV headers the grader expected.

From the task instructions

“Produce three-look O'Brien-Fleming efficacy boundaries at 50%, 75%, and 100% information.”

Agent’s CSVGrader’s expected CSV
information_fractioninfo_fraction
efficacy_zefficacy_z_boundary
futility_zfutility_z_boundary
Read the task instructions
task_instructions.md
# NSCLC Group-Sequential Statistical Analysis Plan

You are given protocol.json for a two-arm Phase III advanced NSCLC trial. Create a complete statistical-analysis-plan deliverable set under output/.

Required files:

  • SAP.md
  • analysis.R
  • sample_size.json
  • boundaries.csv
  • multiple_testing.json
  • power_curve.csv
  • power_curve.png
  • boundary_plot.png

Computational requirements:

  1. Compute log-rank/Schoenfeld sample size for overall survival using the protocol hazard ratio, one-sided alpha, power, and target events.
  2. Produce three-look O'Brien-Fleming efficacy boundaries at 50%, 75%, and 100% information.
  3. Include binding futility only at look 1 using conditional power < 0.20; looks 2 and 3 should record futility as NA.
  4. Report the sample-size inflation factor and adjusted total sample size.
  5. Apply the gated Hochberg procedure to PFS, ORR, and DOR using familywise alpha 0.025, and report subgroup Bonferroni alpha.
  6. Produce a power curve CSV and PNG over plausible hazard ratios including the target HR.
  7. Produce a boundary plot PNG.
  8. Write a full SAP narrative explaining the design assumptions, interim looks, multiplicity, and deliverables.

Rules:

  • Work from protocol.json; do not use internet access.
  • analysis.R should be a reproducible script for generating the outputs.
  • Numeric outputs should be rounded clearly but retain enough precision for verification.
See the sample-size grader excerpt

The same problem appeared in sample_size.json: the grader looked for keys such as per_arm_n and adjusted_total_n, which the task instructions did not specify.

    try:
      sample_size = _load_json(output_files["sample_size.json"])
      pred_n = _to_float(sample_size.get("per_arm_n"))
      ref_n = _to_float(ref["sample_size"]["per_arm_n"])
      if abs(pred_n - ref_n) / ref_n <= 0.03:
          checks["sample_size"] += 0.65
      for key in ["events_required", "total_n", "adjusted_total_n"]:
          if int(round(_to_float(sample_size.get(key)))) == int(ref["sample_size"][key]):
              checks["sample_size"] += 0.08
      inflation = _to_float(sample_size.get("inflation_factor"))
      if abs(inflation - _to_float(ref["sample_size"]["inflation_factor"])) <= 0.01:
          checks["sample_size"] += 0.11
      checks["sample_size"] = min(1.0, checks["sample_size"])
  except Exception as exc:
      issues.append(f"sample_size.json invalid: {exc}")

In one rollout, Sol used information_fraction, efficacy_z, and futility_z in its boundary table. The grader looked for the names on the right. These were reasonable names for the requested quantities, but the task had never told the agent which names it needed to use.

“Structure is flexible”

health_medicine/healthcare_tcga_luad_survival_kras

The survival-analysis task asked the agent to save its proportional-hazards test results in cox_results.json. For the ph_test section, the instructions explicitly allowed a flexible structure.

From the output requirements

“ph_test: structure is flexible”

Sol wroteph_test.tests
A supported layoutph_test.per_variable
0.84The grader treated the tests as missing and capped the score.
See what “flexible” meant to the grader

The grader searched a per_variable object, a table array, or variable entries directly under ph_test. It did not search inside tests.

From the grader
if isinstance(section.get("per_variable"), dict):
    entries.extend(section["per_variable"].items())
if isinstance(section.get("table"), list):
    entries.extend(
        (item.get("variable"), item)
        for item in section["table"]
        if isinstance(item, dict)
    )
if not entries:
    entries.extend(
        (key, value)
        for key, value in section.items()
        if isinstance(value, dict) and _normalize_variable_name(key)
    )
The revised requirement

ph_test must contain a per_variable object keyed by kras_group, age_at_diagnosis_years, and stage_group; each entry must include chi_square and p_value.

Sol put all three test results under ph_test.tests, but the grader did not recognize that layout. I replaced “flexible” with a sentence specifying the expected structure to get a full score rollout.

The key was specified without the value

computing_math/data_pipeline_etl_instance_1

The retail ETL task required a JSON report with a field called schema_drift_columns_filled, but never defined its value. Sol reported how many cells it had filled in each column. The grader expected an ordered list of column names.

Required field

schema_drift_columns_filled

The agent reported fill counts
{
  "discount_pct": 1969,
  "channel": 3883,
  "total_cells": 5852
}
The grader wanted column names
["discount_pct", "channel"]
0.833333 → 1.0One validation rollout after the contract patch.
Show the other format requirements and the fix

The contract also left out the types for completion flags and dates. The grader required Booleans and integer dates even though the agent had supplied counts and ISO date strings.

Agent suppliedGrader required
"timestamps_standardized": 5914"timestamps_standardized": true
"min": "2024-01-01""min": 20240101
From the grader
if transformations.get("timestamps_standardized") is not True:
    return False
if transformations.get("schema_drift_columns_filled") != ["discount_pct", "channel"]:
    return False
Added to the output contract

Standardization flags must be Booleans. schema_drift_columns_filled must be an array of strings in the corresponding column order in target_schema_spec.json. Date-range values must be integers in YYYYMMDD form.

Two original runs scored 0.833333. I added explicit value types and an ordering rule to the visible contract, without supplying the actual counts or dates. A subsequent validation rollout scored 1.0, although later inspection found that the ordering rule still conflicted with the scorer’s expected list order.

I was able to fix some of the tasks by minimally modifying the input on the images while others had to be kept out.

Answer Leakage and network issues

ALE does not restrict network as such. It’s a generally online bench and one of the authors mentioned “we keep it online so agent can sometimes fix its dependencies”.

Some tasks require the agent to look up things online or query APIs and these often came with URLs mentioned in the allow-lists. A few tasks stated network restrictions in their prompts.

Examples:

  1. physical_sciences/lenacapavir_sar_table2_extraction says:

    “Use the staged paper PDF as your source of truth.”
    “Do not rely on web search or external answer sources.”

  2. engineering/chisel_verilog_alignment_seq_1 states:

    “The VM has no internet access during solve time.”

The metadata for network policy (which tasks run online or offline) was not public at the time. Upstream later published a list of 5 tasks that must run offline.

I did 2-3 runs with GPT-5.6 Sol (Medium) without network isolation and found a few tasks that looked up solutions online. These tasks are made to run with network isolation now.

My personal favourite examples of agents looking up stuff on the internet were:

  1. computing_math/go_game_reconstruction_1

    The goal was to reconstruct a Go game from the evidence staged in the sandbox. Instead, the agent realised that the target may be an identifiable public professional game.

    The prompt mentioned “Do not use a Go engine or a web browser.” so the agent proceeded to use curl and download an archive of more than 90,000 professional SGF games, identified the exact historical game, and recovered its 168 moves.

    Screenshot of one of my favourite benchmark cases

    No Go engine or web browser? The agent used curl to find the game in a public archive instead.

  2. health_medicine/wsi_tumor_localization_1

    The task asks for the center point of a tumor in a whole-slide pathology image. The agent started by doing the actual image analysis, zooming into suspicious tissue patches.

    The input had retained its public filename, tumor_001.tif, which probably motivated the agent to look it up online, so the agent searched for the corresponding annotation, tumor_001.xml. It tried GitHub, grep.app, Hugging Face, Google, Bing, DuckDuckGo, the original CAMELYON16 challenge website, and finally Kaggle.

Broken inputs and missing dependencies

Here are a couple of examples.

In computing_math/synthetic_causal_structure_inference, the 40 datasets were arranged in blocks of five, matching the order of the eight scenario families listed in the prompt. The first five were “observed confounder,” the next five “backdoor criterion,” and so on. One agent used SCENARIOS[index // 5] to recover the scenario from the dataset’s position. This also let it derive several related answers, together worth 70% of the rubric, without inferring them from the data.

In engineering/humanoid_wbc_policy_evaluation, the supplied environment depended on a specific MuJoCo development build, 3.7.0.dev886334046. Its download URL 404’ed. The task archive contained the project and its lockfile, but no copy of the required package, so even downloading the complete archive wasn’t enough to recreate the environment. I excluded the task because I couldn’t get the required runtime working from the supplied files.

Missing or corrupt reference data

For some tasks, the reference data was missing or corrupt. This manifested as bad ground truth data, references being questionable or in 1-2 cases as scientifically questionable.

Some of these were fixed upstream. The unresolved ones had to be excluded from the subset.

I reported these issues to the ALE authors on Discord and they were fast to respond. They fixed the reference data for a few tasks in the upstream code and published the list of offline-only tasks (PR 1, commit).

The mitigations mentioned above in the benchmark issues section led us to the formation of ALE-Gold. “Gold” just means tasks I believe are not bad and measure capability well. Nine tasks had to be excluded, leaving us with 96 tasks.

One detail worth mentioning is that I looked at the inputs and traces of problematic tasks and decided on a case-by-case basis what to include. In some cases, the tasks were badly designed or had some issue, but were still included because they seemed to do fine across multiple rollouts.

9 excluded tasks

Swipe to explore
Contract mismatch01 / 09
BPMN governance

The scorer rejects the documented schema and requires undisclosed coordination tasks. Either mismatch can force a zero.

Business & finance

Scorer crash02 / 09
Ranking feature recovery

Scoring crashes after task completion: the evaluator treats a string payload as a mapping.

Computing & mathematics

Answer leakage03 / 09
Causal structure inference

Dataset order reveals scenario families and answers worth 70% of the rubric, bypassing much of the intended inference.

Computing & mathematics

Reference mismatch04 / 09
Chisel–Verilog alignment

The hidden source-location reference contradicts the supplied comparison evidence, putting 80% of credit at risk.

Engineering

Missing dependency05 / 09
Humanoid policy evaluation

A required MuJoCo development wheel is no longer downloadable. The archived payload lacks the replacement wheelhouse.

Engineering

Scorer crash06 / 09
OpenROAD chip signoff

A shell PID-parsing bug crashes scoring before the verifier starts, returning zero even for completed submissions.

Engineering

Unenforced constraint07 / 09
Healthcare bias reproduction

The output-only scorer accepts a direct race bonus prohibited by the prompt; a controlled probe scored 0.861.

Health & medicine

Scoring timeout08 / 09
Prostate radiotherapy planning

The matRad scorer repeatedly exceeds its declared one-hour scoring timeout, preventing reliable evaluation.

Health & medicine

Questionable reference09 / 09
Variant annotation pipeline

Hidden truth appears to mishandle deletion normalization and allele-specific ClinVar, rewarding reference-matching errors.

Health & medicine

Reasons for exclusion refer to the task versions used in this evaluation.

ALE-Gold evaluation results

I ran full evaluations for four models that supported both text and vision. I passed on Opus 5 because of its higher cost and GLM 5.3 because it didn’t support vision.

The ALE CLI taskset used for these evaluations is public on the Prime Environments Hub.

These results are based on the 96-task ALE-Gold subset, with one reported result per model per task. Sol and Luna used xhigh reasoning, while Kimi K3 and GLM 5.3 Flash used max reasoning. Sol, Luna, and GLM used Codex; Kimi used Kimi Code. Rollouts that ended with execution errors were retried, with up to three retries per task.

Aggregate score is the mean task score on a 0–1 scale, including partial credit, with missing scores counted as zero. Pass rate counts only recorded scores exactly equal to 1.0.

ALE-Gold results

Model / configurationAggregate scorePass rate
GPT-5.6 Sol
Xhigh reasoning · Codex
0.6359
32.29%
GPT-5.6 Luna
Xhigh reasoning · Codex
0.5672
27.08%
Kimi K3
Max reasoning · Kimi Code
0.4686
26.04%
GLM 5.3 Flash
Max reasoning · Codex
0.4071
19.79%

Kimi K3 and GLM 5.3 Flash had several rollouts ending mid-turn / partial tool call execution. For Kimi K3, it was an ACP server issue on their side whereas for GLM 5.3 Flash, the issue seemed to be more due to a skill issue with the model. These were retried separately.

Sol (xhigh) vs Luna (xhigh)

Sol vs LunaSol xhighLuna xhigh
Mean score63.59%56.72%
Pass rate32.29%27.08%
Completion tokens2.24M3.03M
Model calls5,0746,691
Recorded cost$139.61$19.19

96 shared tasks. Completion tokens include reasoning tokens. Usage and cost cover the saved traces; earlier retried attempts may not be fully included. Runtime infrastructure is excluded.

Sol scored higher overall, with a mean score of 63.59% against Luna’s 56.72%. It also had the higher pass rate: 32.29% versus 27.08%, corresponding to 31 and 26 full-score results. Across the 96 tasks, Sol scored higher on 24, Luna on 17, and 55 were tied.

Luna used 35% more completion tokens, including reasoning, and made about 32% more model calls. Its recorded model-call cost was 86% lower: $19.19 versus $139.61 for Sol.

The four examples below show different reasons behind Luna’s advantages, from matching a reference verdict to following a numerical specification.

Skull-stripping quality control

Inspect three candidate brain masks in FSLeyes, select the best one, and judge whether it is acceptable.

Luna 1.000Sol 0.000

Both models inspected the candidates and selected the same mask. Luna marked it acceptable, while Sol rejected it. The evaluator required an exact match to the reference mask and verdict, so Luna received full credit and Sol received zero.

Luna’s explanation also mixed up removing brain tissue with leaving non-brain tissue behind, but the evaluator did not check the explanation.

Cell tracking

Segment cells across 30 microscopy frames and track their identities and lineage over time.

Luna 1.000Sol 0.794

Both models produced all 30 segmentation masks and a lineage table. Luna met both the segmentation and tracking thresholds. Sol checked mask connectivity and lineage consistency, but its result still fell below at least one quality threshold.

The evaluation returned only the combined score, so we cannot tell which component accounted for Sol’s lower result. Luna also noticed that some masks looked wrong during its final visual check, but submitted them without further correction.

Fire-detector reconstruction

Reconstruct a room-fire simulation and predict detector responses from supplied heat-release observations.

Luna 0.887Sol 0.626

The task defined ignition as the first positive heat-release observation, at 59.25 seconds. Luna followed that rule. Sol instead moved ignition back to the preceding zero reading, at 18 seconds, reasoning that the grader probably expected the inferred start rather than the literal instruction. This shifted its downstream timing predictions.

Sol’s checks validated the interpretation it had chosen, so they did not catch the departure from the task.

Particle filtering

Track a moving object from noisy observations using a particle filter, across three difficulty levels.

Luna 0.500Sol 0.000

Both models implemented working filters, but calculated a required error measure differently. Luna averaged the distance errors at each time step, matching the evaluator. Sol used the root-mean-square of those distances, then checked its result using the same formula. That mismatch failed a mandatory scoring check and reduced Sol’s score to zero.

Luna earned partial credit, but its hardest level failed because its seeded data-generation sequence differed from the evaluator’s. Both models also approximated a smoothing method that the task required them to implement exactly.

Selected task comparisons

Aggregate scores hide substantial differences between tasks. The five examples below compare all four models across financial analysis, scientific computing, and software tasks. Some show a clear score advantage; others show how similar scores can come from different approaches.

5 task comparisons · Swipe or scroll horizontally to explore →

Business & finance

Reconstruct the Fama–French factors

Rebuild monthly U.S. market, size, value, profitability, and investment factors from public data, covering January 2015 through February 2026.

Sol xhigh0.938
Luna xhigh0.000
Kimi K30.000
GLM Flash0.000

Sol scored higher after more detailed checks of share counts, stock splits, and units. However, it substituted an ETF return series for the market factor, and neither Sol nor Luna obtained the required full paper before starting.

Luna corrected a unit mismatch but retained a fallback to current share counts, weakening its historical market weights. Kimi used narrower data coverage.

Computing & mathematics

Reassemble a sharded checkpoint

Merge model-parallel checkpoint shards into one model file, preserving tensor layouts and matching the supplied reference model’s outputs.

Sol xhigh1.000
Luna xhigh1.000
Kimi K30.000
GLM Flash0.000

Sol and Luna both reconstructed the checkpoint successfully, recovering the unusual tensor layouts and matching the reference outputs. Sol used fewer calls and performed a stronger final check of the saved file; Luna explored more layout variants.

Kimi was still debugging weight ordering when its trace ended. GLM produced a loadable checkpoint with large numerical errors, showing why successfully loading a model is not enough to verify its reconstruction.

Health & medicine

Estimate individual treatment effects

Build a solver that predicts each person’s outcome with and without treatment on the IHDP benchmark, then reports their estimated treatment effect.

Sol xhigh0.959
Luna xhigh0.161
Kimi K30.717
GLM Flash0.286

Sol compared several outcome models and achieved the highest score with a nonlinear model. Kimi compared candidate methods and selected separate regressions on log-transformed outcomes for the treatment and control groups.

Luna closely reproduced the supplied linear baseline, producing a working solver but substantially less accurate predictions than Sol.

Computing & mathematics

Design tests that break flawed code

Write a seeded C++ test generator for a binary-string problem. Its inputs must expose wrong answers and slow solutions across 50 test seeds.

Sol xhigh0.400
Luna xhigh0.700
Kimi K30.800
GLM Flash0.300

Kimi checked its reference implementation against 3,000 brute-force cases and targeted several families of flawed solutions. Sol independently checked the underlying mathematics, while Luna repaired an invalid-interval bug and validated all 50 seeds.

Kimi’s tests earned the most credit. Sol and Luna both produced valid generators, but their tests exposed fewer of the required hidden-solution behaviors.

Health & medicine

Calibrate a CT reconstruction

Tune scanner geometry to reconstruct a phantom from measured projections, achieving at least 0.95 SSIM and the required pixel-error threshold.

Sol xhigh1.000
Luna xhigh1.000
Kimi K30.000
GLM Flash1.000

Sol and Luna generated real reconstructions from the measured projections, debugging geometry and API errors rather than accepting a passing pixel-error check while image similarity still failed. Luna also regenerated its final image from the saved parameter file.

GLM optimized the geometry and checked the saved reconstruction with the intended metric, reaching 0.993 similarity against a 0.95 target. Kimi remained below the required threshold. Sol, Luna, and GLM all received full credit.

Scores include partial credit. Sol and Luna used xhigh reasoning; Kimi K3 and GLM 5.3 Flash used max reasoning. Sol, Luna, and GLM used Codex; Kimi used Kimi Code.

Failure modes

The dominant failure pattern was a skill issue across all models, including Sol and Luna, but it was most pronounced in GLM 5.3 Flash. Models sometimes used an incorrect calculation or implemented the required procedure incorrectly. They also replaced difficult parts of a workflow with simpler approximations without establishing that those approximations met the requirements. The resulting outputs could look plausible and pass basic checks while still drifting away from the requested method.

Verification could repeat the same assumptions as the implementation. Models checked that files existed, outputs had the right shape, or rerunning the code produced the same result, but these checks did not necessarily establish correctness.

In other cases, checks exposed a problem but the model submitted the result anyway. Luna noticed questionable masks in the cell-tracking task but left them unrepaired. GLM produced a loadable checkpoint in checkpoint consolidation despite large numerical errors. Detecting a problem did not always lead to correcting the final output. I believe this often happens when the model is not confident about what it is doing.

Models sometimes produced outputs that did not match the schema or format expected by the scorer, losing partial credit. These mismatches involved details such as column ordering, field names, enum/schema, or file structure.

Model-specific failure modes

More model-specific failure patterns also stood out.

Sol has a tendency to attempt an approximate version of a task when it finds the work difficult. It also sometimes changed its approach based on what it thought the evaluator expected, even when that meant departing from the task’s explicit requirements. Output mismatches with the scorer’s expected schema or format also led to lost credit, although it was difficult to determine their full extent because the direct causes of some score deductions remained unclear.

Luna persisted with weak approaches and sometimes submitted results despite checks exposing unresolved problems. In some tasks, it focused on satisfying the evaluator’s checks without adequately establishing that the underlying work met the requirements.

Kimi K3 substituted shortcuts for the requested workflow without adequately checking that they met the original requirements.

GLM 5.3 Flash often made implementation errors or used the wrong formulas. My interpretation is that these mistakes reflected the capability and knowledge limitations of a smaller model.

Performance by domain

Across the five domains with at least 10 tasks, the four-model average is highest in computing and mathematics, followed by business and finance, then physical sciences. Sol leads in business and finance, physical sciences, life sciences, and health and medicine; Luna leads in computing and mathematics. Sol’s advantages are larger in health and medicine, life sciences, and physical sciences, while they both wiped the floor in computing and mathematics. Among the smaller domains, Luna leads legal, while Kimi K3 leads education and information and narrowly leads engineering. GLM has individual task wins but no outright domain-mean lead. These smaller domains contain between one and five tasks, so their averages are sensitive to individual results.

Performance by domain

Mean task score01
DomainSol xhighLuna xhighKimi K3GLM 5.3 Flash
Health & medicine
19 tasks
0.618
0.513
0.423
0.410
Life sciences
19 tasks
0.665
0.549
0.450
0.426
Computing & mathematics
17 tasks
0.654
0.666
0.612
0.357
Business & finance
13 tasks
0.667
0.619
0.510
0.477
Physical sciences
11 tasks
0.721
0.621
0.498
0.333
Engineering
5 tasks
0.206
0.207
0.207
0.207
Education & information
3 tasks
0.632
0.624
0.634
0.504
Transport & safety
3 tasks
0.875
0.629
0.534
0.541
Legal
2 tasks
0.679
0.710
0.665
0.667
Psychology & neuroscience
2 tasks
0.513
0.430
0.003
0.381
Other
1 task
0.000
0.000
0.000
0.000
Social sciences
1 task
1.000
1.000
0.000
1.000

Engineering has room for improvement

Among domains with at least five tasks, engineering has the lowest scores across all four models. Their mean scores are tightly clustered around 0.21. All four receive full credit on robot-description reconstruction, very little credit on power-feeder reliability, and zero on low-thrust trajectory design, building model-predictive control, and urban traffic calibration.

Engineering

Highest model mean · bars run from 0 to 1. Open the panel to explore the results.

Engineering0.207 5 tasks · Best: Kimi K33 shared zeros · Explore
Sol xhigh0.206
Luna xhigh0.207
Kimi K30.207
GLM Flash0.207

Scored zero for all four models

  • Low-thrust trajectory design
  • Building model-predictive control
  • Urban traffic calibration

The traces show substantive method and implementation problems. In low-thrust trajectory design, Sol and Luna generated real trajectories but replaced the required state/costate shooting method with other approaches while still reporting successful shooting. In building control, Luna substituted thermostat control for the required discrete airflow actions, while GLM substituted scheduled actions for the requested model-predictive policy.

However, these scores also reflect problems with the tasks and their evaluation. The low-thrust task contains conflicting thrust, duration, and velocity-change requirements. The building-control scorer imposes output conventions that are not fully specified in the instructions. Urban traffic calibration awards credit only when all 12 checks pass, so a zero can hide substantial progress on the public calibration targets. Engineering offers room for improvement in both model performance and evaluation quality.

Conclusion

Once you start looking into benchmarks, you will realise that they are not perfect. Often, they are not measuring what they claim to measure, are fully saturated, and/or are unable to differentiate between the capabilities of models. This post was my attempt to highlight some of these issues.

Terminal-Bench 2.1 is a good example: it is no longer very useful for distinguishing the capabilities of leading models, while TB 4.0 still reveals substantial differences. SWE-2 scored 92.8% on TB 2.1, ahead of the other models in the comparison, but only 27.3% on TB 4.0, compared with 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra.

lmao lower TerminalBench 4.0 and DeepSWE than DeepSeek V4.1

— Teortaxes (@teortaxesTex) September 10, 2026

That said, Agents’ Last Exam is not yet saturated (and note that we looked at only its public Linux CLI subset). I expect many domains of ALE to be rapidly hillclimbed in the next six months.

This exploration was an overall great learning experience for me. I hope you got to learn something new too! Thanks for reading! Please share/upvote if you liked it.

Acknowledgements

Xeophon (Florian Brand), for being my peer, discussing ideas with me, and guiding me whenever I was stuck.

Snimu (Sebastian Muller), for coordinating the residency and providing the compute and guidance whenever I requested them.

Thanks to Prime Intellect for the Prime Residency program. Thanks also to the Prime Intellect team for building such great infrastructure. This work would not have been possible without their infrastructure.

#AI #featured