Vibe Code Textbook

Chapter 28 of 28, in Part V, Measurements

SWE-bench Verified task by task: 11 models, one harness, measured

The leaderboard's own per-task records for 11 models on one harness: how task size predicted failure, what a call cap keeps and saves, and what a resolve cost.

Added . 2,568 words, about 12 minutes.

Reviewed by C. B. Zakarian on

Eleven models worked through the same 500 SWE-bench Verified tasks with the same agent, mini-swe-agent 2.0.0, and the leaderboard kept a line for each of the 5,500 attempts: resolved or not, how many model calls it took, and what it cost. Read one task at a time, those lines show three things a score cannot. The size of a task, as human annotators estimated it, moved the resolve rate from 84.7% to 33.1%, while on tasks of the same rated size the best and worst of these models were never more than 28.6 points apart (the three tasks rated over four hours aside). A cap on model calls saves little, because a quarter of the failed attempts ended no later than a typical success. And among the four runs within six tasks of the top score, the recorded cost of a resolved task ranged from $0.097 to $0.982.

What ran: one standard-library Python script (Python 3.13.12, Windows) on 8 October 2026. It read the SWE-bench/experiments repository at commit 40f164d5b8f1, the latest on its main branch that day (dated 3 September 2026), and the 500 Verified tasks from the Hugging Face Hub. The final version fetched everything live at 12:54 UTC, and its output matched, byte for byte, a replay of the responses an earlier live run had saved at 12:31 UTC; that replay is printed near the end. No agent, model or test harness ran for this piece. Every outcome below is a record a submitter committed to the leaderboard, and we re-derived none of them.

Eleven runs with the harness held still

The experiments repository "holds entries, not artifacts": each entry is a folder with a metadata file and a results record, while the predictions, logs and trajectories live wherever the submitter keeps them. Its README says the mini-SWE-agent "bash-only" runs "are now regular Verified submissions under evaluation/verified/, and the leaderboard's Bash Only view is a filter on the Verified board rather than a board of its own." Those entries are the nearest thing the leaderboard has to a controlled comparison. The agent's README says it "Does not have any tools other than bash", and each entry swaps in a different model.

At the pinned commit, 13 of the 182 Verified entries name mini-swe-agent 2.0.0. We could not read two of them task by task. GPT 5.2 Codex's entry, dated 19 February 2026, holds a metadata file and no per-task record. Gemini 3 Pro's, dated 26 February, has a record, but it marks all 500 tasks unresolved while the same entry's metadata reports 69.6%. The other 11 are all dated 17 February 2026, and each record agrees exactly with its own metadata on the number resolved, the total cost and the mean calls per task. Every one made a single attempt per task (attempts: 1), and none carries the leaderboard's verified check (checked: null).

The agent's limits are in its SWE-bench configuration at the v2.0.0 tag: step_limit: 250 and cost_limit: 3., which is 250 model calls or three dollars per task. The agent tests both before every model call, with 0 < self.config.step_limit <= self.n_calls or 0 < self.config.cost_limit <= self.cost, and an attempt that reaches either ends with "exit_status": "LimitsExceeded", "submission": "", an empty patch. Cost is whatever litellm.cost_calculator.completion_cost estimates for each response, added up.

Run, as its entry names it Resolved Under 15 min 15 min to 1 h 1 to 4 h Recorded cost Per resolved task Median calls, resolved / not
Claude 4.5 Opus (high) 76.8% 89.2% 74.7% 35.7% $376.95 $0.982 26.5 / 38.5
Gemini 3 Flash (high) 75.8% 87.6% 73.6% 38.1% $177.98 $0.470 52 / 58
MiniMax M2.5 (high) 75.8% 86.6% 74.3% 38.1% $36.64 $0.097 46 / 67
Claude 4.6 Opus 75.6% 88.1% 72.0% 42.9% $275.76 $0.730 21 / 31
GLM 5 (high) 72.8% 85.1% 70.1% 38.1% $267.19 $0.734 60 / 86.5
GPT 5.2 (high) 72.8% 84.5% 70.1% 40.5% $236.78 $0.651 28 / 37.5
Claude 4.5 Sonnet (high) 71.4% 85.1% 67.8% 33.3% $328.95 $0.921 42 / 53
Kimi K2.5 (high) 70.8% 84.5% 67.4% 31.0% $73.28 $0.207 42 / 56
DeepSeek V3.2 (high) 70.0% 82.5% 67.8% 28.6% $223.92 $0.640 79 / 95
Claude 4.5 Haiku (high) 66.6% 82.0% 62.5% 23.8% $165.46 $0.497 55 / 74
GPT 5 mini 56.2% 76.3% 48.3% 14.3% $23.60 $0.084 16 / 20
All 11 runs 71.3% 84.7% 68.1% 33.1% $2,186.53 $0.557 -

The three tasks rated over four hours are left out of the table. Together they were resolved in 9 of 33 attempts, all 9 on the same task.

Task size moved the rate further than the model did

The labels come from OpenAI's annotation of Verified, which asked annotators "to estimate how much time it would take an experienced software engineer who has had a few hours to familiarize themselves with the codebase to write a patch solving the issue." There are four: "<15 min fix (e.g., a trivial change adding some assertions to a function)", "15 min–1 hour (e.g., a small change that requires a bit of thought)", "1–4 hours (e.g., substantially rewriting a function or editing multiple files)" and ">4 hours". The estimates were "supplementary information (not used for dataset filtering)", and the annotators' answers were combined "by taking the majority choice for a sample, or the median if there is no majority." The Verified rows carry the result in a difficulty column: 194 tasks under 15 minutes, 261 from 15 minutes to an hour, 42 from one to four hours and 3 over four.

Across all 5,500 attempts, a task rated under 15 minutes was resolved 84.7% of the time, one rated 15 minutes to an hour 68.1% of the time, and one rated one to four hours 33.1%. The first step down the scale cost 16.6 points and the second cost 35.0. Within a label, the distance between the best and the worst of the 11 runs was 12.9 points on the smallest tasks, 26.4 on the middle ones and 28.6 on the one-to-four-hour ones, and most of that is GPT 5 mini: without it, the other ten runs sit within 7.2, 12.2 and 19.1 points of each other.

Grouped at the hour, the line is plain. On the 455 tasks rated under an hour, every run resolved between 60.2% (GPT 5 mini) and 80.9% (Claude 4.5 Opus (high)). On the 45 rated an hour or more, the best run, Claude 4.6 Opus, resolved 42.2%, and the pooled rate was 32.7%. That makes the annotators' question a usable filter before a task goes to an agent: how long would this take someone who knows the code a little? The records do not show that cutting a large task into small ones earns small-task success rates, only where the line falls. Spec-first prompting covers how to write the smaller task.

OpenAI attached a caution to the labels that applies here: "Note that this may overestimate the difficulty for a LLM, which may have memorized aspects of codebases and PRs." The tasks are public, and the SWE-bench chapter collects the published evidence that Verified has been seen in training.

Where the eleven runs agree

211 tasks were resolved by all 11 runs, and 59 by none of them. Every difference between the scores in the table comes from the other 230. Another 83 tasks were missed by exactly one run.

Rated Tasks Resolved by all 11 By none Split
Under 15 minutes 194 114 9 71
15 minutes to 1 hour 261 93 34 134
1 to 4 hours 42 4 14 24
Over 4 hours 3 0 2 1

The agreement follows the labels: more than half of the smallest tasks were resolved by every run, and a third of the one-to-four-hour tasks by none. Between them, the 11 runs resolved 441 different tasks, against 384 for the best of them, Claude 4.5 Opus (high). Of the 116 tasks that run missed, 57 were resolved by at least one of the other ten. That is the case for a second attempt with a different model when the first one fails, with the caveat that these records show only that some model succeeded, never which one to choose in advance.

A call cap saves less than it looks

The stopping chapter quotes the FailFast paper's observation that "Failed runs tend to be longer and exhibit redundant exploration or looping". These records agree on the direction and put a size on it: in every run, the median unresolved attempt took more calls than the median resolved one, by 1.12 to 1.48 times.

A cap can be tested on the records after the fact. mini-swe-agent checks its limits before each model call, its SWE-bench prompt at v2.0.0 never mentions them, and an attempt stopped by one submits nothing. So under a cap of k calls, an attempt that resolved its task within k calls would have gone exactly as recorded, and any longer attempt would have ended unresolved after k calls.

Cap on model calls Resolves kept, of 3,923 Calls saved
25 24.1% 54.3%
50 62.5% 25.0%
75 85.4% 10.7%
100 94.9% 4.7%
150 99.4% 1.1%
2 × each run's median resolved attempt 94.1% 4.6%
3 × each run's median resolved attempt 99.1% 1.1%

A cap that keeps 95 resolves in 100 saves under 5% of the calls, and one that saves a quarter of the calls gives up 37.5% of the resolves. Calls are not dollars: the records hold one cost per attempt, so a cap's dollar saving cannot be computed from them. It would probably be somewhat larger than the call saving, because the agent has, in its README's words, "a completely linear history" and sends all of it with every call (self.model.query(self.messages)), so the late calls carry the most context.

The reason the cap does so little is in the failures. 398 of the 1,577 unresolved attempts, 25.2%, used no more calls than their own run's median resolved attempt. Whatever went wrong in them did not show up as length. The configured limits almost never fired: 31 of the 5,500 attempts reached 250 calls or $3.00, and none of the 31 resolved. Failure was still the expensive part. Unresolved attempts were 28.7% of all attempts and took 38.3% of the recorded spend, between 27.6% and 54.4% depending on the run.

FailFast reports that a trained monitor reading the trajectory so far saved "14.6%-20.4% of execution tokens at a target 5% false-positive rate" on its own runs. Those are different runs and a different measure, so the two results do not compare, but the contrast is the useful part: whatever separates a failing attempt from a succeeding one, in these records it is mostly not length. In practice that makes a call cap insurance against the runaway attempt, set at two to three times the length of your own typical successful run, and it leaves the ordinary failure to the checks at the end: a test that fails first and a review of the diff.

What a resolved task cost

The recorded cost of a resolved task, a run's total spend divided by the tasks it resolved, ran from $0.084 for GPT 5 mini, which resolved 281, to $0.982 for Claude 4.5 Opus (high), which resolved 384. Among the four runs within six tasks of the top score, it ran from $0.097 for MiniMax M2.5 (high) to $0.982, about ten times as much for five more resolved tasks. Over all 11 runs it was $0.557.

Those are estimates, not invoices: litellm's price for each response, from the price table in use at the time of the run. Four of the eleven entries mark their model as open-weight (os_model: true): DeepSeek V3.2, GLM 5, Kimi K2.5 and MiniMax M2.5. The records name the model and not the provider that served it, so those prices are whatever litellm listed for the model name. A price table is a snapshot, which is why what one task costs counts tokens and calls first and prices them on the day.

Twenty-six tasks no entry has resolved

Across all 175 Verified entries with a per-task record, from the retrieval baselines of October 2023 to an entry dated 1 September 2026, 474 of the 500 tasks have been resolved at least once, 10 by exactly one entry, and 26 never. The most any one entry resolved is 396, a tie between two December 2025 entries, 20251205_sonar-foundation-agent_claude-opus-4-5 and 20251215_livesweagent_claude-opus-4-5.

The 26 are not all large tasks. 14 carry the 15-minutes-to-an-hour label and 2 the under-15-minutes label (django__django-10999 and matplotlib__matplotlib-25479); 9 are rated one to four hours and 1 over four. Django contributes 13, about its share of the set (231 of 500), while pylint contributes 3 of its 10. Of the 59 tasks none of the 11 runs resolved, 33 have been resolved by some other entry, so most of what defeated all eleven has been done by some system.

A task that nobody resolves is either hard or graded by tests that reject a correct fix. OpenAI's annotation screened for the second ("We check if the FAIL_TO_PASS tests might fail even when a valid solution is provided"), but the Verified rows carry only the time label, so these records cannot tell the two apart.

Run it yourself

The listing in the Code and data box is standard-library Python. Without --cache it fetches everything live; with it, it reads the responses saved in that folder and saves any it has to fetch. This is the run behind every number above, verbatim:

$ python swe-bench-verified-task-by-task.py --cache saved --csv swe-bench-verified-task-by-task.csv
SWE-bench/experiments at 40f164d5b8f1: 182 Verified entries, 175 with a per-task record
SWE-bench Verified: 500 tasks; difficulty labels: <15 min fix 194, 15 min - 1 hour 261, 1-4 hours 42, >4 hours 3

mini-swe-agent 2.0.0 entries: 13
  without a per-task record: 1 (20260219_mini-v2.0.0_gpt-5-2-codex)
  record incomplete or at odds with its own metadata: 1 (20260226_mini-v2.0.0_gemini-3-pro-high (0 resolved in the record, 69.6% in its metadata))
  used: 11 runs x 500 tasks = 5,500 attempts

run                       resolved  <15 min   15m-1h    1-4 h     >4 h   cost $  $/resolve  median calls, resolved / not
Claude 4.5 Opus (high)     384  76.8%    89.2%    74.7%    35.7%    33.3%   376.95      0.982    26.5 / 38.5
Gemini 3 Flash (high)      379  75.8%    87.6%    73.6%    38.1%    33.3%   177.98      0.470      52 / 58
MiniMax M2.5 (high)        379  75.8%    86.6%    74.3%    38.1%    33.3%    36.64      0.097      46 / 67
Claude 4.6 Opus            378  75.6%    88.1%    72.0%    42.9%    33.3%   275.76      0.730      21 / 31
GLM 5 (high)               364  72.8%    85.1%    70.1%    38.1%     0.0%   267.19      0.734      60 / 86.5
GPT 5.2 (high)             364  72.8%    84.5%    70.1%    40.5%     0.0%   236.78      0.651      28 / 37.5
Claude 4.5 Sonnet (high)   357  71.4%    85.1%    67.8%    33.3%    33.3%   328.95      0.921      42 / 53
Kimi K2.5 (high)           354  70.8%    84.5%    67.4%    31.0%    33.3%    73.28      0.207      42 / 56
DeepSeek V3.2 (high)       350  70.0%    82.5%    67.8%    28.6%    33.3%   223.92      0.640      79 / 95
Claude 4.5 Haiku (high)    333  66.6%    82.0%    62.5%    23.8%    33.3%   165.46      0.497      55 / 74
GPT 5 mini                 281  56.2%    76.3%    48.3%    14.3%    33.3%    23.60      0.084      16 / 20
all 11 runs               3923  71.3%    84.7%    68.1%    33.1%    27.3%  2186.53      0.557

Tasks by how many of the 11 runs resolved them
  resolved by:    0    1    2    3    4    5    6    7    8    9   10   11
  tasks:         59   22   15   13   17   11    6   10   25   28   83  211
  <15 min fix      194 tasks: all 11 runs 114, none   9, some  71
  15 min - 1 hour  261 tasks: all 11 runs  93, none  34, some 134
  1-4 hours         42 tasks: all 11 runs   4, none  14, some  24
  >4 hours           3 tasks: all 11 runs   0, none   2, some   1
  resolved by at least one run: 441; by the best run alone: 384

A cap on model calls, applied to the recorded attempts
  cap                              resolves kept   calls saved   kept, lowest and highest run
  25 calls                         944   24.1%       54.3%     0.0% to 84.3%
  50 calls                       2,453   62.5%       25.0%     7.7% to 99.6%
  75 calls                       3,349   85.4%       10.7%     44.6% to 100.0%
  100 calls                      3,724   94.9%        4.7%     77.7% to 100.0%
  150 calls                      3,898   99.4%        1.1%     97.3% to 100.0%
  1.5 x the run's median resolve 3,249   82.8%       10.4%     74.3% to 90.3%
  2 x the run's median resolve   3,690   94.1%        4.6%     87.0% to 98.9%
  3 x the run's median resolve   3,887   99.1%        1.1%     96.0% to 100.0%
  median calls, unresolved attempt / resolved attempt, by run: 1.12x to 1.48x
  unresolved attempts no longer than their run's median resolved attempt: 398 of 1,577 25.2%
  attempts that reached 250 calls or $3.00 (the limits in mini-swe-agent 2.0.0's SWE-bench config): 31, resolved 0
  share of recorded cost spent on unresolved tasks: 27.6% to 54.4% by run, 38.3% over all attempts

All 175 entries with a per-task record
  most resolved by one entry: 396 (20251205_sonar-foundation-agent_claude-opus-4-5, 20251215_livesweagent_claude-opus-4-5)
  tasks resolved by at least one entry: 474; by exactly one: 10; by none: 26
  never resolved, by difficulty: <15 min fix 2, 15 min - 1 hour 14, 1-4 hours 9, >4 hours 1
  never resolved, by repository: django/django 13, pylint-dev/pylint 3, sympy/sympy 3, matplotlib/matplotlib 2, pydata/xarray 2, sphinx-doc/sphinx 2, astropy/astropy 1
  never resolved, <15 min fix: django__django-10999, matplotlib__matplotlib-25479

wrote swe-bench-verified-task-by-task.csv (500 rows)

--harness 1.17.2 holds a different mini-swe-agent version still (five entries from December 2025 name it), --caps takes other caps, and --sha reads another commit of the experiments repository. The commit is pinned in the script, so a live run reads the same records until --sha names a newer one. The task rows are not pinned: the rows API takes no revision, and on 8 October its rows matched the dataset's revision 78f471bf655a.

The dataset

_README  swe-bench-verified-task-by-task.csv  (500 rows, one per task)
instance_id          the SWE-bench Verified task
repo                 its GitHub repository
difficulty           the annotators' time label: <15 min fix | 15 min - 1 hour | 1-4 hours | >4 hours
resolved_by_runs     how many of the 11 mini-swe-agent 2.0.0 runs resolved it (0 to 11)
runs                 11
resolved_by_entries  how many of the 175 Verified entries with a per-task record resolved it
entries              175
median_calls         the median number of model calls across its 11 attempts
median_cost_usd      the median recorded cost across its 11 attempts

Neither the experiments repository nor the dataset card states a licence, so the file holds counts we computed, one row per task, and not a copy of the records; the script reads the records themselves. Sorting it by resolved_by_runs and then median_calls gives a measured difficulty ranking to set beside the annotators' one.

Limits

Eleven runs from one day, one agent and one attempt each. A second attempt by the same model would not reproduce every result, and these records cannot say how often it would differ. The records are the submitters' own: none was re-derived here, and the leaderboard's verified check is empty on all eleven. The labels estimate a person's time, not an agent's. The tasks come from 12 Python repositories, and Django's 231 are almost half of them. The cap result is exact for this agent, whose prompt never mentions its limits; a harness that tells the model its remaining budget could behave differently under a cap. The costs are litellm's estimates from one price table.

A score on this benchmark averages three kinds of task: the 211 every run here resolved, the 59 none did, and the 230 in between where the models differ. The task you are about to hand an agent is one of the three, and the annotators' plain question, how long would this take someone who knows the code a little, moved the odds further than the choice among these eleven models did.