Measurements
Numbers instead of vibes: token costs, time-to-green, benchmark results, and how they were collected.
-
What one coding task costs: tokens, dollars, minutes, three ways
A method for costing an agent task across an interactive session, a headless run, and a cheaper model, with a price calculator and a dataset of published numbers.
-
Coding agent time-to-green: how do you measure ten small tasks fairly?
A method and a first dataset for timing a coding agent from prompt to first passing check: ten defined tasks, a runner script, and only honestly filled columns.
-
SWE-bench explained: what a resolve rate measures and what it misses
How a SWE-bench score is produced, what Verified changed, and five papers on harness effects, lucky passes, contamination, and realistic prompts, with a dataset.