SWE-Bench Verified
TermA framework for evaluating software engineering tools.
article
10 storys
calendar_today
First: 2026-02-03
update
Last: 2026-09-18
open_in_new
Website
menu_book
Wikipedia
Stories
Completed digest stories linked to this service.
-
Coding agent benchmarks harden: SWE-Bench Pro Verified and Real-SWE reset the sc...2026-09-16SWE-Bench Pro Verified and Real-SWE are forcing a reboot of how coding agents are measured in the real world. ...
-
New coding-agent benchmarks raise the bar and cut through leaderboard noise2026-08-12New benchmarks show coding agents still stumble on large-scale refactors and building products from scratch, d...
-
Open Qwen 3.5 narrows the SWE-bench gap with closed models2026-06-29Open Qwen 3.5 is closing the SWE-bench gap with top closed models, which could change your code-agent cost mat...
-
Context beats model: a cheap agent tops SWE-bench Verified2026-05-09A low-cost model paired with richer repo-aware context just topped SWE-bench Verified, showing agent wiring ca...
-
Promptfoo joins OpenAI with a practical playbook for evaluating coding agents2026-04-29Promptfoo is now part of OpenAI and published a hands-on guide that reframes how to evaluate coding agents in ...
-
DeepSeek V4 shows up near the top of SWE‑Bench Verified at lower cost2026-04-26DeepSeek V4 preview models landed high on SWE-Bench Verified, offering near-SOTA scores with 1M context at a l...
-
Agentic coding grows up: domain-grounded agents and verifiable training move fro...2026-04-24Agentic coding is shifting from generic code suggestions to domain-verified systems that generate validated, p...
-
Open-weight coding models surge: Kimi K2.6 hype, Qwen3.6-27B runs local, Meta po...2026-04-23Open-weight coding models jumped forward this week, with Kimi K2.6 hype, a practical Qwen3.6-27B local setup, ...
-
Agentic coding is moving from hype to practice—design for reliability, governanc...2026-04-22Agentic coding is leaving the demo phase, forcing teams to engineer for reliability, governance, and real resu...
-
Anthropic ships Claude Opus 4.7: stronger agentic coding, stricter prompts, and ...2026-04-20Anthropic released Claude Opus 4.7 with big gains in agent coding, tighter instruction-following, and a resear...