Welcome to the 311th edition of Coding Jag brought to you by TestMu AI!๐
What are LLM benchmarks actually measuring? Ai2 built a model called BenchMIRT to find out, running 100 LLMs through 16 benchmarks and more than 34,000 questions. Two dimensions explained most of the scores: safety and general reasoning. Keeping just a tenth of the questions barely changed the rankings. Most suites, it turns out, measure the same thing over and over.
The rest of the week kept circling the same question of what our numbers actually prove. LeadDev asked nearly 600 people whether the AI tools their leaders bought are getting used: Claude Code leads adoption at 78%, but daily use falls to 50%, and just 31% measured the impact at all. GitHub published the receipts on making Copilot cheaper without losing quality, down to the token. And SmartBear paired two numbers that belong together: 93% of teams have adopted AI coding tools, and 92% still test manually.
And the TestMu AI Build vs Buy: AI Testing Agent Whitepaper is out: what an in-house testing agent built on coding agents costs in year one, what buying a purpose-built platform costs at list price, and how to make the call. Written for engineering and QA leaders, free to download.
๐ฌ Come across something useful or interesting? Just reply and let's exchange ideas.
News
12 min
anthropic.com
๐ง Anthropic shipped both models on 1st September at the same price as Fable 5, with cache reads cut by 75%. The number testers should note is AutomationBench: 31.4%, up from 17.1%. Fable 5.1 is on every major cloud today. Mythos 5.1 stays locked behind two verification programmes.
08 min
thehackernews.com
๐ชค Swati Khandelwal reports on GitSpawn, eight flaws Manifold Security found across seven command-line coding agents. A poisoned repository sets one Git setting to a command, and the next status check runs it. Four were still unpatched at publication. Unzip a stranger's repo, and your agent does the rest.
05 min
bleepingcomputer.com
๐ Bill Toulas reports that attackers are hunting a flaw first disclosed in January. VulnCheck's honeypots logged 360 attempts in days, mostly from Russia, all scraping OpenAI and AWS keys out of environment variables. Versions 1.4.2 and older are affected. Upgrade to 1.11.6.
05 min
theregister.com
๐ณ๏ธ Simon Sharwood, APAC Editor at The Register, covers a vote in which almost 450 Debian developers settled. Eight proposals ran, from an outright ban to open use. Proposal E won: Debian neither endorses nor prohibits generative AI, and disclosure is encouraged but not required. Gentoo went the other way.
10 min
surveys.lambdatest.com
๐ Our annual survey on how teams actually use AI in testing is still open. Roughly ten minutes, covering tools, workflows, and what genuinely works. Every response stays confidential, and we publish the results back to the quality community. Honest answers sharpen the picture for everyone, including you.
AI
07 min
forbes.com
๐ Mudit Singh, Co-Founder and Head of Growth at TestMu AI, argues in Forbes that the spec conversation stops one step early. Teams sharpen the brief, then cannot prove the code still matches it. His line to keep: a spec you cannot check against is a way to feel safe, not a safety net.
11 min
github.blog
๐ธ Erik Kristensen and Napalys Klicius publish the receipts on trimming Copilot's agent costs. A meta prompting loop had Copilot rewrite its own prompt to roughly half the length, cutting about 1,300 tokens per turn. Code review costs fell by around 20%, with no quality regression that they could detect.
09 min
huggingface.co
๐ Ai2 ran 100 models through 16 benchmarks and more than 34,000 questions, then asked what the scores actually track. Two dimensions explained most of it: safety and general reasoning. Keeping just 10% of the questions preserved nearly the same ranking. Most suites measure the same thing repeatedly.
05 min
leaddev.com
๐ Chantal Kapani, staff writer at LeadDev, reports on a survey of nearly 600 respondents. Claude Code leads adoption at 78%, but daily use falls to 50%. Only 26% of leaders saw a real gain, and just 31% measured impact at all. Licences bought are not adoption.
Automation
08 min
testmuai.com
๐งฎ Our new report prices both paths across year one, component by component, with every input named and sourced. It starts where an honest analysis has to start: where a do-it-yourself build genuinely fits. Then it costs the limits in infrastructure, tokens, context, and engineering attention.
07 min
smartbear.com
๐ณ๏ธ Klaudia Makiej pairs two numbers from SmartBear's January survey: 93% of teams have adopted AI coding tools, and 92% still test by hand. Forrester supplies the sting. Coding gets 30 to 40% faster, but leave testing manual and team-wide productivity gains average under 10%.
10 min
thegreenreport.blog
๐๏ธ Irfan Mujagic hands the failing test's page state to the model instead of just the stack trace. Playwright's accessibility snapshot renders the page as a compact tree of roles and names. His demo quietly flips a button from Submit Order to Place Order, and the snapshot shows exactly that.
05 min
glebbahmutov.com
โ๏ธ Gleb Bahmutov rewrites one test twice. The explicit version runs 224 characters, the implicit one 179, a 20% trim with identical coverage. His other habit is worth stealing today: name your assertions, so a failure reads "response status: expected 201, got 404" instead of a bare mismatch.
Tools
08 min
cypress.io
โก Jennifer Shehane walks through the major release. Loading 1,000 images took 1,362ms on the new transport against 3,896ms on the old one. Cypress.env() is gone because it pushed every configured variable into the browser. Read the migration guide before you upgrade; the defaults have moved.
05 min
github.com
๐ฌ The headline addition records what you do in the browser and converts it into Playwright code. Turn it on with the devtools capability, click through the flow by hand, and read back a script. Screenshots also return at full size now instead of being quietly shrunk.
09 min
code.visualstudio.com
๐งฐ The September release leads with Agent Merge. Point it at a pull request, and it works through review feedback, failed checks, and merge conflicts, reruns the workflows, and then repeats until the thing is mergeable. VS Code will also ping you when an agent stalls waiting on your input.
Video & Podcast
07 min
testguild.com
๐ค Joe Colantonio talks to Swati Seela, a principal software quality engineer with more than 20 years in testing. Her map splits AI testing into five areas testers should claim. The line that lands: AI opening every reply by agreeing with you is the modern green dashboard.
13 min
youtube.com
๐บ Dave Farley spends 13 minutes on the move from writing individual lines to managing agents. Acceptance testing, TDD, and evals are the tools he reaches for because the loop only works when something can tell the agent it is wrong. Short enough for a lunch break.
Events
06 min
testingmind.com
๐ A one-day, in-person summit at the Emory Conference Center Hotel on 2 October, running from 9 to 5. Speakers include Wasim Haque of OneTrust, Corey Harmon of NRG-CPower, and Sai Rakshit Yerram of Visa. Early bird is $375 until 15 September. We are a Gold Sponsor, and our own Mudassar Syed speaks.