CODING JAG - Issue 311

Welcome to the 311th edition of Coding Jag brought to you by TestMu AI!๐Ÿ‘

What are LLM benchmarks actually measuring? Ai2 built a model called BenchMIRT to find out, running 100 LLMs through 16 benchmarks and more than 34,000 questions. Two dimensions explained most of the scores: safety and general reasoning. Keeping just a tenth of the questions barely changed the rankings. Most suites, it turns out, measure the same thing over and over.

The rest of the week kept circling the same question of what our numbers actually prove. LeadDev asked nearly 600 people whether the AI tools their leaders bought are getting used: Claude Code leads adoption at 78%, but daily use falls to 50%, and just 31% measured the impact at all. GitHub published the receipts on making Copilot cheaper without losing quality, down to the token. And SmartBear paired two numbers that belong together: 93% of teams have adopted AI coding tools, and 92% still test manually.

And the TestMu AI Build vs Buy: AI Testing Agent Whitepaper is out: what an in-house testing agent built on coding agents costs in year one, what buying a purpose-built platform costs at list price, and how to make the call. Written for engineering and QA leaders, free to download.

๐Ÿ“ฌ Come across something useful or interesting? Just reply and let's exchange ideas.

News

Introducing Claude Fable 5.1 and Claude Mythos 5.1

12 minChrome-Extensionanthropic.com

๐Ÿง  Anthropic shipped both models on 1st September at the same price as Fable 5, with cache reads cut by 75%. The number testers should note is AutomationBench: 31.4%, up from 17.1%. Fable 5.1 is on every major cloud today. Mythos 5.1 stays locked behind two verification programmes.

Malicious .git Configs Can Make Claude, Codex, Cursor, and Other AI Agents Run Attacker Code

08 minChrome-Extensionthehackernews.com

๐Ÿชค Swati Khandelwal reports on GitSpawn, eight flaws Manifold Security found across seven command-line coding agents. A poisoned repository sets one Git setting to a command, and the next status check runs it. Four were still unpatched at publication. Unzip a stranger's repo, and your agent does the rest.

Critical Langflow Flaw Exploited to Steal OpenAI and AWS Keys

05 minChrome-Extensionbleepingcomputer.com

๐Ÿ”‘ Bill Toulas reports that attackers are hunting a flaw first disclosed in January. VulnCheck's honeypots logged 360 attempts in days, mostly from Russia, all scraping OpenAI and AWS keys out of environment variables. Versions 1.4.2 and older are affected. Upgrade to 1.11.6.

Debian Votes to Let Contributors Code With AI

05 minChrome-Extensiontheregister.com

๐Ÿ—ณ๏ธ Simon Sharwood, APAC Editor at The Register, covers a vote in which almost 450 Debian developers settled. Eight proposals ran, from an outright ban to open use. Proposal E won: Debian neither endorses nor prohibits generative AI, and disclosure is encouraged but not required. Gentoo went the other way.

TestMu AI State of AI in Testing Survey 2026

10 minChrome-Extensionsurveys.lambdatest.com

๐Ÿ“Š Our annual survey on how teams actually use AI in testing is still open. Roughly ten minutes, covering tools, workflows, and what genuinely works. Every response stays confidential, and we publish the results back to the quality community. Honest answers sharpen the picture for everyone, including you.

AI

Spec-Driven Development Fixed How AI Writes Code, Nobody's Fixed How Humans Verify It

07 minChrome-Extensionforbes.com

๐Ÿ“ Mudit Singh, Co-Founder and Head of Growth at TestMu AI, argues in Forbes that the spec conversation stops one step early. Teams sharpen the brief, then cannot prove the code still matches it. His line to keep: a spec you cannot check against is a way to feel safe, not a safety net.

How We Make AI Coding More Cost-Efficient Without Sacrificing Task Quality

11 minChrome-Extensiongithub.blog

๐Ÿ’ธ Erik Kristensen and Napalys Klicius publish the receipts on trimming Copilot's agent costs. A meta prompting loop had Copilot rewrite its own prompt to roughly half the length, cutting about 1,300 tokens per turn. Code review costs fell by around 20%, with no quality regression that they could detect.

BenchMIRT: What Are LLM Benchmarks Actually Measuring?

09 minChrome-Extensionhuggingface.co

๐Ÿ“Š Ai2 ran 100 models through 16 benchmarks and more than 34,000 questions, then asked what the scores actually track. Two dimensions explained most of it: safety and general reasoning. Keeping just 10% of the questions preserved nearly the same ranking. Most suites measure the same thing repeatedly.

You Bought the AI Tool. Are Your Engineers Using It?

05 minChrome-Extensionleaddev.com

๐Ÿ“‰ Chantal Kapani, staff writer at LeadDev, reports on a survey of nearly 600 respondents. Claude Code leads adoption at 78%, but daily use falls to 50%. Only 26% of leaders saw a real gain, and just 31% measured impact at all. Licences bought are not adoption.

Automation

Build vs Buy: Your AI Testing Agent Decision

08 minChrome-Extensiontestmuai.com

๐Ÿงฎ Our new report prices both paths across year one, component by component, with every input named and sourced. It starts where an honest analysis has to start: where a do-it-yourself build genuinely fits. Then it costs the limits in infrastructure, tokens, context, and engineering attention.

Your AI Investment Has a Governance Gap, and It's Called Testing

07 minChrome-Extensionsmartbear.com

๐Ÿ•ณ๏ธ Klaudia Makiej pairs two numbers from SmartBear's January survey: 93% of teams have adopted AI coding tools, and 92% still test by hand. Forrester supplies the sting. Coding gets 30 to 40% faster, but leave testing manual and team-wide productivity gains average under 10%.

Give Your AI Agent Eyes: Capturing DOM Snapshots for LLM-Assisted Test Debugging

10 minChrome-Extensionthegreenreport.blog

๐Ÿ‘๏ธ Irfan Mujagic hands the failing test's page state to the model instead of just the stack trace. Playwright's accessibility snapshot renders the page as a compact tree of roles and names. His demo quietly flips a button from Submit Order to Place Order, and the snapshot shows exactly that.

Implicit Assertions Are More Readable

05 minChrome-Extensionglebbahmutov.com

โœ๏ธ Gleb Bahmutov rewrites one test twice. The explicit version runs 224 characters, the implicit one 179, a 20% trim with identical coverage. His other habit is worth stealing today: name your assertions, so a failure reads "response status: expected 201, got 404" instead of a bare mismatch.

Tools

Cypress 16: Faster Tests, Starting With HTTP/2 Support

08 minChrome-Extensioncypress.io

โšก Jennifer Shehane walks through the major release. Loading 1,000 images took 1,362ms on the new transport against 3,896ms on the old one. Cypress.env() is gone because it pushed every configured variable into the browser. Read the migration guide before you upgrade; the defaults have moved.

Playwright MCP v0.0.80

05 minChrome-Extensiongithub.com

๐ŸŽฌ The headline addition records what you do in the browser and converts it into Playwright code. Turn it on with the devtools capability, click through the flow by hand, and read back a script. Screenshots also return at full size now instead of being quietly shrunk.

Visual Studio Code 1.136

09 minChrome-Extensioncode.visualstudio.com

๐Ÿงฐ The September release leads with Agent Merge. Point it at a pull request, and it works through review feedback, failed checks, and merge conflicts, reruns the workflows, and then repeats until the thing is mergeable. VS Code will also ping you when an agent stalls waiting on your input.

Video & Podcast

AI Testing Is Bigger Than You Think: 5 Areas Testers Must Own with Swati Seela

07 minChrome-Extensiontestguild.com

๐ŸŽค Joe Colantonio talks to Swati Seela, a principal software quality engineer with more than 20 years in testing. Her map splits AI testing into five areas testers should claim. The line that lands: AI opening every reply by agreeing with you is the modern green dashboard.

The Crucial AI Shift You Need to Master Right Now

13 minChrome-Extensionyoutube.com

๐Ÿ“บ Dave Farley spends 13 minutes on the move from writing individual lines to managing agents. Acceptance testing, TDD, and evals are the tools he reaches for because the loop only works when something can tell the agent it is wrong. Short enough for a lunch break.

Events

Software Quality Summit | Atlanta

06 minChrome-Extensiontestingmind.com

๐Ÿ‘ A one-day, in-person summit at the Emory Conference Center Hotel on 2 October, running from 9 to 5. Speakers include Wasim Haque of OneTrust, Corey Harmon of NRG-CPower, and Sai Rakshit Yerram of Visa. Early bird is $375 until 15 September. We are a Gold Sponsor, and our own Mudassar Syed speaks.