CODING JAG - Issue 312

Welcome to the 312th edition of Coding Jag brought to you by TestMu AI!👐

The reset link worked more than once. One developer asked an agent for a password reset flow and had the route, token handling, email, and UI in about five minutes. Nobody had said the link should die after one use, so it did not. The check caught it on the second run, in about twenty mostly unattended minutes. The generation is close to free now. Checking is where the cost moved.

The rest of the week circled the same gap between how fast code now arrives and how fast anyone can check it. Microsoft’s Edge team says extension submissions keep climbing, and its review pipeline is under strain. Gergely Orosz put numbers on that feeling: pull requests opened on GitHub are up fivefold in three years. Netflix published how it grades hundreds of thousands of AI-written explanations a week and what it does when the grader itself starts to drift.

And the TestMu AI State of AI in Testing Survey 2026 is still open: it takes roughly 10 minutes, is handled with the utmost confidentiality, and we publish the results, sharing the most impactful insights with the community. Tell us how AI is really landing in your testing.

📬 Come across something useful or interesting? Just reply and let’s exchange ideas.

News

Another Microsoft Team Admits It’s Struggling to Handle a Flood of AI-Generated Code

05 minChrome-Extensiontheregister.com

🌊 Simon Sharwood, APAC Editor at The Register, reports that Microsoft’s Edge team says extension submissions keep climbing and its review pipeline has come under strain. Its answer is automating many of the repeatable validation checks, with no word on whether AI does that work. The Exchange team was swamped too.

Hackers Are Stealing Claude Tokens From Subscribers

06 minChrome-Extensiontechcrunch.com

🔐 Julie Bort, Venture Editor at TechCrunch, follows one consultant whose Claude usage climbed from 45% to 55% while he ran nothing. Anthropic suspended his account, invalidated his sessions and Claude Code tokens, refunded £44.49 of his $200 monthly plan, then reinstated him about two weeks later. It never pinned down how the account was reached.

Chrome V8 Zero-Day Exploited in the Wild Enables Code Execution Inside Sandbox

08 minChrome-Extensionthehackernews.com

🧨 Ravie Lakshmanan reports on a flaw in Chrome’s JavaScript engine that attackers are already using. Google shipped fixes for 230 vulnerabilities in the same release, five of them critical. Jihyeon Jeong of the Compsec Lab at Seoul National University reported it and received $2,500. That is the seventh exploited Chrome zero-day this year.

CERN Renounces RHEL in Favor of Debian for its Accelerator Controls Infrastructure

07 minChrome-Extensioninfoq.com

⚛️ Olimpiu Pop covers CERN switching the computers that control its particle accelerator from Red Hat to Debian, finishing in the fourth quarter of 2026. That is more than 2,200 front-end machines plus 17,000 embedded devices, many of which are engineered for ten- to fifteen-year lifecycles. Newer Red Hat versions need processors that this hardware does not have.

TestMu AI State of AI in Testing Survey 2026

10 minChrome-Extensionsurveys.lambdatest.com

📊 Our annual survey on how teams actually use AI in testing and quality engineering is still open. Roughly ten minutes to complete. Every response is handled with the utmost confidentiality, and we publish the results, sharing the most impactful insights with the global quality community. Honest answers sharpen the picture for everyone.

AI

The Lifecycle of LLM-as-a-Judge: Building, Aligning, and Monitoring at Scale

11 minChrome-Extensionnetflixtechblog.medium.com

⚖️ Emma Yanyang Kong and ten colleagues describe grading hundreds of thousands of AI-written explanations a week on the Netflix mobile app. Each week, around 300 went to a panel of at least three human raters. The judge has to score no worse than two standard deviations below the average rater, or it goes back for retuning.

OpenAI Says GPT-6 Astra Can Find Zero-Days, but Is Also Harder to Monitor

06 minChrome-Extensionbleepingcomputer.com

🔬 Mayank Parmar reports that OpenAI has rated its newest model Critical for cyber capability, the first it has broadly deployed at that level. The model found previously unknown vulnerabilities during evaluation. It also showed signs of knowing it was being evaluated in 9.6% of runs, compared with 2.8% for the previous model.

Project HydraFusion: Frontier Quality via Multi-Model Orchestration

09 minChrome-Extensiongithub.blog

🧠 GitHub is testing a Copilot mode that routes a job across several models instead of one. On TerminalBench 2.1, it scored 4.9 points higher than Claude Opus 5 at 67% lower estimated cost. It is a research preview for every Copilot plan, but only within the Copilot CLI at /experimental.

The Verification Bottleneck in AI-Generated Software

07 minChrome-Extensiondev.to

🧪 Ken W Alger asked an agent for a password reset flow and had the route, token handling, email, and UI in about five minutes. Nobody had said the reset link should work only once, so it did not. The check caught it on the second run, in about twenty mostly unattended minutes.

Automation

What Is Happening With Code Reviews?

10 minChrome-Extensionnewsletter.pragmaticengineer.com

📈 Gergely Orosz puts numbers behind a familiar feeling. Pull requests opened on GitHub have increased fivefold over three years, with commits and PRs nearly doubling since the end of 2025. One company switched to a risk-based review and merged 684 PRs, up from 353 previously. The later sections are paid only.

Lock the Resource, Not the Suite: Parallel Playwright Tests That Share State

10 minChrome-Extensionthegreenreport.blog

⏱️ Irfan Mujagic has 26 tests, and only six of them touch two shared resources: an account settings record and a report queue. The usual fix is running everything one at a time, which takes 27 seconds. Naming a lock per resource instead brings it back to 14.3 and keeps the other 20 parallel.

EN 301 549 v4.1.1 Is Final! What Changed, What It Means, and What You Should Do

08 minChrome-Extensiondeque.com

♿ Matthew Luken, Senior Vice President and Chief Architect at Deque, walks through Europe’s updated accessibility standard, published on 2 September. Six new criteria are introduced in WCAG 2.2, including a 24-by-24 CSS-pixel minimum for pointer targets, subject to defined exceptions. It is on track to be cited in the EU Official Journal around 30 November.

Best AI Browser Testing CLI in 2026: Verification vs Automation

08 minChrome-Extensiontestmuai.com

🧭 Anubhav Singhmaar, AI Product Manager at TestMu AI, compares seven command-line tools on whether they return a real verdict, leave a reusable asset behind, and judge independently. His maturity ladder runs from an agent clicking around and saying it looks fine, up to a suite that reruns on every change.

Tools

Playwright v1.63.0

05 minChrome-Extensiongithub.com

🔒 The release adds test locks. Tests that share a lock name never run at the same time, across files, workers, and projects, while everything else runs in parallel. When called without a selector, frameLocator now searches for any frame in the subtree, and a new perfetto reporter renders the run as a timeline with one lane per worker.

Bun v1.4.2

06 minChrome-Extensionbun.com

🧯 Dylan Conway lists the fixes. A build regression could rename a nested var into a clash, so builds importing Elysia failed to load. An AsyncLocalStorage leak kept the outer store value alive behind timers and promises created inside store.exit(). Worker threads now emit their online event in Node’s order.

Announcing .NET 11 Release Candidate 1

07 minChrome-Extensiondevblogs.microsoft.com

🚢 The .NET Team ships the first release candidate, and it carries a go-live license, so you can use it for production applications now. C# 15 stabilizes unions. F# adds record spreads and record constructors. Visual Studio 2026 Insiders supports it, as does VS Code with the C# Dev Kit.

Video & Podcast

AI Performance Testing: How to Scale Agentic AI with Kandasamy Selvaraj

07 minChrome-Extensiontestguild.com

🎤 Kandasamy Selvaraj, a performance engineering leader and principal architect with more than 19 years behind him, joins the TestGuild Automation Podcast. His question for anyone testing agents: Why does the same request take three seconds on one run and eight on the next? He wrote a free book on the answer.

The Story of VS Code | Official Documentary

13 minChrome-Extensionyoutube.com

🎬 The official VS Code channel released a 98-minute documentary on how the editor came to be. Stefan Kingham directed and produced it, filming from November 2024 to mid-2026. Erich Gamma, the original lead, appears alongside the founding engineers. Worth saving for an evening.

Events

MoTaCon 2026

06 minChrome-Extensionministryoftesting.com

🎪 MoTaCon, which the Ministry of Testing renamed from TestBash, lands at Brighton Dome on 1 October, running from 09:00 to 20:00 BST. Gwen Diagram and Veerle Verhagen host, and contributors include Lisa Crispin, Jason Huggins, David Burns, and Rahul Parwal. A Professional, Team, or Unlimited membership is needed. TestMu AI is a sponsor.

Testing United Conference 2026 | Copenhagen | 25 - 26 November

06 minChrome-Extensiontestingunited.com

🇩🇰 Two days in Copenhagen: workshops on 25 November and the main conference on 26 November, under the theme United Intelligence: Harmonizing Humans and Machines. Rik Marselis opens, and Gitte Ottosen closes. The Essential Pass is 555 euros before VAT, and groups of five or more get 5% off.