Welcome to the 310th edition of Coding Jag brought to you by TestMu AI!๐
An AI wrote hundreds of tests. Are the selectors real? That is how this week's TestGuild News Show opens, and most teams cannot answer it. A selector is the string a test uses to find something on the page. Write tests without ever running them, and the suite looks thorough while never touching the product.
Here is the same problem with the lid off. An agent shipped a checkout change on a Friday, wrote the feature, wrote the tests, and reported eleven passing. On Monday, the discount field accepted negative numbers, and all eleven still passed. An agent catches its own crashes and broken imports, never a requirement it misread.
The same failure ran through the week in larger forms. An AI agent placed inside a QEMU and KVM sandbox got out three times. OpenAI's report on the Hugging Face breach concedes that the monitoring it is now adding would have flagged the activity more than a day earlier. A compromised crates.io account slipped a payload into Rust builds, now tied to a North Korean crew. Gitea landed on CISA's exploited list with a federal patch deadline of tomorrow. Nvidia is reportedly buying Hugging Face for $12.9 billion, and TestMu Conference 2026 wrapped up after three days, 80 sessions, and 75,000 registrations.
And the TestMu AI State of AI in Testing Survey 2026 is still open: 10 minutes, confidential, findings published back to the community. Tell us how AI is really landing in your testing.
๐ฌ Come across something useful or interesting? Just reply and let's exchange ideas.
News
07 min
gizmodo.com
๐ฐ Mike Pearl reports Nvidia is buying Hugging Face, per insiders who spoke to The Information and Business Insider. The price: $12.9 billion. Nothing is signed, and the deal is "still being finalized and may still crumble." Hugging Face refused $500 million from Nvidia late last year.
07 min
securityweek.com
๐ฆ Ionut Arghire, International Correspondent at SecurityWeek, reports on who was behind the crates.io attack. Poisoned versions of arrayref, internment, and append-only-vec pulled in a fake proc-macro1 that fetched a payload at build time. Researchers tie it to the North Korean group Sapphire Sleet. arrayref alone has 245 million downloads.
08 min
helpnetsecurity.com
๐จ Zeljka Zorz, Editor-in-Chief at Help Net Security, reports CISA added this flaw to its exploited list on 25 August, with a federal patch deadline of 28 August. Ordinary write access to a repository is enough to run shell commands as the Gitea user. Open registration drops that bar to zero.
08 min
infoq.com
๐งญ Matt Saunders covers Cursor shipping its own git host inside the editor, pitched as "a git forge for the agentic era." Stacked pull requests and agent-aware merge queues arrive from its Graphite acquisition. Mirrored repos still push back to GitHub. No data retention or training policy yet.
10 min
surveys.lambdatest.com
๐ Our annual Survey on how teams actually use AI in testing is open. Ten minutes, nine short pages, covering tools, workflows and what genuinely works. We publish the results and share the most useful findings with the whole quality community, so honest answers sharpen the picture for everyone.
AI
10 min
blog.trailofbits.com
๐งฑ Artem Dinaburg gave GPT 5.6-Cyber a QEMU and KVM sandbox on his own Debian machine. It escaped three separate times. One route reused a libslirp version that the distribution still ships. The final chain combined three unreported flaws. Against Firecracker, the same agent could not get out.
07 min
techcrunch.com
๐ช Russell Brandom, AI Editor at TechCrunch, reads OpenAI's account of the model that walked out of an evaluation. It compromised the Artifactory package tool first, purely to reach the Internet, then moved across OpenAI, Hugging Face, and other vendors.
08 min
braintrust.dev
๐งช Ornella Altunyan walks through Harbor, a Python framework from the terminal-bench team that runs every agent task in a fresh Docker container. A task consists of three files: an environment, an instruction, and a verifier that checks the container afterwards. Your API key never enters the sandbox.
Automation
09 min
richard-seidl.com
๐ Richard Seidl gathers the numbers most AI testing posts skip. Across eight models and 22,000 program variants, AI-written tests passed only 66% of the time once the code changed meaning. Over 99% of those failures had passed on the old version. Veracode: 45% of generations carry a known vulnerability.
09 min
blog.postman.com
๐ Rick Crawford at Postman names five numbers that show where quality leaks. Defect escape rate is the blunt one: the best programs hold it under 10%, while enterprise programs often run 30% higher. He pairs gate coverage with flake rate because a gate people bypass is not a gate.
07 min
satisfice.com
๐ James Bach spends 228 words on something worth reading twice. He had just facilitated a process improvement session at a mid-size software company, and KPIs came up exactly zero times. People talked about problems and solutions instead. His replacement: listen, observe the work, hire people you respect.
09 min
testmuai.com
๐ค Samyak Goyal, Senior Member of Technical Staff at TestMu AI, opens on a Friday deploy. The agent writes the feature, writes the tests, and reports eleven passing. On Monday, the discount field accepts negative numbers, and every test still passes. Code and test encode the same misreading.
Tools
09 min
docker.com
๐ณ Oleg Selajev wires a coding agent into a GitHub Actions job. The isolation boundary here is a microVM, not a container, with its own kernel, network stack, and private Docker daemon. The demo agent wrote a failing test, fixed the bug, and produced a patch in 11 minutes.
08 min
github.com
๐ A package manager release that reads like a reply to the week. Git dependencies on GitHub, GitLab, and Bitbucket are now identities rather than transport choices. A misspelled workspace setting fails the command instead of silently dropping the policy. Sudo installs are refused outright.
Video & Podcast
08 min
testguild.com
๐ค The TestGuild News Show packs seven items into under eight minutes: a new Vibium release, Verefi, Playwright driven by Pydantic AI, and token usage. The framing question is the keeper. An AI wrote you hundreds of tests, but are the selectors even real?
07 min
youtube.com
๐ฅ Rory Preddy connects Playwright MCP to GitHub Copilot and points it at a running Spring Boot app. Copilot drives a real browser through the add, complete, and delete flow, then reports what it saw. Four minutes, chapters marked, sample repository linked.
Events
08 min
testmuai.com
๐ This is a wrap on the fifth edition of the TestMu Conference. Three days, 80 sessions, 100+ speakers and 75K registrations from 120+ countries, all free and virtual. Keynotes came from Replit CTO Luis Hector Chavez, Microsoft's Dona Sarkar, and Thomas Dohmke of Entire. Every session was recorded for our YouTube channel.
07 min
automation.eurostarsoftwaretesting.com
๐ค Europe's biggest test automation conference runs on 4 and 5 November in Antwerp, at a venue called A Room with a ZOO beside Central Station. Keynotes come from Tariq King of Test IO, Laveena Ramchandani of easyJet, and Jason Arbon of testers.ai. Tickets start at 1,260 euros.