Next-Gen App & Browser Testing Cloud
Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Your Agent Still Doesnt Know What Matters [Testμ 2026]
Your Agent Still Doesnt Know What Matters [Testμ 2026]
Angela Cutulenco on the four CMDB layers an AI agent cannot read, why an empty field looks like a low rating, and what to ask instead of fixing your data.
Published on:
On This Page
- Two Databases, Twenty Years Apart
- The Compliment About Discovery
- Complete Is Not Trustworthy
- Service Mapping And Ownership
- Criticality And Dependency
- Silent Drift In Four Layers
- Four Seconds Inside A Head
- The Knowledge That Leaves
- The Access Request Walkthrough
- The Blind Spot In Agent Evals
- Autonomy Levels, Not Data Projects
Four seconds. That is how long an experienced engineer spends silently correcting a configuration record before carrying on with their day, and none of those corrections is ever written down.
At Testμ Conf 2026, Angela Cutulenco, Senior Director of Governance at the Toronto Transit Commission, asked what happens when that engineer’s work is handed to an agent. The ticket handling transfers. The corrections do not, because only the work was ever documented.
If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.
TL;DR
Accurate but not trustworthy is Angela Cutulenco’s verdict on the modern CMDB, or configuration management database: automated discovery keeps it accurate about what exists, while the human-entered layers describing what those things mean to the business have nothing forcing them to stay current. It matters because an AI agent cannot tell the two halves apart.
- Is your CMDB out of date? - No, and Angela Cutulenco says so as a technical statement rather than a courtesy. The asset layer of a modern CMDB is frequently the most current record in the entire organisation, more accurate than architecture diagrams, personal spreadsheets and often the documentation vendors hand over at go-live. She hedges it with frequently and offers no evidence.
- Does accurate mean trustworthy? - No. Angela Cutulenco separates the two: complete is about how much of the environment you have recorded, which discovery delivered, and trustworthy is about whether the records say what those things mean to the business, which was never finished.
- Do the business-meaning layers update themselves? - No. Service mapping, ownership, criticality and dependency behaviour are all entered by people, and Angela Cutulenco says none of the four has a mechanism forcing it to stay correct. Discovery has a schedule and you notice when it stops; ownership has no schedule at all.
- Can an AI agent tell an empty field from a deliberate low rating? - No. Angela Cutulenco’s point is that an empty criticality field and a value somebody set look identical to an agent, and at the Toronto Transit Commission tiers one and two are populated and accurate while tiers three and four were never marked at all.
- Can automated discovery see a failover pair? - No. Angela Cutulenco describes a never-fail architecture where a replacement server with the same name and IP address is created at the disaster recovery site, so discovery only ever sees one server because only one exists at a time. Every dependency on that server therefore reads as critical when the system actually survives losing it.
- Is the CMDB’s unreliability already causing failures? - No, and Angela Cutulenco says something has been protecting everyone: experienced engineers correct the record in their heads. She estimates four seconds for an analyst to know the assignment group is the pre-reorganisation one, that an empty rating means tier three rather than unimportant, and that a failover makes a critical-looking dependency survivable. The four seconds is rhetoric with no measurement behind it.
- Does the AI agent make a mistake in the failure scenario? - No. Angela Cutulenco walks through an access request where every step is carried out correctly on true data and the outcome is still wrong, and says that given that data she would defend the agent’s decision herself.
- Did that access failure actually happen at the Toronto Transit Commission? - No. Angela Cutulenco frames it as conditional, describing what would happen if an access administration agent went into production. The production agent today handles password resets on one incident type and completes that work with no people involved.
- Will better agent evaluations fix this? - No. Angela Cutulenco says her team tested reasoning, instruction-following, consistency and explainability, and the agent passed everything. The tests examine how the agent thinks and never examine whether what it was told is true.
- Should you fix the CMDB first? - No, and Angela Cutulenco calls that advice useless from twenty years on this work. A fix-first plan is a multi-year programme needing funding every year and ending as a project that never finishes, while the agents keep arriving.
- What should teams ask instead of how good their CMDB is? - What the agent is allowed to do. Angela Cutulenco offers three autonomy levels: the agent reads anything but recommends nothing, the agent recommends an action, or the agent acts on its own. She gives no criteria for promoting an agent between the three.
- What is the context integrity test? - Four questions Angela Cutulenco delivers verbally and never shows: does this service have a complete map, do we know who actually answers, does the criticality rating mean anything, and would the map have predicted our last ten incidents. She says there is a way to calculate the responses and never describes the calculation.
Two Databases, Twenty Years Apart
Twenty years ago she built a configuration management database by hand in Microsoft Access, and stresses she means that literally. Every asset was created manually and every relationship entered one at a time, because no tool in her organisation could find them on the network.
Everyone in the department could read it. Nobody but her could change it. Every update arrived through change management and she entered each one herself.
She names the objection before the audience can, agreeing that the arrangement sounds like a bottleneck and confirming it was one. She then calls it the most trustworthy configuration database she has ever worked with.
Today discovery runs on schedule with service mapping in place, and in one night the system does what took her months by hand.
Her agent implementation runs in two places at once. A non-production environment is where the team learns what it is capable of, and production is restricted to one incident type on one service, where it completes the work on its own with no people involved.
Early on she promises to explain how they chose that one incident type, saying the choice is the reason she is giving the talk. The payoff arrives 26 minutes later as a single clause.
The Compliment About Discovery
She flags the compliment as a technical statement rather than a courtesy, and names the genre she is rejecting: talks that open by telling you your configuration data is out of date and full of errors. That starting point, she says, was correct 15 years ago and is not correct today. No source, sample or method accompanies the claim.
What nightly discovery finds is servers, virtual machines, databases, installed software and its versions, and network devices, without asking anyone for permission and without waiting for anyone to remember it needs doing.
Her Access database was accurate only because she kept it accurate, which works while the environment stays small. Discovery has no such limit, finds far more systems than she ever could, and never forgets to look.
The direct version is that the asset layer of a modern CMDB is frequently the most current record in the entire organisation, ahead of architecture diagrams, well ahead of personal spreadsheets, and in many organisations ahead of the documentation vendors produce at go-live. Both hedges are hers, and no evidence is offered for either.
She then names the industry habit she blames: solving the part of the problem we know how to solve, and treating the whole problem as finished.
Complete Is Not Trustworthy
Her definitional split is the argument in one move. Complete is about how much of the environment you have recorded, and discovery delivered that. Trustworthy is about whether the records tell you what those things mean to the business, and that half was never finished.
What discovery can answer is whatever announces itself. A server responds, a database responds, a service is listening on the port; discovery asks and the system answers, and the answer is a fact.
What it cannot answer are the questions people ask when something breaks. Does the customer depend on this server. Who do I call at three in the morning when it stops. How bad is it if it stops. What else stops working when it fails.
Her concrete version of that gap comes from her own sector. A server cannot tell you that a rider is standing on a platform in the cold, waiting for the arrival information that server provides. The connection exists in the business world and is nowhere on the network.
Hence the structural claim: a CMDB is two databases sitting in the same tables, an asset list that discovery maintains and a description of what those assets mean that a person maintains. Everyone looks at one screen and assumes one quality of data.
Discovery spots the hardware. Humans define the context.
— TestMu AI (@testmuai) August 21, 2026
At #TestMuConf, Angela Cutulenco breaks down the ultimate tag team:
• Discovery finds: Servers, databases, software, and devices.
• Humans supply: Ownership, criticality, service mapping, and dependency behavior.… pic.twitter.com/uBpR6q3G59
Note: An agent cannot tell a deliberate rating from a field nobody filled in. Try TestMu AI now!
Service Mapping And Ownership
Service mapping, in her definition, is the path from a piece of infrastructure to a business service somebody outside the organisation would recognise. The limit is that discovery can tell you two systems talk to each other and cannot tell you what those systems let a customer do.
Her example is Wheel-Trans, which she describes as the scheduling transportation system for people with disabilities, taking people to dialysis appointments.
The service depends on background scripts that carry the workload, and those scripts are not configuration items. They are records inside the workload system that discovery never sees, so the map shows servers and applications while the most important part of the service stays invisible to anyone reading it.
Ownership is the second layer. Every configuration item stores a name and a support group, and somebody populated those on the day the system was built. That build-time observation is about ownership; the published description attaches it to criticality instead, which reverses her point about that field.
Her case is a reorganisation. During the first phase of an IT transformation, one large end-user group became three teams. Asset life cycle and inventory management stayed in the same section, and the third became a new field services team under a different manager, with responsibilities that had never existed before.
Nobody updated the CMDB, so tickets sat with teams that no longer owned the work while the records looked complete the entire time. No dates, ticket counts or duration are given, and no record was shown.
Criticality And Dependency
Criticality is the field that says how much damage is done when something stops. Her organisation records tier one for mission-critical systems and tier two for business-critical ones.
Tiers three and four are not marked at all, which she attributes to not having had time. Tier three covers operating support systems and tier four admin systems, and both sit with an empty criticality field.
The agent-facing consequence is the sharpest line in this part of the talk. When an agent reads an empty field, it treats it the way it would treat a low rating, because the empty field and the value somebody set look exactly the same to it. The failure here is absence rather than staleness, and tiers one and two are populated and accurate.
Dependency behaviour is the fourth layer and the one she calls the most important. Discovery sees that two systems are connected and cannot see what happens to the second when the first fails, and that difference is what tells you how far damage travels.
Her example is an internal mission-critical system on a never-fail architecture. When a server at the main site goes down, a replacement with the same name and the same IP address is created automatically at the disaster recovery site. Discovery sees one server because only one exists at a time, and the CMDB has no way to record that the failover exists.
The result is an inverted signal. Every dependency on that server reads as critical, while the system carries on working even when the server is lost.
Silent Drift In Four Layers
All four layers depend on a person to enter them and keep them current, and none has a mechanism that forces them to stay correct.
The asymmetry she draws is with discovery itself. Discovery has a schedule, and if it stops running you find out quickly. Ownership has no schedule at all.
When a service map stops matching reality, nothing tells you. When the last person who understood a dependency leaves the organisation, nothing tells you either.
These layers do not fail in a way anyone notices. They drift, and nothing in the system raises its hand to say a map is out of date, an assignment group is wrong or a rating was never set.
She concedes the obvious objection immediately: that unreliability has caused very little trouble so far.
Four Seconds Inside A Head
She puts the challenge to herself directly. Twenty years running on this data, thousands of changes made with it, most of them fine, so either she is overstating the problem or something has been protecting everyone. Her answer is that something has, and that it is about to be removed.
What an experienced engineer does with a change record is three corrections in sequence. Read the assignment group and know it is the pre-reorganisation group, so route to field services instead. Read the empty criticality field and know that empty here means tier three rather than unimportant. See every dependency on that system reading as critical, know about the failover, and decline to escalate.
All of that happens in roughly four seconds inside a person’s head, and then they carry on with their day. The figure is an estimate with no measurement behind it.
Ask that engineer later what they did and they will tell you they checked the CMDB. They will not mention ignoring three of the fields they read, because they do not think of it that way.
They would not call it expertise either. To them it is knowing how things work in their own organisation, which is exactly why nobody has ever written it down.
The Knowledge That Leaves
She is explicit that her hand-built database did not work because of the tool. Access was never designed for it and she knew that at the time. Nor was it her own precision, since she made mistakes like everybody else.
It worked because of where the information came from. Nothing entered that database unless it came from an approved change, and by the time a change reached her, somebody had already decided what the system was, who owned it and what it connected to. She was not inventing any of it.
She endorses the shift away from that model anyway, saying letting discovery take over the writing was the right decision and she would make it again. The cost she names is that discovery records what is there and not what any of it means, so the facts kept arriving and the meaning stopped.
Give the same record to an agent and every silent correction disappears. It routes to a team that no longer does the work, and treats an absent criticality rating as a low one because it has no way to know the rating was never set.
Her framing of that outcome is careful. The agent is not making a mistake; it is doing what it was told using the information it was given. It follows the logic correctly, names the records it read, and will explain its decisions clearly if asked. And it will be wrong, because what the analyst knows was never written down.
The part she says concerns her most is what leaves with the work. Analysts supply information the CMDB does not hold every time they touch a ticket, from experience rather than instructions, and it appears in no job description and no report. Hand the work to an agent and their knowledge goes with it, because only the work was ever recorded.
The Access Request Walkthrough
She rules out the dramatic failure first. Hallucination, an invented server, changing the wrong environment: those are real, and they are the failures teams are already good at catching, which discovery makes rare. If a system is not there, the agent will not find a record for it.
The scenario that follows has not happened and she says so. Their production agent today handles password resets, and she introduces the walkthrough as what would happen if they decided to put an access administration agent into production.
The setup is a user reporting lost access to a business-functional system, where approval is not an IT prerogative but is granted by the product owner on the business side, whose name is recorded in the CMDB.
The first two steps are clean. The agent identifies the system in the CMDB and the record is correct, confirmed by discovery the previous night. It reads criticality and finds tier one, a field that is both populated and accurate. Everything it has read so far is true.
The next two steps introduce the gap. The agent looks up the product owner to route the approval and finds a name. The record does not say when that name was entered, or whether the person is still in the role. It sends the approval request there.
The final step branches two ways, and both are bad. Either the request sits with someone who no longer holds the role and nobody approves it, so the user waits and cannot work. Or that person approves it because it looks routine and they used to do exactly this, and the user receives access to confidential information from someone who no longer has the authority to grant it.
The Blind Spot In Agent Evals
Her challenge to the audience is to look back at those steps and say where the agent made a mistake. It did not. Every step was carried out correctly.
It read the fields it was supposed to read, followed the logic it was supposed to follow, and reached the conclusion that follows from the information available. Given that data, she says she would defend the decision.
Her team did test the agent before it went near production, and it passed everything. They tested whether it reasons correctly and it does. They tested whether it follows instructions and it did. They tested consistency and explainability too.
That absolute is worth noting as hers: no source is offered for nobody.
Her conclusion from it is that better evaluations will not solve the problem. The tests examine how the agent thinks and never examine whether what it was told is true. This is the argument the session rests on; the CMDB anecdotes are illustration for it.
Autonomy Levels, Not Data Projects
She rejects the obvious remedy outright: the answer is not to fix the CMDB.
From twenty years on this work, her description of the fix-first plan is a programme that runs for several years, needs funding in every one of them, and leaves you with a project that never finishes. Meanwhile the agents keep arriving, which is why she calls the advice useless.
Her reframe is to stop asking how good your CMDB is and start asking what the agent is allowed to do, then decide next steps from that.
The three autonomy levels are read-only, where the agent can search, summarise and report but cannot recommend; recommendation, where it can propose an action; and action, where it acts on its own. She gives no criteria for promoting an agent between them.
Her candid admission about her own programme is the payoff of the promise made 26 minutes earlier, and it lands in one clause. Their first agent was chosen on volume; the second will go through the approach she is recommending.
The context integrity test is four questions delivered verbally and never shown: does this service have a complete map, do we know who actually answers, does the criticality rating mean anything, and would the map have predicted our last ten incidents. She says there is a way to calculate the responses and never describes it.
Her closing reframe is that the question is not whether the agent is good enough, but whether the picture your organisation holds is useful.
This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.
Author
TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




