Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Keeping Quality Transparency Throughout the Organization
Keeping Quality Transparency Throughout the Organization
Testers understand product quality best. Learn how to report test depth, keep open bug counts visible, and share a quality score card with the whole team.
Last Updated on:
Quality transparency means the whole organization can see how deeply a build was tested and how good it looked to the testers. Testers pair a one-to-five depth of testing score with a high, medium, or poor quality rating, and publish the open bug count daily where the team already looks. This guide covers subjective quality and test depth, reporting new feature quality, maintaining visible open bug counts, scoring the overall quality of the product, and reporting quality when AI writes part of the code.
Key Takeaways
- Testers usually understand best what quality means for a product, which puts testers in the strongest position to keep the whole team aware of quality standards.
- A one-to-five depth-of-testing score reported next to a subjective quality rating says more about product quality than a defect count on its own.
- When testers reveal their depth scores at the same time and still disagree after discussion, the team records the most pessimistic evaluation.
- Publishing the open bug count daily where the team already looks, such as the team chat channel or the delivery dashboard, keeps the count a shared signal instead of a report someone has to request.
- Splitting a large product into functional areas and scoring each area for depth and quality produces a score card the whole team can read.
- Recording how much of a change an AI assistant produced, and showing that record next to the depth score, keeps quality reporting honest because AI-generated code carries different risk at the same depth score.
Subjective Quality and Test Depth
We understand that it is not easy to express if the quality is excellent or bad. For example, if the team only had time to execute a basic first pass at testing a certain product area and discovered no defects, was the quality good or low? Of course, it's hard to say. So, a good start can be to generate a report on both the "depth" of testing they've been able to accomplish and their subjective quality evaluation.
Example:
POOR Quality ------> Average Quality ------> Great Quality
To describe testing depth, use a value between one and five, with one indicating a shallow initial pass and five indicating extensive testing of all parts of the software, including borders and extreme failure conditions. If you're testing in a group, have each colleague think of a number between one and five to signify how thoroughly he believes the application has been tested. When everyone has a number, have them all demonstrate it at the same time by raising that number of fingers on a hand.
This is similar to a game of rock-paper-scissors, except no one loses. If you discover that not everyone agrees, which is likely, debate the anomalies as a unit. For example, if I pick one and you pick three, we may talk about why you think the level of testing was average and why I think it was shallow. You can vote again after some debate to see where we all stand. If we still can't agree after a certain number of votes, we'll go with the most pessimistic evaluation.
Use a similar mechanism to assess quality, except this time use your thumbs up for good quality, down for bad quality, and sideways for middling. This brief, collaborative activity also provides everyone on the team with a shared notion of depth and quality. Report the pair of evaluations to the team as a one-to-five score for depth and a high, medium, or poor-quality assessment for your testing team. Will use a number of stars to represent depth and a happy, neutral, or frowning face to represent quality.
Key Takeaway: Rating testing depth from one for a shallow first pass to five for extensive testing of boundaries and extreme failure conditions, and pairing that number with a high, medium, or poor quality rating, gives a team a shared way to describe how well a product was tested.
Report the New Feature Quality
At the end of a development cycle, it is typical practice in agile development to organize a product demo. Each new feature is displayed for the entire team to view during the session.
In the past, the testing team could have grumbled about not having enough time to adequately test. There is no need to complain anymore. The extent of the tests is clearly stated. It is up to the entire team to determine what to do. After observing this technique, I've heard engineers agree that they should finish coding earlier in the sprint to allow testers more time. I'm not exaggerating.
Key Takeaway: Stating the depth of testing openly at the end-of-cycle product demo turns a complaint about insufficient test time into a decision for the whole team, and some developers respond by finishing their code earlier in the sprint.
Maintain visible open bug counts
The volume of bugs discovered isn't necessarily a good indicator of software quality. Low testing depth, for example, may yield a small number of bugs, whereas excessive testing depth is likely to provide a greater number of defects. Even if there are few defects, a subjective quality evaluation might reveal poor quality: "I didn't do much testing, but everything I tried broke."
You may then publish a new version of the graph every morning and keep it somewhere the team already looks. The graph depicts a curve that gradually climbs over time and occasionally lurches downward as the development team repairs defects in an effort to lower the bug count.
Most teams no longer share a room, so the chart needs a home the team already visits every day. Pin it in the team chat channel, add it as a panel on the delivery dashboard, or have the build pipeline post the current count as a comment on each release branch. The format matters less than the access. If someone has to request the number, it has stopped being a shared signal and has become a report.
Agree on the counting rules before you publish the first chart. Write down which states in your bug tracking tool count as open, which severities are included, and what happens to a defect the team decides to defer. Deferred defects should stay in the count until someone closes them deliberately, because dropping them quietly is how a low number stops meaning anything. Record the date whenever you change a rule, so an old reading is never compared against a new definition.
Keeping the total low has become a source of pride for the development team, which now saves a few days at the end of each development cycle to concentrate on defects. They don't want the cycle to conclude with a high defect count since they know this figure is visible to everyone, and a large number of defects will damage the team's quality evaluation.
Key Takeaway: An open bug count only means something when the counting rules are written down first and the open bug chart is published every day somewhere the team already visits, because a count that has to be requested is no longer a shared signal.
The overall quality of the product
More and more features are added as the product's development progresses. Testing of the newly added features and the entire product is still being done diligently by testers. They want to increase their test coverage across the product and learn how all these features work together. The testers evaluate the complete product quality at the conclusion of each cycle to update the team on their opinion of the quality of the product as they perceive it.
Because the product being developed is vast, providing a single evaluation for the complete product would be insufficient. Fortunately, this product is separated into a dozen different functional areas. The test team provides a depth and quality assessment for each functional area. The end result is a type of "score card" made up of depth scores and quality assessments. Pushing both depth and quality as high as possible is the team's overall objective.
The team should be aware of any defects and do everything possible to address them as soon as they are discovered. You may keep the number of defects discovered visible at all times by implementing a graphic that resembles an agile burn chart. The vertical axis of this two-dimensional chart indicates the number of known bugs, while the horizontal axis is time.
Key Takeaway: Giving every functional area of a large product its own depth score and quality assessment at the end of each cycle builds a score card the team works to push higher on both depth and quality.
How Should You Report Quality When AI Writes Part of the Code?
Record how much of a change was produced with an AI assistant, and show that record next to the depth score. The 2025 DORA State of AI-assisted Software Development report found that 90% of survey respondents use AI at work and that 30% report little or no trust in the code AI generates. The same research observed a positive relationship between AI adoption and software delivery throughput, and a negative relationship between AI adoption and software delivery stability. A depth of 3 over code a developer wrote and a depth of 3 over code an assistant generated do not carry the same risk. A score card that hides the difference stops telling the team the truth.
Capture the attribution in the commit rather than in someone's memory. Git supports a Co-authored-by trailer, written as Co-authored-by: NAME <NAME@EXAMPLE.COM> after an empty line at the end of the commit message. Require that trailer on any commit produced with an assistant. The pipeline that publishes your open bug count can then also report what share of the cycle's commits carry it, so the depth vote argues over a fact instead of an impression.
Publish the rule as well as the number. DORA lists a clear and communicated AI stance among the capabilities that amplify the effect of AI, and it describes what happens without one: developers either hide their usage or avoid the tools entirely. Write down which assistants are approved, what code review an AI-assisted change needs before it reaches the depth vote, and who decides. A team that cannot say how a change was produced cannot honestly rate its quality.
Key Takeaway: Marking AI-assisted commits with a Co-authored-by trailer and publishing a written AI stance lets a team report quality on AI-generated code honestly, because a depth score means something different when an assistant wrote the code.
Closing
The product manager ultimately decides whether to ship or not at time of release, but the entire group is knowledgeable of the quality of the product being shipped. The choice to deploy a product with a low degree of testing and poor quality would be dangerous, but that risk is now transparent to anyone and everyone.
Author
David Tzemach is a software quality and engineering leader with 19+ years of experience in software testing, quality assurance, and large-scale R&D operations. He specializes in building QA organizations from scratch, defining quality frameworks, and implementing agile and shift-left testing practices across enterprise environments. David has served as Head of QA and QA Architect, authored multiple books on agile quality and testing, and actively contributes to the testing community through his QualityBreach platform and publications.
Quality Transparency FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests



