Root cause analysis (RCA) is a technique for identifying and addressing the underlying or fundamental causes of a defect rather than only its symptoms. RCA supports a structured approach to quality improvement, and its primary objective is to prevent the recurrence of defects. The TA uses various techniques to identify root causes of defects and failures (e.g., defect taxonomies, the five whys technique, cause-effect diagrams, and Pareto analysis).
The classical RCA involves subject matter experts studying a defect in considerable detail after resolving it. However, there are usually many defects that need to be analyzed. Therefore, it would be very inefficient and time-consuming to have preventive action planning for each defect. One way to approach this problem is to classify defects and then perform the RCA for the defect types occurring.
Defect classification is based on the recognition that individual defects capture a great deal of information about the development process and the system under test. Defect classification allows the TA to extract information about various aspects of the development process from the defect and turn it into a process measurement. This, in turn, gives an insight into the types of errors made during development, which is helpful for process improvement. Defect classification bridges the gap between quantitative defect statistics and qualitative RCA. To effectively support RCA, the defects should be uniformly classified throughout the entire SDLC, from early testing to production.
The TA should support their organization in standardizing software defect classification. This will improve communication and the exchange of information regarding defects among developers and organizations, facilitating the RCA.
Examples of defect classification methods are:
Defects can also be mapped to quality attributes using software quality models such as ( ISO/IEC 25010 , 2023) or the FURPS model (Grady et al., 1987).
More information on RCA can be found in (ISTQB-ITP, v1.0).
When multiple agile teams need to collaborate in order to implement a system or a solution, some of the quality assurance (QA) and test activities will span multiple teams and the responsibility for delivering a working solution is shared between the teams. If a single team tries to fix a problem the solution may cause new problems for the other agile teams.
In a value stream, bottlenecks are a root cause for waste. Some typical bottlenecks in development value streams are:
It requires a flexible set of root cause analysis techniques to discover many potentially relevant root causes using systems thinking. If not used, there is a risk of concluding too quickly that there is just one single root cause. A basic root cause analysis technique in lean is “Five Whys.” Causal loop diagram (CLD) is a method that can help if the feedback structure of human interaction or of the technical system needs to be identified.
It is necessary to understand there are degrees of failure. Just because a security vulnerability is not detected and resolved, it doesn’t necessarily mean the security testing approach failed. There are too many possible security vulnerabilities, with new ones being discovered daily. However, there are other cases where security test approaches have been inadequate to effectively identify security risks, which have led to sensitive data and other digital assets being compromised.
Root cause analysis can help identify why a security testing approach may have failed. Possible causes include:
Using a model-based improvement approach, as described in the previous section, improvements are introduced by comparing the test approach of a project or team to external best practices. Analytical approaches identify problems based on data from the project or team itself. Appropriate improvements can be derived from an analysis of these problems. Analytical approaches can be used together with a model-based approach to verify results and provide diversity.
Problems can be identified by using quantitative and qualitative data. Section 1.5.3 of this syllabus, Analytical-Based Test Process Improvement Approach, introduces analytical approaches that mainly use quantitative data from the test process and data from defects to assess the current approach. Section 1.5.4 of this syllabus, Retrospectives, introduces retrospectives, in which qualitative data regarding what works well and what does not work well, is collected from development and test team members.
Data analysis is important for objective test process improvement and a valuable support to purely qualitative assessments, which may otherwise result in imprecise recommendations that are not supported by data. Applying an analytical approach to improvement most often involves a quantitative analysis of the test process to identify problem areas and set project-specific goals. The definition and measurement of key parameters is required to assess the test process and evaluate whether improvements are successful.
Examples of analytical approaches are:
Root cause analysis is the study of problems to identify their root causes. This allows the identification of solutions that remove the causes of problems rather than merely addressing the immediate obvious symptoms. A possible analysis procedure would involve selecting an appropriate set of defects, identifying clusters in this data, and using cause-effect diagrams (also called Ishikawa or fishbone diagrams) to identify the root causes of important defect clusters. Improvements are then derived to prevent similar defects from occurring.
Measures, metrics, and indicators are used in a quantitative manner to assess how well the test process in the project or team is performed. Key attributes of the test process to be considered are effectiveness, efficiency, and predictability. For each of these attributes, one or several metrics can be selected. By collecting and analyzing corresponding data, the key areas requiring improvement can be identified.
The GQM approach (Basili, et al., 2014) (van Solingen & Berghout, 1999) provides a framework to define and analyze metrics that are tailored to the information needs of relevant project stakeholders. Measurement goals define a quality aspect of an object that needs to be measured for a particular purpose, perspective, and context. These goals are refined into questions that define the quality aspect from the stakeholders’ viewpoint. Metrics are then selected that provide the necessary information to answer the question. Data collected for the metrics answer the questions, to assess the measurement goal and satisfy stakeholders’ information needs.
More information on these analytical-based test process improvement approaches can be found in ISTQB® Expert Level Improving the Test Process Syllabus.
Systems thinking and root cause analysis are important disciplines that provide many different techniques to analyze complex problems. An agile test leader needs to participate in and facilitate analysis of complex problems to help the organization grow and optimize its value streams.
Human beings make errors (mistakes), which produce defects (faults, bugs), which in turn may result in failures. Humans make errors for various reasons, such as time pressure, complexity of work products, processes, infrastructure or interactions, or simply because they are tired or lack adequate training.
Defects can be found in documentation, such as a requirements specification or a test script, in source code, or in a supporting work product such as a build file. Defects in work products produced earlier in the SDLC, if undetected, often lead to defective work products later in the lifecycle. If a defect in code is executed, the system may fail to do what it should do, or do something it shouldn’t, causing a failure. Some defects will always result in a failure if executed, while others will only result in a failure in specific circumstances, and some may never result in a failure.
Errors and defects are not the only cause of failures. Failures can also be caused by environmental conditions, such as when radiation or electromagnetic fields cause defects in firmware.
A root cause is a fundamental reason for the occurrence of a problem (e.g., a situation that leads to an error). Root causes are identified through root cause analysis, which is typically performed when a failure occurs or a defect is identified. It is believed that further similar failures or defects can be prevented or their frequency reduced by addressing the root cause, such as by removing it.
Solution analysis is used to identify potential solutions to problems and then to choose between those solutions. Any chosen improvement(s) or solution(s) may be decided in a number of ways:
The solution analysis process includes one or more of the following, depending on the method chosen:
Causal analysis is the study of problems to identify their possible root causes. This allows identification of solutions which will remove the causes of problems and not just address the immediately obvious symptoms. If causal analysis is not used, attempts to improve test processes may fail because the actual root causes are not addressed and the same or similar problems recur.
Many software process improvement models emphasize the use of causal analysis as a means of continually improving the maturity of the software process.
The following systematic methods for causal analysis are described below as examples:
Note: other methods are available for causal analysis (see Advanced syllabus) and also checklists of common causes may be used as an input to the causal analysis, for example when carrying out causal analysis on defects.
Cause-Effect diagrams (also known as Ishikawa fishbone diagrams: Ishikawa fishbone diagrams) were developed for the manufacturing and other industries [Ishikawa 91] and have been adopted in the IT industry [Juran]. These diagrams provide a mechanism to identify and discuss root causes under a number of headings.
The steps to apply, (according to [Robson 95]), are described as:
4. Use the brainstorming method – possible causes are brainstormed and added to the ribs of the diagram. For each first level cause, checklists are used to identify the underlying root causes, which may be on a different part of the diagram.
5. Incubate the ideas for a period of time
6. Analyze the diagram to look for clusters of causes and symptoms. Use the Pareto idea (80% of the gain from 20% of the effort) to identify clusters that are candidates to solve.
Cause Effect diagrams may also be used to work from the effects back to the causes, as noted in [BS7925-2] and [Copeland 03].
Security auditing is a manual examination and evaluation that identifies weaknesses in an organization’s security processes and infrastructure. Security audits at the procedural level (e.g., to review internal controls) may be performed manually. Security audits at the architectural level are often performed with security audit tools, which may be aligned with a particular vendor solution for networking, server architecture and workstations.
Just like security testing, a security audit does not guarantee all vulnerabilities will be found. However, the audit is one more activity in the security process to identify problem areas and indicate where remediation is needed.
In some security auditing approaches, testing is performed as part of the auditing process. However, the scope of security auditing is much larger than security testing. Security auditing often investigates areas such as procedures, policies and controls that are difficult to test in a direct way. Security testing is more involved with the technologies to support security, such as firewall configuration, correct application of authentication and encryption, and application of user rights.
There are five pillars to security auditing [Jackson, 2010]:
Assessment – Assessments document and identify potential threats, key assets, policies and procedures, and management’s tolerance for risk. Assessments are not one-time events. Since the environment and business are constantly in flux, assessments must be performed on a regular basis. This also provides the opportunity to know if security policies are still relevant and effective.
Prevention – This extends beyond technology and includes administrative, operational, and technical controls. Prevention is not accomplished just through technology, but also through policies, procedures, and awareness. While the prevention of any and all attacks is unrealistic, the combination of defenses can help make it much more difficult for an attacker to succeed.
Detection – Detection is how a security breach or intrusion is identified. Without adequate detection mechanisms, there is the risk of not knowing whether the network has been compromised. Detective controls help to identify security incidents and provide visibility into activities on the network. Early incident detection enables an appropriate reaction to recover services quickly.
Reaction – Reaction time is greatly reduced with good security defenses and detection mechanisms. While security breaches are bad news, it is important to know if one has occurred. Fast reaction time is critical to minimize the exposure to the incident. Fast reaction requires good preventive defenses and detection mechanisms to provide the data and context needed for response. The speed and efficiency of incident response is a key indicator of the effectiveness of an organization’s security efforts.
Recovery – Recovery starts with determining what occurred so that systems can be recovered without recreating the same vulnerability or condition that caused the incident in the first place. The recovery phase does not end with restoring the system. There is also the root cause analysis that determines what changes need to be made to processes, procedures, and technologies to reduce the likelihood of the same type of vulnerability in the future. An auditor must ensure that the organization has a plan for recovery that includes ways to prevent future similar incidents.
The inspection process is described in the ISTQB Foundation syllabus and expanded in the ISTQB Advanced syllabus. Using the software inspection process [Gilb & Graham] suggests a different approach to causal analysis.
The causal analysis meeting is a facilitated discussion which lasts two hours and follows a set timescale and format.
In the defect analysis, each defect is categorized with:
When a test script fails or passes unexpectedly, root cause analysis must be performed. This will include inspecting test logs, performance data, setup, and teardown of the test script.
It is also helpful to execute a few isolated tests. Intermittent failures are more difficult to analyze. The defect can be in the test case, the SUT, the TAF, the hardware or the network. Monitoring system resources may yield clues for the root cause. Test log file analysis of the test case, the SUT and the TAF can help identify the root cause of the defect. Debugging may also be necessary. To aid in identifying the root cause may require support from a test analyst, business analyst, developer, or system engineer.
Verify if all the assertions are in place. Missing assertions may result in inconclusive test results.
Data can be collected from the following sources:
Since a TAS has automated testware at its core, the automated testware can be enhanced to record information about its use. Testware enhancements made to the underlying testware can be used by all the higher-level automated test scripts. For example, enhancing the underlying testware to record the start and end time of test execution may well apply to all tests.
Features of test automation that support measurement and test report generation
The scripting languages of many test tools support measurement and reporting through facilities that can be used to record and log information before, during, and after test execution of individual tests, and entire test suites.
Test reporting on each of a series of test runs needs to have an analysis feature to consider the test results of the previous test runs so it can highlight trends, such as changes in the test success rate.
Test automation typically requires automating both the test execution and the test verification, the latter being achieved by comparing specific elements of the actual results with the expected results. This comparison is best done by a test tool using assertions. The level of information that is reported as a result of this comparison must be considered. It is important that the test status be determined correctly (i.e., passed or failed). In the case of failed status, more information about the cause of the failure will be required (e.g., screen shots).
Differences between actual results and expected results of a test are not always clear, and tool support can help greatly in defining comparisons that ignore differences that are expected, such as dates and times, while highlighting any unexpected differences.
Test logging
Test logs are a source that is frequently used to analyze potential defects within the TAS and the SUT. In the following section are examples of test logging, categorized by TAS and SUT.
TAS logging
The context determines whether the TAF or the test execution is responsible for logging information that should include the following:
SUT logging
Correlation of the test automation results with SUT logs to helps identify the root cause of defects in the SUT and the TAS.
Integration with other third-party tools (e.g., spreadsheets, XML, documents, databases, and report tools)
When information from the execution of automated test cases is used in other tools for tracking and reporting (e.g., updating traceability information), it is possible to provide the information in a format that is suitable for third-party tools. This is often achieved through existing test tool functionality (e.g., export formats for test reporting) or by creating customized reporting that is output in a format consistent with other software.
Visualization of test results
Test results can be made visible using charts. Consider using colored icons such as traffic lights to indicate the overall status of the test execution/test automation so that decisions can be made based on reported information. Management is particularly interested in visual summaries to see the test results, which aides in decision making. If more information is needed, they can still drill down into the details.
The application of methods described above may require the selection of specific defects for analysis from a potentially large collection. The following approaches can be taken in combination to selecting defects for analysis:
A TA can proactively help minimize the recurrence of defects into the software. This syllabus discusses two approaches related to mitigating the recurrence of defects: analyzing test results to improve test analysis and test design and supporting root cause analysis with defect classification.
A project retrospective meeting is a team meeting that assesses the degree to which the team effectively and efficiently achieved its objectives for the project or activities just concluded, and team satisfaction with the way in which those objectives were achieved. In agile projects, retrospectives are crucial to the overall success of the team. By having a retrospective at the end of each iteration, the agile team has an opportunity to review processes, adapt to changing conditions and interact in a way that helps to identify and resolve any interpersonal issues.
A Test Manager should be able to organize and moderate an effective retrospective meeting. Because these meetings are sometimes discussing controversial points, the following rules should be applied:
According to [Derby&Larsen 06] a retrospective should follow a structure as outlined here:
Of course, the objective of the retrospective is to deliver outcomes that result in actual change. Successful retrospectives involve the following:
A Test Manager may participate in project-wide retrospectives. There is also value in having a test retrospective with the test team alone focusing on those things the test team can control and improve.
The TA can contribute to defect prevention in several ways, using their domain knowledge, test expertise, and analytical skills. Examples include:
In addition to participating in defect prevention, the TA assesses (usually in consultation with the test manager) whether the proposed measures have resulted in the desired effect. Examples of metrics that help to evaluate the effectiveness of these measures include:
Retrospectives are meetings in which a team reviews its methods and collaboration, captures lessons learned (good and bad), and decides on changes and actions to achieve improvements (both for testing and non-testing issues). Retrospectives address topics such as the process, people, organization, collaboration, and tools.
Retrospectives are used in both sequential development models and Agile software development. In sequential development models they are a part of test completion. In this context, retrospectives aim at generating lessons learned in order to better manage future projects. In Agile software development retrospectives are generally held at the end of each iteration to discuss what was successful and what needs to be improved, and how those improvements can be incorporated in the next iteration. Retrospectives are performed by the entire team and thus support the whole team approach and foster continuous improvement. Note that dedicated retrospectives are sometimes required to address testing issues.
A typical retrospective consists of the following steps:
Introduction: The goal and agenda of the retrospective are reviewed, and an atmosphere of mutual trust is created so that problems can be discussed without placing blame or judgment.
Collect data: Data is collected regarding what happened during the iteration or project. It is possible to collect qualitative data, such as a timeline of key events that identifies issues and lists how each team member feels about those issues. In addition, quantitative data from metrics can be presented, for example, data for test progress, defect detection, test effectiveness, test efficiency, and predictability can provide an objective insight into the testing of the project or iteration.
Derive improvements: The collected data is analyzed to understand the current situation and to generate improvement ideas. For example, root cause analysis can be applied to identify root causes of identified problems and a brainstorming session can be held in order to generate ideas on how to resolve the root causes.
Decide on improvement actions: Actions to implement the improvement ideas are derived and prioritized. An improvement plan and responsibilities are defined. Goals and associated metrics can be defined to evaluate the impact of the actions on the identified problems. Implementing too many improvements at once is difficult to manage with verifiable steps.
Close retrospective: In this last step, the retrospective itself is reviewed to identify strengths and improvements in the retrospective process. A retrospective is performed regularly, especially in Agile software development. Continuous improvement is also applied to the retrospective itself.
It is important to appropriately document the results of a retrospective. In a sequential development model, findings, conclusions, and recommendations need to be distributed and communicated in an understandable way to members of the organization. In Agile software development, problems and actions should also be documented to allow the review of actions and their potential impact on the problems in the next iteration.
Testers, being a part of the (project) team, bring in their unique perspective. They can raise testingrelated problems (and others) and stimulate the team to think about possible improvements.
Further information can be found in (Derby & Larsen, 2006).
Metrics-based test process improvement in Agile software development requires careful interpretation of both product and process indicators, aligning with continuous improvement principles and drawing on team capabilities and tools. Except for genuine key performance indicators, metrics should serve as feedback mechanisms rather than performance targets, to avoid misuse that could distort behavior and lead to local optimization.
When the defect detection percentage declines unexpectedly, a root cause analysis may reveal ineffective exploratory testing or insufficient test automation. In response, the team may consider refining their test chartering techniques, improving session-based test management, or investing in tooling for observability or exploratory testing support.
If tests repeatedly fail due to environmental instability, investing in test environment management, service virtualization, or containerized solutions may improve stability and confidence in test results.
High cycle times for resolving defects might indicate weak collaboration across roles or siloed responsibilities. A possible improvement could involve adopting whole team approaches to defect triage, involving testers earlier in refinement sessions, and embedding Agile test leaders to reinforce shared ownership of quality. Pairing metrics such as mean time to failure and mean time to repair with retrospectives enables teams to recognize bottlenecks in feedback loops and adapt accordingly.
Low coverage on critical paths, as reported by code coverage tools, may prompt efforts to strengthen automated coverage at the component testing or component integration testing levels, potentially rebalancing the test pyramid toward more granular, fast-executing tests. However, teams must assess whether these tests provide meaningful insights or simply inflate metrics.
Redundant or fragile tests indicated by high failure rates with low defect yields should be refactored or replaced using design patterns such as page objects or test data builders. For more information on design patterns in test automation, see (ISTQB-TAE, v2.0).
A declining participation in retrospectives or quality discussions may reflect a cultural issue. Coaching may be required to facilitate psychological safety and reinforce testing as a shared responsibility. In such a context, qualitative metrics, such as perceived trust and openness in collaboration, can be just as important as numerical indicators.
When test debt accumulates, as evidenced by an increasing backlog of flaky tests or manual regression test time, a test strategy may schedule time in each iteration to maintain tests, rotate test ownership, and apply test automation refactoring practices. Measurement of test automation effectiveness should include not just pass/fail rates, but also maintenance effort and test execution lead times.
Monitoring-based metrics, such as mean time to recovery or production defect escape rates, might initiate test process improvements. A high MTTR may indicate missing system alerts, resource constraints, or insufficient exploratory testing of operational behaviors. This could lead to incorporating chaos testing or software development techniques such as canary releases, feature toggles, or dark launches.
For people development, metrics such as pair testing frequency, defect learning session counts, or crossfunctional contribution ratios may indicate whether knowledge sharing and capability growth are occurring. A lack of cross-functional collaboration across roles may necessitate mentoring, role rotation, or structured learning pathways such as community of practice initiatives.
Tool-related improvements can also be informed by lead time metrics from code commit to feedback, flaky test analysis, and integration failure rates. Tools that enable short feedback loops, such as dashboards, automated root cause analysis, or smart test selection, can directly support team learning and enable faster responses.
Any improvement action should be measurable and validated. For this, retrospectives can play a vital role. Metrics must be traceable to test objectives and risk priorities, rather than arbitrary benchmarks. Models such as GQM (Goal-Question-Metric) are helpful for structuring the use of metrics in a goal-oriented manner. Finally, teams should maintain critical thinking and contextual sensitivity to prevent metrics from becoming disconnected from the core test mission.
The IDEAL model describes the following high-level activities for the “Diagnosing” phase:
The end result of this phase is typically a test assessment report.
Based on the high-level activities of the IDEAL model, the following considerations must be taken into account in this phase:
The activities performed in this phase depend on the approach to be taken to test process improvement (see Chapter 5).
If an analytical-based approach is to be adopted (see Chapter 4) the various causal analysis techniques (Section 4.2) may be applied and metrics, measures and indicators (Section 4.4) analyzed.
If a model-based approach is to be used (see Chapter 3) then an assessment will be planned and performed. Sections 6.3.1 to 6.3.3 cover these aspects in more detail.
Control flow analysis is the static technique where the steps followed through a program are analyzed through the use of a control flow graph, usually with the use of a tool. There are a number of anomalies which can be found in a system using this technique, including loops that are badly designed (e.g., having multiple entry points or that do not terminate), ambiguous targets of function calls in certain languages, incorrect sequencing of operations, code that cannot be reached, uncalled functions, etc.
Control flow analysis can be used to determine cyclomatic complexity. The cyclomatic complexity is a positive integer which represents the number of independent paths in a strongly connected graph.
The cyclomatic complexity is generally used as an indicator of the complexity of a component. Thomas McCabe's theory [McCabe76] was that the more complex the system, the harder it would be to maintain and the more defects it would contain. Many studies have noted this correlation between complexity and the number of contained defects. Any component that is measured with a higher complexity should be reviewed for possible refactoring, for example division into multiple components.
Testers collect and report discrepancies between actual and expected outcome through defect reports. A defect report contains all relevant information the tester can provide to help the business analyst understand what happened and to assess the deviation.
Defect analysis is a joint activity of testers and business analysts. Usually, the tester identifies the acceptance criteria that are not satisfied. The business analyst may then be asked to analyze its impact on the related business processes. This includes determining the priority of the defect (e.g., low, medium, high, critical) with respect to its potential business impact on system usage.
To analyze the business impact of a defect, the business analyst and tester may do the following:
The impact analysis and the resulting decision regarding further actions to be taken are documented in the defect report.