Human beings make errors (mistakes), which produce defects (faults, bugs), which in turn may result in failures. Humans make errors for various reasons, such as time pressure, complexity of work products, processes, infrastructure or interactions, or simply because they are tired or lack adequate training.
Defects can be found in documentation, such as a requirements specification or a test script, in source code, or in a supporting work product such as a build file. Defects in work products produced earlier in the SDLC, if undetected, often lead to defective work products later in the lifecycle. If a defect in code is executed, the system may fail to do what it should do, or do something it shouldn’t, causing a failure. Some defects will always result in a failure if executed, while others will only result in a failure in specific circumstances, and some may never result in a failure.
Errors and defects are not the only cause of failures. Failures can also be caused by environmental conditions, such as when radiation or electromagnetic fields cause defects in firmware.
A root cause is a fundamental reason for the occurrence of a problem (e.g., a situation that leads to an error). Root causes are identified through root cause analysis, which is typically performed when a failure occurs or a defect is identified. It is believed that further similar failures or defects can be prevented or their frequency reduced by addressing the root cause, such as by removing it.
It is necessary to understand there are degrees of failure. Just because a security vulnerability is not detected and resolved, it doesn’t necessarily mean the security testing approach failed. There are too many possible security vulnerabilities, with new ones being discovered daily. However, there are other cases where security test approaches have been inadequate to effectively identify security risks, which have led to sensitive data and other digital assets being compromised.
Root cause analysis can help identify why a security testing approach may have failed. Possible causes include:
Availability is typically specified in terms of the amount of time a system (or software) is available to users and other systems under normal operating conditions. Systems may have a low maturity, but still have a high availability. For instance, a phone network may fail to connect several calls (and thus have low maturity), but as long as the system recovers quickly and allows the next attempts to connect most users will be content. However, a single failure that caused a phone network outage for several hours would represent an unacceptable level of availability. Availability is often specified as part of an SLA and measured for operational systems, such as websites and software as a service (SaaS) applications. The availability of a system may be described as 99.999% (‘five nines’), in which case it should be unavailable no more than 5 minutes per year, alternatively system availability may be specified in terms of unavailability (e.g., the system shall not be down for more than 60 minutes per month).
Measuring availability prior to operation (e.g., as part of making the release decision) is often performed using the same tests used for measuring maturity; tests are based on an operational profile of expected use over a prolonged period and performed in a test environment as close to the operational environment as possible. Availability can be measured as MTTF/(MTTF + MTTR), where MTTF is the mean time to failure and MTTR is the mean time to repair (MTTR), which is often measured as part of maintainability testing. Where a system is high-reliability and incorporates recoverability (see section 4.4.5) then we can substitute mean time to recover for MTTR in the equation when the system takes some time to recover from a failure.
The benefits of testing are offset by quality costs. A means of quantifying the total cost of quality-related efforts and defects is called cost of quality. Cost of quality involves classifying project and operational costs into four categories related to product defect costs:
The total appraisal costs and internal failure costs are usually significantly less than the external failure costs. This therefore makes testing extremely valuable. By determining the costs in these four categories, test managers can create a convincing business case for testing.
There are more approaches that can be considered for defining the cost of quality. The ISTQB® syllabus supports two of them. This syllabus is based on the Feigenbaum’s approach, and the ISTQB® Foundation Level Syllabus V.4 presents Boehm’s approach (see the ISTQB® Foundation Level Syllabus V.4, Section 1.3, Testing Principles). These two approaches have been selected to reach a broader understanding of the cost of quality. Feigenbaum’s approach (Feigenbaum, Nov/Dec 1956) considers quality as a customer-oriented and company-wide process, while Boehm’s approach (Boehm, 1979) focuses on the trade-off between the cost of prevention and the cost of failure in software development (Hadjicostas, 2004).
Piloting a proposed improvement is an effective way of reducing the risk of failure, gaining experience, building support and reducing the risk of implementation failure. This is especially important where those improvements involve major changes to working practices or place a heavy demand on resources.
Selection of a pilot should balance the following factors:
While there certainly are many different performance failure modes that can be found during dynamic testing, the following are some examples of common failures (including system crashes), along with typical causes:
Slow response under all load levels
In some cases, response is unacceptable regardless of load. This may be caused by underlying performance issues, including, but not limited to, bad database design or implementation, network latency, and other background loads. Such issues can be identified during functional and usability testing, not just performance testing, so test analysts should keep an eye open for them and report them.
Slow response under moderate-to-heavy load levels
In some cases, response degrades unacceptably with moderate-to-heavy load, even when such loads are entirely within normal, expected, allowed ranges. Underlying defects include saturation of one or more resources and varying background loads.
Degraded response over time
In some cases, response degrades gradually or severely over time. Underlying causes include memory leaks, disk fragmentation, increasing network load over time, growth of the file repository, and unexpected database growth.
Inadequate or graceless error handling under heavy or over-limit load
In some cases, response time is acceptable but error handling degrades at high and beyond-limit load levels. Underlying defects include insufficient resource pools, undersized queues and stacks, and too rapid time-out settings.
Specific examples of the general types of failures listed above include: