Software Fault Reporting Processes in
Business-Critical Systems
Jon Arvid Børretzen
Doctoral Thesis
Submitted for the partial fulfilment of the requirements for the degree ofPhilosophiae Doctor
Department of Computer and Information Science
Faculty of Information Technology, Mathematics and Electrical Engineering Norwegian University for Science and Technology
Copyright © 2007 Jon Arvid Børretzen ISBN 82-471-xxxx-x (printed)
ISBN 82-471-xxxx-x (electronic) ISSN 1503-8181
Abstract
Today’s society is crucially dependent on software systems. The number of areas where functioning software is at the core of operation is growing steadily. Both financial systems and e-business systems relies on increasingly larger and more complex computer and software systems. To increase e.g. the reliability and performance of such systems we rely on a plethora of methods, techniques and processes specifically aimed at improving the development, operation and maintenance of such software. The BUCS project (BUsiness-Critical Systems) is seeking to develop and evaluate methods to improve the support for development, operation and maintenance of business-critical software and systems. Improving software processes relies on the ability to analyze previous projects and derive concrete improvement proposals. The research in this thesis is based on empirical studies performed in several Norwegian companies that develop business-critical software. The work specifically aims to assess the use of fault reporting approaches, and describe how improvement in this area can benefit process and product quality.
Some specific software methods will be adopted from safety-critical software engineering practices, while others will be taken from general software engineering. Together they will be tuned and refined for this particular context. A specific goal in the BUCS project has been to facilitate the use of traditional Software Criticality Analysis techniques for the development of business-critical software. This encompasses techniques used to evaluate and explore potential risks and hazards in a system. The thesis describes six studies of software development technology for business-critical systems. The main goal is to attain a better understanding of business-critical systems, as well as to adapt and improve relevant methods and processes. Through data mining of historical software project data and other studies of relevant projects, we have gathered information to be evaluated with the goal of improving business-critical systems development. The BUCS project has been involved in investigation of development projects for business-critical systems, investigations that have been continued in the EVISOFT user-driven project. The main goal was to study the effects of revised development methods for business-critical software, in order to improve important quality aspects of these systems.
The main research questions in this work are:
• RQ1. What is the role of fault reporting in existing industrial software development? • RQ2. How can we improve existing fault reporting processes?
• RQ3. What are the most common and severe fault types, and how can we reduce them in number and severity?
• RQ4. How can we use safety analysis techniques together with failure report analysis to improve the development process?
The main contributions of this thesis are:
• C1. Describing how to utilize safety criticality techniques to improve the development process for business-critical software.
• C2. Identification of typical shortcomings in fault reporting.
Preface
This thesis is submitted to the Norwegian University of Science and Technology (NTNU) in partial fulfilment of the requirements for the degree Philosophiae Doctor. The work has been performed at the Department of Computer and Information Science, NTNU, Trondheim, with Professor Reidar Conradi as the main advisor, and Professor Tor Stålhane and Professor Torbjørn Skramstad as co-advisors.
The thesis is part of the BUCS project (BUsiness-Critical Systems) and has been financed for three years by the Norwegian Research Council through the IKT’2010 basic IT Programme under NFR grant number 152923/V30. In addition comes one year as a teaching assistant paid by NTNU. The BUCS project has been lead by Professor Tor Stålhane. Some of the work in this thesis has also partly been financed by the EVISOFT user-driven R&D project under NFR grant number 174390/I40.
Acknowledgements
During the work on this thesis, I have been lucky to been in contact with many people who have provided help, inspiration and motivation. First of all, I want to thank my supervisor, Professor Reidar Conradi, for giving valuable feedback and comments on many drafts and ideas during the last four years. Also, I want to thank Professor Tor Stålhane, my co-advisor, for being the source of a lot of good advice and many bad jokes. I also want to thank the present and former members of the software engineering group at IDI, NTNU for giving me a good working environment. A special thanks to my BUCS colleagues Torgrim Lauritsen and Per Trygve Myhrer for collaboration in our research and daily work.
Parts of the work for this thesis have been done in collaboration with people from several industrial organizations. I am very grateful to these companies and the people I have been in touch with from these organizations who have been helpful and accommodating when sharing their information and experience with me. Also I want to thank master student Jostein Dyre-Hansen for helping me analyze a great deal of data material.
Finally, I want to thank my family and friends for their encouragement and inspiration, and I would especially express my thanks to Ingvild for her love and patience.
Trondheim, Nov 1, 2007 Jon Arvid Børretzen
Table of contents
1 Introduction ... 1 1.1 Motivation ... 1 1.2 Research Context... 2 1.3 Research design ... 41.4 Research questions and contributions... 4
1.5 Included research papers ... 5
1.6 Thesis structure... 7
2 State-of-the-art... 9
2.1 Introduction ... 9
2.2 Software engineering... 9
2.3 Software Quality... 11
2.4 Anomalies: Faults, errors, failures and hazards... 12
2.5 Current methods and practices ... 16
2.6 Business-critical software... 20
2.6.1 Criticality definitions... 20
2.7 Techniques and methods used to develop safety-critical systems... 22
2.8 Empirical Software Engineering ... 26
2.9 Main challenges in business-critical software engineering ... 29
3 Research Context and Design... 31
3.1 BUCS Context ... 31
3.2 Research Focus ... 32
3.3 Research approach and research design ... 35
3.4 Overview of the studies ... 41
4 Results ... 43
4.1 Study 1: Preliminary Interviews with company representatives (used in P1)... 43
4.2 Study 2: Combining safety methods in the BUCS project (Paper P1) ... 44
4.3 Study 3: Fault report analysis (Papers P2, P3, P5) ... 45
4.4 Study 4: Fault report analysis (Paper P4) ... 48
4.5 Study 5: Interviewing practitioners about fault management (Paper P6)... 50
4.6 Study 6: Using hazard identification to identify faults (Paper P7)... 51
4.7 Study 7: Experiences from fault report studies (Technical Report P8)... 52
5 Evaluation and Discussion ... 55
5.1 Contributions ... 55
5.2 Contribution of this thesis vs. literature... 57
5.3 Revisiting the Thesis Research Questions, RQ1-RQ4 ... 58
5.4 Evaluation of validity ... 59
5.5 Industrial relevance of results... 60
5.6 Reflection: Research cooperation with industry... 61
6 Conclusions and future work... 63
6.2 Future Work... 64 Glossary ... 67 Term definitions ... 67 References ... 73 Appendix A: Papers... 81
List of Figures
Figure 1-1 The studies with their related papers and contributions 5
Figure 1-2 The structure of this thesis 8
Figure 2-1 Relationship between faults, errors, failures and reliability 13 Figure 2-2 Relationship between hazards, accidents and safety 14 Figure 2-3 Faults, Hazards, Reliability and Safety 15
Figure 2-4 The Rational Unified Process 18
Figure 2-5 Relationship of business-critical and other types of criticality 21 Figure 2-3 Relationship between faults, errors and failures 22
Figure 4-1 Combining PHA/HazOp and Safety Case 45
Figure 4-2 Percentage of high severity faults in some fault categories 47 Figure 4-3 Quality views associated to defect data, and their relations 48 Figure 4-4 Distribution of severity with respect to fault types for all projects 50 Figure 4-5 Distribution of hazards represented as fault types (%) 51
List of Tables
Table 2-1 Examples of different systems’ criticality 22 Table 2-2 Properties of some safety criticality analysis techniques 25 Table 2-3 12 ways of studying technology, from [Zelkowitz98] 26
Table 2-4 Empirical research approaches 28
Table 3-1 Description of our studies 33
Table 3-2 Type of studies in this thesis 41
Table 3-3 Relation between main and local research questions 41 Table 4-1 Distribution of all faults in fault type categories 47 Table 4-2 Distribution of all faults in fault type categories 47 Table 4-3 Fault type distribution across all projects 49 Table 5-1 Relationship of contributions and research questions 56
Abbreviations
BUCS Business-Critical Software (project)
CBD Component-Based Development
CBSE Component-Based Software Engineering
CCA Cause-Consequence Analysis
COTS Commercial Off The Shelf
DBMS Data Base Management System
GQM Goal Question Metric
GUI Graphical User Interface
ETA Event Tree Analysis
EVISOFT EVidence based Improvement of SOFTware engineering (project) FMEA Failure Mode and Effects Analysis
FMECA Failure Mode Effects and Criticality Analysis
FTA Fault Tree Analysis
HAZOP Hazard and Operability Analysis
IEEE Institute of Electrical and Electronics Engineers
INCO Incremental and component-based software development (project) ISO International Organization for Standardization
NFR Norwegian Research Council
NS-ISO Norwegian Standard
NTNU Norwegian University of Science and Technology
OMG Object Management Group
OS Operating System
OSS Open Source Software
PHA Preliminary Hazard Analysis
QA Quality Assurance
RUP Rational Unified Process (by Rational)
SPI Software Process Improvement
UML Unified Modelling Language (by Rational, later OMG)
1
Introduction
In this chapter the background and research context for this thesis is presented. The chapter also introduces the research design, the research questions and the contributions. Finally, the list of papers and the outline of the thesis is presented.
1.1 Motivation
The technological development in our society has lead to software systems being introduced into an increasing number of different business domains. In many of these areas we become more or less dependent on these systems, and their potential weaknesses could have grave consequences. In this respect, we can coarsely divide software products into three categories: safety-critical software (e.g. controlling traffic signals), business-critical software (e.g. for banking) and non-critical software (e.g. for word processing).
Evidently, the definition of business-critical versus the other two categories may be difficult to state precisely, and would in many cases depend on the particular viewpoint of the business and users. To clarify the distinction between business-critical and safety-critical, we can consider what consequences operation failure (observable and erroneous behaviour of the system compared to the requirements) will have in the two different cases. For safety-critical applications, the result of a failure could easily be a physical accident or an action leading to physical harm for one or more human beings. In the case of business-critical systems, the consequences of failures are not that grave, in the sense that accidents do not mean real physical damage, but that the negative implications may be of a more financial or trust-threatening nature.
Ian Sommerville states that business-criticality signifies the ability of core computer and other support systems of a business to have sufficient QoS to preserve the stability of the business [Sommerville04]. Thus business-critical systems are those whose failure could threaten the stability of a business.
The overall goal for the BUCS project is to better understand and thus sensibly improve software technologies, including processes used for developing business-critical software. In order to do this, empirical studies of projects have been performed in cooperation with Norwegian ICT industry.
Specific BUCS goals as presented in the BUCS project proposal [BUCS02] are the
BG1 To obtain a better understanding of the problems encountered by Norwegian industry during development, operation and maintenance of business-critical software.
BG2 Study the effects of introducing safety-critical methods and techniques into the development of business-critical software, to reduce the number of system failures (increased reliability).
BG3 Provide adapted and annotated methods and processes for development of business-critical software
BG4 Package and disseminate the effective methods into Norwegian software industry.
In this thesis, we aim to study how software faults and software fault reporting practises affects business-critical software, and also if techniques (e.g. PHA, Hazop, etc) from the area of safety-critical systems development can have a positive effect on other quality attributes (e.g. reliability) than safety. The relation between faults and failures is explained in Section 2.4.
1.2 Research Context
This thesis is a part of the work done in the BUCS basic research and development project (BUsiness-Critical Software). The BUCS project was funded by the Norwegian Research Council as a basic R&D project in IT, and was run in 2003-2007. Some parts of the work in this thesis were also financed by the EVISOFT project, a national, user-driven R&D project on software process improvement funded by the Norwegian Research Council [EVISOFT06].
Within the BUCS project, this thesis will focus on fault reporting processes in
business-critical systems. Some important research issues we want to study are the following:
• How do software faults affect the reliability and safety of business-critical systems? • What are the common fault types in business-critical systems?
• How can we use system safety methods in business-critical application development?
1.2.2 The BUCS project
The goal of the BUCS project is not to help developers to finish their development on schedule and budget. We are not particularly interested in the delivered functionality or how to identify or avoid process and project risk. This is not because we think that these properties are not important – it is just that we have defined them out of the BUCS project.
The goal of the BUCS project is to help developers, users and other stakeholders to develop software whose later use is less prone to critical problems, i.e. has sufficient reliability and safety. In a business environment this means that the system seldom behaves in such a way that it causes the customer or his users to lose money, important
Another term is business-safe, which means that a system fulfils the criteria for business-safety in a business-critical system.
That a system is business-safe does not mean that the system is fault-free, i.e. cannot possibly fail. What this means is that the system will have a low probability of entering a state where it will cause serious losses. In this respect, the system characteristic is close to the term “safe”. This term is, however, wider, since it is concerned with all activities that can cause damage to people, equipment, the environment or severe economic losses. Just as with general safety, business-safety is not a characteristic of the system alone – it is a characteristic of the system’s interactions with its usage environment.
BUCS are considering two groups of stakeholders and wants to help them both: • The customers and their users. They need methods that enables them to:
o Understand the dangers that can occur when they start to use the system as part of their business.
o Write or state requirements to the developers so that they can take care of the risks incurred when operating the system.
• The developers. They need help to implement the system so that: o It can be made business-safe.
o They can support their claims with analysis and documentation.
o It is possible to change the systems in such a way that when the operating environment or profile changes, the systems are still business-safe.
BUCS aim to help the developers to build a business-safe system without large increases in development costs or schedule. This is achieved by the following contributions from BUCS:
BC1 A set of methods for analysing business-safety concerns. These methods are adapted to the software development process in general and – for the first version – especially to the Rational Unified Process (RUP).
BC2 A systematic approach for analysing, understanding, and protecting against business-safety related events.
BC3 A method for testing that the customers’ business-safety concerns are adequately taken care of in the implementation.
Why should development organizations do something that costs extra, i.e. is this a smart business proposition? We definitively mean that the answer is “Yes”, and for the following reasons:
•••• The only solution most companies have to offer to customers with business-safety concerns today is that the developers will be more careful and test more – this is not a good enough solution.
•••• By building a business-safe system, the developers will help the customer to achieve an efficient operation of their business and thus build an image of a company that have their customers’ interest in focus. Applying new methods to increase the products’ business-safety must thus be viewed as an investment. The return on the investment will come as more business from large, important customers.
BUCS will not invent entirely new methods. What we will do, is to take commonly used methods, especially from the area of systems safety such as Hazard Analysis and FMEA, and adapt them to more mainstream software development. This is done by extending the methods, making them:
• More practical to use in a software development environment.
• Suitable to fit into the ways developers work in a software project environment – concerning both process and related software tools and methods.
1.3 Research design
As stated in the BUCS project proposal, “The principal goal is through empirical studies to understand and improve the software technologies and processes used for developing business-critical software” [BUCS02]. This entails both quantitative and qualitative studies, and in some cases a combination. Several aspects have to be considered when performing such studies, and particularly:
• Deciding on the metrics used in the investigations.
• Deciding on the process of retrieving information (data mining, observation, surveys).
Members of the BUCS project have conducted interviews, experiments, data analysis, surveys, and case studies. The methods employed in this part of the BUCS project are structured interviews, historical data mining and analysis, and case studies.
1.4 Research questions and contributions
The goal of this research is to explore quality issues of business-critical software, with focus on fault reporting and management, as well as the use of safety analysis techniques for this type of software development. In this thesis, four overall research questions have been defined:
RQ1. What is the role of fault reporting in existing industrial software development?
RQ2. How can we improve on existing fault reporting processes?
RQ3. What are the most common and severe fault types, and how can we reduce them in number and severity?
RQ4. How can we use safety analysis techniques together with fault report analysis to improve the development process?
First fault report analysis study
2005
Second fault report analysis study 2006 Assessing hazard analysis vs. fault report analysis 2007 Preliminary Interview Study 2003 Literature study of safety methods 2004 Interviews on fault reports 2007 Experiences on fault reporting 2007
Phase 1 Phase 2 Phase 3
C P June 2007 June 2003 Quantitative study Qualitative study Contribution Paper Input P1 P2 P3 P4 P7 P6 C1 C1 C2 C2 C2 C1 C2 C3 C3 C3 C3 P5 Industrial cooperation Study 1 Study 2 Study 3 Study 4 Study 5 Study 7 Study 6
Figure 1-1 The studies with their related papers and contributions
The research questions together with the studies performed have resulted in the following contributions:
C1. Describing how to utilize safety criticality techniques to improve the development process for business-critical software.
C2. Identification of typical shortcomings in fault reporting.
C3. Improved model of fault origins and types for business-critical software. Figure 1-1 illustrates how the studies, contributions and research papers are connected. It also shows the time and sequence of the studies and how the different studies have influenced each other with input and experience. The background cloud shows which studies were performed with industrial cooperation.
1.5 Included research papers
This thesis includes seven papers numbered P1 to P7, whose full text is included verbatim in Appendix A. The papers are briefly described in the following:
P1. Jon Arvid Børretzen, Tor Stålhane, Torgrim Lauritsen, and Per Trygve Myhrer:
"Safety activities during early software project phases", In Proc. Norwegian Informatics Conference (NIK'04), pp. 180-191, Stavanger, 29. Nov. - 1. Dec. 2004.
Relevance to the thesis: This paper describes the introduction and use of safety
criticality analysis techniques in early project phases. It presents several relevant techniques and how they can be combined with a common development methodology like RUP.
My contribution: I was the leading author and contributed 80% of the work, including
literature review and paper writing.
P2. Jon Arvid Børretzen and Reidar Conradi: "A study of Fault Reports in
Commercial Projects", In Jürgen Münch and Matias Vierimaa (Eds): Proc. 7th International Conference on Product Focused Software Process Improvement (PROFES'2006), pp. 389-394, Amsterdam, the Netherlands, 12-14 June 2006.
Relevance to the thesis: This paper presents work done in the area of fault report
analysis, and describes how using a fault categorization scheme can help identify problem areas in the development process.
My contribution: I was the leading author and contributed 80% of the work, including
research design, data collection, data analysis and paper writing.
P3. Parastoo Mohagheghi, Reidar Conradi, and Jon A. Børretzen: "Revisiting the
Problem of Using Problem Reports for Quality Assessment", In Kenneth Anderson (Ed.): Proc. the 4th Workshop on Software Quality, held at ICSE'06, 21 May 2006 - as part of Proc. 28th International Conference on Software Engineering & Co-Located Workshops, 21-26 May 2006, Shanghai, P. R. China, ACM Press 2006, ISBN 1-59593-085-X, ISSN 0270-5257
Relevance to the thesis: This paper describes experience with working with problem
reports from industry. It discusses several problems with using this type of data and how they can be used for assessing software quality.
My contribution: I contributed on 30% of the work, including data collection and
analysis, commenting the data material and draft paper.
P4. Jon Arvid Børretzen and Jostein Dyre-Hansen: Investigating the Software Fault
Profile of Industrial Projects to Determine Process Improvement Areas: An Empirical Study, Proceedings of the European Systems & Software Process Improvement and
Innovation Conference 2007 (EuroSPI07), pp. 212-223, Potsdam, Germany, 26-28 Sept.
2007.
Relevance to the thesis: This paper continues the fault report study focus, refining the
design and execution of the previous study and confirming several of our findings.
My contribution: I was the leading author and contributed 80% of the work, including
P5. Jingyue Li, Anita Gupta, Jon Arvid Børretzen, and Reidar Conradi: "The
Empirical Studies on Quality Benefits of Reusing Software Components" Proc. The First IEEE International Workshop on Quality Oriented Reuse of Software (QUORS'2007), held in conjunction with IEEE COMPSAC 2007, 5 p, Beijing, July 23-27, 2007.
Relevance to the thesis: This paper uses data from our first fault report study, and
presents a study where defect types are compared in reusable components with non-reusable components.
My contribution: I contributed 20% of the work, including theory definition, data
collection, data analysis and commenting on data material, results and draft paper.
P6. Jon Arvid Børretzen: “Fault classification and fault management: Experiences
from a software developer perspective”. 14 pages, submitted to Journal of Systems and Software.
Relevance to the thesis: This paper presents findings from a series of interviews
performed with developers involved in fault reporting, and seeks to describe problems and issues in fault management and reporting as seen from the practitioners’ viewpoint.
My contribution: I contributed 95% of the work, including interviews, transcription,
coding, analysis and paper writing.
P7. Jon Arvid Børretzen: “Using Hazard Identification to Identify Potential Software
Faults: A Proposed Method and Case Study”. 10 pages, submitted to the First International Conference on Software Testing, Verification and Validation, Lillehammer, Norway, April 9-11, 2008.
Relevance to the thesis: This paper seeks to combine the knowledge gained from fault
reports analysis with the potential of hazard analysis techniques, and proposes a novel method for doing this.
My contribution: I contributed 85% of the work, including Hazard analysis, Fault
report analysis, data analysis and paper writing.
1.6 Thesis structure
Chapter 2 deals with issues about software engineering in general and state-of-the-art, including software criticality and especially business-critical software and an overview of the most important challenges in these areas. Chapter 3 presents the context for the BUCS project and research, the methods used and the research questions for this thesis. Chapter 4 presents the results of the studies performed. An evaluation of the contributions and results are made in Chapter 5. Chapter 6 sums up the thesis work and present relevant issues for further work. Figure 1-2 illustrates how the thesis is composed.
Theory and state-of-the-art Chapter 2 Research Context and design Chapter 3 Conclusions and future work Chapter 6 Evaluation Chapter 5 Results Chapter 4
Figure 1-2 The structure of this thesis
In this thesis I have used the term “we” when presenting the work, both when presenting my description of the work in Chapters 1-6 and in the collaborative work from the papers P1 to P7 – in Appendix A.
2
State-of-the-art
This chapter describes the challenges in software engineering that are the motivation behind improving approaches for business-critical software development. Then, there is a presentation of literature related to business-critical software development. The definitions of these subjects are discussed and research challenges are described for each of them. Finally, the chapter is summarized and the research challenges are described related to the studies in this thesis.
2.1 Introduction
In the engineering of business-critical software systems, as in the engineering of other software systems, there is a multitude of different methods, techniques and processes being employed by industry. Since business-critical applications are not really an established phrase or topic within the software engineering community, it is therefore difficult to point out specific methods and techniques that are being used when business-critical systems are being developed. Instead, the most common methods of software engineering will be presented in the following, with additional comments on how they may be best utilized to aid the development of business-critical systems. Also, a presentation of methods from the development of safety-critical applications will be made as these are relevant for use in the BUCS project in general.
2.2 Software engineering
Software Engineering is an engineering discipline dealing with all aspects of software development from the early stages of system specification to maintaining the system after it has gone into use. Software engineering is the profession concerned with creation and maintenance of software by applying computer technology, project management, domain knowledge, and other skills and technologies. Fairley says that:
"Software engineering is the technological and managerial discipline concerned with systematic production and maintenance of software products" of required functionality
and quality "on time and within cost estimates" [Fairley85].
On the other hand, software systems have social and economic value, by making people more productive, improving their quality of life, and making them able to perform work and actions that would otherwise be impossible, like controlling a modern aeroplane. Software engineering technologies and practices help developers by improving productivity and quality. The field has been evolving continuously from its early days
in the 1940s until today in the 2000s. The ongoing goal is to improve technologies and practices, seeking to improve the productivity of practitioners and the quality of applications for the users.
The effort in software engineering technology was stepped up due to the “software crisis” (a term coined in 1968), which identified many problems of software development [Glass94]. Many software projects ran over budget and schedule. Some projects caused property damage, and a few projects actually caused loss of life. The software crisis was originally defined in terms of productivity, but evolved to emphasize quality. The most common result of failed software development projects are projects that overrun their schedule and budget, but more serious consequences may also be the result of poorly executed software projects.1
Cost and Budget Overruns: A survey conducted at the Simula Research Laboratory in
2003 showed that 37% of the investigated projects used more effort than estimated. The average effort overrun was 41%, with 67% in projects with a public client, and 21% for projects with a private client [Moløkken04].
Property Damage: Software defects can cause property damage. Poor software
security allows hackers to steal identities, and defective control systems can damage the physical systems the software is controlling. The result is lost time, money, and damaged reputation. The expensive European Ariane 5 rocket exploded on its virgin voyage in 1996, because its software operated under different flight conditions than the software was designed for [Kropp98].
Life and Death: Defects in software can be lethal. Some software systems used in
radiotherapy machines failed so gravely that they administered lethal doses of radiation to patients [Leveson95].
The use of the term “software crisis” has been slowly fading out; perhaps because the software engineering community have come to the understanding that it is unrealistic and unproductive to remain in crisis mode for this many years. Software engineers are accepting that the problems of software engineering are truly difficult and only hard work over a long period of time can solve them. Processes and methods have become major parts of software engineering, e.g. object-orientation (OO) and the Rational Unified Process (RUP). Studies have however shown, that many practitioners resist formalized processes, which often treats them impersonally like machines, rather than creative people [Thomas96]. The profession of software engineering is important, and has made big improvements since 1968, even though it is not perfect. Software engineering is a relatively young field, and practitioners and researchers continually work to improve the technologies and practices, in order to improve the final products and to better comply with the needs of the users and customers.
2.3 Software Quality
The word quality can have several meanings and definitions, even though most of these definitions try to communicate practically the same idea. Often, the context in which the quality is to be judged, decides which definition that will be used. The context could be user-orientated, product-oriented, production-oriented or even emotionally oriented. ISO defines quality as “The totality of features and characteristics of a product/service
that bears upon its ability to satisfy stated or implied needs” [ISO 8402]. Another ISO
definition is “Quality: ability of a set of inherent characteristics of a product, system or
process to fulfil requirements of customers and other interested parties” [ISO 9001].
Aune presents the following simplified definitions from the ISO 8402 standard [Aune00]:
1. Quality: Conformity with requirements (or needs, expectations, specifications) 2. Quality: The satisfaction of the customer
3. Quality: Suitability for use (at a given time)
Software quality in terms of reliability is often related to faults and failures, e.g. in number of faults found, or failure rate over a period of time during use. Added to this, as the before mentioned definitions imply, there are other quality factors that are important, e.g. the software’s ability to be used in an effective way (i.e. its usability). There is a multitude of concepts that together can be used to define quality, where the importance of a given factor or characteristic depends on the software context. Reliability, Usability, Safety, Security, Availability and Performance are common examples. The glossary in Appendix A describes some of the relevant quality attributes.
2.3.1 Software Quality practices
Quality Assurance (QA)
QA is the planned and systematic efforts needed to gain sufficient confidence in that a product or a service will satisfy stated requirements to quality (e.g. degree of safety/reliability). Alternatively, QA is control of product and process throughout software development, so that we increase the probability that we manage to fulfil the requirements specifications. Software QA involves the entire software development process, monitoring and improving the process, making sure that any agreed-upon standards and procedures are followed, and ensuring that problems are found and dealt with. QA work is oriented towards problem “prevention”. Solving problems is a high-visibility process; preventing problems is low-high-visibility.
Among the duties of a QA team are certification and standardization work, as well as internal inspections and reviews. Other relevant QA tasks are inspections, testing, verification and validation, some of which are presented further in section 2.5.4.
Software Process Improvement (SPI)
Software Process Improvement is basically systematic improvement of the work processes used in a software-producing organization, based on organizational goals and
backed by empirical studies and results. Capability Maturity Model Integration (CMMI) and ISO 9000 are examples of ways to assess and certify software processes. Statistical Process Control (SPC) and the Goals/Question/Metric (GQM) paradigm are examples of methods used to implement Software Process Improvement [Dybå00], but these require a certain level of stability in an organization to be applicable.
To be able to measure improvement, we have to introduce measurement into software development processes. SPI initiatives are generally based on measurement of processes, followed by results and information feedback into the process under study. The work in this thesis is directed towards measurement of faults in software, and how this information may be used to improve the software process and product.
2.4 Anomalies: Faults, errors, failures and hazards
Improving software quality is a goal of most software development organizations. This is not a trivial task, and different stakeholders will have different views on what software quality is. In addition, the character of the actual software will influence what is considered the most important quality attributes of that software. For many organizations, analyzing routinely collected data could be used to improve their process and product quality. Fault reports is one possible source of such data, and research shows that fault analysis can be a viable approach to certain parts of software process improvement [Grady92]. One important issue in developing business-critical software is to remove possible causes for failure, which may lead to wrong operations of the system. In our studies we will investigate fault reports from business-critical industrial software projects.
Software quality encompasses a great number of properties or attributes. The ISO 9126 standard defines many of these attributes as sub-attributes of the term “quality of use” [ISO91]. When speaking about business-critical systems, the critical quality attribute is often experienced as the dependability of the system. In [Laprie95], Laprie states that “a
computer system’s dependability is the quality of the delivered service such that reliance can justifiably be placed on this service.” According to [Avizienis04] and
[Littlewood00], dependability is a software quality attribute that encompasses several other attributes, especially reliability, availability, safety, integrity and maintainability2.
The term dependability can also be regarded subjectively as the “amount of trust one has in the system”. Quality-of-Service (QoS) is the dependability plus performance, usability and certain provision aspects [Emstad03].
Much effort has been put into reducing the probability of software failures, but this has not removed the need for post-release fault-fixing. Faults in the software are detrimental to the software’s quality, to a greater or lesser extent dependent on the nature and severity of the fault. Therefore, one way to improve the quality of developed software is to reduce the number of faults introduced into the system during initial development.
Faults are potential flaws (i.e. incorrect versus explicitly stated requirements) in a
software system, that later may be activated to produce an error (as incorrect internal dynamic state). An error is the execution of a "passive fault", and my lead to a failure (for incorrect external dynamic state). This relationship is illustrated in Figure 2-1. A
failure results in observable and incorrect external behaviour and system state. The
remedies for errors and failures are to limit the consequences of an active error or failure, in order to resume service. This may be in the form of duplication, repair, containment etc. These kinds of remedies do work, but studies have shown that this kind of downstream (late) protection is more expensive than preventing the faults from being introduced into the code [Leveson95].
Figure 2-1 Relationship between faults, errors, failures and reliability
Faults that unintentionally have been introduced into the system during some lifecycle phase can be discovered either by formal proof or manual inspections before the system is run, by testing during development or when the application is run on site. The discovered faults are then reported in some fault reporting system, to be candidates for later correction. Software may very well have faults that do not lead to failure, since they may never be executed, given the actual context and usage profile. Many such faults will remain in the system unknown to the developers and users. That is, a system with few discovered faults is not necessarily the same as a system with few faults. Indeed, many reported faults may be deemed too “exotic” or irrelevant to correct. Inversely, a system with many reported faults may be a very reliable system, since most relevant faults can have been eliminated. Faults are also commonly known as defects or
bugs, while a more extensive concept is anomaly, used in the IEEE 1044 standard
[IEEE 1044].
The relationship between faults, errors and failures concerns the reliability dimension. If we look at the safety dimension, we have a relationship between hazards and accidents. A hazard is a state or set of conditions of a system or an object that, together with other conditions in the environment of the system or object, may lead to an accident (safety dimension) [Leveson95]. Leveson defines an accident as “an undesired and unplanned
(but not necessatily unexpected) event that results in at least a specified level of loss.” Fault (static) Potential flaw, erroneous program. Error (dynamic) Erroneous internal system state. Failure (dynamic) Erroneous external behaviour. Reliability
The connection between hazards and safety is defined through Leveson’s definition of safety: “Safety is freedom form accidents or losses”. Figure 2-2 illustrates this relationship.
Figure 2-2 Relationship between hazards, accidents and safety
To reduce the chances of critical faults existing in a software system, the latter should be analyzed in the context of its environment and operation to identify possible hazardous events [Leveson95]. Hazard analysis techniques like Failure and Effect Analysis (FMEA) and Hazard and Operability Study (Hazop) can help us to reduce the product risk stemming from such accident. Hazards encompass a greater scope than faults, because a system can be prone to many hazards even if it has no faults. Hazards are related to the system’s environment, not just to the software itself. Therefore they may be present even though the system fulfils the requirements specifications completely, i.e. has no faults.
The full lines in Figure 2-3 show the common view of how faults are related to reliability and hazards are related to safety. In parts of the thesis we also suggest that faults may influence safety and hazards may influence reliability, as shown by the dotted lines. Literature searches shows little work that have been done in this specific area, but the fact that faults and hazards do share some characteristics make plausible connections between faults and safety and hazards and reliability also, at least from a pragmatic viewpoint.
Avizienis et al. emphasize that fault prevention and fault tolerance aim to provide the ability to deliver a service that can be trusted, while fault removal and fault forecasting aim to reach confidence in that ability, by justifying that the functional, dependability and security specifications are adequate, and that the system is likely to meet them [Avizienis04]. Hence, by working towards techniques that can prevent faults and reduce the number and severity of faults in a system, the quality of the system can be improved in the area of reliability (and thus dependability).
Hazards (static) Potential negative event Accident/loss (dynamic) Negative effect of event occuring Safety
Figure 2-3 Faults, Hazards, Reliability and Safety
A usual course of events leading to a fault report is that someone reports a failure through testing or operation, whereupon a report is logged. This report could initially be classed as a failure report, as it describes what happened when the system failed. As developers examine the report, they will eliminate reported “problems” that were not real failures (often caused by wrong user commands) or duplicates of previously reported ones. Primarily, they work to establish what caused the failure, i.e. the originall fault. When they identify the fault, they can choose to repair the fault and report what the fault was and how it was repaired. The failure report has thus become a fault report. When looking at a large collection of fault/failure reports in a system in testing or operation, some faults have been repaired, while others have not (and may never be). Still, we choose to refer to a report of a software failure as a fault report, even if the fault has not yet been identified, since it is stored with the other fault reports, and work is usually being done to identify the fault that caused the failure.
2.4.1 Reflection and challenges
As stated in Section 1.1, the terminology from the literature, although clear and concise in each individual field and source, gets confusing and conflicting when you compare definitions. In our work, we have not tried to redefine the terms and definitions to make them smoothly fit together, we merely want to explain some of our understanding about faults and fault reporting, to the degree it is relevant for the thesis.
We still see a need for work unifying concepts, especially in the reliability area. There is great diversity in the literature on the terminology used to report software or system related problems. The possible differences between problems, troubles, bugs, anomalies, defects, errors, faults or failures are discussed in books (e.g., [Fenton97]), standards and classification schemes such as IEEE Std. 1044-1993 [IEEE 1044] the United Kingdom Software Metrics Association (UKSMA)’s scheme [UKSMA], and papers; e.g., [Freimut01]. Until there is agreement on the terminology used in reporting problems, we must be aware of these differences and answer the above questions when using a term.
Faults
Hazards
Reliability
2.5 Current methods and practices
2.5.1 General software engineering paradigms
In software engineering there have been many different paradigms or life-cycle models. The most common and well known paradigms are presented in the following.
The traditional software process (waterfall): The waterfall model was the first widely
used software development model. It was first proposed in 1970 by W. W. Royce [Royce70], in which software development is seen as flowing steadily through the phases of requirements analysis, design, implementation, testing (validation), integration and maintenance. In the original article, Royce advocated using the model repeatedly, in an iterative way. However, many people do not know that, and some have unjustly discredited this paradigm for real use. In practice, the process rarely proceeds in a purely linear fashion. Iterations, by going back to or adapting results of previous stages, are common.
The spiral model: The spiral model was defined by Barry Boehm [Boehm88], and
combines elements of both design and prototyping in stages, so it's a mix of top-down and bottom-up concepts. This model was not the first model to discuss iteration, but it was the first model to explain why iteration is important. As originally envisioned, the iterations were typically 6 months to 2 years long. This persisted until around 2000. Increasingly, development has turned towards shorter iteration periods, because of higher time-to-market demand. In her doctoral thesis, Parastoo Mohagheghi reports iterations of 2-3 months being common [Mohagheghi04b].
Prototyping, iterative and incremental development: The prototyping model is a
software development process that starts with (incomplete) requirements gathering, followed by prototyping and user evaluation. Often the customer/user may not be able to provide a complete set of application objectives, detailed input, processing, or output requirements at the start. After user evaluation, another prototype will be built based on feedback from users, and again the cycle returns to customer evaluation.
Agile methods: The benefits of agile methods for small teams working with rapidly
changing requirements have been documented [Beck99]. However, both by proponents and critics, the applicability of agile methods to larger projects is hotly debated. Large-scale projects, with high QA requirements, have traditionally been seen as the home-ground for plan-driven software development methods. Deciding when to use agile methods also depends on the values and principles that a developer wishes to be reflected in her/his work. Extreme Programming (XP) [Beck99], one of the more popular of the agile methods, is explicit in its demand for developers to follow a "code of software conduct" that transmits these values and principles to the project at-hand. In keeping with the philosophy of agile methods, there is no rigid structure defining when to use any particular feature of these approaches(!).
2.5.2 Software Reuse
Reuse in software development is a term describing development that includes
systematic activities for creation and later incorporation ("reuse") of common, domain-specific artifacts. Reuse can lead to profound technological, practical, economic, and legal obstacles, but the benefits may be substantial. It mostly concerns program artifacts in the form of components. In the SEI’s report [Bachmann00] on technical aspects of CBD, a component is defined as:
• An opaque implementation of functionality. • Subject to third-party composition.
• Conformant to component model.
Software development that systematically develops domain-specific and generalized software artifacts for possible, later reuse is called software development for reuse. Software development that systematically makes use of such pre-made, reusable artefacts, is called software development with reuse.
Component-based software engineering (CBSE)
Component-based software engineering is a field of study within software engineering, building on prior theories of software objects, software architectures, software frameworks and software design patterns, and on extensive theory of object-oriented (OO) programming and design of all these. It claims that software components, like the idea of a hardware component used e.g. in telecommunication systems, can be ultimately made interchangeable and reliable. CBSE is often said to be mostly software development with reuse, and with emphasis on reusing components developed outside the actual project.
Commercial Off-The-Shelf (COTS)
COTS components are external executable software components being sold, leased, or licensed to the general public; offered by a vendor trying to profit from it; supported and evolved by the vendor, and used by the customers normally without source code access or modification ("black box"). Different ways of incorporating COTS-based activities is described by Li et al. in [Li06].
Open Source Software (OSS)
Open Source Software is software released following the principles of the open source movement. In particular, it must be released under an Open Source license as defined by the Open Source Definition, with there being over 50 license types. The Open Source movement is a result of the free software movement, that advocates the term "Open Source Software" as an alternative term for free software, and primarily makes its arguments on pragmatical rather than philosophical grounds. Nearly all Open Source Software is also "Free Software". An OSS component is an external component for which the source code is available ("white box"), and the source code can be acquired either free of charge or for a nominal fee, and with a possible obligation to report back any changes done.
2.5.3 Specific software development methods
The two following methods are well-known and commonly used in software development.
Rational Unified Process (RUP)
The Rational Unified Process (RUP) is a software process, design and development method created by the Rational Software Corporation [Rational], and is described in [Kruchten00] and [Kroll03]. It describes how to effectively deploy software using commercially proven techniques. It is really a heavyweight process, and therefore particularly applicable to larger software development teams working on large projects. It is essentially an incremental development process which centers on the Unified Modelling Language (UML) [Fowler04]. It divides a project into four distinct phases; Inception, Elaboration, Construction and Transition. Figure 2-4 shows the overall architecture of the RUP.
Figure 2-4 The Rational Unified Process
Patterns and Architecture-driven methods
Design patterns are recurring solutions to problems in object-oriented design. The phrase was introduced to computer science in the 1990s by the text “Design Patterns: elements of reusable object-oriented software” [Gamma95]. The scope of the term remained a matter of dispute into the next decade. Algorithms are not thought of as design patterns, since they solve implementation problems rather than design problems. Typically, a design pattern is thought to encompass a tight interaction of a few classes and objects. Three major terms have been proposed: pattern languages, pattern catalogs and pattern systems [Riehle96].
The architect Christopher Alexander's work on a pattern language, for designing buildings and communities, was the inspiration for the design patterns of software [Price99]. Interest in sharing patterns in the software community has led to a number of books and symposia. The goal of the pattern literature is to make the experience of past designers accessible to beginners and others in the field. Design patterns thus presents
different solutions in a common format, to provide a language for discussing design issues.
2.5.4 Techniques for increasing trust in software systems
In addition to the general practices of QA and SPI for improving quality in software systems, there are also some specific verification techniques that are commonly used in software development to increase the trust in software. Software verification is a discipline whose goal is to assure that software fully satisfies all the expected requirements, and the following are some well known techniques in use:
Testing: Dynamic verification is performed during the execution of software, and
dynamically checks its behaviour; it is commonly known as testing. Testing is part of more or less all software development processes, and can be performed at many levels, for instance unit level, interface level and system level.
Inspections: An inspection is also a very common sort of review used in software
development projects. The goal of the inspection is for all of the inspectors to reach consensus on a work product and approve it for use in the project. Commonly inspected work products include software requirements specifications, design documentation and test plans. In an inspection, a work product is selected for review and a team is gathered for an inspection meeting to review the work product. In an inspection, a defect is any part of the work product that will keep an inspector from approving it. For example, if the team is inspecting a software requirements specification, each defect will be text in the document which an inspector disagrees with. Basili et al. describes an investigation of an inspection technique called perspective-based testing in [Basili00].
Formal methods: Formal methods are mathematically-based techniques for the
specification, development and verification of software and hardware systems. The use of formal methods for software and hardware design is motivated by the expectation that, as in other engineering disciplines, performing appropriate mathematical analyses can contribute to the reliability and robustness of a design. However, the high cost of using formal methods means that they are usually only used in the development of high-integrity systems, where safety or security is important. Heimdahl and Heitmayer present some issues concerning formal methods in [Heimdahl98].
2.5.5 Business-Critical computing and related terms
At first glance, there is little evidence of work on business-critical computing, when searching the literature. The term “mission-critical” is much more commonly used, and can be interpreted to include many of the characteristics of “business-critical”. The key similarity is that both terms are related to the core activity of an organization, and that the computer systems supporting this activity should not fail. Another term that comes from Software Engineering Institute (SEI) is “performance-critical” [SEI], and has much of the same meaning as “business-critical”.
“Safety-critical” systems are closely connected to these former terms, but this term has a more severe meaning. Nonetheless, most of the main characteristics of these terms are the same; i.e. that reliability, availability and similar quality attributes are deemed very important. Safety-critical systems have been much more thoroughly researched than the other types of “-critical” systems, simply because of the seriousness of failure and the potential effects of failure in safety-critical systems.
2.6 Business-critical software
As mentioned, our societies’ dependency on timely and well-functioning software systems is increasing. Banking systems, train control systems, airport landing systems, automatic teller machines and industrial process control systems are but examples of the systems many of us are directly or indirectly critically dependent on. Of these, some are highly critical to our safety (e.g. traffic control), while others are critical only in the sense that we are able to perform operations that we want or need to carry out our work/business (for instance cinema ticket sales).
That a software-intensive system is business-critical means that:
If and when a system failure occurs, the consequences are restricted to financial or financially related negative implications, not including physical harm to humans, animals or physical objects. The consequences are severe enough to mean a considerable loss of money if the fault or failure is not corrected or averted swiftly enough.
2.6.1 Criticality definitions
Business-critical software systems have a lot in common with safety-critical systems, but there are also quite telling differences. A simplistic way to distinguish them is to put them into classes according to the effects that software anomalies (faults or hazards) may have on the environment. The classes are safety-critical, mission-critical, performance-critical, business-critical, and non-critical:
Safety-critical: A safety-critical system could be a computer, electronic or
electromechanical system where a hazardous event may cause injury or even death to human beings, or physical harm to other objects that interact with the system. Examples are aircraft control systems and nuclear power-station control systems, where an accident in most cases will lead to economic losses as well as injury and other physical damage. Common tools to design safety-critical systems are redundancy and formal methods, and a spectrum of specialized technologies exist for safety-critical systems (Hazop, Fault-tree analysis etc). The IEC 61508 standard is intended to be a basic functional safety standard applicable to all kinds of industry, and is also used to define the safety standards of some safety-critical systems [IEC 61508].
Mission-critical: The term mission-critical system reflects military usage and is used to
describe activities, processing etc., that are deemed vital to the organization's business success and, possibly, its very existence. Some major software systems are described as mission-critical if such a system, product or service experiences a failure or is otherwise unavailable to the organization, it will have a significant negative impact upon the organization. Such systems typically include support for accounts/billing, customer balances, computer-controlled machinery and production lines, just-in-time ordering, and delivery scheduling. Examples of related technologies are Enterprise Resource Planning tools, such as SAP [SAP].
Performance-critical: The SEI defines performance-criticality as the ability of
software-intensive systems to perform successfully under adverse circumstances, e.g., under heavy or unexpected load or in the presence of subsystem failures. One trivial example of this is the performance of the SMS telecom services during New Years Eve. Some services like this can have critical functions, and yet, the behaviour of systems under such circumstances is often less than acceptable [SEI].
Business-critical: The difference between a business-critical and a regular commercial
software system is really defined by the business. There is no established general definition telling us which software applications are critical to an operation. In a retail business, a Customer Relationship Management (CRM) system may be the most important. On the other hand, it may be the manufacturing or supplier management software that is the most important. We need to consider the impact of relevant services from software on the business operations, and determine how much value each brings to the business and the impacts of such software parts being unavailable. The impact can be lost revenue, corrupted data or lost user time, as well as indirect and more elusive losses in customer reputation, goodwill, slipped deadlines, and increased levels of stress among employees and customers.
Non-critical: Although important enough, some types of software will simply not be
classified as critical. Word processors, spreadsheets and graphical design software are examples of such software. Of course it is expected that such tools are reasonably fault-free and stable, but should they fail, the damage will usually be limited, typically a person-day of effort in the worst case scenario.
Figure 2-5 shows the relationship between business-criticality and the other types of criticality defined here. As we see, safety-, performance-, and mission-critical systems can also be business-critical, but a business-critical system need not be one of the others. Table 2-1 illustrate the overlap between the different categories.
Figure 2-5 Relationship of business-critical and other types of criticality
Table 2-1 Examples of different systems’ criticality
Criticality category
Example
Safety-critical Nuclear reactor control system.
Performance-critical Electronic toll collection in traffic, must process and transfer information quickly enough to keep up with traffic.
Mission-critical Software handling financial transactions between banks. Functional and non-functional aspects of such applications are considered.
Business-critical Software handling financial transactions between banks. As mission-critical, but wider consequences are also considered. Non-critical Computer games, word processor application.
2.7 Techniques and methods used to develop safety-critical
systems
There are a number of methods and techniques that are commonly employed when making safety-critical systems. Some of them will be presented here and related to business-critical computing. According to [Leveson95] and [Rausand91], the most common ones are the following:
o PHA (Preliminary Hazard analysis): Preliminary Hazard Analysis (PHA) is used
in the early project life cycle stages to identify critical system functions and broad system hazards, so as to enable hazard elimination, reduction or control further on in the project. The identified hazards are assessed and prioritized, and safety design