Better risk assessment can identify root causes for potential catastrophes before they occur
STANDARDIZED APPROACH TO RISK
An RA tool was constructed to evaluate the risk represented by process releases resulting from catastrophic pump failures. The guideline was developed to be consistent with, and borrows heav-ily from, the approach defined in API Publication 581 Risk-Based TABLE 1. No. 2 fuel oil pump failure history
Failure Date Problem
1 Aug. 18, 1993 Leaking head
2 Sept. 3, 1993 Failed coupling
3 Jan. 27, 1994 Failed coupling
4 Feb. 22, 1994 Leaking head
5 Nov. 22,1994 Cavitating noise
6 April 27, 1998 Leaking seal
7 May 29, 1998 Failed coupling
8 Nov. 20, 1998 Failed thrust bearings
9 April 6, 1999 Leaking seal
10 March 30, 2000 Leaking seal
11 Aug. 3, 2000 Leaking seal
12 Feb. 20, 2002 Failed coupling
13 Oct. 21, 2002 Failed thrust bearings 14 Sept. 9, 2003 Failed thrust bearings
PLANT/PROCESS OPTIMIZATION SPECIALREPORT
48
I
JUNE 2011 HydrocarbonProcessing.comInspection (RBI) Base Resource Document.4 RBI is a widely accepted method currently practiced across the refining industry.
Although API 581 applies primarily to fixed equipment, the approach has many parallels that apply to failure RA for rotating machines. Accordingly, the standard RBI components are supple-mented with data, methods and tools more specific to centrifugal pumps when needed.
A standardized RA approach reduces the inconsistency that different PHA teams may encounter at different times. More importantly, a standardized approach adds value by connecting the reliability of a specific pump installation to process safety risk tolerance. The benefit comes from determining a realistic target for the MI program to achieve, instead of motivating reliability professionals to achieve their safety goals with nonspecific targets like “work harder,” or “do better” or “fail less.” Setting a tangible reliability target allows a responsible decision to be made as to whether or not risk tolerance can be achieved through the MI program alone. If the MI program cannot realistically achieve the desired level of risk control, then additional layers of protection must be added to manage the risk to an acceptable level.
In some cases, the MI program may adequately drive risk to an acceptable level without requiring any additional safeguards or improvements. At such a time, the PHA team has a basis to conclude that no further actions are needed to mitigate the poten-tial hazards associated with a catastrophic pump failure. In short, the existing safeguards have been evaluated and are considered adequate. The process of evaluating the potential risk associ-ated with catastrophic pump failures begins with determining an acceptable level of risk. This prevents the PHA process from defeating its purpose by creating action items that consume
avail-able resources that should be working on resolving more impor-tant process safety risks.
Method overview. Fig. 3 shows the basic process used to evaluate the risk represented by a catastrophic centrifugal pump seal failure. The analysis begins with a technical pump risk assess-ment—risk-based pump analysis ( RBPA). This step is performed according to the consequence analysis and likelihood analysis methods described in API 581 Sections 7 and 8. Afterward, a quantitative layer of protection analysis (LOPA) is used to com-pare the specific pump risk against an acceptable risk tolerance.
This makes it possible to develop a reliability plan to operate the pumps within risk tolerance.
Catastrophic pump-seal failure and consequences.
OSHA data from 1992 to 2009 contains a record of 36 cata-strophic releases of highly hazardous chemicals that resulted in fatalities.5 These incidents are responsible for 52 fatalities and 250 employee injuries. Ninety-eight of these injuries were severe enough to require hospitalization. One of these incidents involved a process release that occurred while steaming-out a pump casing.
The pump casing split open, resulting in a hot oil release that immediately exploded (Jan. 19, 2005, Kern Oil Refinery, Bakers-field, California). The conditions present during this failure are similar to those that the CSB documents in the first case history.
However, none of the fatal incidents contained in the OSHA database resulted from a pump-reliability issue.
It would not be responsible to conclude that a catastrophic pump seal failure could not result in a fatality based on these his-torical statistics. The second case history illustrates the potential for pump-failure mechanisms to be directly involved in a process safety incident capable of causing severe consequences. Although there is insufficient data for a straightforward fatality frequency calculation, enough statistical information exists to estimate a minimum frequency based on site-specific data and industry aver-ages. A frequency/consequence diagram, such as the one shown in Fig. 4, can be constructed using this information along with these facts and assumptions:
• A total estimated 2009 refining capacity of 17.67 million bpd6
• The relationship of approximately one fire for every one thousand repairs, as cited by an industry reliability authority. 7,8 This was corroborated by a large US refinery in 2009.
According to this analysis, the frequency for a fatality (highest severity consequence) is estimated to be lower than 1x10-6 (1/1 million) years. This frequency suggests that a fatality caused by a pump-reliability issue is probably more likely than an airline fatal-ity but less likely than other typical US workplace fatalfatal-ity causes.9 Based on the industry workplace fatality statistics contained in the OSHA database, this relative ranking seems reasonable.
This information makes it possible to define risk tolerance.
Risk tolerance (or literally “tolerance to risk”) implies that the choice has been made to operate equipment in a responsible man-Risk-based
pump analysis Risk-based
reliability plan
Catastrophic pump-seal failure risk assessment method overview. Failure per year frequency
b
Catastrophic pump failure frequency/consequence plot.
FIG. 4
PLANT/PROCESS OPTIMIZATION SPECIALREPORT
HYDROCARBON PROCESSING JUNE 2011
I
49ner rather than shutting it down to mitigate a process safety failure risk. Risk tolerance will vary between different organizations. It is a decision that should be made under the direction of legal counsel and supported by industry statistics.
Risk-based pump analysis. Fig. 5 shows the primary steps involved in the RBPA. In the RBPA, results from the conse-quence analysis are combined with the likeli-hood analysis to determine the risk associated with a catastrophic pump failure. Compar-ing actual operatCompar-ing risk against a designated risk tolerance makes it possible to assess risk reduction options that may adequately con-trol the process safety hazard. To be effec-tive, the risk reduction options must directly address the factors governing process safety.
Consequence analysis. The conse-quence analysis is covered extensively in API
581 RBI Base Resource Document Section 7. It is used to cal-culate the release area that would develop upon a loss of process containment caused by a catastrophic equipment failure. In this case, the RBI principles of API 581 Section 7 are being applied to potential releases caused by a catastrophic pump failure. Fig.
6 outlines the recommended approach for working through the consequence analysis using the methods described in API 581 Section 7.
The analysis should be based on a representative fluid and should assume that typical refinery pump service is constantly changing and the process material properties being evaluated may be best described as an estimate of average operating condi-tions over a time period. API 581 breaks process fluids down to a discrete number of representative fluids. This level of detail is sufficient for the consequence analysis.
The flow area for a major leak is represented by an annular area between the shaft sleeve and the closest fixed dimension of the pump casing or packing gland. The OD of the shaft sleeve and the ID of the closest fixed dimension of the pump casing or pack-ing gland are determined from the seal manufacturer’s detailed drawing as illustrated in Fig. 7. These dimensions are then used to calculate a major seal failure leak rate.
Likelihood analysis. The likelihood analysis is described in detail by API 581 RBI Base Resource Document Section 8.
Its purpose is to generate an initiating event frequency for both the major and full bore leak scenarios. The likelihood analysis described in this study makes use of generic initiating event fre-quencies (IEFg) that are based on the empirical data shown in Fig.
8. This figure is based on catastrophic pump failure data placed into the public domain by multiple sources.10–18 This information covers a wide range of leak rates from minor leaks (low severity) to full bore leaks (high severity). The middle area of the chart represents the major leak range.
The likelihood analysis is performed by 1) selecting an appro-priate IEFg based on the analysis represented in Fig. 9 then 2) adjusting the IEFg based on the specific pump’s actual reliability history (MTBFa —see Eq. 1) compared with the standard reliabil-ity of a generic refinery process pump (MTBFg). This adjustment
is made according to Eq. 2, which produces the initiating event frequency for a specific pump installation (IEFa).
(1)MTBF Years of pump repair history
Repairs for a
a
¦
lll pumps in the functional group(2)IEF IEF MTBF
a g MTBFg
a
u
Risk analysis. The risk analysis takes place as the LOPA that assesses the pump operating risk against the designated risk toler-ance. Its purpose is to determine if a pump installation meets its reliability expectations. This is true if the frequency of mitigated consequences is less than the designated risk tolerance. If the frequency of mitigated consequences is more than the designated risk tolerance, then guidance should be suggested to improve performance to meet process safety objectives.
Consequence
Calculate and plot the impact area for each scenario:
t.BKPSMFBL t'VMMCPSFMFBL Consequence analysis steps.
FIG. 6
Simplified seal sketch—Major leak flow path.
FIG. 7
PLANT/PROCESS OPTIMIZATION SPECIALREPORT
50
I
JUNE 2011 HydrocarbonProcessing.comProbability of personnel in affected area. The prob-ability of personnel in the affected area, Pp, is a function of the size of the affected area, Aa, and the amount of time personnel are
likely to be in this area. There are causes of catastrophic pump failures that increase the probability of personnel being in the affected area at the time of the event. An example may be an abnormal process condition (such as flow loss) where the console operator calls for the outside personnel to respond. There are also causes that are random in nature where the probability of person-nel in the affected area is based on the average amount of time that people are in the area on any given day.
Failure cause distribution estimates for centrifugal pumps in US process plants indicate that approximately 12% of failures are caused by improper operation.10–19 Some of these causes result from chronic poor operating practices that reduce pump reliability. They may have been normalized over time and do not result in an operator response. An example may include cavitation noises caused by low NPSHa operation or long-term flow outside of recommended reliability limits.20 An estimate of 10% of the causes of major releases that result in increased occupancy of the affected area is assumed for this analysis. The remaining random occupancy that does not increase the probability of personnel in the affected area would therefore be 90%.
An estimated random occupancy of 1 hr/day/1,000 ft2 is assumed for normal process areas. This estimate should be modi-fied if there is evidence of higher or lower occupancy. Remote areas that are not frequented with multiple rounds a shift will be less. Affected areas that include known high-occupancy zones will be greater. Any basis for choosing a different random occupancy should be documented. This random occupancy is further simpli-fied to a probability of 0.04/1,000 ft2. By combining cause gener-ated occupancy with random occupancy an overall probability of personnel in the affected area can be determined by Eq. 3.
Pp = 0.10 + Aa (0.04/1,000 ft2) (3)
Probability of ignition. API 581 reports probability of ignition, Pi, for five potential outcomes in tables. The proper table in API 581 Section 7 should be selected based upon the process leak assessment made during the consequence analysis.
Risk-based reliability plan. The out-put from the RBPA feeds into a process for managing the risk of a process safety failure.
The frequency of mitigated consequences, Fm, is the product of the frequency of the specific pump’s initiating event frequency, IEFa, the total probability of failure on demand for each independent layer of pro-tection, PFDt, the probability of personnel in the affected area and the probability for ignition as calculated in Eq. 4. If the fre-quency of mitigated consequences, Fm, is higher than the designated risk tolerance, then a risk-based reliability plan must be developed to manage the risk for a process safety failure. This can be accomplished by either increasing the pump’s reliabil-ity, MTBF, or by applying safeguards suf-ficient to mitigate the consequences of a catastrophic pump failure to an acceptable
1 Event frequency per year
Release area, in.2
Generic centrifugal pump leak area/frequency.
FIG. 8
Establish safe design objectives (based on LOPA results) t.5#'
t*OEFQFOEFOUMBZFSTPGQSPUFDUJPO
Identify, evaluate and select alternatives (for meeting safe design objectives)
Implement risk-based reliability plan (to meet safe design objectives) Identify alternatives
Design review
Evaluate and select alternatives '.&"
NBKPSMFBLGBJMVSF NPEFT
3FMJBCJMJUZTUSBUFHZ
QMBO 1SPKFDUQMBO "VEJUQMBO 'BJMVSFBOBMZTJT QMBO
Creating a risk-based reliability plan.
FIG. 9
PLANT/PROCESS OPTIMIZATION
level. The basic process used to develop a risk-based reliability plan is shown in Fig. 9.
Fm = IEFa (PFDt)(Pp)(Pi) (4) The LOPA results designate a target MTBF for meeting a des-ignated risk tolerance. MTBF improvements have a number of advantages. For example, they reduce both maintenance costs and the potential to introduce some major leak failure modes dur-ing repairs like the one described in the first case history. MTBF improvements are typically preventive instead of reactive. However, it may be difficult to quantify the expected MTBF improvement available through failure analysis and investigation. Failure analysis skills, training and methods are involved in developing an effective set of corrective actions to increase MTBF. This depends greatly upon the failure investigator’s individual capabilities.
Machinery engineers must be consulted to determine if the MTBF improvement is realistically achievable. Consideration should be given to proven technology and both industry and
per-sonal experience with the process requirements. MTBF improve-ments can be applied together with additional safeguards to meet the overall risk tolerance criteria. If MTBF alternatives are selected as a part of the strategy to meet the risk tolerance, MTBF becomes a part of the process safety risk management for the pump group under consideration. It should be managed with the same diligence and priority as defined by safe operating limits.
Case history in preventive risk mitigation. An investi-gation was used to determine the cause for a series of recurring seal and thrust bearing failures on two heavy vacuum gasoil (HVGO) service pumps operating side-by-side in a refinery vacuum crude TABLE 2. HVGO pump WO history
Date Problem Cost
Dec. 13, 2006 Need to replace inboard and outboard bearings $1,445 April 9, 2007 Replace inboard and outboard pump bearings $1,974
March 31, 2008 Inboard seal leak $4,771
Nov. 1, 2008 Inboard pump seal leaking $3,853
April 7, 2009 Outboard bearing leak $3,487
$15,530
Catastrophic HVGO pump seal failure consequences.
FIG. 10
Select 162 at www.HydrocarbonProcessing.com/RS
PLANT/PROCESS OPTIMIZATION SPECIALREPORT
52
I
JUNE 2011 HydrocarbonProcessing.comunit. The maintenance history of these pumps is shown in Table 2.
The investigation determined that high frequency vibration caused by vortex cavitation suction recirculation (VCSR) was responsible for the low MTBF. Based on this diagnosis, an action item was cre-ated to increase the pumps’ NPSH margin ratio to reduce the cavi-tation forces responsible for excessive stress on the thrust bearings.
Addressing this action item would require either redesigning or replacing the pumps at considerable expense. Based on the result-ing maintenance expenses, other competresult-ing reliability improve-ment projects offered a greater return on investimprove-ment. Therefore, it was decided that the risk of catastrophic HVGO pump failure should continue to be managed by repairs. The repairs were to be triggered by condition monitoring until the higher priority reli-ability improvement projects could be completed.
The potential consequences of HVGO leaks in, vacuum crude unit service are not comforting (Fig. 10). A leak of sufficient size would likely autoignite upon contacting air. The consequence for a catastrophic pump failure represents a potential PSM incident in addition to property damage and business interruption. But condition monitoring seemed to be an acceptable approach to managing the risk for a catastrophic HVGO pump failure based on previous operating history.
Upon developing the RBPA guidance, the HVGO pumps were reevaluated to verify that the reliability strategy was in agreement with refinery risk tolerance. The analysis showed that the pump group was one protective layer short at its present MTBF, and its reliability would have to be increased to at least six years MTBF to operate within refinery risk tolerance. This immediately changed the basis for the project from a reliability improvement opportunity to
a process safety risk mitigation project. The priority of the HVGO pump project was elevated and an execution date was scheduled.
Disclaimer and conclusions. The guideline and method-ology discussed in this article attempts to be generally applicable to all centrifugal pumps. However, good engineering judgment must prevail while applying this guideline. The approach can be modified as appropriate following a recommended peer-review and documenting the technical basis for deviations.
Tolerating repeat failures on machinery that contains potentially hazardous process materials can have disappointing consequences.
However, it is not uncommon for equipment failures to be accepted without comparing actual reliability performance against a desig-nated risk tolerance. In cases where breakdown maintenance is the option selected to manage the risk for catastrophic process releases, a definitive and objective basis is needed to expose a potentially unacceptable process safety hazard before an incident occurs.
A standard RA method can be developed to evaluate pump reliability on the basis of managing its failure frequency sufficiently low to realistically avoid a process safety incident. However, the MI program by itself may not sufficiently elevate equipment reli-ability to a level where process safety consequences can confidently be prevented. Risk tolerance ultimately determines the complete plan needed to fully address a process safety risk. In many cases, a complete plan represents a combination of reliability (MI program) improvements and safeguards (layers of protection).
The RBI method described in API 581 Sections 7 and 8 pro-vides a sound engineering basis to do risk analysis on equip-ment whose failure may represent process safety consequences.
Engineering advanced
© 2011 Chemstations, Inc. All rights reserved. | CMS-322-1 5/11