• No results found

Principles of Big Data Algorithms and Application for Unconventional Oil and Gas Resources

N/A
N/A
Protected

Academic year: 2021

Share "Principles of Big Data Algorithms and Application for Unconventional Oil and Gas Resources"

Copied!
9
0
0

Loading.... (view fulltext now)

Full text

(1)

Principles of Big Data Algorithms and Application for Unconventional Oil

and Gas Resources

Avi Lin, Halliburton

Copyright 2014, Society of Petroleum Engineers

This paper was prepared for presentation at the SPE Large Scale Computing and Big Data Challenges in Reservoir Simulation Conference and Exhibition held in Istanbul, Turkey, 15–17 September 2014.

This paper was selected for presentation by an SPE program committee following review of information contained in an abstract submitted by the author(s). Contents of the paper have not been reviewed by the Society of Petroleum Engineers and are subject to correction by the author(s). The material does not necessarily reflect any position of the Society of Petroleum Engineers, its officers, or members. Electronic reproduction, distribution, or storage of any part of this paper without the written consent of the Society of Petroleum Engineers is prohibited. Permission to reproduce in print is restricted to an abstract of not more than 300 words; illustrations may not be copied. The abstract must contain conspicuous acknowledgment of SPE copyright.

Abstract

This paper presents challenges encountered during reservoir modeling and simulation, while demonstrat-ing that the ultimate solution for these challenges can be approached in a Big Data (BD) framework usdemonstrat-ing a combination of physics- and analytics-based approaches.

Introduction: Insufficient Resources (ISR) Computing and BD

BD, previously the concern of database enthusiasts or a marketing technique among retailers, is currently part of mainstream consciousness, jargon, and vocabulary. BD is profoundly impacting the upstream exploration and production (E&P) industry, particularly unconventional production enhancement (PE) (Chelmis et al. 2012). As it becomes a more significant computational subject, it is necessary to better define BD and treat it carefully in terms of computational methods and algorithms (Abramson et al. 2014), particularly in the oil and gas (OG) industry, which heavily relies on these algorithms to increase productivity. The goal is to better define, precisely understand, and set proper expectations regarding the impact of BD on the unconventional PE OG industry, particularly during drilling and completion phases. According to the classical definition, BD refers to datasets with sizes beyond the ability of typical database software tools to handle (capture, store, manage, and analyze) (Chen et al. 2014). As BD technological advancements are made, the size of these datasets capable of being examined also increases. The important issue is to properly meet or engage the concepts and analytics of BD in the OG space in addition to real problems and challenges currently and in the very near future being considered. Among the major challenges in the unconventional OG operations space BD-related, the following can be determined:

● Help improve exploration.

● Help improve production, optimize extraction, and effectively manage well operations using real-time interpretation of drilling information.

● Help improve reservoir performance by analyzing and acting on information from all collected data, such as microseismicity images, seismic collections, and production systems.

(2)

● Obtain better site specific physical and chemical models. Help optimize operational efficiencies in terms of fluids used, proppants used, and the correct number and location of the hydro fractures. The overall expectations of the OG operators from the BD-like solutions can be defined as the large amount of data associated with a predefined well or job (signals and historical data) and deriving or extracting information and knowledge necessary to properly support the operator’s decisions. Therefore, the issue is to identify the “the set of computational operators (COs)” that, when applied to the data, generate the necessary relevant supporting information.

Traditionally, there are two options for the “CO:”

Physics-related computational operators (PCO):the multiphysics phenomena are well-defined by the various physical domains and their governing partial-differential equations. Examples of these operators include equation systems governing the flow in the wellbore and perforation regions, physics of the fluid streams and proppants in the fractures network and reservoirs, and structural dynamics of the rock blocks, among others. The datasets are necessary for properly setting up the initial conditions and boundary conditions, as well as assisting during tuning the various models’ coefficients, such as those related to the materials’ physical and chemical properties.

Analytics-related computational operators (ACO): this is used for the case where there is no knowledge on the specific formation structure and the relative impact, for example, of the hydro-fractures on the shale, and/or to estimate the chances that it will affect the underground water reservoir in case they exist. The “analytical COs” are more of the data analysis, data mining and data fusion type of algorithms (Loshin 2013). Here, the OG operators’ objectives are often times to find or sense the abnormalities, as they are affecting the operators’ decisions.

Combined computational operators: this is a combination of the previous two. Part of the multiphysics is governed by physical partial differential equations and data analytics.

Therefore, simply stated, the key differences between the ACO and PCO with respect to BD are: ● The ACO attempts to discover trends, rules, and abnormalities in the BD, while the PCO attempts

to obtain additional information based on the provided data and governing physical equations. ● The PCO uses the modeling of the CO in the discrete space, while the ACO mainly uses the

available data, thus making use of the BD concepts. ● Results quality.

X The PCO result quality primarily depends on the discrete mapping of the operators. For ACO, it is commonly agreed upon that, the more data used, the better the results obtained from the analysis; therefore, its quality primarily depends on the amount of data considered.

X Moreover, the quality of the PCO is well-controlled, while the quality of the ACO is not; thus, the effect of not using all of the available data is not well-defined.

● The results for both COs depend on the amount of data, its occupied subspaces, and the way the analytic operators and physics discrete operators are obtained, as well as on the quality of data.

Physics and BD

It is assumed that an elliptic equation for ␾ is being solved over a domain⍀ with a boundary⵲⍀. The value of ␾ is given at each point along ⵲⍀. For engineering purposes, the reservoir domain and its boundaries are mathematically not uniquely defined; thus, for a given task, several elliptic problems must be solved. Considering that ␾is a vector of variables, the amount of data generated is large, particularly when requiring a reasonable confidence level. Additionally, the solution quality is determined by the finite mesh spread over ⍀. To reduce uncertainty, a suitably fine mesh should be used to help reduce the error

(3)

(or uncertainty). This can significantly increase the amount of the information obtained from this simulator. These facts classify the overall solution as a high-level BD-type situation. This BD system contains at least three categories of data— data along the boundary⵲⍀, initial time data, and information derived over the mesh. Mathematically, the low confidence level of the simulated results can be caused by the small amount of data in one of the classes. Thus, the quality of the results depends on the category that contains the smallest dataset.

Analytics and BD

Usually, the analytic algorithms consider all relevant data available to them, either post-job static data or quasi-real-time data. The outcomes of these algorithms are usually a mix of a set of rules or patterns on the data and some additional information. For example, in the OG space, the data fusion algorithms can point out the existence of a fault in the vicinity of a specific fracturing stage based on real-time unstructured data originating from the seismicity measurements, microseismicity images, bottomhole pressure (BHP) values, image logs from the previous stage, and production zone of a nearby well. Additionally, it predicts the proppant distribution along the primary hydraulic fractures during the current and previous stages. The major issues are related to the real requirement of these algorithms to use all of the available data. Usually, it is almost impossible to know whether all this data or which parts of the data are necessary to draw the same rule and additional information. It is almost impossible to assess what the uncertainty levels of the analytics are when less data are used (He et al. 2014). This issue is important when accelerated real-time results are necessary, while using, for example, the waterfall approach. It is also important to exclude part of the data linearly dependent on other parts, even for unstructured data types. For example, for gamma ray measurements and magnetic measurements in close vicinity to the wellbore casing (approximately within 30 ft away from the wellbore), the formation properties inferred from both datasets are oftentimes very similar. Therefore, it can be understood that, in this wellbore vicinity, these two datasets are linearly dependent. Of course, these datasets might be linearly independent in other formation zones.

Problem Definition

Let⍀be an OG reservoir domain with a boundary⭸⍀, where⍀僆ᑬ2(and in general can be inᑬ3). The domain ⍀ can be convex or concave, but it is assumed that it is well connected. There are several parameters characterizing this formation domain to be of the reservoir type, and the objective is to determine the size and domain extent, as well as the quality level of this reservoir, where 100% quality means that all of the characterization parameters are positive to their maximum extent. The qualification of the domain under consideration is performed throughout the job’s workflow.

Before, during, and after the fracturing job, multiple sensors incorporating various technologies collect and record measurements on the formation throughout the job duration, and specifically on the reservoir domain. These technologies include, for example, seismic measurements, microseismic (MS) measure-ments, image logs, gamma rays, and magnetic measurements. In this paper, a set of r variables characterizes the reservoir using the measurementsM ⬅{m1,m2, . . .mr}. The variableTi,i ⫽1, 2, . . . p denotes a typical measurement technology incorporated during the job, with p being the total number of technologies used. Each of these technologies measures or senses at a specific point in time and space, {X,t} and a set of scalar quantities si,j, j ⫽ 1, 2, . . . ni, where niis the number of scalar measurements performed by the technology Ti, andsi⫽{si,j}, where each si,jis one of the members ofM, thussiM. Examples of these measurements are the event strength for MS measurements and pressure values reflected by the magnetic measurements.

The differentrmeasurements under consideration (e.g., pressure, temperature, event strength, and local flow velocity) can be calculated for a typical situation using Table 1.

The set of variables in Table 1 refers to the measurements of the technology Ti. The length of this is a vector that is the number of samples this technology has collected, where each entry

(4)

has 4 ⫹ 2ni entries, and e is the appropriate error component for each of the measurements with this technology.

One of the primary objectives is to calculate the reservoir characteristics and properties based on the values of M given on the boundaries of the reservoir and throughout the reservoir neighborhood. This includes fluid variables (such as velocity, pressure, component concentrations, etc.), which are groups under Class V, and reservoir properties (such as porosity, permeability, and fractures planes), which are grouped under ClassI. Oftentimes, the values for a subset of the variables inIcan be obtained by applying some nonlinear operators on a subset of the M measured data over ⍀ (e.g., statistical or stochastic operators operating on the seismic and/or MS events to generate the fractures planes or operators applied on the image logs and/or gamma rays data to approximate the permeability). These operators are denoted by P, and usually carry specific parameters (e.g., P might map the post-job MS events for a given unconventional horizontal completion stage to a set of fractures planes with the parameter of average distance between two neighboring planes of the same family, the minimum number of MS supporting events per plane, etc.). At the same time, the physics in the domain ⍀ can be simulated by solving a set of governing equations (GE). These equations solve for a set of variablesV, whereVM. The solvability of V using the GE depends on the availability of necessary inputs, including material coefficient data, boundary conditions (on ⭸⍀), and initial conditions. Some of these inputs can be obtained directly from the Table 1. However, oftentimes, the necessary information might be hidden in Table 1, a combination of several measurements, or even completely unknown. For example:

● It might be necessary to know the location of the various fractures and faults in ⍀. For that, it might be necessary to combine several mi from Table 1, while applying some (non-linear) operators, to obtain several options for the possible fractures in this domain and their uncertainty tag. The GE simulation must be applied for all permutations of the fracture field to determine the best matching the detailed measurements from Table 1.

● It might be necessary to know the formation properties (i.e., porosity and permeability), which are completely unknown for most of the domain, while some technologies (image logs) can offer values for part of the domain (e.g., close to the wellbore).

● Oftentimes, the domain ⍀ is not precisely defined, and calculations must be performed for the ranges of ⍀s and ⭸⍀s.

These unknowns are grouped into the set U ⬅ {u1,u2 . . . }, and the various options for the domain

are captured by the set D ⬅ {d1,d2, . . . }.

The last consideration is user opinion. The set C ⬅ {c1,c2, . . . } is a set of user opinions acting as constraints on the description of anything occurring in the specific reservoir. For example, the user can dictate that, most likely, the direction or orientation of the fractures is within a specified range, the

Table 1—DIFFERENTrMEASUREMENTS

T1 T2 T3 T4 Tp S s1⫽m2,s2⫽m3 m1,m3,m6 – – – – – m7,m9,mr X,t X1,t1,s1,e1 X2,t2,s2,e2 X3,t3,s3,e3 X4,t4,s4,e4 – – – Xp,tp,sp,ep m1 – ⫻ – ⫻ ⫻ – – – m2 ⫻ – ⫻ ⫻ – – – – m3 ⫻ ⫻ – – – ⫻ – – m4 – – ⫻ – – – – ⫻ – – – – – – – – – – – ⫻ – – – – – ⫻ mr – – ⫻ ⫻ – ⫻ – –

(5)

pressure values are upper bounded by some values, or the fracture propagation can be bounded. These constraints must be considered during all phases of characterizing the OG reservoir domain.

The objective is to describe the physics of the reservoir throughout the range of possible⍀s by means of the variablesMandUand the uncertainty tags, subject toC, while the classVshould be properly fitted to the relevant part of M at all of the given discrete times.

Solution Process

The solution is composed of several levels. At the core level, the reservoir is numerically simulated throughout a time span or similar to the time extent that the various technologies inTable 1offer in terms of measurement data. This simulation is based on boundary and initial conditions derived from the data in Table 1 with initial assumptions for the values or geometric structure for the necessary unknown variables.

At the next computational level, the simulation results, which are the computed variable values over the reservoir’s discrete grid for all of the time steps, are compared to the measured values of the appropriate technologies inTable 1. When attempting to obtain the best fit, considering the uncertainties, both in the computing values and measured variables, the values of the unknowns are properly changed to help achieve a global best fit (Bermudez et al. 2014 and Zadeh et al. 2012).

At the Level 3, the algorithm attempts to sense or better estimate the appropriate values for the reservoir domain and its boundaries, as well as the parameters for the transformation operatorP. Here, solutions are sought for the various ranges of the entire domain and operator’s parameters, attempting to obtain clear and well-defined sweet spots.

Level 1: the Core Problem Solver

The core problem is to obtain a solution for the evolution of the reservoir characteristics and properties at any location inside the reservoir domain. Generally, before considering any solution to this problem, the following steps are taken:

1. A specific domain d in D is chosen.

2. Based on Step 1, using the values for the entries ofVrelevant to the chosen solution domain, a set of parameters for the preprocessing operator P is chosen.

3. Based on Step 2, the operatorP is applied onMsubject toC. The result is a set of values for the members of I, each tagged with an appropriate confidence level.

4. A specific set of values ␯ for the unknowns in V is now assumed.

5. Based on these choices, the initial conditions, as well as the boundary conditions, are derived from M.

Using the information and values generated in the preprocessing steps discussed, the time-dependent problem is solved over dfor the time extent as defined by the sensor measurements provided inTable 1. For the time integration, an appropriate second order implicit scheme (a modified Crank Nicholson scheme) is used. At each time-step, the set of nonlinear equations at the discrete grid points over the entire domaindare solved in a quick and stable manner. This process estimates the values forVat the grid points and specific discrete time levels, which is accompanied by a well-defined error.

The time complexity of obtaining the solution for the basic problem is Eq. 1:

(1) wheretis the number of time steps,qis the average number of nonlinear iterations, andnis the number of grid points in the domain.

Usually, t⫽20,000, q⫽ 6, |v|⫽10, and n ⫽1000, thus time complexity per time step ⫽103(with some mathematical manipulations on the zero entries elements, it might decrease to 1011). The space

(6)

complexity is approximately 109. Both complexities reflect a well-defined confidence level (or accuracy) of the numerical simulation. Roughly speaking, the simulation error is proportional to . This fact inserts some flexibility into the choice of n so that the error level can be maintained within the user’s predefined limits. Thus, adaptive and dynamic numerical schemes are used (Pop 2014).

The value discussed fortrepresents usually 4 to 7 hr of a completion job. A regular sensor technology collects data at a rate five to six times slower. Therefore, to prepare for the optimal fit at the other levels, it is necessary to save the spatial computation results at all of the times that the available technologies offer relevant measurements. Actual experience demonstrates that 800 to 1,000 spatial reservoir fields are usually necessary for such a fit.

The effects of the temporal MS events are embedded in the calculations by assuming that matrix properties at and around the MS event are dramatically changed with respect to its neighborhood. That is, the porosity and permeability are different in a factor of K compared to these values in the closed vicinity. It is assumed that a low-path filter is controlling these properties; thus, for the permeabilityk, for example, the following model is assumed in Eq. 2:

(2) Where K ⫽ 0 if no MS event exists and K-5 is there if an MS event in the said geometric location.

Level 2: Optimal Parameters Fit

The calculated values ofVcan now be compared to the measured values defined inTable 1at the specific time the measurements are available, considering the calculated values are accompanied by an uncertainty originating from the numerical schemes errors, as well as the uncertainty of some of the formation properties (fracture location).

The fitting process uses the Newton approach with acceleration (thus iterative) using Krylov spaces, which amounts to correcting iteratively the Jacobian and Hessian matrices. The Jacobian matrix is the size of approximately (|V||M|)2, while the size of the Hessian is its square. The Jacobian inverse and Hessian products were iteratively approximated (using Kyrlov). These are the matrices usually handled in the BD environment, and thus, must be handled iteratively so that distributed algorithms can be implemented.

Moreover, the system is formulated so adaptive schemes can be naturally implemented. The adaptation is performed in three major domains:

● Number of grid points distributed inside the computational reservoir domain: this number might be different from time step to time step, and the grid concentration centers might move from location to location over time.

● Fit process onto which the number of time slices is implemented.

● Number of subsets of the parameters to be fitted: it can be small at the beginning of the process and grow as the process converges to the final ultimate set.

The adaptation approach can cause an increased speed of approximately 6 compared to the static case.

Level 3: Optimal Reservoir and Mapping Operator Parameters

This is the external optimization loop containing variables that are the reservoir domain’s geometry, its boundaries, and four of the mapping operator parameters. For that process, the ACO of the BD is used, where all of the computations are performed out of core and, in additional approximations, use the following concepts:

● The Lagrange multipliers for the constraints appear as a “quasi-time” in the iteration direction. ● The domain’s shape and size evolves in time at the boundary’s grid point level with the limitation

that the number of grid points on the boundary is constant.

(7)

Results and Discussion

All of the data used can be accessed within the public domain.

Tables 2 through 4 summarize the results for 1,200 MS events, 615 collected seismic events, and gamma ray signatures (all for 3 hr of the job time). For the core problem solution, 2,500 grid points were used with five variables.

For Level 2, five different formation parameters were fitted, and two fracture families were considered; each has between five to seven high-confidence fracture planes. The optimization procedure was implemented throughout 380 time slices; each had measurement values and error estimation for all of the variables (“no holes” condition).

For Level 3, the reservoir domain was expanded normally to its boundaries. For the mapping operator, only two parameters were chosen (distance between adjacent fractures planes and number of supporting events).

The AWS of Amazon (AWS Documentation 2014) was used for parallel and distributed computing (Zhang et al. 2014; Facchinei et al. 2014). Tables 2 through 4 summarize the process of obtaining the solution, performance, and efficiency of the overall strategy.

Solution Workflow

The solution workflow in these tables is as follows.

Level 1: Computational Domains and Numerical Schemes

All of the calculations and simulations began after approximately 5 min into the job; by then, more than 10 MS events were collected. The reason for that is only then can the solution domain, ⍀, be better defined. Two options for the solution domain were considered:

Dynamic Option This option will be presented in the forthcoming paper.

Table 2—SOLUTION, PERFORMANCE, AND EFFICIENCY OF OVERALL STRATEGY Global

Iteration 1

Norm of Error when Iterations Stopped

Number of Iterations Overall

Number of Central Processing Units (CPUs) Used

CPU Time Spent at Each Level (sec)

Memory Usage per CPU (gigabyte [GB])

Level 1 5 10^(⫺3) 35 120 800 8 Level 2 3 10^(⫺1) 8 20 110 51 Level 3 5 10^(⫺1) 3 5 40 250

Table 3—SOLUTION, PERFORMANCE, AND EFFICIENCY OF OVERALL STRATEGY

Global Iteration 2

Norm of Error when Iterations Stopped Number of Iterations Overall Number of CPUs Used

CPU Time Spent at Each Level (sec)

Memory Usage per CPU (GB)

Level 1 2 10^(⫺6) 30 120 650 8 Level 2 1 10^(⫺3) 6 20 95 55 Level 3 2 10^(⫺3) 3 5 43 270

Table 4 —SOLUTION, PERFORMANCE, AND EFFICIENCY OF OVERALL STRATEGY

Global Iteration 3

Norm of Error when Iterations Stopped Number of Iterations Overall Number of CPUs Used

CPU Time Spent at Each Level (sec)

Memory Usage per CPU (GB)

Level 1 2 10^(⫺9) 45 120 980 8 Level 2 7 10^(⫺7) 11 30 95 58 Level 3 5 10^(⫺6) 9 10 43 280

(8)

Here, the algorithm takes care of the possible changes in the computational domain in time because of the possibility that more fluid in the reservoir gets into motion in time.

When the computational domain changes, it is using a full similarity mapping, meaning that the boundaries of the new domain are parallel to those of the previous one; yet, the new domain has more grid points. The pressure values at the new computational points at the previous time step are those as of the pore pressure at “infinity.”

Static Option Presented in Table 1, here, the algorithm calculates the initial domain similar to the

dynamic option, after which it is increased in a factor of ‘f’, and then it is kept static throughout the computations. The factor of f ⫽ 8 was shown to be large enough in the present case.

A reduced version of the automatic fractures matching (Ma and Lin 2014) was used, where only the high confidence fracture of the primary family of fractures was considered with a static nonuniform histogram for the planes orientation.

The smallest rectangular containing all of the MS events was considered the medium computational domain, where another two with the size of ⫹20 and -20% were considered. Altogether, three compu-tational domains were considered. This is important for the iterations at Level 3. For the case of this paper, these computational domains increased by a factor of 8.

The calculations over each domain begin with a 51 ⫻51 grid, where the finite difference solution of the nonlinear time-dependent Darcy equation was used. As previously discussed, it is assumed that the MS event represents a fault (a fracture) in the matrix, and thus, the material properties close to the location of the MS event should be significantly different from the neighborhood. The current model assumes for the porosity and permeability a change of x Q, where usually Q⫽5 was chosen (this model is quite simple and used to exemplify the current computational algorithm). In the case there is another MS in the neighborhood, its influence (weighted function) reduces at the same amount of Q.

The Darcy equation was solved over the chosen domain with the chosen formation properties in parallel for the three chosen domains. They were distributed over 120 processors in the cloud, where each processor obtains approximately 90 equations for the elimination phase. These equations were quasi-linearized beforehand.

Level 2: Fitting of Pressure Values

During the time integration, there is some additional information available with respect to the pressure values. Often, the pressure drop through the reservoir is given. Sometimes, the leakoff is given, and thus, the pressure gradient on the domain’s boundaries can also be calculated. Sometimes, specific pressure values at a given location and at a specific time are given. Using ACO algorithms for optimal discrete fitting are used, as previously discussed.

The optimal fit is using a discrete version of the Newton algorithms in addition to Krylov iterations to obtain a second order rate of convergence.

This is executed for all of the domains under consideration

Level 3: Domain and Boundary Conditions Convergence

This level attempts to sense the optimal computational domain, which is defined to be located between the current computational domains under consideration and for which the formation properties do not change significantly with the change of the computational domain geometry.

Tables 2through4intend to show the computational performance of the mixing of PCO and ACO for the BD numerical computing of reservoir simulation. Each of these tables shows a different strategy to solve the BD numerical computing problem.

In Table 3, for example, the computational performed three Level 3 iterations, each performing six Level 2 iterations, where, for each at each time step, Level 1 performed 30 nonlinear iterations at most (converging the pressure and formation properties). 120 processors have been used at Level 1 (primarily

(9)

for elimination), 20 at Level 2 (primarily for the Jacoby inversion and obtaining the second order rate of convergence for parallel Krylov), and five processors at Level 3 for optimal change of the domains.

The intent of this paper is not to present the relevant results for this problem, but to present only the BD computational strategy. The results will be presented in a future paper.

References

Abramson, D., Lees, M., Krzhizhanovskaya, V.V., et al. 2014. Big Data Meets Computational Science, Preface for ICCS 2014. Procedia Computer Science 29: 1–7.

AWS Documentation. 2014. Amazon Elastic Compute Cloud, API Version 2014-06-15 User Guide, http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html.

Bermudez, G.M.T., Espada, J.P., and Nuñez-Valdez, E.R. 2014. An Application for Recommender Systems in the Contents Industry. Proc., The 8th International Conference on Knowledge Management in Organizations, The Netherlands.

Chelmis, C., Sorathia, V.S., Agarwal, S., et al. 2012. Semiautomatic Semantic Assistance to Manual Curation of Data in Smart Oil Fields. Paper SPE 153271 presented at the SPE Western Regional Meeting, Bakersfield, California, USA, 21–23 March. 10.2118/153271-MS.

Chen, C.L.P. and Zhang, C. 2014. Data-Intensive Applications, Challenges, Techniques and Tech-nologies: A survey on Big Data. Information Sciences275(2014): 314 –347. 10.1016/j.ins.2014.01.015. Facchinei, F., Sagratella, S., and Scutari, G. 2014. “Parallel Algorithms for Big Data Optimization.” arXiv preprint arXiv:1402.5521 http://arxiv.org/pdf/1402.5521.pdf(accessed 1 July 2014).

He, Q., Wang, H., Zhuang, F., et al. 2014. In press. Parallel Sampling from Big Data with Uncertainty Distribution. Fuzzy Sets and Systems (online). 10.1016/j.fss.2014.01.016.

Loshin, D. 2013. Big Data Analytics: from Strategic Planning to Enterprise Integration with Tools, Techniques, NoSQL, and Graph, Burlington, Massachusetts: Gulf Professional Publishing/Elsevier.

Ma, J. and Lin, A. 2014. In press, Stimulated Rock Information in Multistage Hydraulic Fracturing Treatment. SPE Journal SJ-0414-0011.

Pop, F. 2014. High Performance Numerical Computing for High Energy Physics: A New Challenge for Big Data Science. Advances in High Energy Physics2014. Article ID 507690. 10.1155/2014/507690 Zadeh, R., Bosagh, Z., and Goel, A. 2012. Dimension Independent Similarity Computation. The Journal of Machine Learning Research 14 (1): 1605–1626.

Zhang, X., Liu, C., Nepal, S., et al. 2014. A Hybrid Approach for Scalable Sub-Tree Anonymization over Big Data using MapReduce on Cloud.Journal of Computer and System Sciences80(5): 1008 –1020.

http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html. http://arxiv.org/pdf/1402.5521.pdf

References

Related documents

This paper describes our experiences using Active Learning in four first-year computer science and industrial engineering courses at the School of Engineering of the Universidad

This study is aimed to determine malnutrition of children under five years old as well as to identify the correlation between risk factors and malnutrition on the area of

This essay asserts that to effectively degrade and ultimately destroy the Islamic State of Iraq and Syria (ISIS), and to topple the Bashar al-Assad’s regime, the international

Under the stress of three inhibitors, the metabolites and key enzymes/proteins involved in glycolysis, reductive tricarboxylic acid (TCA) cycle, acetone–butanol synthesis and

Requisitos: Java Runtime Environment, OpenOffice.org/LibreOffice Versión más reciente: 1.5 (2011/09/25) Licencia: LGPL Enlaces de interés Descarga:

A: The extended warranty provides the same coverage as the standard warranty; replacement parts, standard shipping charges, labor and travel expenses for onsite Field Service

Powerful, simple and user-friendly software that makes total energy supervision possible for power analyzers, energy meters, earth leakages and com- plete control of

 There are disparities between figures published by PHE and the TIIG data presented in this report; the PHOF reported Allerdale, Eden and South Lakeland local authorities to