In this thesis, we have developed different control architectures to manage the power and performance of virtualized server cluster in data centers. First, we design a hierar- chical architecture which dynamically tunes the cluster’s processing capacity to match the incoming workload intensity so that the cluster’s power consumption is minimized and the desired response times for multiple applications are satisfied. Specifically, in this hierar- chy, fully distributed local controllers optimize the CPU share of VMs under their control, and a supervisory controller on top shuts down the unneeded machines during periods of light workload. Two different strategies: receding horizon control and neural network based control, are compared for the local controllers. We validate the framework on a clus- ter supporting three online services, showing that our scheme adapts quickly to dynamic workload changes, and is scalable and quite flexible in that servers can be added/removed anytime while maintaining overall system performance. Also, when managed using our control scheme, the cluster saves, on average, 20% in power-consumption costs over a three hour period when compared to a system operating without dynamic control.
Second, we propose a decentralized control architecture to further enhance the scala- bility of the hierarchical design. Here, each controller manages one server: its inner loop constantly optimizes the per-VM computing resources to guarantee the service level agree- ment (SLA), and its outer loop appropriately switches the processor cores on/off so that dynamic workload is consolidated onto the fewest number of active servers to reduce power consumption. In addition, we organize the controllers in different fashions and analyze how the organizations affect the overall performance of large clusters with up to a thousand servers. Our studies indicate that the control structure, when organized as a causal system in which a precedence relation exists among the individual controllers, achieves a high de- gree of SLA satisfaction (> 98%) while significantly reducing the corresponding switching cost.
sumption of multiple geographically distributed data centers, so that data centers can curtail power consumption as requested by electric utilities and earn financial reward as a return. The idea is to integrate the demand response (DR) program into data center operations and reduce power consumption in a data center by migrating the workload in the form of live VM migration to other centers. The optimizer aims to maximize the expected profit by trading off among reward, costs, VM migration time/distance, and risks from bandwidth and reward uncertainties. A set of case studies involving data centers participating in an economic DR program is used to validate the framework.
So far, regarding the integration of DR program into data center operations as de- scribed in Section 4, we have only proposed the idea and conducted some fundamental work in this direction. Some important issues need to be further studied in the future. First, the mathematical model could be further improved. The bandwidth cost should be taken into consideration, as it may add a large part to the total cost. Instead of a linear function, the relationship between migration time and distance can be replaced with a more practical quadratic function. In the current MINLP optimizer, the weighting factors of re- ward, costs, distance, and risks, are chosen prior to the test scenarios, based on exhaustive simulation results. Autonomic computation of these factors need to be realized to simply the configuration process and yield optimal performance. Second, currently we assume that all distributed data centers belong to one owner so that a centralized optimizer is adopted. For data centers belonging to different owners, a distributed framework need to be designed so that they can negotiate with each other regarding VM migration. Finally, the idea in Section 4 is only evaluated via simulation so far, and a real testbed across multiple data centers needs to be set up to further validate it.
Bibliography
[1] A glance at data centers. http://healthyhug.com/facebooks-green-data-center. html.
[2] Intel and core i7 (nehalem) dynamic power management. http://impact.asu.edu/ cse591sp11/Nahelempm.pdf.
[3] Zahra Abbasi, Tridib Mukherjee, Georgios Varsamopoulos, and Sandeep K. S. Gupta. DAHM: A green and dynamic web application hosting manager across geographically distributed data centers. ACM Journal on emerging technology, 2012.
[4] T. F. Abdelzaher, K. G. Shin, and N. Bhatti. Performance guarantees for web server end–systems: A control theoretic approach. IEEE Trans. Parallel & Distributed Syst., 13(1):80–96, January 2002.
[5] M. Arlitt and T. Jin. A workload characterization study of the 1998 world cup web site. IEEE Network, 14(3):30–37, May/Jun 2000.
[6] AST. Data center trends, October 2008.
[7] T. Atwood. Right architecture for the right workload: The application tier. Technical report, Sun Microsystems Report, Products, Jul. 2004.
[8] Dennis Bouley. Estimating a data center’s electrical carbon footprint, April 2010. [9] K. Brammer and G. Siffling. Kalman-Bucy Filters. Artec House: Norwood, MA, 1989. [10] E. F. Camacho and C. Bordons. Model Predictive Control. Springer-Verlag, London,
1999.
[11] Changbing Chen, Bingsheng He, and Xueyan Tang. Green-aware workload schedul- ing in geographically distributed data centers. In Intl. Conf. on Cloud Computing
Technology and Science, 2012.
[12] Yuan Chen, D. Gmach, C. Hyser, Z. Wang, C. Bash, C. Hoover, and S. Singhal. Integrated management of application performance, power and cooling in data centers. In Network Operations and Mgmt. Symposium, 2010.
[13] Cisco. Data center interconnect design guide for virtualized workload mobility with Cisco, NetApp and VMware, 2011.
[14] Cisco. Data center interconnect implementation guide for virtualized workload mobility with Cisco, EMC and VMware, 2011.
[15] Rajarshi Das, Jeffrey O. Kephart, Charles Lefurgy, Gerald Tesauro, David W. Levine, and Hoi Chan. Autonomic multi-agent management of power and performance in data centers. In Conf. Autonomous agents and multiagent systems, 2008.
[17] W. B. Dunbar and R. M. Murray. Distributed receding horizon control for multi-vehicle formation stabilization. Automatica, 42(4):549–558, 2006.
[18] EIA. Electric Power Monthly. http://www.eia.gov/electricity/monthly/index. cfm.
[19] EPA. Report to congress on server and data center energy efficiency, July 2007. [20] F5. Deploying the BIG-IP v10.2 to enable long distance live migration with VMware
vSphere vMotion, 2010.
[21] A. G. Ganek and T. A. Corbi. The dawn of the autonomic computing era. IBM Systems
Journal, 42(1):5–18, 2003.
[22] Jesse Goellner et al. Demand dispatchintelligent demand for a more efficient grid.
Technical Report DOE/NETL- DE-FE0004001, August 2011.
[23] A. Guez, I. Rusnak, and I. Bar Kana. Multiple objectives optimization approach to adaptive and learning control. Intl. Journal of Control, 56(2):469–482, September 1992. [24] A. C. Harvey. Forecasting, Structural Time Series Models and the Kalman Filter.
Cambridge University Press, Cambridge, UK, 2001.
[25] J. Hellerstein, S. Singhal, and Q. Wang. Research challenges in control engineering of computing systems. IEEE Trans. Network & Service Mgmt., 6(4):206–211, Dec. 2009. [26] J. L. Hellerstein, Y. Diao, S. Parekh, and D. M. Tilbury. Feedback Control of Computing
Systems. Wiley-IEEE Press, 2004.
[27] Yu-Chi Ho and K’ai-Ching Chu. Team decision theory and information structures in optimal control problems–Part I. Automatic Control, IEEE Transactions on, 17(1):15 – 22, Feb. 1972.
[28] IBM. Trade6 Performance-Characterizing Application for WebSphere.
http://www.ibm.com/developerworks/edu/dm-dw-dm-0506lau.html, 2005.
[29] David Irwin, Navin Sharma, and Prashant Shenoy. Towards continuous policy-driven demand response in data centers. In ACM Workshop on Green Networking, Aug. 2011. [30] ISO-New-England. Real-time price response program, 2012.
[31] Gueyoung Jung, M.A. Hiltunen, K.R. Joshi, R.D. Schlichting, and C. Pu. Mistral: Dy- namically managing power, performance, and adaptation cost in cloud infrastructures. In IEEE Intl. Conf. on Distributed Computing Systems, June 2010.
[32] Pavlo Krokhmal, Michael Zabarankin, and Stan Uryasev. Modeling and optimization of risk. Surveys in Operations Research and Management Science, 16(2):49 – 66, 2011. [33] D. Kusic, J.O. Kephart, J.E. Hanson, N. Kandasamy, and G. Jiang. Power and perfor- mance management of virtualized computing environments via lookahead control. In
[34] Dara Kusic, Jeffrey Kephart, James Hanson, Nagarajan Kandasamy, and Guofei Jiang. Power and performance management of virtualized computing environments via looka- head control. Cluster Computing, 12:1–15, 2009.
[35] Kien Le, Jingru Zhang, Jiandong Meng, R. Bianchini, Y. Jaluria, and T.D. Nguyen. Reducing electricity cost through virtual machine placement in high performance com- puting clouds. In Intl. Conf. for High Performance Computing, Networking, Storage
and Analysis, 2011.
[36] Alberto Leon-Garcia. Probability, statistics, and random processes for electrical engi-
neering. Prentice Hall, 2008.
[37] X. Liu, X. Zhu, S. Singhal, and M. Arlitt. Adaptive entitlement control of resource containers on shared servers. In 9th IFIP/IEEE Int’l Symp. Integrated Network Man-
agement (IM), pages 163–176, 2005.
[38] Zhenhua Liu, Minghong Lin, Adam Wierman, Steven H. Low, and Lachlan L.H. An- drew. Geographical load balancing with renewables. SIGMETRICS Perform. Eval.
Rev., 39(3):62–66, December 2011.
[39] Zhenhua Liu, Minghong Lin, Adam Wierman, Steven H. Low, and Lachlan L.H. An- drew. Greening geographical load balancing. In ACM SIGMETRICS, 2011.
[40] C. Lu, G. A. Alvarez, and J. Wilkes. Aqueduct: Online data migration with perfor- mance guarantees. In Proc. USENIX Conf. File Storage Tech., pages 219–230, 2002. [41] J. M. Maciejowski. Predictive Control with Constraints. Prentice Hall, London, 2002. [42] Spyros G. Makridakis, Steven C. Wheelwright, and Rob J. Hyndman. Forecasting:
methods and applications. Wiley, 1998.
[43] McKinsey. Revolutionizing data center efficiency, July 2008.
[44] Xiaoqiao Meng, Canturk Isci, Jeffrey Kephart, Li Zhang, Eric Bouillet, and Dimitrios Pendarakis. Efficient resource provisioning in compute clouds via VM multiplexing. In
Intl Conf. on Autonomic computing, New York, NY, USA, 2010.
[45] M. Morari and J. Lee. Model predictive control: Past, present and future. Computers
and Chemical Engineering, 23:667–682, 1999.
[46] D. Mosberger and T. Jin. httperf: A tool for measuring web server performance. Perf.
Eval. Review, 26:31–37, Dec. 1998.
[47] K. Narendra and S. Mukhopadhyay. Adaptive control using neural networks and ap- proximate models. IEEE Trans. Neural Networks, 8(3):475–485, May 1997.
[48] PJM. Introduction to PJM demand response, Oct. 2008. [49] PJM. PJM demand side response overview, May 2012.
[50] PJM. Retail electricity consumer opportunities for demand response in PJM’s whole- sale markets, 2012.
[51] S. Qin and T. Badgewell. An overview of industrial model predictive control technology.
Chemical Process Control, 93(316):232–256, 1997.
[52] Asfandyar Qureshi. Plugging into energy market diversity. In 7th ACM Workshop on
Hot Topics in Networks, Oct. 2008.
[53] Ramya Raghavendra, Parthasarathy Ranganathan, Vanish Talwar, Zhikui Wang, and Xiaoyun Zhu. No ”power” struggles: coordinated multi-level power management for the data center. In Intl. Conf. on Architectural support for programming languages and
operating systems, pages 48–59, 2008.
[54] Jia Rao, Xiangping Bu, Cheng-Zhong Xu, Leyi Wang, and George Yin. Vconf: a reinforcement learning approach to virtual machines auto-configuration. In Intl. Conf.
on Autonomic computing, pages 137–146, 2009.
[55] Lei Rao, Xue Liu, Le Xie, and Wenyu Liu. Minimizing electricity cost: Optimization of distributed internet data centers in a multi-electricity-market environment. In IEEE
INFOCOM, 2010.
[56] RUBBoS. RUBBoS: Bulletin Board Benchmark. http://jmob.ow2.org/rubbos.html. [57] RUBiS. RUBiS: Rice University Bidding System. http://rubis.ow2.org/.
[58] N. Sandell, P. Varaiya, M. Athans, and M. Safonov. Survey of decentralized control methods for large scale systems. IEEE Trans. Automatic Control, 23(2):108–128, April 1978.
[59] T. Simunic and S. Boyd. Managing power consumption in networks on chips. In Proc.
Design, Automation & Test Europe, pages 110–6, Mar. 2002.
[60] Sabrina Spatari. Average and marginal reductions in greenhouse gas emissions from data center power optomization, 2010.
[61] M. Steinbach. Markowitz revisited: Mean-variance models in financial portfolio anal- ysis. SIAM Review, 43(1):31–85, 2001.
[62] Gerald Tesauro. Reinforcement learning in autonomic computing: A manifesto and case studies. IEEE Internet Computing, 11:22–30, 2007.
[63] VMware. VMware vSphere — the best platform for cloud infrastructures, 2011. [64] VMware. VMware vSphere vMotion: 5.4 times faster than Hyper-V live migration,
2011.
[65] VMware and Cisco. Virtual machine mobility with VMware VMotion and Cisco data center interconnect technologies, 2009.
[66] Rahul Walawalkar, Stephen Fernands, Netra Thakur, and Konda Reddy Chevva. Evo- lution and current status of demand response (DR) in electricity markets: Insights from PJM and NYISO. Energy, 35(4):1553 – 1560, 2010.
[67] Rui Wang and Nagarajan Kandasamy. A distributed control framework for perfor- mance management of virtualized computing environments: some preliminary results. In 1st Workshop on Automated control for datacenters and clouds, 2009.
[68] Rui Wang and Nagarajan Kandasamy. On the design of decentralized control archi- tectures for workload consolidation in large-scale server clusters. In Intl. Conf. on
Autonomic computing, 2012.
[69] Rui Wang and Nagarajan Kandasamy. Workload consolidation in virtualized comput- ing systems via hierarchical control. Intel Technology Journal, 16, June 2012.
[70] Rui Wang, Nagarajan Kandasamy, and Chika Nwankpa. Data centers as demand re- sponse resources in the electricity market: Some preliminary results. In Intl. Workshop
on Feedback Computing, 2012.
[71] Rui Wang, Nagarajan Kandasamy, Chika Nwankpa, and David R. Kaeli. Data centers as controllable load resources in the electricity market. In Intl. Conf. on Distributed
Computing Systems, 2013.
[72] Rui Wang, Dara Marie Kusic, and Nagarajan Kandasamy. A distributed control frame- work for performance management of virtualized computing environments. In Intl.
Conf. on Autonomic computing, 2010.
[73] Xiaorui Wang and Ming Chen. Cluster-level feedback power control for performance optimization. In High Performance Computer Architecture, 2008. HPCA 2008. IEEE
14th International Symposium on, pages 101 –110, feb. 2008.
[74] Xiaorui Wang, Ming Chen, C. Lefurgy, and T.W. Keller. Ship: Scalable hierarchical power control for large-scale data centers. In Intl. Conf. on Parallel Architectures and
Compilation Techniques, sept. 2009.
[75] Yefu Wang, Xiaorui Wang, Ming Chen, and Xiaoyun Zhu. Power-efficient response time guarantees for virtualized enterprise servers. In Real-Time Systems Symp., 2008. [76] Z. Wang et al. Appraise: Application-level performance management in virtualized server environments. IEEE Trans. Network & Service Mgmt., 6(4):240–254, Dec. 2009. [77] Wikipedia. Normal distribution. http://en.wikipedia.org/wiki/Normal_
distribution.
[78] J. Xu, M. Zhao, J. Fortes, R. Carpenter, and M. Yousif. On the use of fuzzy modeling in virtualized data center management. pages 25–35, Jun. 2007.
VITA
Rui Wang was born in Yinchuan, China. He received his Bachelor of Science in Automa- tion from the Department of Automation at University of Science and Technology Beijing, China, in 2007. He joined the Department of Electrical and Computer Engineering at Drexel University in September 2008, and since then has been under the supervision of Dr. Nagarajan Kandasamy. His research falls into the big area of autonomic computing and cloud computing. Specifically, he designs fully decentralized control architectures, using control theory and optimization techniques, to manage the power and performance of large- scale virtualized server clusters. He also works as a teaching assistant in the Department of Electrical and Computer Engineering from 2008 to 2013. He is a student member of IEEE.