The Experts below are selected from a list of 2811 Experts worldwide ranked by ideXlab platform

Qi Liu - One of the best experts on this subject based on the ideXlab platform.

  • An optimized Speculative Execution Strategy Based on Local Data Prediction in Heterogeneous Hadoop Environment
    Zhonghua Minguo Diannao Xuehui, 2018
    Co-Authors: Liu Xiaodong, Qi Liu, Jin Dan-dan, Linge Nigel
    Abstract:

    Hadoop is a famous parallel computing framework that is applied to process large-scale data, but there exists such a task in Hadoop framework, which is called “Straggling task” and has a serious impact on Hadoop. Speculative execution (SE) is an effective way to deal with the “Straggling task” by monitoring the real-time rate of running tasks and back up the “Straggler” on another node to increase the opportunity of completing backup task ahead of original. There are many problems in the proposed SE strategies, such as “Straggling task” misjudgment, improper selection of backup nodes, which will result in inefficient implementation of SE. In this paper, we propose an optimized SE strategy based on local data prediction, it collects task execution information in real time and uses Local regression to predict remaining time of the current task, and selects the appropriate backup task node according to the actual requirements, at the same time, it uses the consumption and benefit model to maximizes the effectiveness of SE. Finally, the strategy is implemented in Hadoop-2.6.0, the experiment proves that the optimized strategy not only enhances the accuracy of selecting the “Straggler” task candidates, but also shows better performance in heterogeneous Hadoop Environment

  • an optimized speculative execution strategy based on local data prediction in a heterogeneous Hadoop Environment
    Computational Science and Engineering, 2017
    Co-Authors: Xiaodong Liu, Qi Liu
    Abstract:

    Hadoop is a famous distributed computing framework that is applied to process large-scale data. "Straggling tasks" have a serious impact on Hadoop performance due to imbalance of slow tasks distribution. Speculative execution (SE) presents a way to deal with Straggling tasks by monitoring the real-time progress of running tasks and replicating potential "Stragglers" on another node to increase the opportunity of completing backup tasks ahead of original. Current proposed SE strategies meet their challenges such as misjudgment of "Straggling tasks", improper selection of backup nodes, etc., which result in inefficient performance of the SE and its Hadoop system. In this paper, we propose an optimized SE strategy based on local data prediction, which collects task execution information in real time and uses Locally Weighted Regression (LWR) to predict remaining time of each running tasks, and selects an appropriate backup task node according to the actual requirements. It also combines a cost-benefit model to maximize the effectiveness of SE. According to the results, the proposed SE strategy implemented in Hadoop-2.6.0 enhances the accuracy of selecting potential Straggler task candidates, and shows better performance in various situations in a heterogeneous Hadoop Environment.

  • vpch a consistent hashing algorithm for better load balancing in a Hadoop Environment
    International Conference on Advanced Cloud and Big Data, 2015
    Co-Authors: Qi Liu, Weidong Cai, Jian Shen, Baowei Wang, Nigel Linge
    Abstract:

    MapReduce (MR) is a popular programming model for the purposes of processing large data sets among data clusters or grids, e.g. a Hadoop Environment. Load balancing as a key factor affecting the performance of map resource distribution, has recently gained high concerns to optimize. Current MR processes in the realization of distributing tasks to clusters use hashing with random modulo operations, which can lead to uneven data distribution and inclined loads, thereby obstruct the performance of the entire distribution system. In this paper, a virtual partition consistent hashing (VPCH) algorithm is proposed for the reduce stage of MR processes, in order to achieve such a trade-off on job allocation. According to the results, using our method can reduce task execution time with or without MJR (mapreduce.job.reduce.slowstart.completedmaps) parameter set.

Xiangrong Wang - One of the best experts on this subject based on the ideXlab platform.

  • parallel cellular automata markov model for land use change prediction over mapreduce framework
    ISPRS international journal of geo-information, 2019
    Co-Authors: Junfeng Kang, Lei Fang, Shuang Li, Xiangrong Wang
    Abstract:

    The Cellular Automata Markov model combines the cellular automata (CA) model’s ability to simulate the spatial variation of complex systems and the long-term prediction of the Markov model. In this research, we designed a parallel CA-Markov model based on the MapReduce framework. The model was divided into two main parts: A parallel Markov model based on MapReduce (Cloud-Markov), and comprehensive evaluation method of land-use changes based on cellular automata and MapReduce (Cloud-CELUC). Choosing Hangzhou as the study area and using Landsat remote-sensing images from 2006 and 2013 as the experiment data, we conducted three experiments to evaluate the parallel CA-Markov model on the Hadoop Environment. Efficiency evaluations were conducted to compare Cloud-Markov and Cloud-CELUC with different numbers of data. The results showed that the accelerated ratios of Cloud-Markov and Cloud-CELUC were 3.43 and 1.86, respectively, compared with their serial algorithms. The validity test of the prediction algorithm was performed using the parallel CA-Markov model to simulate land-use changes in Hangzhou in 2013 and to analyze the relationship between the simulation results and the interpretation results of the remote-sensing images. The Kappa coefficients of construction land, natural-reserve land, and agricultural land were 0.86, 0.68, and 0.66, respectively, which demonstrates the validity of the parallel model. Hangzhou land-use changes in 2020 were predicted and analyzed. The results show that the central area of construction land is rapidly increasing due to a developed transportation system and is mainly transferred from agricultural land.

Devi Anjali - One of the best experts on this subject based on the ideXlab platform.

  • Big Data Security Issues and Challenges in Cloud Computing
    ASIAN SOCIETY FOR SCIENTIFIC RESEARCH, 2018
    Co-Authors: Gangawane, Aarti A, Devi Anjali
    Abstract:

    In this paper we are going to discuss Big Data security issues and challenges in cloud computing, and also discuss Map Reduce and Hadoop Environment used for data processing file management. In this paper we discuss Need of security in cloud computing with big data. Could Computing Environments gaining popularity now a days, along with this the security issues through use of this technology are increasing. We will discuss various possible solutions for the security issues in cloud computing and Big Data. Many organizations, business, companies and many industries use the big data applications. Cloud computing security includes computer security, network security, information security, and data privacy. Cloud computing plays a very important role in protecting data, applications with the help of policies, technologies, controls, and big data tools

Sr Ravuri - One of the best experts on this subject based on the ideXlab platform.

  • Security Issues Associated With Big Data in Cloud Computing
    International Journal of Network Security & Its Applications, 2014
    Co-Authors: Vn Inukollu, S Arsi, Sr Ravuri
    Abstract:

    In this paper, we discuss security issues for cloud computing, Big data, Map Reduce and Hadoop Environment. The main focus is on security issues in cloud computing that are associated with big data. Big data applications are a great benefit to organizations, business, companies and many large scale and small scale industries.We also discuss various possible solutions for the issues in cloud computing security and Hadoop. Cloud computing security is developing at a rapid pace which includes computer security, network security, information security, and data privacy. Cloud computing plays a very vital role in protecting data, applications and the related infrastructure with the help of policies, technologies, controls, and big data tools. Moreover, cloud computing, big data and its applications, advantages are likely to represent the most promising new frontiers in science.

Roberto Cerchione - One of the best experts on this subject based on the ideXlab platform.

  • extracting knowledge from big data for sustainability a comparison of machine learning techniques
    Sustainability, 2019
    Co-Authors: Raghu Garg, Himanshu Aggarwal, Piera Centobelli, Roberto Cerchione
    Abstract:

    At present, due to the unavailability of natural resources, society should take the maximum advantage of data, information, and knowledge to achieve sustainability goals. In today’s world condition, the existence of humans is not possible without the essential proliferation of plants. In the photosynthesis procedure, plants use solar energy to convert into chemical energy. This process is responsible for all life on earth, and the main controlling factor for proper plant growth is soil since it holds water, air, and all essential nutrients of plant nourishment. Though, due to overexposure, soil gets despoiled, so fertilizer is an essential component to hold the soil quality. In that regard, soil analysis is a suitable method to determine soil quality. Soil analysis examines the soil in laboratories and generates reports of unorganized and insignificant data. In this study, different big data analysis machine learning methods are used to extracting knowledge from data to find out fertilizer recommendation classes on behalf of present soil nutrition composition. For this experiment, soil analysis reports are collected from the Tata soil and water testing center. In this paper, Mahoot library is used for analysis of stochastic gradient descent (SGD), artificial neural network (ANN) performance on Hadoop Environment. For better performance evaluation, we also used single machine experiments for random forest (RF), K-nearest neighbors K-NN, regression tree (RT), support vector machine (SVM) using polynomial function, SVM using radial basis function (RBF) methods. Detailed experimental analysis was carried out using overall accuracy, AUC–ROC (receiver operating characteristics (ROC), and area under the ROC curve (AUC)) curve, mean absolute prediction error (MAE), root mean square error (RMSE), and coefficient of determination (R2) validation measurements on soil reports dataset. The results provide a comparison of solution classes and conclude that the SGD outperforms other approaches. Finally, the proposed results support to select the solution or recommend a class which suggests suitable fertilizer to crops for maximum production.