The Experts below are selected from a list of 10731 Experts worldwide ranked by ideXlab platform
Ji Liu - One of the best experts on this subject based on the ideXlab platform.
-
a comprehensive Linear Speedup analysis for asynchronous stochastic parallel optimization from zeroth order to first order
arXiv: Optimization and Control, 2016Co-Authors: Xiangru Lian, Huan Zhang, Chojui Hsieh, Yijun Huang, Ji LiuAbstract:Asynchronous parallel optimization received substantial successes and extensive attention recently. One of core theoretical questions is how much Speedup (or benefit) the asynchronous parallelization can bring us. This paper provides a comprehensive and generic analysis to study the Speedup property for a broad range of asynchronous parallel stochastic algorithms from the zeroth order to the first order methods. Our result recovers or improves existing analysis on special cases, provides more insights for understanding the asynchronous parallel behaviors, and suggests a novel asynchronous parallel zeroth order method for the first time. Our experiments provide novel applications including model blending problems using the proposed asynchronous parallel zeroth order method.
-
a comprehensive Linear Speedup analysis for asynchronous stochastic parallel optimization from zeroth order to first order
Neural Information Processing Systems, 2016Co-Authors: Xiangru Lian, Huan Zhang, Chojui Hsieh, Yijun Huang, Ji LiuAbstract:Asynchronous parallel optimization received substantial successes and extensive attention recently. One of core theoretical questions is how much Speedup (or benefit) the asynchronous parallelization can bring to us. This paper provides a comprehensive and generic analysis to study the Speedup property for a broad range of asynchronous parallel stochastic algorithms from the zeroth order to the first order methods. Our result recovers or improves existing analysis on special cases, provides more insights for understanding the asynchronous parallel behaviors, and suggests a novel asynchronous parallel zeroth order method for the first time. Our experiments provide novel applications of the proposed asynchronous parallel zeroth order method on hyper parameter tuning and model blending problems.
-
asynchronous stochastic coordinate descent parallelism and convergence properties
arXiv: Optimization and Control, 2014Co-Authors: Ji Liu, Stephen J WrightAbstract:We describe an asynchronous parallel stochastic proximal coordinate descent algorithm for minimizing a composite objective function, which consists of a smooth convex function plus a separable convex function. In contrast to previous analyses, our model of asynchronous computation accounts for the fact that components of the unknown vector may be written by some cores simultaneously with being read by others. Despite the complications arising from this possibility, the method achieves a Linear convergence rate on functions that satisfy an optimal strong convexity property and a subLinear rate ($1/k$) on general convex functions. Near-Linear Speedup on a multicore system can be expected if the number of processors is $O(n^{1/4})$. We describe results from implementation on ten cores of a multicore processor.
Dries Harnie - One of the best experts on this subject based on the ideXlab platform.
-
scaling machine learning for target prediction in drug discovery using apache spark
Future Generation Computer Systems, 2017Co-Authors: Dries Harnie, Mathijs Saey, Alexander Vapirev, Jorg K Wegner, Andrey Gedich, Marvin Steijaert, Hugo CeulemansAbstract:Abstract In the context of drug discovery, a key problem is the identification of candidate molecules that affect proteins associated with diseases. Inside Janssen Pharmaceutica, the Chemogenomics project aims to derive new candidates from existing experiments through a set of machine learning predictor programs, written in single-node C++. These programs take a long time to run and are inherently parallel, but do not use multiple nodes. We show how we reimplemented the pipeline using Apache Spark, which enabled us to lift the existing programs to a multi-node cluster without making changes to the predictors. We have benchmarked our Spark pipeline against the original, which shows almost Linear Speedup up to 8 nodes. In addition, our pipeline generates fewer intermediate files while allowing easier checkpointing and monitoring.
-
scaling machine learning for target prediction in drug discovery using apache spark
IEEE ACM International Symposium Cluster Cloud and Grid Computing, 2015Co-Authors: Dries Harnie, Alexander Vapirev, Jorg K Wegner, Andrey Gedich, Marvin Steijaert, Roel Wuyts, Wolfgang De MeuterAbstract:In the context of drug discovery, a key problem is the identification of candidate molecules that affect proteins associated with diseases. Inside Janssen Pharmaceutical, the Chemo genomics project aims to derive new candidates from existing experiments through a set of machine learning predictor programs, written in single-node C++. These programs take a long time to run and are inherently parallel, but do not use multiple nodes. We show how we reimplementation the pipeline using Apache Spark, which enabled us to lift the existing programs to a multi-node cluster without making changes to the predictors. We have benchmarked our Spark pipeline against the original, which shows almost Linear Speedup up to 8 nodes. In addition, our pipeline generates fewer intermediate files while allowing easier check pointing and monitoring.
Hugo Ceulemans - One of the best experts on this subject based on the ideXlab platform.
-
scaling machine learning for target prediction in drug discovery using apache spark
Future Generation Computer Systems, 2017Co-Authors: Dries Harnie, Mathijs Saey, Alexander Vapirev, Jorg K Wegner, Andrey Gedich, Marvin Steijaert, Hugo CeulemansAbstract:Abstract In the context of drug discovery, a key problem is the identification of candidate molecules that affect proteins associated with diseases. Inside Janssen Pharmaceutica, the Chemogenomics project aims to derive new candidates from existing experiments through a set of machine learning predictor programs, written in single-node C++. These programs take a long time to run and are inherently parallel, but do not use multiple nodes. We show how we reimplemented the pipeline using Apache Spark, which enabled us to lift the existing programs to a multi-node cluster without making changes to the predictors. We have benchmarked our Spark pipeline against the original, which shows almost Linear Speedup up to 8 nodes. In addition, our pipeline generates fewer intermediate files while allowing easier checkpointing and monitoring.
Alexander Vapirev - One of the best experts on this subject based on the ideXlab platform.
-
scaling machine learning for target prediction in drug discovery using apache spark
Future Generation Computer Systems, 2017Co-Authors: Dries Harnie, Mathijs Saey, Alexander Vapirev, Jorg K Wegner, Andrey Gedich, Marvin Steijaert, Hugo CeulemansAbstract:Abstract In the context of drug discovery, a key problem is the identification of candidate molecules that affect proteins associated with diseases. Inside Janssen Pharmaceutica, the Chemogenomics project aims to derive new candidates from existing experiments through a set of machine learning predictor programs, written in single-node C++. These programs take a long time to run and are inherently parallel, but do not use multiple nodes. We show how we reimplemented the pipeline using Apache Spark, which enabled us to lift the existing programs to a multi-node cluster without making changes to the predictors. We have benchmarked our Spark pipeline against the original, which shows almost Linear Speedup up to 8 nodes. In addition, our pipeline generates fewer intermediate files while allowing easier checkpointing and monitoring.
-
scaling machine learning for target prediction in drug discovery using apache spark
IEEE ACM International Symposium Cluster Cloud and Grid Computing, 2015Co-Authors: Dries Harnie, Alexander Vapirev, Jorg K Wegner, Andrey Gedich, Marvin Steijaert, Roel Wuyts, Wolfgang De MeuterAbstract:In the context of drug discovery, a key problem is the identification of candidate molecules that affect proteins associated with diseases. Inside Janssen Pharmaceutical, the Chemo genomics project aims to derive new candidates from existing experiments through a set of machine learning predictor programs, written in single-node C++. These programs take a long time to run and are inherently parallel, but do not use multiple nodes. We show how we reimplementation the pipeline using Apache Spark, which enabled us to lift the existing programs to a multi-node cluster without making changes to the predictors. We have benchmarked our Spark pipeline against the original, which shows almost Linear Speedup up to 8 nodes. In addition, our pipeline generates fewer intermediate files while allowing easier check pointing and monitoring.
Xiangru Lian - One of the best experts on this subject based on the ideXlab platform.
-
a comprehensive Linear Speedup analysis for asynchronous stochastic parallel optimization from zeroth order to first order
arXiv: Optimization and Control, 2016Co-Authors: Xiangru Lian, Huan Zhang, Chojui Hsieh, Yijun Huang, Ji LiuAbstract:Asynchronous parallel optimization received substantial successes and extensive attention recently. One of core theoretical questions is how much Speedup (or benefit) the asynchronous parallelization can bring us. This paper provides a comprehensive and generic analysis to study the Speedup property for a broad range of asynchronous parallel stochastic algorithms from the zeroth order to the first order methods. Our result recovers or improves existing analysis on special cases, provides more insights for understanding the asynchronous parallel behaviors, and suggests a novel asynchronous parallel zeroth order method for the first time. Our experiments provide novel applications including model blending problems using the proposed asynchronous parallel zeroth order method.
-
a comprehensive Linear Speedup analysis for asynchronous stochastic parallel optimization from zeroth order to first order
Neural Information Processing Systems, 2016Co-Authors: Xiangru Lian, Huan Zhang, Chojui Hsieh, Yijun Huang, Ji LiuAbstract:Asynchronous parallel optimization received substantial successes and extensive attention recently. One of core theoretical questions is how much Speedup (or benefit) the asynchronous parallelization can bring to us. This paper provides a comprehensive and generic analysis to study the Speedup property for a broad range of asynchronous parallel stochastic algorithms from the zeroth order to the first order methods. Our result recovers or improves existing analysis on special cases, provides more insights for understanding the asynchronous parallel behaviors, and suggests a novel asynchronous parallel zeroth order method for the first time. Our experiments provide novel applications of the proposed asynchronous parallel zeroth order method on hyper parameter tuning and model blending problems.