The Experts below are selected from a list of 61728 Experts worldwide ranked by ideXlab platform

Mahmut Kandemir - One of the best experts on this subject based on the ideXlab platform.

  • quantifying and optimizing data access parallelism on manycores
    Modeling Analysis and Simulation On Computer and Telecommunication Systems, 2018
    Co-Authors: Jihyun Ryoo, Xulong Tang, Orhan Kislal, Mahmut Kandemir
    Abstract:

    Data access parallelism (DAP) indicates how well available hardware resources are utilized by data accesses. This paper investigates four complementary components of data access parallelism in detail: cache-Level parallelism (CLP), bank-Level parallelism (BLP), network-Level parallelism (NLP), and memory Controller-Level parallelism (MLP). Specifically, we first quantify these four components for a set of 20 multi-threaded benchmark programs, and show that, when executed on a state-of-the-art manycore platform, their original values are quite low compared to the maximum possible values they could take. We next perform a limit study, which indicates that significant performance improvements are possible if the values of these four components of DAP could be maximized. Building upon our observations from this limit study, we then present two practical computation and network access scheduling schemes. Both these schemes make use of profile data, but, while the compiler-based strategy uses fixed priorities of CLP, BLP, NLP, and MLP, the machine learning-based one employs a predictive machine learning model. Our experiments indicate 30.8% and 36.9% performance improvements with the compiler-based and learning-based schemes, respectively. Our results also show that the proposed schemes consistently achieve significant improvements under different values of the major experimental parameters.

  • Compiler Support for Optimizing Memory Bank-Level Parallelism
    2014 47th Annual IEEE ACM International Symposium on Microarchitecture, 2014
    Co-Authors: Wei Ding, Diana Guttman, Mahmut Kandemir
    Abstract:

    Many prior compiler-based optimization schemes focused exclusively on cache data locality. However, cache locality is only one part of the overall performance of applications running on emerging multicores or many cores. For example, memory stalls could constitute a very large fraction of execution time even in cache-optimized codes, and one of the main reasons for this is lack of memory-Level parallelism. Motivated by this, we propose a compiler-based Bank-Level Parallelism (BLP) optimization scheme that uses loop tile scheduling. More specifically, we first use Cache Miss Equations to predict where the last-Level cache miss will happen in each tile, and then identify the set of memory banks that will be accessed in each tile. Using this information, two tile scheduling algorithms are proposed to maximize BLP, each targeting a different scenario. We further discuss how our compiler-based scheme can be enhanced to consider memory Controller-Level parallelism and row-buffer locality. Our experimental evaluation using 11 multithreaded applications shows that the proposed BLP optimization can improve average BLP by 17.1% on average, resulting in a 9.2% reduction in average memory access latency. Furthermore, considering memory Controller-Level parallelism and row-buffer locality (in addition to BLP) takes our average improvement in memory access latency to 22.2%.

Wei Ding - One of the best experts on this subject based on the ideXlab platform.

  • Compiler Support for Optimizing Memory Bank-Level Parallelism
    2014 47th Annual IEEE ACM International Symposium on Microarchitecture, 2014
    Co-Authors: Wei Ding, Diana Guttman, Mahmut Kandemir
    Abstract:

    Many prior compiler-based optimization schemes focused exclusively on cache data locality. However, cache locality is only one part of the overall performance of applications running on emerging multicores or many cores. For example, memory stalls could constitute a very large fraction of execution time even in cache-optimized codes, and one of the main reasons for this is lack of memory-Level parallelism. Motivated by this, we propose a compiler-based Bank-Level Parallelism (BLP) optimization scheme that uses loop tile scheduling. More specifically, we first use Cache Miss Equations to predict where the last-Level cache miss will happen in each tile, and then identify the set of memory banks that will be accessed in each tile. Using this information, two tile scheduling algorithms are proposed to maximize BLP, each targeting a different scenario. We further discuss how our compiler-based scheme can be enhanced to consider memory Controller-Level parallelism and row-buffer locality. Our experimental evaluation using 11 multithreaded applications shows that the proposed BLP optimization can improve average BLP by 17.1% on average, resulting in a 9.2% reduction in average memory access latency. Furthermore, considering memory Controller-Level parallelism and row-buffer locality (in addition to BLP) takes our average improvement in memory access latency to 22.2%.

Linpeng Huang - One of the best experts on this subject based on the ideXlab platform.

  • a pure hardware driven scheduler for enhancing bank Level parallelism in a persistent memory Controller
    Future Generation Computer Systems, 2020
    Co-Authors: Dongliang Xue, Linpeng Huang
    Abstract:

    Abstract Researchers are attempting to exploit the non-volatile nature of Persistent Memory (PM) in various applications. To utilize persistence, many solutions enforce strict, sequential write orderings that are propagated through the cache hierarchy, inevitably resulting in extra write requests. Other solutions introduce persistence logging operations to ensure transactional consistency, which also introduces extra writes. To rationally schedule these write requests, current work classifies the sources of write requests into subtypes at the software Level based on the characteristics of the application, subsequently transferring the classified requests as hint messages to the memory Controller Level for guiding scheduling. However, categorizing diverse applications requires incompatible modifications to software and can degrade both the performance and fairness of the results with respect to low bank-Level parallelism and low row-buffer locality. To address these problems, we bypass software intervention and propose PHD-scheduler, a P ure H ardware- D riven scheduler for enhancing bank-Level parallelism in PM Controller. PHD-scheduler is composed of three main ideas: (1) PHD-scheduler moves the classification into the PM’s Controller Level rather than intervening in the compatibility of the software, (2) it introduces a dynamic computing mechanism to ensure bank access from centralization to distribution, and (3) it maximizes row buffer locality by reshuffling memory requests with three novel criteria to achieve an appropriate batch scheduling. The experimental results show that PHD-scheduler achieves an average bank-Level parallelism improvement of 11.8%. Moreover, performance and fairness are improved by 4.2% and 3.3%, respectively.

I Yellowley - One of the best experts on this subject based on the ideXlab platform.

  • a new approach to contour error control in high speed machining
    International Journal of Machine Tools & Manufacture, 2015
    Co-Authors: Mostafizur Rahaman, Rudolf Seethaler, I Yellowley
    Abstract:

    Abstract High speed machining technology attempts to maximize productivity through the use of high spindle speeds and axis traverse rates. The technology is dependent upon the development of suitable mechanical hardware, electrical drives and associated control software to ensure that all components are used to maximum advantage. The role of the control software is particularly demanding since one needs to maximize traverse rates while providing the necessary accuracy, and indeed providing a margin of safety to deal with unexpected changes in process, or system parameters. There have been relatively few improvements in commercial CAD or CAM systems that would help machine tool users to take maximum advantage of high speed machining; rather the majority of the approaches have been undertaken at the machine tool Controller Level. This paper uses circular interpolation and corner tracking to compare several such control techniques, (Cross Coupled Control (CCC), Zero Phase Error Tracking Control (ZPETC), and Realtime Frequency Modulated Interpolation (FMI)), each of which have been proposed in the literature order to improve machining accuracy. None of these approaches are found to be universally successful when used alone and the authors, in this paper, examine the use of these systems in combination. Particular attention is focused upon an extension of a simplified version of cross coupled control together with Frequency Modulated Interpolation. It is shown that the combined system performs extremely well, and is easily actuated at high frequencies with conventional hardware. A custom built high speed x-y table is used to confirm system performance with multiple constraints present.

Heiko Claussen - One of the best experts on this subject based on the ideXlab platform.

  • industrial robot grasping with deep learning using a programmable logic Controller plc
    Conference on Automation Science and Engineering, 2020
    Co-Authors: Eugen Solowjow, Ines Ugalde, Yash Shahapurkar, Juan Aparicio, Jeffrey Mahler, Vishal Satish, Ken Goldberg, Heiko Claussen
    Abstract:

    Universal grasping of a diverse range of previously unseen objects from heaps is a grand challenge in e-commerce order fulfillment, manufacturing, and home service robotics. Recently, deep learning based grasping approaches have demonstrated results that make them increasingly interesting for industrial deployments. This paper explores the problem from an automation systems point-of-view. We develop a robotics grasping system using Dex-Net, which is fully integrated at the Controller Level. Two neural networks are deployed on a novel industrial AI hardware acceleration module close to a PLC with a power footprint of less than 10 W for the overall system. The software is tightly integrated with the hardware allowing for fast and efficient data processing and real-time communication. The success rate of grasping an object form a bin is up to 95% with more than 350 picks per hour, if object and receptive bins are in close proximity. The system was presented at the Hannover Fair 2019 (world’s largest industrial trade fair) and other events, where it performed over 5,000 grasps per event.

  • industrial robot grasping with deep learning using a programmable logic Controller plc
    arXiv: Robotics, 2020
    Co-Authors: Eugen Solowjow, Ines Ugalde, Yash Shahapurkar, Juan Aparicio, Jeffrey Mahler, Vishal Satish, Ken Goldberg, Heiko Claussen
    Abstract:

    Universal grasping of a diverse range of previously unseen objects from heaps is a grand challenge in e-commerce order fulfillment, manufacturing, and home service robotics. Recently, deep learning based grasping approaches have demonstrated results that make them increasingly interesting for industrial deployments. This paper explores the problem from an automation systems point-of-view. We develop a robotics grasping system using Dex-Net, which is fully integrated at the Controller Level. Two neural networks are deployed on a novel industrial AI hardware acceleration module close to a PLC with a power footprint of less than 10 W for the overall system. The software is tightly integrated with the hardware allowing for fast and efficient data processing and real-time communication. The success rate of grasping an object form a bin is up to 95 percent with more than 350 picks per hour, if object and receptive bins are in close proximity. The system was presented at the Hannover Fair 2019 (world s largest industrial trade fair) and other events, where it performed over 5,000 grasps per event.