The Experts below are selected from a list of 52845 Experts worldwide ranked by ideXlab platform

Stephen W Keckler - One of the best experts on this subject based on the ideXlab platform.

  • a real time energy efficient superpixel hardware accelerator for mobile Computer Vision Applications
    Design Automation Conference, 2016
    Co-Authors: Injoon Hong, Iuri Frosio, Jason Clemons, Brucek Khailany, Rangharajan Venkatesan, Stephen W Keckler
    Abstract:

    Superpixel generation is a common preprocessing step in Vision processing aimed at dividing an image into non-overlapping regions. Simple Linear Iterative Clustering (SLIC) is a commonly used superpixel algorithm that offers a good balance between performance and accuracy. However, the algorithm's high computational and memory bandwidth requirements result in performance and energy efficiency that do not meet the requirements of real-time embedded Applications. In this work, we explore the design of an energy-efficient superpixel accelerator for real-time Computer Vision Applications. We propose a novel algorithm, Subsampled SLIC (S-SLIC), that uses pixel subsampling to reduce the memory bandwidth by 1.8×. We integrate S-SLIC into an energy-efficient superpixel accelerator and perform an in-depth design space exploration to optimize the design. We completed a detailed design in a 16nm FinFET technology using commercially-available EDA tools for high-level synthesis to map the design automatically from a C-based representation to a gate-level implementation. The proposed S-SLIC accelerator achieves real-time performance (30 frames per second) with 250× better energy efficiency than an optimized SLIC software implementation running on a mobile GPU.

  • DAC - A real-time energy-efficient superpixel hardware accelerator for mobile Computer Vision Applications
    Proceedings of the 53rd Annual Design Automation Conference on - DAC '16, 2016
    Co-Authors: Injoon Hong, Iuri Frosio, Jason Clemons, Brucek Khailany, Rangharajan Venkatesan, Stephen W Keckler
    Abstract:

    Superpixel generation is a common preprocessing step in Vision processing aimed at dividing an image into non-overlapping regions. Simple Linear Iterative Clustering (SLIC) is a commonly used superpixel algorithm that offers a good balance between performance and accuracy. However, the algorithm's high computational and memory bandwidth requirements result in performance and energy efficiency that do not meet the requirements of real-time embedded Applications. In this work, we explore the design of an energy-efficient superpixel accelerator for real-time Computer Vision Applications. We propose a novel algorithm, Subsampled SLIC (S-SLIC), that uses pixel subsampling to reduce the memory bandwidth by 1.8×. We integrate S-SLIC into an energy-efficient superpixel accelerator and perform an in-depth design space exploration to optimize the design. We completed a detailed design in a 16nm FinFET technology using commercially-available EDA tools for high-level synthesis to map the design automatically from a C-based representation to a gate-level implementation. The proposed S-SLIC accelerator achieves real-time performance (30 frames per second) with 250× better energy efficiency than an optimized SLIC software implementation running on a mobile GPU.

Rajesh Gupta - One of the best experts on this subject based on the ideXlab platform.

  • multi tenant mobile offloading systems for real time Computer Vision Applications
    International Conference of Distributed Computing and Networking, 2019
    Co-Authors: Zhou Fang, Jenghau Lin, Mani Srivastava, Rajesh Gupta
    Abstract:

    Offloading techniques enable many emerging Computer Vision Applications on mobile platforms by executing compute-intensive tasks on resource-rich servers. Although there have been a significant amount of research efforts devoted in optimizing mobile offloading frameworks, most previous works are evaluated in a single-tenant setting, that is, a server is assigned to a single client. However, in a practical scenario that servers must handle tasks from many clients running diverse Applications, contention on shared server resources may degrade application performance. In this work, we study scheduling techniques to improve serving performance in multi-tenant mobile offloading systems, for Computer Vision algorithms running on CPUs and deep neural networks (DNNs) running on GPUs. For CPU workloads, we present methods to mitigate resource contention and to improve delay using a Plan-Schedule approach. The planning phase predicts future workloads from all clients, estimates contention, and adjusts future task start times to remove or reduce contention. The scheduling phase dispatches arriving offloaded tasks to the server that minimizes contention. For DNN workloads running on GPUs, we propose adaptive batching algorithms using information of batch size, model complexity and system load to achieve the best Quality of Service (QoS), which are measured from accuracy and delay of DNN tasks. We demonstrate the improvement of serving performance using several real-world Applications with different server deployments.

  • ICDCN - Multi-tenant mobile offloading systems for real-time Computer Vision Applications
    Proceedings of the 20th International Conference on Distributed Computing and Networking, 2019
    Co-Authors: Zhou Fang, Jenghau Lin, Mani Srivastava, Rajesh Gupta
    Abstract:

    Offloading techniques enable many emerging Computer Vision Applications on mobile platforms by executing compute-intensive tasks on resource-rich servers. Although there have been a significant amount of research efforts devoted in optimizing mobile offloading frameworks, most previous works are evaluated in a single-tenant setting, that is, a server is assigned to a single client. However, in a practical scenario that servers must handle tasks from many clients running diverse Applications, contention on shared server resources may degrade application performance. In this work, we study scheduling techniques to improve serving performance in multi-tenant mobile offloading systems, for Computer Vision algorithms running on CPUs and deep neural networks (DNNs) running on GPUs. For CPU workloads, we present methods to mitigate resource contention and to improve delay using a Plan-Schedule approach. The planning phase predicts future workloads from all clients, estimates contention, and adjusts future task start times to remove or reduce contention. The scheduling phase dispatches arriving offloaded tasks to the server that minimizes contention. For DNN workloads running on GPUs, we propose adaptive batching algorithms using information of batch size, model complexity and system load to achieve the best Quality of Service (QoS), which are measured from accuracy and delay of DNN tasks. We demonstrate the improvement of serving performance using several real-world Applications with different server deployments.

Injoon Hong - One of the best experts on this subject based on the ideXlab platform.

  • a real time energy efficient superpixel hardware accelerator for mobile Computer Vision Applications
    Design Automation Conference, 2016
    Co-Authors: Injoon Hong, Iuri Frosio, Jason Clemons, Brucek Khailany, Rangharajan Venkatesan, Stephen W Keckler
    Abstract:

    Superpixel generation is a common preprocessing step in Vision processing aimed at dividing an image into non-overlapping regions. Simple Linear Iterative Clustering (SLIC) is a commonly used superpixel algorithm that offers a good balance between performance and accuracy. However, the algorithm's high computational and memory bandwidth requirements result in performance and energy efficiency that do not meet the requirements of real-time embedded Applications. In this work, we explore the design of an energy-efficient superpixel accelerator for real-time Computer Vision Applications. We propose a novel algorithm, Subsampled SLIC (S-SLIC), that uses pixel subsampling to reduce the memory bandwidth by 1.8×. We integrate S-SLIC into an energy-efficient superpixel accelerator and perform an in-depth design space exploration to optimize the design. We completed a detailed design in a 16nm FinFET technology using commercially-available EDA tools for high-level synthesis to map the design automatically from a C-based representation to a gate-level implementation. The proposed S-SLIC accelerator achieves real-time performance (30 frames per second) with 250× better energy efficiency than an optimized SLIC software implementation running on a mobile GPU.

  • DAC - A real-time energy-efficient superpixel hardware accelerator for mobile Computer Vision Applications
    Proceedings of the 53rd Annual Design Automation Conference on - DAC '16, 2016
    Co-Authors: Injoon Hong, Iuri Frosio, Jason Clemons, Brucek Khailany, Rangharajan Venkatesan, Stephen W Keckler
    Abstract:

    Superpixel generation is a common preprocessing step in Vision processing aimed at dividing an image into non-overlapping regions. Simple Linear Iterative Clustering (SLIC) is a commonly used superpixel algorithm that offers a good balance between performance and accuracy. However, the algorithm's high computational and memory bandwidth requirements result in performance and energy efficiency that do not meet the requirements of real-time embedded Applications. In this work, we explore the design of an energy-efficient superpixel accelerator for real-time Computer Vision Applications. We propose a novel algorithm, Subsampled SLIC (S-SLIC), that uses pixel subsampling to reduce the memory bandwidth by 1.8×. We integrate S-SLIC into an energy-efficient superpixel accelerator and perform an in-depth design space exploration to optimize the design. We completed a detailed design in a 16nm FinFET technology using commercially-available EDA tools for high-level synthesis to map the design automatically from a C-based representation to a gate-level implementation. The proposed S-SLIC accelerator achieves real-time performance (30 frames per second) with 250× better energy efficiency than an optimized SLIC software implementation running on a mobile GPU.

Leonidas J Guibas - One of the best experts on this subject based on the ideXlab platform.

  • supervised earth mover s distance learning and its Computer Vision Applications
    European Conference on Computer Vision, 2012
    Co-Authors: Fan Wang, Leonidas J Guibas
    Abstract:

    The Earth Mover's Distance (EMD) is an intuitive and natural distance metric for comparing two histograms or probability distributions. It provides a distance value as well as a flow-network indicating how the probability mass is optimally transported between the bins. In traditional EMD, the ground distance between the bins is pre-defined. Instead, we propose to jointly optimize the ground distance matrix and the EMD flow-network based on a partial ordering of histogram distances in an optimization framework. Our method is further extended to accept information from general labeled pairs. The trained ground distance better reflects the cross-bin relationships, hence produces more accurate EMD values and flow-networks. Two Computer Vision Applications are used to demonstrate the effectiveness of the algorithm: first, we apply the optimized EMD value to face verification, and achieve state-of-the-art performance on the PubFig and the LFW data sets; second, the learned EMD flow-network is used to analyze face attribute changes, obtaining consistent paths that demonstrate intuitive transitions on certain facial attributes.

  • ECCV (1) - Supervised earth mover's distance learning and its Computer Vision Applications
    Computer Vision – ECCV 2012, 2012
    Co-Authors: Fan Wang, Leonidas J Guibas
    Abstract:

    The Earth Mover's Distance (EMD) is an intuitive and natural distance metric for comparing two histograms or probability distributions. It provides a distance value as well as a flow-network indicating how the probability mass is optimally transported between the bins. In traditional EMD, the ground distance between the bins is pre-defined. Instead, we propose to jointly optimize the ground distance matrix and the EMD flow-network based on a partial ordering of histogram distances in an optimization framework. Our method is further extended to accept information from general labeled pairs. The trained ground distance better reflects the cross-bin relationships, hence produces more accurate EMD values and flow-networks. Two Computer Vision Applications are used to demonstrate the effectiveness of the algorithm: first, we apply the optimized EMD value to face verification, and achieve state-of-the-art performance on the PubFig and the LFW data sets; second, the learned EMD flow-network is used to analyze face attribute changes, obtaining consistent paths that demonstrate intuitive transitions on certain facial attributes.

Zhou Fang - One of the best experts on this subject based on the ideXlab platform.

  • multi tenant mobile offloading systems for real time Computer Vision Applications
    International Conference of Distributed Computing and Networking, 2019
    Co-Authors: Zhou Fang, Jenghau Lin, Mani Srivastava, Rajesh Gupta
    Abstract:

    Offloading techniques enable many emerging Computer Vision Applications on mobile platforms by executing compute-intensive tasks on resource-rich servers. Although there have been a significant amount of research efforts devoted in optimizing mobile offloading frameworks, most previous works are evaluated in a single-tenant setting, that is, a server is assigned to a single client. However, in a practical scenario that servers must handle tasks from many clients running diverse Applications, contention on shared server resources may degrade application performance. In this work, we study scheduling techniques to improve serving performance in multi-tenant mobile offloading systems, for Computer Vision algorithms running on CPUs and deep neural networks (DNNs) running on GPUs. For CPU workloads, we present methods to mitigate resource contention and to improve delay using a Plan-Schedule approach. The planning phase predicts future workloads from all clients, estimates contention, and adjusts future task start times to remove or reduce contention. The scheduling phase dispatches arriving offloaded tasks to the server that minimizes contention. For DNN workloads running on GPUs, we propose adaptive batching algorithms using information of batch size, model complexity and system load to achieve the best Quality of Service (QoS), which are measured from accuracy and delay of DNN tasks. We demonstrate the improvement of serving performance using several real-world Applications with different server deployments.

  • ICDCN - Multi-tenant mobile offloading systems for real-time Computer Vision Applications
    Proceedings of the 20th International Conference on Distributed Computing and Networking, 2019
    Co-Authors: Zhou Fang, Jenghau Lin, Mani Srivastava, Rajesh Gupta
    Abstract:

    Offloading techniques enable many emerging Computer Vision Applications on mobile platforms by executing compute-intensive tasks on resource-rich servers. Although there have been a significant amount of research efforts devoted in optimizing mobile offloading frameworks, most previous works are evaluated in a single-tenant setting, that is, a server is assigned to a single client. However, in a practical scenario that servers must handle tasks from many clients running diverse Applications, contention on shared server resources may degrade application performance. In this work, we study scheduling techniques to improve serving performance in multi-tenant mobile offloading systems, for Computer Vision algorithms running on CPUs and deep neural networks (DNNs) running on GPUs. For CPU workloads, we present methods to mitigate resource contention and to improve delay using a Plan-Schedule approach. The planning phase predicts future workloads from all clients, estimates contention, and adjusts future task start times to remove or reduce contention. The scheduling phase dispatches arriving offloaded tasks to the server that minimizes contention. For DNN workloads running on GPUs, we propose adaptive batching algorithms using information of batch size, model complexity and system load to achieve the best Quality of Service (QoS), which are measured from accuracy and delay of DNN tasks. We demonstrate the improvement of serving performance using several real-world Applications with different server deployments.