The Experts below are selected from a list of 11946 Experts worldwide ranked by ideXlab platform

Hoijun Yoo - One of the best experts on this subject based on the ideXlab platform.

  • hnpu an adaptive dnn training processor utilizing stochastic dynamic fixed point and active Bit Precision searching
    IEEE Journal of Solid-state Circuits, 2021
    Co-Authors: Donghyeon Han, Gwangtae Park, Youngwoo Kim, Seokchan Song, Juhyoung Lee, Hoijun Yoo
    Abstract:

    This article presents HNPU, which is an energy-efficient deep neural network (DNN) training processor by adopting algorithm-hardware co-design. The HNPU supports stochastic dynamic fixed-point representation and layer-wise adaptive Precision searching unit for low-Bit-Precision training. It additionally utilizes slice-level reconfigurability and sparsity to maximize its efficiency both in DNN inference and training. Adaptive bandwidth reconfigurable accumulation network enables reconfigurable DNN allocation and maintains its high core utilization even in various Bit-Precision conditions. Fabricated in a 28-nm process, the HNPU accomplished at least 5.9x higher energy efficiency and 2.5x higher area efficiency in actual DNN training compared with the previous state-of-the-art on-chip learning processors.

  • hnpu an adaptive dnn training processor utilizing stochastic dynamic fixed point and active Bit Precision searching
    IEEE Journal of Solid-state Circuits, 2021
    Co-Authors: Donghyeon Han, Gwangtae Park, Youngwoo Kim, Seokchan Song, Juhyoung Lee, Hoijun Yoo
    Abstract:

    This article presents HNPU, which is an energy-efficient deep neural network (DNN) training processor by adopting algorithm-hardware co-design. The HNPU supports stochastic dynamic fixed-point representation and layer-wise adaptive Precision searching unit for low-Bit-Precision training. It additionally utilizes slice-level reconfigurability and sparsity to maximize its efficiency both in DNN inference and training. Adaptive bandwidth reconfigurable accumulation network enables reconfigurable DNN allocation and maintains its high core utilization even in various Bit-Precision conditions. Fabricated in a 28-nm process, the HNPU accomplished at least $5.9\times $ higher energy efficiency and $2.5\times $ higher area efficiency in actual DNN training compared with the previous state-of-the-art on-chip learning processors.

  • 1b-16b Variable Bit Precision DNN Processor for Emotional HRI System in Mobile Devices
    2020
    Co-Authors: Changhyeon Kim, Jinmook Lee, Sang Hoon Kang, Sangyeob Kim, Hoijun Yoo
    Abstract:

    We propose an energy-efficient DNN processor with the proposed look-up-table-based processing engine (LPE) and near-zero skipper. A CNN-based facial emotion recognition model and an RNN-based emotional dialogue generation model are integrated for the natural human-robot interaction (HRI) system, and it is evaluated by the proposed processor. LPE supports 1 to 16 Bit variable weight Bit Precision, and it achieves 57.6% and 28.5% lower energy consumption than the conventional multiplier-accumulator (MAC) units in 1-16 Bit weight Precision. Furthermore, the near-zero skipper reduces 36% of MAC operations and consumes 28% lower energy consumption in facial emotion recognition tasks. Implemented in 65 nm CMOS process, the proposed processor occupies 1784×1784 μm2 areas and dissipates 0.28 mW and 34.4 mW at 1 frame-per-second (fps) and 30 fps facial emotion recognition tasks.

  • unpu an energy efficient deep neural network accelerator with fully variable weight Bit Precision
    IEEE Journal of Solid-state Circuits, 2019
    Co-Authors: Jinmook Lee, Changhyeon Kim, Sang Hoon Kang, Dongjoo Shin, Sangyeob Kim, Hoijun Yoo
    Abstract:

    An energy-efficient deep neural network (DNN) accelerator, unified neural processing unit (UNPU), is proposed for mobile deep learning applications. The UNPU can support both convolutional layers (CLs) and recurrent or fully connected layers (FCLs) to support versatile workload combinations to accelerate various mobile deep learning applications. In addition, the UNPU is the first DNN accelerator ASIC that can support fully variable weight Bit Precision from 1 to 16 Bit. It enables the UNPU to operate on the accuracy-energy optimal point. Moreover, the lookup table (LUT)-based Bit-serial processing element (LBPE) in the UNPU achieves the energy consumption reduction compared to the conventional fixed-point multiply-and-accumulate (MAC) array by 23.1%, 27.2%, 41%, and 53.6% for the 16-, 8-, 4-, and 1-Bit weight Precision, respectively. Besides the energy efficiency improvement, the unified DNN core architecture of the UNPU improves the peak performance for CL by 1.15 $\times$ compared to the previous work. It makes the UNPU operate on the lower voltage and frequency for the given DNN to increase energy efficiency. The UNPU is implemented in 65-nm CMOS technology and occupies the $4 \times 4$ mm2 die area. The UNPU can operates from 0.63- to 1.1-V supply voltage with maximum frequency of 200 MHz. The UNPU has peak performance of 345.6 GOPS for 16-Bit weight Precision and 7372 GOPS for 1-Bit weight Precision. The wide operating range of UNPU makes the UNPU achieve the power efficiency of 3.08 TOPS/W for 16-Bit weight Precision and 50.6 TOPS/W for 1-Bit weight Precision. The functionality of the UNPU is successfully demonstrated on the verification system using ImageNet deep CNN (VGG-16).

  • unpu a 50 6tops w unified deep neural network accelerator with 1b to 16b fully variable weight Bit Precision
    International Solid-State Circuits Conference, 2018
    Co-Authors: Jinmook Lee, Changhyeon Kim, Sang Hoon Kang, Dongjoo Shin, Sangyeob Kim, Hoijun Yoo
    Abstract:

    Deep neural network (DNN) accelerators [1-3] have been proposed to accelerate deep learning algorithms from face recognition to emotion recognition in mobile or embedded environments [3]. However, most works accelerate only the convolutional layers (CLs) or fully-connected layers (FCLs), and different DNNs, such as those containing recurrent layers (RLs) (useful for emotion recognition) have not been supported in hardware. A combined CNN-RNN accelerator [1], separately optimizing the computation-dominant CLs, and memory-dominant RLs or FCLs, was reported to increase overall performance, however, the number of processing elements (PEs) for CLs and RLs was limited by their area and consequently, performance was suboptimal in scenarios requiring only CLs or only RLs. Although the PEs for RLs can be reconfigured into PEs for CLs or vice versa, only a partial reconfiguration was possible resulting in marginal performance improvement. Moreover, previous works [1-2] supported a limited set of weight Bit Precisions, such as either 4b or 8b or 16b. However, lower weight Bit-Precisions can achieve better throughput and higher energy efficiency, and the optimal Bit-Precision can be varied according to different accuracy/performance requirements. Therefore, a unified DNN accelerator with fully-variable weight Bit-Precision is required for the energy-optimal operation of DNNs within a mobile environment.

Kaushik Roy - One of the best experts on this subject based on the ideXlab platform.

  • sbsnn stochastic Bits enabled binary spiking neural network with on chip learning for energy efficient neuromorphic computing at the edge
    IEEE Transactions on Circuits and Systems I-regular Papers, 2020
    Co-Authors: Minsuk Koo, Gopalakrishnan Srinivasan, Yong Shim, Kaushik Roy
    Abstract:

    In this work, we propose stochastic Binary Spiking Neural Network (sBSNN) composed of stochastic spiking neurons and binary synapses (stochastic only during training) that computes probabilistically with one-Bit Precision for power-efficient and memory-compressed neuromorphic computing. We present an energy-efficient implementation of the proposed sBSNN using ‘stochastic Bit’ as the core computational primitive to realize the stochastic neurons and synapses, which are fabricated in 90nm CMOS process, to achieve efficient on-chip training and inference for image recognition tasks. The measured data shows that the ‘stochastic Bit’ can be programmed to mimic spiking neurons, and stochastic Spike Timing Dependent Plasticity (or sSTDP) rule for training the binary synaptic weights without expensive random number generators. Our results indicate that the proposed sBSNN realization offers possibility of up to $32\times $ neuronal and synaptic memory compression compared to full Precision (32-Bit) SNN and energy efficiency of 89.49 TOPS/Watt for two-layer fully-connected SNN.

  • sbsnn stochastic Bits enabled binary spiking neural network with on chip learning for energy efficient neuromorphic computing at the edge
    arXiv: Emerging Technologies, 2020
    Co-Authors: Minsuk Koo, Gopalakrishnan Srinivasan, Yong Shim, Kaushik Roy
    Abstract:

    In this work, we propose stochastic Binary Spiking Neural Network (sBSNN) composed of stochastic spiking neurons and binary synapses (stochastic only during training) that computes probabilistically with one-Bit Precision for power-efficient and memory-compressed neuromorphic computing. We present an energy-efficient implementation of the proposed sBSNN using 'stochastic Bit' as the core computational primitive to realize the stochastic neurons and synapses, which are fabricated in 90nm CMOS process, to achieve efficient on-chip training and inference for image recognition tasks. The measured data shows that the 'stochastic Bit' can be programmed to mimic spiking neurons, and stochastic Spike Timing Dependent Plasticity (or sSTDP) rule for training the binary synaptic weights without expensive random number generators. Our results indicate that the proposed sBSNN realization offers possibility of up to 32x neuronal and synaptic memory compression compared to full Precision (32-Bit) SNN and energy efficiency of 89.49 TOPS/Watt for two-layer fully-connected SNN.

  • dynamic Bit width adaptation in dct an approach to trade off image quality and computation energy
    IEEE Transactions on Very Large Scale Integration Systems, 2010
    Co-Authors: Jongsun Park, Jung Hwan Choi, Kaushik Roy
    Abstract:

    This paper presents a dynamic Bit-width adaptation scheme for applications using discrete cosine transform (DCT). The technique can efficiently trade off image quality and computation energy. Based on sensitivity differences of 64 DCT coefficients, separate operand Bit-widths are used for different frequency components to reduce computation energy. To select the appropriate operand Bit-widths that achieve significant reduction of power consumption with minimum image quality degradation, we also propose a Bit-width selection algorithm. The proposed variable Bit Precision DCT algorithm can be efficiently implemented using carry save adder trees. The reconfigurable DCT architecture can achieve power savings ranging from 36% to 75% compared to normal operation at the expense of minor image quality degradation.

Juhyoung Lee - One of the best experts on this subject based on the ideXlab platform.

  • hnpu an adaptive dnn training processor utilizing stochastic dynamic fixed point and active Bit Precision searching
    IEEE Journal of Solid-state Circuits, 2021
    Co-Authors: Donghyeon Han, Gwangtae Park, Youngwoo Kim, Seokchan Song, Juhyoung Lee, Hoijun Yoo
    Abstract:

    This article presents HNPU, which is an energy-efficient deep neural network (DNN) training processor by adopting algorithm-hardware co-design. The HNPU supports stochastic dynamic fixed-point representation and layer-wise adaptive Precision searching unit for low-Bit-Precision training. It additionally utilizes slice-level reconfigurability and sparsity to maximize its efficiency both in DNN inference and training. Adaptive bandwidth reconfigurable accumulation network enables reconfigurable DNN allocation and maintains its high core utilization even in various Bit-Precision conditions. Fabricated in a 28-nm process, the HNPU accomplished at least 5.9x higher energy efficiency and 2.5x higher area efficiency in actual DNN training compared with the previous state-of-the-art on-chip learning processors.

  • hnpu an adaptive dnn training processor utilizing stochastic dynamic fixed point and active Bit Precision searching
    IEEE Journal of Solid-state Circuits, 2021
    Co-Authors: Donghyeon Han, Gwangtae Park, Youngwoo Kim, Seokchan Song, Juhyoung Lee, Hoijun Yoo
    Abstract:

    This article presents HNPU, which is an energy-efficient deep neural network (DNN) training processor by adopting algorithm-hardware co-design. The HNPU supports stochastic dynamic fixed-point representation and layer-wise adaptive Precision searching unit for low-Bit-Precision training. It additionally utilizes slice-level reconfigurability and sparsity to maximize its efficiency both in DNN inference and training. Adaptive bandwidth reconfigurable accumulation network enables reconfigurable DNN allocation and maintains its high core utilization even in various Bit-Precision conditions. Fabricated in a 28-nm process, the HNPU accomplished at least $5.9\times $ higher energy efficiency and $2.5\times $ higher area efficiency in actual DNN training compared with the previous state-of-the-art on-chip learning processors.

  • Z-PIM: A Sparsity-Aware Processing-in-Memory Architecture With Fully Variable Weight Bit-Precision for Energy-Efficient Deep Neural Networks
    IEEE Journal of Solid-state Circuits, 2021
    Co-Authors: Ji-hoon Kim, Juhyoung Lee, Jinsu Lee, Jaehoon Heo, Joo-young Kim
    Abstract:

    We present an energy-efficient processing-in-memory (PIM) architecture named Z-PIM that supports both sparsity handling and fully variable Bit-Precision in weight data for energy-efficient deep neural networks. Z-PIM adopts the Bit-serial arithmetic that performs a multiplication Bit-by-Bit through multiple cycles to reduce the complexity of the operation in a single cycle and to provide flexibility in Bit-Precision. To this end, it employs a zero-skipping convolution SRAM, which performs in-memory and operations based on custom 8T-SRAM cells and channel-wise accumulations, and a diagonal accumulation SRAM that performs Bit- and spatial-wise accumulation on the channel-wise accumulation results using diagonal logic and adders to produce the final convolution outputs. We propose the hierarchical Bitline structure for energy-efficient weight Bit pre-charging and computational readout by reducing the parasitic capacitances of the Bitlines. Its charge reuse scheme reduces the switching rate by 95.42% for the convolution layers of VGG-16 model. In addition, Z-PIM's channel-wise data mapping enables sparsity handling by skip-reading the input channels with zero weight. Its read-operation pipelining enabled by a read-sequence scheduling improves the throughput by 66.1%. The Z-PIM chip is fabricated in a 65-nm CMOS process on a 7.568-mm² die, while it consumes average 5.294-mW power at 1.0-V voltage and 200-MHz frequency. It achieves 0.31-49.12-TOPS/W energy efficiency for convolution operations as the weight sparsity and Bit-Precision vary from 0.1 to 0.9 and 1 to 16 Bit, respectively. For the figure of merit considering input Bit-width, weight Bit-width, and energy efficiency, the Z-PIM shows more than 2.1 times improvement over the state-of-the-art PIM implementations.

Minsuk Koo - One of the best experts on this subject based on the ideXlab platform.

  • sbsnn stochastic Bits enabled binary spiking neural network with on chip learning for energy efficient neuromorphic computing at the edge
    IEEE Transactions on Circuits and Systems I-regular Papers, 2020
    Co-Authors: Minsuk Koo, Gopalakrishnan Srinivasan, Yong Shim, Kaushik Roy
    Abstract:

    In this work, we propose stochastic Binary Spiking Neural Network (sBSNN) composed of stochastic spiking neurons and binary synapses (stochastic only during training) that computes probabilistically with one-Bit Precision for power-efficient and memory-compressed neuromorphic computing. We present an energy-efficient implementation of the proposed sBSNN using ‘stochastic Bit’ as the core computational primitive to realize the stochastic neurons and synapses, which are fabricated in 90nm CMOS process, to achieve efficient on-chip training and inference for image recognition tasks. The measured data shows that the ‘stochastic Bit’ can be programmed to mimic spiking neurons, and stochastic Spike Timing Dependent Plasticity (or sSTDP) rule for training the binary synaptic weights without expensive random number generators. Our results indicate that the proposed sBSNN realization offers possibility of up to $32\times $ neuronal and synaptic memory compression compared to full Precision (32-Bit) SNN and energy efficiency of 89.49 TOPS/Watt for two-layer fully-connected SNN.

  • sbsnn stochastic Bits enabled binary spiking neural network with on chip learning for energy efficient neuromorphic computing at the edge
    arXiv: Emerging Technologies, 2020
    Co-Authors: Minsuk Koo, Gopalakrishnan Srinivasan, Yong Shim, Kaushik Roy
    Abstract:

    In this work, we propose stochastic Binary Spiking Neural Network (sBSNN) composed of stochastic spiking neurons and binary synapses (stochastic only during training) that computes probabilistically with one-Bit Precision for power-efficient and memory-compressed neuromorphic computing. We present an energy-efficient implementation of the proposed sBSNN using 'stochastic Bit' as the core computational primitive to realize the stochastic neurons and synapses, which are fabricated in 90nm CMOS process, to achieve efficient on-chip training and inference for image recognition tasks. The measured data shows that the 'stochastic Bit' can be programmed to mimic spiking neurons, and stochastic Spike Timing Dependent Plasticity (or sSTDP) rule for training the binary synaptic weights without expensive random number generators. Our results indicate that the proposed sBSNN realization offers possibility of up to 32x neuronal and synaptic memory compression compared to full Precision (32-Bit) SNN and energy efficiency of 89.49 TOPS/Watt for two-layer fully-connected SNN.

Donghyeon Han - One of the best experts on this subject based on the ideXlab platform.

  • hnpu an adaptive dnn training processor utilizing stochastic dynamic fixed point and active Bit Precision searching
    IEEE Journal of Solid-state Circuits, 2021
    Co-Authors: Donghyeon Han, Gwangtae Park, Youngwoo Kim, Seokchan Song, Juhyoung Lee, Hoijun Yoo
    Abstract:

    This article presents HNPU, which is an energy-efficient deep neural network (DNN) training processor by adopting algorithm-hardware co-design. The HNPU supports stochastic dynamic fixed-point representation and layer-wise adaptive Precision searching unit for low-Bit-Precision training. It additionally utilizes slice-level reconfigurability and sparsity to maximize its efficiency both in DNN inference and training. Adaptive bandwidth reconfigurable accumulation network enables reconfigurable DNN allocation and maintains its high core utilization even in various Bit-Precision conditions. Fabricated in a 28-nm process, the HNPU accomplished at least 5.9x higher energy efficiency and 2.5x higher area efficiency in actual DNN training compared with the previous state-of-the-art on-chip learning processors.

  • hnpu an adaptive dnn training processor utilizing stochastic dynamic fixed point and active Bit Precision searching
    IEEE Journal of Solid-state Circuits, 2021
    Co-Authors: Donghyeon Han, Gwangtae Park, Youngwoo Kim, Seokchan Song, Juhyoung Lee, Hoijun Yoo
    Abstract:

    This article presents HNPU, which is an energy-efficient deep neural network (DNN) training processor by adopting algorithm-hardware co-design. The HNPU supports stochastic dynamic fixed-point representation and layer-wise adaptive Precision searching unit for low-Bit-Precision training. It additionally utilizes slice-level reconfigurability and sparsity to maximize its efficiency both in DNN inference and training. Adaptive bandwidth reconfigurable accumulation network enables reconfigurable DNN allocation and maintains its high core utilization even in various Bit-Precision conditions. Fabricated in a 28-nm process, the HNPU accomplished at least $5.9\times $ higher energy efficiency and $2.5\times $ higher area efficiency in actual DNN training compared with the previous state-of-the-art on-chip learning processors.