The Experts below are selected from a list of 36 Experts worldwide ranked by ideXlab platform

Xin Dong - One of the best experts on this subject based on the ideXlab platform.

  • training for multi Resolution Inference using reusable quantization terms
    Architectural Support for Programming Languages and Operating Systems, 2021
    Co-Authors: Sai Qian Zhang, Bradley Mcdanel, H T Kung, Xin Dong
    Abstract:

    Low-Resolution uniform quantization (e.g., 4-bit bitwidth) for both Deep Neural Network (DNN) weights and data has emerged as an important technique for efficient Inference. Departing from conventional quantization, we describe a novel training approach to support Inference at multiple Resolutions by reusing a single set of quantization terms (the same set of nonzero bits in values). The proposed approach streamlines the training and supports dynamic selection of Resolution levels during Inference. We evaluate the method on a diverse range of applications including multiple CNNs on ImageNet, an LSTM on Wikitext-2, and YOLO-v5 on COCO. We show that models resulting from our multi-Resolution training can support up to 10 Resolutions with only a moderate performance reduction (e.g., ≤ 1%) compared to training them individually. Lastly, using an FPGA, we compare our multi-Resolution multiplier-accumulator (mMAC) against other conventional MAC designs and evaluate the Inference performance. We show that the mMAC design broadens the choices in trading off cost, efficiency, and latency across a range of computational budgets.

Sai Qian Zhang - One of the best experts on this subject based on the ideXlab platform.

  • training for multi Resolution Inference using reusable quantization terms
    Architectural Support for Programming Languages and Operating Systems, 2021
    Co-Authors: Sai Qian Zhang, Bradley Mcdanel, H T Kung, Xin Dong
    Abstract:

    Low-Resolution uniform quantization (e.g., 4-bit bitwidth) for both Deep Neural Network (DNN) weights and data has emerged as an important technique for efficient Inference. Departing from conventional quantization, we describe a novel training approach to support Inference at multiple Resolutions by reusing a single set of quantization terms (the same set of nonzero bits in values). The proposed approach streamlines the training and supports dynamic selection of Resolution levels during Inference. We evaluate the method on a diverse range of applications including multiple CNNs on ImageNet, an LSTM on Wikitext-2, and YOLO-v5 on COCO. We show that models resulting from our multi-Resolution training can support up to 10 Resolutions with only a moderate performance reduction (e.g., ≤ 1%) compared to training them individually. Lastly, using an FPGA, we compare our multi-Resolution multiplier-accumulator (mMAC) against other conventional MAC designs and evaluate the Inference performance. We show that the mMAC design broadens the choices in trading off cost, efficiency, and latency across a range of computational budgets.

Feng Cao - One of the best experts on this subject based on the ideXlab platform.

  • a multi clause dynamic deduction algorithm based on standard contradiction separation rule
    Information Sciences, 2021
    Co-Authors: Feng Cao, Jun Liu, Shuwei Chen
    Abstract:

    Abstract in the past decades, automated theorem proving (ATP) for first-order logic has made good progress, in which binary Resolution Inference rule plays a crucial role. However, as shown in the latest benchmark library of the ATP system, there are still many practical problems that have not been resolved or cannot be effectively resolved. Recently, in order to overcome the limitations of ATP based on binary Resolution Inference rules, a novel multi-clause dynamic standard contradiction separation (S-CS) Inference rule and its automated deduction theory have been proposed. Based on this theory, this paper first clarifies the generality of this S-CS rule by comparing it with some well-known variants of the binary Resolution rule, and then focuses on how to design a specific and effective algorithm along with search strategies to realize the S-CS based deductive theory with its implementation. Specifically, the present work proposes a novel S-CS dynamic deduction algorithm (in short SDDA) based on different strategies and summarizes its implementation procedures. In addition, we focus on evaluating whether SDDA, as a novel perspective multi-clause dynamic automatic deduction algorithm, can be applied on top of the current leading ATP system architectures to further improve their performances. Therefore, SDDA is applied to the current leading first-order ATP systems, i.e., Vampire and E, respectively forming two integrated APT systems, denoted as SDDA_V and SDDA_E. Then the capabilities of SDDA_V and SDDA_E are evaluated on the latest benchmark database TPTP, such as the CASC-J9 problems (FOF division) as well as the hard problems with a rating of 1 in the TPTP benchmark database. The experimental results show the effectiveness of SDDA: SDDA_V outperforms Vampire itself, and SDDA_E, outperforms E itself, and the two improved ATP systems have solved a number of hard problems with the rating of 1 in TPTP, that is, some problems in the latest benchmark database TPTP which have not yet been solved by other current first-order ATP systems.

Bradley Mcdanel - One of the best experts on this subject based on the ideXlab platform.

  • training for multi Resolution Inference using reusable quantization terms
    Architectural Support for Programming Languages and Operating Systems, 2021
    Co-Authors: Sai Qian Zhang, Bradley Mcdanel, H T Kung, Xin Dong
    Abstract:

    Low-Resolution uniform quantization (e.g., 4-bit bitwidth) for both Deep Neural Network (DNN) weights and data has emerged as an important technique for efficient Inference. Departing from conventional quantization, we describe a novel training approach to support Inference at multiple Resolutions by reusing a single set of quantization terms (the same set of nonzero bits in values). The proposed approach streamlines the training and supports dynamic selection of Resolution levels during Inference. We evaluate the method on a diverse range of applications including multiple CNNs on ImageNet, an LSTM on Wikitext-2, and YOLO-v5 on COCO. We show that models resulting from our multi-Resolution training can support up to 10 Resolutions with only a moderate performance reduction (e.g., ≤ 1%) compared to training them individually. Lastly, using an FPGA, we compare our multi-Resolution multiplier-accumulator (mMAC) against other conventional MAC designs and evaluate the Inference performance. We show that the mMAC design broadens the choices in trading off cost, efficiency, and latency across a range of computational budgets.

H T Kung - One of the best experts on this subject based on the ideXlab platform.

  • training for multi Resolution Inference using reusable quantization terms
    Architectural Support for Programming Languages and Operating Systems, 2021
    Co-Authors: Sai Qian Zhang, Bradley Mcdanel, H T Kung, Xin Dong
    Abstract:

    Low-Resolution uniform quantization (e.g., 4-bit bitwidth) for both Deep Neural Network (DNN) weights and data has emerged as an important technique for efficient Inference. Departing from conventional quantization, we describe a novel training approach to support Inference at multiple Resolutions by reusing a single set of quantization terms (the same set of nonzero bits in values). The proposed approach streamlines the training and supports dynamic selection of Resolution levels during Inference. We evaluate the method on a diverse range of applications including multiple CNNs on ImageNet, an LSTM on Wikitext-2, and YOLO-v5 on COCO. We show that models resulting from our multi-Resolution training can support up to 10 Resolutions with only a moderate performance reduction (e.g., ≤ 1%) compared to training them individually. Lastly, using an FPGA, we compare our multi-Resolution multiplier-accumulator (mMAC) against other conventional MAC designs and evaluate the Inference performance. We show that the mMAC design broadens the choices in trading off cost, efficiency, and latency across a range of computational budgets.