The Experts below are selected from a list of 5121 Experts worldwide ranked by ideXlab platform

Chiouyng Lee - One of the best experts on this subject based on the ideXlab platform.

  • embracing systolic super systolization of large scale Circulant Matrix vector multiplication on fpga with subquadratic space complexity
    Field Programmable Gate Arrays, 2019
    Co-Authors: Jiafeng Xie, Chiouyng Lee
    Abstract:

    The recent advance in artificial intelligence (AI) technology has led to a new round of systolic structure innovation. Many AI accelerators have employed systolic structure to realize the core large-scale Matrix-vector multiplication for high-performance processing, which has a complexity of $o(n^2)$ for Matrix size of $n\times n$ (difficult to be implemented on the field-programmable gate array (FPGA) platform). To overcome this drawback, in this paper, we propose a super systolization strategy to implement the core Circulant Matrix-vector multiplication into a systolic structure with subquadratic space complexity. The proposed effort is carried out through two stages of coherent interdependent efforts: (i) a novel Matrix-vector multiplication algorithm based on Toeplitz Matrix-vector product (TMVP) approach is proposed to obtain subquadratic space complexity; (ii) a series of optimization techniques are introduced to map the proposed algorithm into desired systolic structure. Finally, detailed complexity analysis and comparison have been conducted to prove the efficiency of the proposed strategy. The proposed strategy is highly efficient and can be extended in many neural network based hardware implementation platforms.

  • FPGA - Embracing Systolic: Super Systolization of Large-Scale Circulant Matrix-vector Multiplication on FPGA with Subquadratic Space Complexity
    Proceedings of the 2019 ACM SIGDA International Symposium on Field-Programmable Gate Arrays, 2019
    Co-Authors: Jiafeng Xie, Chiouyng Lee
    Abstract:

    The recent advance in artificial intelligence (AI) technology has led to a new round of systolic structure innovation. Many AI accelerators have employed systolic structure to realize the core large-scale Matrix-vector multiplication for high-performance processing, which has a complexity of $o(n^2)$ for Matrix size of $n\times n$ (difficult to be implemented on the field-programmable gate array (FPGA) platform). To overcome this drawback, in this paper, we propose a super systolization strategy to implement the core Circulant Matrix-vector multiplication into a systolic structure with subquadratic space complexity. The proposed effort is carried out through two stages of coherent interdependent efforts: (i) a novel Matrix-vector multiplication algorithm based on Toeplitz Matrix-vector product (TMVP) approach is proposed to obtain subquadratic space complexity; (ii) a series of optimization techniques are introduced to map the proposed algorithm into desired systolic structure. Finally, detailed complexity analysis and comparison have been conducted to prove the efficiency of the proposed strategy. The proposed strategy is highly efficient and can be extended in many neural network based hardware implementation platforms.

Jiafeng Xie - One of the best experts on this subject based on the ideXlab platform.

  • embracing systolic super systolization of large scale Circulant Matrix vector multiplication on fpga with subquadratic space complexity
    Field Programmable Gate Arrays, 2019
    Co-Authors: Jiafeng Xie, Chiouyng Lee
    Abstract:

    The recent advance in artificial intelligence (AI) technology has led to a new round of systolic structure innovation. Many AI accelerators have employed systolic structure to realize the core large-scale Matrix-vector multiplication for high-performance processing, which has a complexity of $o(n^2)$ for Matrix size of $n\times n$ (difficult to be implemented on the field-programmable gate array (FPGA) platform). To overcome this drawback, in this paper, we propose a super systolization strategy to implement the core Circulant Matrix-vector multiplication into a systolic structure with subquadratic space complexity. The proposed effort is carried out through two stages of coherent interdependent efforts: (i) a novel Matrix-vector multiplication algorithm based on Toeplitz Matrix-vector product (TMVP) approach is proposed to obtain subquadratic space complexity; (ii) a series of optimization techniques are introduced to map the proposed algorithm into desired systolic structure. Finally, detailed complexity analysis and comparison have been conducted to prove the efficiency of the proposed strategy. The proposed strategy is highly efficient and can be extended in many neural network based hardware implementation platforms.

  • FPGA - Embracing Systolic: Super Systolization of Large-Scale Circulant Matrix-vector Multiplication on FPGA with Subquadratic Space Complexity
    Proceedings of the 2019 ACM SIGDA International Symposium on Field-Programmable Gate Arrays, 2019
    Co-Authors: Jiafeng Xie, Chiouyng Lee
    Abstract:

    The recent advance in artificial intelligence (AI) technology has led to a new round of systolic structure innovation. Many AI accelerators have employed systolic structure to realize the core large-scale Matrix-vector multiplication for high-performance processing, which has a complexity of $o(n^2)$ for Matrix size of $n\times n$ (difficult to be implemented on the field-programmable gate array (FPGA) platform). To overcome this drawback, in this paper, we propose a super systolization strategy to implement the core Circulant Matrix-vector multiplication into a systolic structure with subquadratic space complexity. The proposed effort is carried out through two stages of coherent interdependent efforts: (i) a novel Matrix-vector multiplication algorithm based on Toeplitz Matrix-vector product (TMVP) approach is proposed to obtain subquadratic space complexity; (ii) a series of optimization techniques are introduced to map the proposed algorithm into desired systolic structure. Finally, detailed complexity analysis and comparison have been conducted to prove the efficiency of the proposed strategy. The proposed strategy is highly efficient and can be extended in many neural network based hardware implementation platforms.

Zhao-lin Jiang - One of the best experts on this subject based on the ideXlab platform.

Moon Ho Lee - One of the best experts on this subject based on the ideXlab platform.

  • one bit feedback for quasi orthogonal space time block codes based on Circulant Matrix
    IEEE Transactions on Wireless Communications, 2009
    Co-Authors: Zhu Chen, Moon Ho Lee
    Abstract:

    During the last few years, a number of Quasi-Orthogonal Space-Time Block Codes (QOSTBC) have been proposed for using in multiple transmit antennas systems. In this letter, based on Circulant Matrix, we propose a novel method of extending any QOSTBC constructed for 4 transmit antennas to a closed-loop scheme. We show that with the aid of multiplying the entries of QOSTBC code words by the appropriate phase factors which depend on the channel information, the proposed scheme can improve its transmit diversity with one bit feedback. The performances of the proposed scenario extended from Jafarkhani's QOSTBC as well as its optimal constellation rotated scheme are analyzed. The simulation results suggest that there is a significant Eb/No advantage in the proposed scheme which is able to be designed easily.

Patrice Abry - One of the best experts on this subject based on the ideXlab platform.

  • Smoothing Windows for the Synthesis of Gaussian Stationary Random Fields Using Circulant Matrix Embedding
    Journal of Computational and Graphical Statistics, 2014
    Co-Authors: Hannes Helgason, Vladas Pipiras, Patrice Abry
    Abstract:

    When generating Gaussian stationary random fields, a standard method based on Circulant Matrix embedding usually fails because some of the associated eigenvalues are negative. The eigenvalues can be shown to be nonnegative in the limit of increasing sample size. Computationally feasible large sample sizes, however, rarely lead to nonnegative eigenvalues. Another solution is to extend suitably the covariance function of interest so that the eigenvalues of the embedded Circulant Matrix become nonnegative in theory. Though such extensions have been found for a number of examples of stationary fields, the method depends on nontrivial constructions in specific cases.In this work, the embedded Circulant Matrix is smoothed at the boundary by using a cutoff window or overlapping windows over a transition region. The windows are not specific to particular examples of stationary fields. The resulting method modifies the standard Circulant embedding, and is easy to use. It is shown that this straightforward approach...

  • Synthesis of multivariate stationary series with prescribed marginal distributions and covariance using Circulant Matrix embedding
    Signal Processing, 2011
    Co-Authors: Hannes Helgason, Vladas Pipiras, Patrice Abry
    Abstract:

    The problem of synthesizing multivariate stationary series Y[n]=(Y"1[n],...,Y"P[n])^T, [email protected]?Z, with prescribed non-Gaussian marginal distributions, and a targeted covariance structure, is addressed. The focus is on constructions based on a memoryless transformation Y"p[n]=f"p(X"p[n]) of a multivariate stationary Gaussian series X[n]=(X"1[n],...,X"P[n])^T. The mapping between the targeted covariance and that of the Gaussian series is expressed via Hermite expansions. The various choices of the transforms f"p for a prescribed marginal distribution are discussed in a comprehensive manner. The interplay between the targeted marginal distributions, the choice of the transforms f"p, and on the resulting reachability of the targeted covariance, is discussed theoretically and illustrated on examples. Also, an original practical procedure warranting positive definiteness for the transformed covariance at the price of approximating the targeted covariance is proposed, based on a simple and natural modification of the popular Circulant Matrix embedding technique. The applications of the proposed methodology are also discussed in the context of network traffic modeling. Matlab codes implementing the proposed synthesis procedure are publicly available at http://www.hermir.org.