The Experts below are selected from a list of 117135 Experts worldwide ranked by ideXlab platform
Daisuke Takahashi - One of the best experts on this subject based on the ideXlab platform.
-
PPAM (1) - Implementation of Parallel 3-D Real FFT with 2-D Decomposition on Intel Xeon Phi Clusters
Parallel Processing and Applied Mathematics, 2020Co-Authors: Daisuke TakahashiAbstract:In this paper, we propose an implementation of a parallel 3-D real fast Fourier transform (FFT) with 2-D decomposition on Intel Xeon Phi clusters. The proposed implementation of the parallel 3-D real FFT is based on the conjugate symmetry property of the discrete Fourier transform (DFT) and the row-column FFT algorithm. We vectorized FFT kernels using the Intel Advanced Vector Extensions 512 (Intel AVX-512) instructions. Performance results of parallel 3-D real FFTs on Intel Xeon Phi clusters are reported. We successfully achieved a level of performance over 10 TFlops on 2048 nodes of Fujitsu PRIMERGY CX1640 M1 cluster for an \(8192^3\)-point FFT.
-
Mixed-Radix FFT Algorithms
High-Performance Computing Series, 2019Co-Authors: Daisuke TakahashiAbstract:This chapter presents Mixed-Radix FFT Algorithms. First, two-dimensional formulation of DFT is given. Next, radix-3, 4, 5, and 8 FFT algorithms are described.
-
High-Performance FFT Algorithms
High-Performance Computing Series, 2019Co-Authors: Daisuke TakahashiAbstract:This chapter presents high-performance FFT algorithms. First, the four-step FFT algorithm and five-step FFT algorithm are described. Next, the six-step FFT algorithm and blocked six-step FFT algorithm are explained. Then, nine-step FFT algorithm and recursive six-step FFT, and blocked multidimensional FFT algorithms are described. Finally, FFT algorithms suitable for fused multiply–add instructions and FFT algorithms for SIMD instructions are explained.
-
ICCSA (1) - An Implementation of Parallel 1-D Real FFT on Intel Xeon Phi Processors
Computational Science and Its Applications – ICCSA 2017, 2017Co-Authors: Daisuke TakahashiAbstract:In this paper, we propose an implementation of a parallel one-dimensional real fast Fourier transform (FFT) on Intel Xeon Phi processors. The proposed implementation of the parallel one-dimensional real FFT is based on the conjugate symmetry property for the discrete Fourier transform (DFT) and the six-step FFT algorithm. We vectorized FFT kernels using the Intel Advanced Vector Extensions 512 (AVX-512) instructions, and parallelized the six-step FFT by using OpenMP. Performance results of one-dimensional FFTs on Intel Xeon Phi processors are reported. We successfully achieved a performance of over 91 GFlops on an Intel Xeon Phi 7250 (1.4 GHz, 68 cores) for a \(2^{29}\)-point real FFT.
Asmita Haveliya - One of the best experts on this subject based on the ideXlab platform.
-
design and simulation of 32 point FFT using radix 2 algorithm for fpga implementation
International Conference on Advanced Computing, 2012Co-Authors: Asmita HaveliyaAbstract:The Fast Fourier Transform (FFT) is one of the rudimentary operations in field of digital signal and image processing. Some of the very vital applications of the fast fourier transform include Signal analysis, Sound filtering, Data compression, Partial differential equations, Multiplication of large integers, Image filtering etc. Fast Fourier transform (FFT) is an efficient implementation of the discrete Fourier transform (DFT). This paper concentrates on the development of the Fast Fourier Transform (FFT), based on Decimation-In-Time (DIT) domain, Radix-2 algorithm, this paper uses VHDL as a design entity, and their Synthesis by Xilinx Synthesis Tool on Vertex kit has been done. The input of Fast Fourier transform has been given by a PS2 KEYBOARD using a test bench and output has been displayed using the waveforms on the Xilinx Design Suite 12.1. The synthesis results show that the computation for calculating the 32-point Fast Fourier transform is efficient in terms of speed.
Mohammad Rastegari - One of the best experts on this subject based on the ideXlab platform.
-
Butterfly Transform: An Efficient FFT Based Neural Architecture Design
2020 IEEE CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020Co-Authors: Keivan Alizadeh Vahid, Anish Prabhu, Ali Farhadi, Mohammad RastegariAbstract:In this paper, we show that extending the butterfly operations from the FFT algorithm to a general Butterfly Transform (BFT) can be beneficial in building an efficient block structure for CNN designs. Pointwise convolutions, which we refer to as channel fusions, are the main computational bottleneck in the state-of-the-art efficient CNNs (e.g. MobileNets). We introduce a set of criterion for channel fusion, and prove that BFT yields an asymptotically optimal FLOP count with respect to these criteria. By replacing pointwise convolutions with BFT, we reduce the computational complexity of these layers from O(n^2) to O(n log n) with respect to the number of channels. Our experimental evaluations show that our method results in significant accuracy gains across a wide range of network architectures, especially at low FLOP ranges. For example, BFT results in up to a 6.75% absolute Top-1 improvement for MobileNetV1, 4.4 % for ShuffleNet V2 and 5.4% for MobileNetV3 on ImageNet under a similar number of FLOPS. Notably, ShuffleNet-V2+BFT outperforms state-of-the-art architecture search methods MNasNet, FBNet and MobilenetV3 in the low FLOP regime.
Keivan Alizadeh Vahid - One of the best experts on this subject based on the ideXlab platform.
-
Butterfly Transform: An Efficient FFT Based Neural Architecture Design
2020 IEEE CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020Co-Authors: Keivan Alizadeh Vahid, Anish Prabhu, Ali Farhadi, Mohammad RastegariAbstract:In this paper, we show that extending the butterfly operations from the FFT algorithm to a general Butterfly Transform (BFT) can be beneficial in building an efficient block structure for CNN designs. Pointwise convolutions, which we refer to as channel fusions, are the main computational bottleneck in the state-of-the-art efficient CNNs (e.g. MobileNets). We introduce a set of criterion for channel fusion, and prove that BFT yields an asymptotically optimal FLOP count with respect to these criteria. By replacing pointwise convolutions with BFT, we reduce the computational complexity of these layers from O(n^2) to O(n log n) with respect to the number of channels. Our experimental evaluations show that our method results in significant accuracy gains across a wide range of network architectures, especially at low FLOP ranges. For example, BFT results in up to a 6.75% absolute Top-1 improvement for MobileNetV1, 4.4 % for ShuffleNet V2 and 5.4% for MobileNetV3 on ImageNet under a similar number of FLOPS. Notably, ShuffleNet-V2+BFT outperforms state-of-the-art architecture search methods MNasNet, FBNet and MobilenetV3 in the low FLOP regime.
J.v. Mccanny - One of the best experts on this subject based on the ideXlab platform.
-
A 64-point Fourier transform chip for video motion compensation using phase correlation
IEEE Journal of Solid-State Circuits, 1996Co-Authors: C. Chiu, Hui, Tiong Jiu Ding, J.v. MccannyAbstract:Details of a new low power fast Fourier transform (FFT) processor for use in digital television applications are presented. This has been fabricated using a 0.6-/spl mu/m CMOS technology and can perform a 64 point complex forward or inverse FFT on real-time video at up to 18 Megasamples per second. It comprises 0.5 million transistors in a die area of 7.8/spl times/8 mm/sup 2/ and dissipates 1 W. The chip design is based on a novel VLSI architecture which has been derived from a first principles factorization of the discrete Fourier transform (DFT) matrix and tailored to a direct silicon implementation.