The Experts below are selected from a list of 225 Experts worldwide ranked by ideXlab platform

Jack Dongarra - One of the best experts on this subject based on the ideXlab platform.

  • soft error resilient qr factorization for hybrid system with gpgpu
    Journal of Computational Science, 2013
    Co-Authors: Piotr Luszczek, Jack Dongarra, Stanimire Tomov
    Abstract:

    Abstract The general purpose graphics processing units (GPGPUs) are increasingly deployed for scientific computing due to their performance advantages over CPUs. What followed is the fact that fault tolerance has become a more serious concern compared to the period when GPGPUs were used exclusively for graphics applications. Using GPUs and CPUs together in a hybrid computing system increases flexibility and performance but also increases the possibility of the computations being affected by soft errors, for example, in the form of bit flips. In this work, we propose a soft error resilient algorithm for QR factorization on such hybrid systems. Our contributions include: (1) a checkpointing and recovery mechanism for the left-factor Q whose performance is scalable on hybrid systems; (2) optimized Givens Rotation utilities on GPGPUs to efficiently reduce an upper Hessenberg matrix to an upper triangular form for the protection of the right factor R ; and (3) a recovery algorithm based on QR update on GPGPUs. Experimental results show that our fault tolerant QR factorization can successfully detect and recover from soft errors in the entire matrix with little overhead on hybrid systems with GPGPUs.

  • soft error resilient qr factorization for hybrid system with gpgpu
    Proceedings of the second workshop on Scalable algorithms for large-scale systems, 2011
    Co-Authors: Piotr Luszczek, Stanimire Tomov, Jack Dongarra
    Abstract:

    The general purpose graphics processing units (GPGPU) are increasingly deployed for scientific computing due to their performance advantages over CPUs. As a result, fault tolerance has become a more serious concern compared to the period when GPGPUs were used exclusively for graphics applications. Using GPUs and CPUs together in a hybrid computing system increases flexibility and performance but also increases the possibility of the computations being affected by soft errors. In this work, we propose a soft error resilient algorithm for QR factorization on such hybrid systems. Our contributions include (1) a checkpointing and recovery mechanism for the left-factor Q whose performance is scalable on hybrid systems; (2) optimized Givens Rotation utilities on GPGPUs to efficiently reduce an upper Hessenberg matrix to an upper triangular form for the protection of the right factor R, and (3) a recovery algorithm based on QR update on GPGPUs. Experimental results show that our fault tolerant QR factorization can success- fully detect and recover from soft errors in the entire matrix with little overhead on hybrid systems with GPGPUs.

  • soft error resilient qr factorization for hybrid system
    University of Tennessee Computer Science Technical Report, 2011
    Co-Authors: Piotr Luszczek, Stanimire Tomov, Jack Dongarra
    Abstract:

    As the general purpose graphics processing units (GPGPU) are increasingly deployed for scientific computing for its raw performance advantages compared to CPUs, the fault tolerance issue has started to become more of a concern than before when they were exclusively used for graphics applications. The pairing of GPUs with CPUs to form a hybrid computing systems for better flexibility and performance creates a massive amounts of computations that have a higher possibility to be affected by transient error – a soft error that silently modifies data causing errors to pass unnoticed. This is despite the fact that the newest Fermi generation of GPUs from NVIDIA are equipped with error correcting units to protect their memories. This problem is particularly serious for applications that employ numerical linear algebra since large sections of data are often modified between steps, and therefore even a single error could eventually propagate into a large area of result. In order to give protection to dense linear algebra computations on such hybrid systems, we developed an algorithm that is resilient to soft errors. We chose the right-looking Householder QR factorization as a demonstration of our algorithm for a hybrid system that features both GPUs and CPUs. Algorithm based fault tolerance (ABFT) is used to protect from errors in the trailing matrix and the right factor, while a checkpointing method is used to ensure the left factor is error-free. This work is based on a previous study of fault tolerance in matrix factorizations. Our contribution includes (1) a stable multiple-error checkpointing and recovery mechanism for the left-factor, which is also scalable in performance in the hybrid execution environment and does not cause severe performance degradation. (2) optimized Givens Rotation utilities on the GPU to efficiently reduce an upper Hessenberg matrix to upper triangular form, and (3) a recovery algorithm based on QR update inside a hybrid system. Experimental results show that, our fault tolerant QR factorization can successfully detect and correct data altered by soft errors in both the left and right factors and we observe a decreasing percentage of overhead as the matrix size grows.

Yi Long - One of the best experts on this subject based on the ideXlab platform.

  • A Scalable Limited Feedback Design for Network MIMO Using Per-Cell Product Codebook
    IEEE Transactions on Wireless Communications, 2010
    Co-Authors: Yong Cheng, Yi Long
    Abstract:

    In network MIMO systems, channel state information is required at the transmitter side to multiplex users in the spatial domain. Since perfect channel knowledge is difficult to obtain in practice, limited feedback is a widely accepted solution. The dynamic number of cooperating BSs and heterogeneous path loss effects of network MIMO systems pose new challenges on limited feedback design. In this paper, we propose a scalable limited feedback design for network MIMO systems with multiple base stations, multiple users and multiple data streams for each user. We propose a limited feedback framework using per-cell product codebooks, along with a low-complexity feedback indices selection algorithm. We show that the proposed per-cell product codebook limited feedback design can asymptotically achieve the same performance as the joint-cell codebook approach. We also derive an asymptotic per-user throughput loss due to limited feedback with per-cell product codebooks. Based on that, we show that when the number of per-user feedback-bits Bk is {O}( NnTnR log2(ρgksum)), the system operates in the noise-limited regime in which the per-user throughput is {O} ( nR log2 (nRρgksum/NnT)). On the other hand, when the number of per-user feedback-bits Bk does not scale with the system SNR ρ, the system operates in the interference-limited regime where the per-user throughput is {O}(nRBk/(NnT)2). Numerical results show that the proposed design is very flexible to accommodate dynamic number of cooperating BSs and achieves much better performance compared with other baselines (such as the Givens Rotation approach).

  • A Scalable Limited Feedback Design for Network MIMO using Per-Cell Product Codebook
    arXiv: Information Theory, 2010
    Co-Authors: Yong Cheng, Yi Long
    Abstract:

    In network MIMO systems, channel state information is required at the transmitter side to multiplex users in the spatial domain. Since perfect channel knowledge is difficult to obtain in practice, \emph{limited feedback} is a widely accepted solution. The {\em dynamic number of cooperating BSs} and {\em heterogeneous path loss effects} of network MIMO systems pose new challenges on limited feedback design. In this paper, we propose a scalable limited feedback design for network MIMO systems with multiple base stations, multiple users and multiple data streams for each user. We propose a {\em limited feedback framework using per-cell product codebooks}, along with a {\em low-complexity feedback indices selection algorithm}. We show that the proposed per-cell product codebook limited feedback design can asymptotically achieve the same performance as the joint-cell codebook approach. We also derive an asymptotic \emph{per-user throughput loss} due to limited feedback with per-cell product codebooks. Based on that, we show that when the number of per-user feedback-bits $B_{k}$ is $\mathcal{O}\big( Nn_{T}n_{R}\log_{2}(\rho g_{k}^{sum})\big)$, the system operates in the \emph{noise-limited} regime in which the per-user throughput is $\mathcal{O} \left( n_{R} \log_{2} \big( \frac{n_{R}\rho g_{k}^{sum}}{Nn_{T}} \big) \right)$. On the other hand, when the number of per-user feedback-bits $B_{k}$ does not scale with the \emph{system SNR} $\rho$, the system operates in the \emph{interference-limited} regime where the per-user throughput is $\mathcal{O}\left( \frac{n_{R}B_{k}}{(Nn_{T})^{2}} \right)$. Numerical results show that the proposed design is very flexible to accommodate dynamic number of cooperating BSs and achieves much better performance compared with other baselines (such as the Givens Rotation approach).

  • a scalable limited feedback design for network mimo using per cell codebook
    Wireless Communications and Networking Conference, 2010
    Co-Authors: Yong Cheng, Vincent K N Lau, Yi Long
    Abstract:

    In network MIMO systems, channel state information is required at the transmitter side to multiplex users in the spatial domain. Since perfect channel knowledge is difficult to obtain in practice, \emph{limited feedback} is a widely accepted solution. The {\em dynamic number of cooperating BSs} and {\em heterogeneous path loss effects} of network MIMO systems pose new challenges on limited feedback design. In this paper, we propose a scalable {\em limited feedback framework using per-cell codebooks}, along with a {\em low-complexity feedback indices selection algorithm}. We show that the proposed per-cell codebook limited feedback design can asymptotically achieve the same performance as the joint-cell codebook approach. We also derive an asymptotic \emph{per-user throughput loss} due to limited feedback with per-cell codebooks. Based on that, we show that when the number of per-user feedback-bits $B_{k}$ is $O\big( Nn_{T}n_{R}\log_{2}(\rho g_{k}^{sum} )\big)$, the system operates in the \emph{noise-limited} regime in which the per-user throughput is $O \left( n_{R} \log_{2} \big( \frac{n_{R}\rho g_{k}^{sum}}{Nn_{T}} \big) \right)$. On the other hand, when the number of per-user feedback-bits does not scale with the \emph{system SNR} $\rho$, the system operates in the \emph{interference-limited} regime where the per-user throughput is $O\left( \frac{n_{R}B_{k}}{(Nn_{T})^{2}} \right)$. Numerical results show that the proposed design is very flexible to accommodate dynamic number of cooperating BSs and achieves much better performance compared with other baselines (such as the Givens Rotation approach).

Yong Cheng - One of the best experts on this subject based on the ideXlab platform.

  • A Scalable Limited Feedback Design for Network MIMO Using Per-Cell Product Codebook
    IEEE Transactions on Wireless Communications, 2010
    Co-Authors: Yong Cheng, Yi Long
    Abstract:

    In network MIMO systems, channel state information is required at the transmitter side to multiplex users in the spatial domain. Since perfect channel knowledge is difficult to obtain in practice, limited feedback is a widely accepted solution. The dynamic number of cooperating BSs and heterogeneous path loss effects of network MIMO systems pose new challenges on limited feedback design. In this paper, we propose a scalable limited feedback design for network MIMO systems with multiple base stations, multiple users and multiple data streams for each user. We propose a limited feedback framework using per-cell product codebooks, along with a low-complexity feedback indices selection algorithm. We show that the proposed per-cell product codebook limited feedback design can asymptotically achieve the same performance as the joint-cell codebook approach. We also derive an asymptotic per-user throughput loss due to limited feedback with per-cell product codebooks. Based on that, we show that when the number of per-user feedback-bits Bk is {O}( NnTnR log2(ρgksum)), the system operates in the noise-limited regime in which the per-user throughput is {O} ( nR log2 (nRρgksum/NnT)). On the other hand, when the number of per-user feedback-bits Bk does not scale with the system SNR ρ, the system operates in the interference-limited regime where the per-user throughput is {O}(nRBk/(NnT)2). Numerical results show that the proposed design is very flexible to accommodate dynamic number of cooperating BSs and achieves much better performance compared with other baselines (such as the Givens Rotation approach).

  • A Scalable Limited Feedback Design for Network MIMO using Per-Cell Product Codebook
    arXiv: Information Theory, 2010
    Co-Authors: Yong Cheng, Yi Long
    Abstract:

    In network MIMO systems, channel state information is required at the transmitter side to multiplex users in the spatial domain. Since perfect channel knowledge is difficult to obtain in practice, \emph{limited feedback} is a widely accepted solution. The {\em dynamic number of cooperating BSs} and {\em heterogeneous path loss effects} of network MIMO systems pose new challenges on limited feedback design. In this paper, we propose a scalable limited feedback design for network MIMO systems with multiple base stations, multiple users and multiple data streams for each user. We propose a {\em limited feedback framework using per-cell product codebooks}, along with a {\em low-complexity feedback indices selection algorithm}. We show that the proposed per-cell product codebook limited feedback design can asymptotically achieve the same performance as the joint-cell codebook approach. We also derive an asymptotic \emph{per-user throughput loss} due to limited feedback with per-cell product codebooks. Based on that, we show that when the number of per-user feedback-bits $B_{k}$ is $\mathcal{O}\big( Nn_{T}n_{R}\log_{2}(\rho g_{k}^{sum})\big)$, the system operates in the \emph{noise-limited} regime in which the per-user throughput is $\mathcal{O} \left( n_{R} \log_{2} \big( \frac{n_{R}\rho g_{k}^{sum}}{Nn_{T}} \big) \right)$. On the other hand, when the number of per-user feedback-bits $B_{k}$ does not scale with the \emph{system SNR} $\rho$, the system operates in the \emph{interference-limited} regime where the per-user throughput is $\mathcal{O}\left( \frac{n_{R}B_{k}}{(Nn_{T})^{2}} \right)$. Numerical results show that the proposed design is very flexible to accommodate dynamic number of cooperating BSs and achieves much better performance compared with other baselines (such as the Givens Rotation approach).

  • a scalable limited feedback design for network mimo using per cell codebook
    Wireless Communications and Networking Conference, 2010
    Co-Authors: Yong Cheng, Vincent K N Lau, Yi Long
    Abstract:

    In network MIMO systems, channel state information is required at the transmitter side to multiplex users in the spatial domain. Since perfect channel knowledge is difficult to obtain in practice, \emph{limited feedback} is a widely accepted solution. The {\em dynamic number of cooperating BSs} and {\em heterogeneous path loss effects} of network MIMO systems pose new challenges on limited feedback design. In this paper, we propose a scalable {\em limited feedback framework using per-cell codebooks}, along with a {\em low-complexity feedback indices selection algorithm}. We show that the proposed per-cell codebook limited feedback design can asymptotically achieve the same performance as the joint-cell codebook approach. We also derive an asymptotic \emph{per-user throughput loss} due to limited feedback with per-cell codebooks. Based on that, we show that when the number of per-user feedback-bits $B_{k}$ is $O\big( Nn_{T}n_{R}\log_{2}(\rho g_{k}^{sum} )\big)$, the system operates in the \emph{noise-limited} regime in which the per-user throughput is $O \left( n_{R} \log_{2} \big( \frac{n_{R}\rho g_{k}^{sum}}{Nn_{T}} \big) \right)$. On the other hand, when the number of per-user feedback-bits does not scale with the \emph{system SNR} $\rho$, the system operates in the \emph{interference-limited} regime where the per-user throughput is $O\left( \frac{n_{R}B_{k}}{(Nn_{T})^{2}} \right)$. Numerical results show that the proposed design is very flexible to accommodate dynamic number of cooperating BSs and achieves much better performance compared with other baselines (such as the Givens Rotation approach).

Piotr Luszczek - One of the best experts on this subject based on the ideXlab platform.

  • soft error resilient qr factorization for hybrid system with gpgpu
    Journal of Computational Science, 2013
    Co-Authors: Piotr Luszczek, Jack Dongarra, Stanimire Tomov
    Abstract:

    Abstract The general purpose graphics processing units (GPGPUs) are increasingly deployed for scientific computing due to their performance advantages over CPUs. What followed is the fact that fault tolerance has become a more serious concern compared to the period when GPGPUs were used exclusively for graphics applications. Using GPUs and CPUs together in a hybrid computing system increases flexibility and performance but also increases the possibility of the computations being affected by soft errors, for example, in the form of bit flips. In this work, we propose a soft error resilient algorithm for QR factorization on such hybrid systems. Our contributions include: (1) a checkpointing and recovery mechanism for the left-factor Q whose performance is scalable on hybrid systems; (2) optimized Givens Rotation utilities on GPGPUs to efficiently reduce an upper Hessenberg matrix to an upper triangular form for the protection of the right factor R ; and (3) a recovery algorithm based on QR update on GPGPUs. Experimental results show that our fault tolerant QR factorization can successfully detect and recover from soft errors in the entire matrix with little overhead on hybrid systems with GPGPUs.

  • soft error resilient qr factorization for hybrid system with gpgpu
    Proceedings of the second workshop on Scalable algorithms for large-scale systems, 2011
    Co-Authors: Piotr Luszczek, Stanimire Tomov, Jack Dongarra
    Abstract:

    The general purpose graphics processing units (GPGPU) are increasingly deployed for scientific computing due to their performance advantages over CPUs. As a result, fault tolerance has become a more serious concern compared to the period when GPGPUs were used exclusively for graphics applications. Using GPUs and CPUs together in a hybrid computing system increases flexibility and performance but also increases the possibility of the computations being affected by soft errors. In this work, we propose a soft error resilient algorithm for QR factorization on such hybrid systems. Our contributions include (1) a checkpointing and recovery mechanism for the left-factor Q whose performance is scalable on hybrid systems; (2) optimized Givens Rotation utilities on GPGPUs to efficiently reduce an upper Hessenberg matrix to an upper triangular form for the protection of the right factor R, and (3) a recovery algorithm based on QR update on GPGPUs. Experimental results show that our fault tolerant QR factorization can success- fully detect and recover from soft errors in the entire matrix with little overhead on hybrid systems with GPGPUs.

  • soft error resilient qr factorization for hybrid system
    University of Tennessee Computer Science Technical Report, 2011
    Co-Authors: Piotr Luszczek, Stanimire Tomov, Jack Dongarra
    Abstract:

    As the general purpose graphics processing units (GPGPU) are increasingly deployed for scientific computing for its raw performance advantages compared to CPUs, the fault tolerance issue has started to become more of a concern than before when they were exclusively used for graphics applications. The pairing of GPUs with CPUs to form a hybrid computing systems for better flexibility and performance creates a massive amounts of computations that have a higher possibility to be affected by transient error – a soft error that silently modifies data causing errors to pass unnoticed. This is despite the fact that the newest Fermi generation of GPUs from NVIDIA are equipped with error correcting units to protect their memories. This problem is particularly serious for applications that employ numerical linear algebra since large sections of data are often modified between steps, and therefore even a single error could eventually propagate into a large area of result. In order to give protection to dense linear algebra computations on such hybrid systems, we developed an algorithm that is resilient to soft errors. We chose the right-looking Householder QR factorization as a demonstration of our algorithm for a hybrid system that features both GPUs and CPUs. Algorithm based fault tolerance (ABFT) is used to protect from errors in the trailing matrix and the right factor, while a checkpointing method is used to ensure the left factor is error-free. This work is based on a previous study of fault tolerance in matrix factorizations. Our contribution includes (1) a stable multiple-error checkpointing and recovery mechanism for the left-factor, which is also scalable in performance in the hybrid execution environment and does not cause severe performance degradation. (2) optimized Givens Rotation utilities on the GPU to efficiently reduce an upper Hessenberg matrix to upper triangular form, and (3) a recovery algorithm based on QR update inside a hybrid system. Experimental results show that, our fault tolerant QR factorization can successfully detect and correct data altered by soft errors in both the left and right factors and we observe a decreasing percentage of overhead as the matrix size grows.

Peiyun Tsai - One of the best experts on this subject based on the ideXlab platform.

  • efficient implementation of qr decomposition for gigabit mimo ofdm systems
    IEEE Transactions on Circuits and Systems, 2011
    Co-Authors: Zhengyu Huang, Peiyun Tsai
    Abstract:

    This paper presents a VLSI architecture of QR decomposition for 4×4 MIMO-OFDM systems. A real-value decomposed MIMO system model is handled and thus the channel matrix to be processed is extended to the size of 8×8. Instead of direct factorization, a QR decomposition scheme by cascading one complex-value and one real-value Givens Rotation stages is proposed, which can save 44% hardware complexity. Besides, the requirement of skewed inputs in the conventional QR-decomposition systolic array is eliminated and 36% of delay elements are removed. The real-value Givens Rotation stage is also constructed in a form of a stacked triangular systolic array to match with the throughput of the complex-value one. Hardware sharing is considered to enhance the utilization. The proposed design is implemented in 0.18-μm CMOS technology with 152K gates. From measurement, the maximum operating frequency is 100 MHz. It generates QR decomposition results every four clock cycles and accomplishes continuous projection every clock cycle to support MIMO detection up to 2.4 Gb/s. The measured power consumption is 318.6 mW and 219.6 mW for QR decomposition and projection, respectively, at the highest operating frequency. From the comparison, our proposed design achieves the highest throughput with high efficiency.