The Experts below are selected from a list of 63 Experts worldwide ranked by ideXlab platform

Masami Saeki - One of the best experts on this subject based on the ideXlab platform.

  • An Efficient GPU Implementation of Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    International Journal of Networking and Computing, 2017
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present an efficient GPU implementation of bulk computation of eigenvalues for many small, non-symmetric, real matrices. This work is motivated by the necessity of such bulk computation in designing of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. Several efforts have been devoted to accelerating the eigenvalue computation including computer languages, systems, environments supporting matrix manipulation offering specific libraries/function calls. Some of them are optimized for computing the eigenvalues of a very large matrix by parallel processing. However, such libraries/function calls are not aimed at accelerating the eigenvalues computation for a lot of small matrices. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, utilization of the shared memory, and so forth. In particular, we present two types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. Furthermore, to hide CPU-GPU data transfer latency, overlapping computation on the GPU with the transfer is employed. Experimental results on NVIDIA TITAN~X show that our GPU implementation attains a speed-up factor of up to 83.50 and 17.67 over the sequential CPU implementation and the parallel CPU implementation with eight threads on Intel Core i7-6700K, respectively.

  • gpu accelerated bulk computation of the eigenvalue problem for many small real non symmetric matrices
    International Symposium on Computing and Networking, 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

  • CANDAR - GPU-Accelerated Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    2016 Fourth International Symposium on Computing and Networking (CANDAR), 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

Koji Nakano - One of the best experts on this subject based on the ideXlab platform.

  • PPAM (1) - A GPU Implementation of Bulk Execution of the Dynamic Programming for the Optimal Polygon Triangulation
    Parallel Processing and Applied Mathematics, 2018
    Co-Authors: Kohei Yamashita, Yasuaki Ito, Koji Nakano
    Abstract:

    The optimal polygon triangulation problem for a convex polygon is an optimization problem to find a triangulation with minimum total weight. It is known that this problem can be solved using the dynamic programming technique in \(O(n^3)\) time. The main contribution of this paper is to present an efficient parallel implementation of this \(O(n^3)\)-time algorithm for a lot of instances on the GPU (Graphics Processing Unit). In our proposed GPU implementation, we focused on the computation for a lot of instances and considered programming issues of the GPU architecture such as Coalesced Access of the global memory, warp divergence. Our implementation solves the optimal polygon triangulation problem for 1024 convex 1024-gons in 4.77 s on the NVIDIA TITAN X, while a conventional CPU implementation runs in 241.53 s. Thus, our GPU implementation attains a speedup factor of 50.6.

  • An Efficient GPU Implementation of Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    International Journal of Networking and Computing, 2017
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present an efficient GPU implementation of bulk computation of eigenvalues for many small, non-symmetric, real matrices. This work is motivated by the necessity of such bulk computation in designing of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. Several efforts have been devoted to accelerating the eigenvalue computation including computer languages, systems, environments supporting matrix manipulation offering specific libraries/function calls. Some of them are optimized for computing the eigenvalues of a very large matrix by parallel processing. However, such libraries/function calls are not aimed at accelerating the eigenvalues computation for a lot of small matrices. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, utilization of the shared memory, and so forth. In particular, we present two types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. Furthermore, to hide CPU-GPU data transfer latency, overlapping computation on the GPU with the transfer is employed. Experimental results on NVIDIA TITAN~X show that our GPU implementation attains a speed-up factor of up to 83.50 and 17.67 over the sequential CPU implementation and the parallel CPU implementation with eight threads on Intel Core i7-6700K, respectively.

  • gpu accelerated bulk computation of the eigenvalue problem for many small real non symmetric matrices
    International Symposium on Computing and Networking, 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

  • CANDAR - GPU-Accelerated Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    2016 Fourth International Symposium on Computing and Networking (CANDAR), 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

Hiroki Tokura - One of the best experts on this subject based on the ideXlab platform.

  • An Efficient GPU Implementation of Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    International Journal of Networking and Computing, 2017
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present an efficient GPU implementation of bulk computation of eigenvalues for many small, non-symmetric, real matrices. This work is motivated by the necessity of such bulk computation in designing of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. Several efforts have been devoted to accelerating the eigenvalue computation including computer languages, systems, environments supporting matrix manipulation offering specific libraries/function calls. Some of them are optimized for computing the eigenvalues of a very large matrix by parallel processing. However, such libraries/function calls are not aimed at accelerating the eigenvalues computation for a lot of small matrices. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, utilization of the shared memory, and so forth. In particular, we present two types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. Furthermore, to hide CPU-GPU data transfer latency, overlapping computation on the GPU with the transfer is employed. Experimental results on NVIDIA TITAN~X show that our GPU implementation attains a speed-up factor of up to 83.50 and 17.67 over the sequential CPU implementation and the parallel CPU implementation with eight threads on Intel Core i7-6700K, respectively.

  • gpu accelerated bulk computation of the eigenvalue problem for many small real non symmetric matrices
    International Symposium on Computing and Networking, 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

  • CANDAR - GPU-Accelerated Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    2016 Fourth International Symposium on Computing and Networking (CANDAR), 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

Yasuaki Ito - One of the best experts on this subject based on the ideXlab platform.

  • PPAM (1) - A GPU Implementation of Bulk Execution of the Dynamic Programming for the Optimal Polygon Triangulation
    Parallel Processing and Applied Mathematics, 2018
    Co-Authors: Kohei Yamashita, Yasuaki Ito, Koji Nakano
    Abstract:

    The optimal polygon triangulation problem for a convex polygon is an optimization problem to find a triangulation with minimum total weight. It is known that this problem can be solved using the dynamic programming technique in \(O(n^3)\) time. The main contribution of this paper is to present an efficient parallel implementation of this \(O(n^3)\)-time algorithm for a lot of instances on the GPU (Graphics Processing Unit). In our proposed GPU implementation, we focused on the computation for a lot of instances and considered programming issues of the GPU architecture such as Coalesced Access of the global memory, warp divergence. Our implementation solves the optimal polygon triangulation problem for 1024 convex 1024-gons in 4.77 s on the NVIDIA TITAN X, while a conventional CPU implementation runs in 241.53 s. Thus, our GPU implementation attains a speedup factor of 50.6.

  • An Efficient GPU Implementation of Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    International Journal of Networking and Computing, 2017
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present an efficient GPU implementation of bulk computation of eigenvalues for many small, non-symmetric, real matrices. This work is motivated by the necessity of such bulk computation in designing of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. Several efforts have been devoted to accelerating the eigenvalue computation including computer languages, systems, environments supporting matrix manipulation offering specific libraries/function calls. Some of them are optimized for computing the eigenvalues of a very large matrix by parallel processing. However, such libraries/function calls are not aimed at accelerating the eigenvalues computation for a lot of small matrices. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, utilization of the shared memory, and so forth. In particular, we present two types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. Furthermore, to hide CPU-GPU data transfer latency, overlapping computation on the GPU with the transfer is employed. Experimental results on NVIDIA TITAN~X show that our GPU implementation attains a speed-up factor of up to 83.50 and 17.67 over the sequential CPU implementation and the parallel CPU implementation with eight threads on Intel Core i7-6700K, respectively.

  • gpu accelerated bulk computation of the eigenvalue problem for many small real non symmetric matrices
    International Symposium on Computing and Networking, 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

  • CANDAR - GPU-Accelerated Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    2016 Fourth International Symposium on Computing and Networking (CANDAR), 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

Yushiro Hirota - One of the best experts on this subject based on the ideXlab platform.

  • An Efficient GPU Implementation of Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    International Journal of Networking and Computing, 2017
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present an efficient GPU implementation of bulk computation of eigenvalues for many small, non-symmetric, real matrices. This work is motivated by the necessity of such bulk computation in designing of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. Several efforts have been devoted to accelerating the eigenvalue computation including computer languages, systems, environments supporting matrix manipulation offering specific libraries/function calls. Some of them are optimized for computing the eigenvalues of a very large matrix by parallel processing. However, such libraries/function calls are not aimed at accelerating the eigenvalues computation for a lot of small matrices. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, utilization of the shared memory, and so forth. In particular, we present two types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. Furthermore, to hide CPU-GPU data transfer latency, overlapping computation on the GPU with the transfer is employed. Experimental results on NVIDIA TITAN~X show that our GPU implementation attains a speed-up factor of up to 83.50 and 17.67 over the sequential CPU implementation and the parallel CPU implementation with eight threads on Intel Core i7-6700K, respectively.

  • gpu accelerated bulk computation of the eigenvalue problem for many small real non symmetric matrices
    International Symposium on Computing and Networking, 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.

  • CANDAR - GPU-Accelerated Bulk Computation of the Eigenvalue Problem for Many Small Real Non-symmetric Matrices
    2016 Fourth International Symposium on Computing and Networking (CANDAR), 2016
    Co-Authors: Hiroki Tokura, Yasuaki Ito, Koji Nakano, Takumi Honda, Mitsuya Nishino, Yushiro Hirota, Masami Saeki
    Abstract:

    The main contribution of this paper is to present a very efficient GPU implementation of bulk computation of eigenvalues for a large number of small non-symmetric real matrices. This work is motivated by the necessity of such bulk computation in design of control systems, which requires to compute the eigenvalues of hundreds of thousands non-symmetric real matrices of size up to 30x30. In our GPU implementation, we considered programming issues of the GPU architecture including warp divergence, Coalesced Access of the global memory, bank conflict of the shared memory, etc. In particular, we present three types of assignments of GPU threads to matrices and introduce three memory arrangements in the global memory. The experimental results on NVIDIA GeForce GTX TITAN X show that our GPU implementation for 500000 matrices of size 5x5 to 30x30 attains a speed-up factor of approximately 15 over the CPU implementation on Intel Core i7-4790.