The Experts below are selected from a list of 9 Experts worldwide ranked by ideXlab platform
Nong Xiao - One of the best experts on this subject based on the ideXlab platform.
-
a specialized low cost Vectorized Loop buffer for embedded processors
Design Automation and Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
-
DATE - A specialized low-cost Vectorized Loop buffer for embedded processors
2011 Design Automation & Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
Libo Huang - One of the best experts on this subject based on the ideXlab platform.
-
a specialized low cost Vectorized Loop buffer for embedded processors
Design Automation and Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
-
DATE - A specialized low-cost Vectorized Loop buffer for embedded processors
2011 Design Automation & Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
Zhiying Wang - One of the best experts on this subject based on the ideXlab platform.
-
a specialized low cost Vectorized Loop buffer for embedded processors
Design Automation and Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
-
DATE - A specialized low-cost Vectorized Loop buffer for embedded processors
2011 Design Automation & Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
Li Shen - One of the best experts on this subject based on the ideXlab platform.
-
a specialized low cost Vectorized Loop buffer for embedded processors
Design Automation and Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
-
DATE - A specialized low-cost Vectorized Loop buffer for embedded processors
2011 Design Automation & Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
Hongyi Lu - One of the best experts on this subject based on the ideXlab platform.
-
a specialized low cost Vectorized Loop buffer for embedded processors
Design Automation and Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.
-
DATE - A specialized low-cost Vectorized Loop buffer for embedded processors
2011 Design Automation & Test in Europe, 2011Co-Authors: Libo Huang, Zhiying Wang, Li Shen, Hongyi Lu, Nong XiaoAbstract:Current Loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting Loop buffer is to obtain its performance benefit. In this paper, we propose an application specific Loop buffer organization for Vectorized processing kernels, to achieve low-power and high-performance goals. The Vectorized Loop buffer (VLB) is simplified with single Loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices.