The Experts below are selected from a list of 2772 Experts worldwide ranked by ideXlab platform
N. Yano - One of the best experts on this subject based on the ideXlab platform.
-
A fully pipelined single-precision floating-point unit in the synergistic processor element of a CELL processor
IEEE Journal of Solid-State Circuits, 2006Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit (FPU) in the synergistic processor element (SPE) of a CELL processor is a fully pipelined 4-way single-instruction multiple-data (SIMD) unit designed to accelerate media and data streaming with 128-bit operands. It supports 32-bit single-precision floating-point and 16-bit integer operands with two different latencies, six-cycle and seven-cycle, with 11 FO4 delay per stage. The FPU optimizes the performance of critical single-precision multiply-add operations. Since exact rounding, exceptions, and de-norm number handling are not important to multimedia applications, IEEE correctness on the single-precision floating-point numbers is sacrificed for performance and simple design. It employs fine-grained clock gating for power saving. The design has 768K transistors in 1.3 mm/sup 2/, fabricated SOI in 90-nm technology. Correct operations have been observed up to 5.6 GHz with 1.4 V and 56/spl deg/C, delivering 44.8 GFlops. Architecture, logic, circuits, and integration are codesigned to meet the performance, power, and area goals.
-
IEEE Symposium on Computer Arithmetic - The vector floating-point unit in a synergistic processor element of a CELL processor
17th IEEE Symposium on Computer Arithmetic (ARITH'05), 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
the vector floating point unit in a synergistic processor element of a cell processor
Symposium on Computer Arithmetic, 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
A fully-pipelined single-precision floating point unit in the synergistic processor element of a CELL processor
Digest of Technical Papers. 2005 Symposium on VLSI Circuits 2005., 2005Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating point unit in the synergistic processor element of a CELL processor is a fully-pipelined 4-way SIMD unit designed to accelerate media and data streaming. It supports 32-bit single-precision floating point and 16-bit integer operands with two different latencies, optimizing the performance of critical single-precision multiply-add operations. It employs fine-grained clock gating for power saving. Architecture, logic, circuits and integration are co-designed to meet the performance, power, and area goals.
Hwa-joon Oh - One of the best experts on this subject based on the ideXlab platform.
-
A fully pipelined single-precision floating-point unit in the synergistic processor element of a CELL processor
IEEE Journal of Solid-State Circuits, 2006Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit (FPU) in the synergistic processor element (SPE) of a CELL processor is a fully pipelined 4-way single-instruction multiple-data (SIMD) unit designed to accelerate media and data streaming with 128-bit operands. It supports 32-bit single-precision floating-point and 16-bit integer operands with two different latencies, six-cycle and seven-cycle, with 11 FO4 delay per stage. The FPU optimizes the performance of critical single-precision multiply-add operations. Since exact rounding, exceptions, and de-norm number handling are not important to multimedia applications, IEEE correctness on the single-precision floating-point numbers is sacrificed for performance and simple design. It employs fine-grained clock gating for power saving. The design has 768K transistors in 1.3 mm/sup 2/, fabricated SOI in 90-nm technology. Correct operations have been observed up to 5.6 GHz with 1.4 V and 56/spl deg/C, delivering 44.8 GFlops. Architecture, logic, circuits, and integration are codesigned to meet the performance, power, and area goals.
-
IEEE Symposium on Computer Arithmetic - The vector floating-point unit in a synergistic processor element of a CELL processor
17th IEEE Symposium on Computer Arithmetic (ARITH'05), 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
the vector floating point unit in a synergistic processor element of a cell processor
Symposium on Computer Arithmetic, 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
A fully-pipelined single-precision floating point unit in the synergistic processor element of a CELL processor
Digest of Technical Papers. 2005 Symposium on VLSI Circuits 2005., 2005Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating point unit in the synergistic processor element of a CELL processor is a fully-pipelined 4-way SIMD unit designed to accelerate media and data streaming. It supports 32-bit single-precision floating point and 16-bit integer operands with two different latencies, optimizing the performance of critical single-precision multiply-add operations. It employs fine-grained clock gating for power saving. Architecture, logic, circuits and integration are co-designed to meet the performance, power, and area goals.
S.m. Mueller - One of the best experts on this subject based on the ideXlab platform.
-
A fully pipelined single-precision floating-point unit in the synergistic processor element of a CELL processor
IEEE Journal of Solid-State Circuits, 2006Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit (FPU) in the synergistic processor element (SPE) of a CELL processor is a fully pipelined 4-way single-instruction multiple-data (SIMD) unit designed to accelerate media and data streaming with 128-bit operands. It supports 32-bit single-precision floating-point and 16-bit integer operands with two different latencies, six-cycle and seven-cycle, with 11 FO4 delay per stage. The FPU optimizes the performance of critical single-precision multiply-add operations. Since exact rounding, exceptions, and de-norm number handling are not important to multimedia applications, IEEE correctness on the single-precision floating-point numbers is sacrificed for performance and simple design. It employs fine-grained clock gating for power saving. The design has 768K transistors in 1.3 mm/sup 2/, fabricated SOI in 90-nm technology. Correct operations have been observed up to 5.6 GHz with 1.4 V and 56/spl deg/C, delivering 44.8 GFlops. Architecture, logic, circuits, and integration are codesigned to meet the performance, power, and area goals.
-
IEEE Symposium on Computer Arithmetic - The vector floating-point unit in a synergistic processor element of a CELL processor
17th IEEE Symposium on Computer Arithmetic (ARITH'05), 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
the vector floating point unit in a synergistic processor element of a cell processor
Symposium on Computer Arithmetic, 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
A fully-pipelined single-precision floating point unit in the synergistic processor element of a CELL processor
Digest of Technical Papers. 2005 Symposium on VLSI Circuits 2005., 2005Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating point unit in the synergistic processor element of a CELL processor is a fully-pipelined 4-way SIMD unit designed to accelerate media and data streaming. It supports 32-bit single-precision floating point and 16-bit integer operands with two different latencies, optimizing the performance of critical single-precision multiply-add operations. It employs fine-grained clock gating for power saving. Architecture, logic, circuits and integration are co-designed to meet the performance, power, and area goals.
T. Namatame - One of the best experts on this subject based on the ideXlab platform.
-
A fully pipelined single-precision floating-point unit in the synergistic processor element of a CELL processor
IEEE Journal of Solid-State Circuits, 2006Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit (FPU) in the synergistic processor element (SPE) of a CELL processor is a fully pipelined 4-way single-instruction multiple-data (SIMD) unit designed to accelerate media and data streaming with 128-bit operands. It supports 32-bit single-precision floating-point and 16-bit integer operands with two different latencies, six-cycle and seven-cycle, with 11 FO4 delay per stage. The FPU optimizes the performance of critical single-precision multiply-add operations. Since exact rounding, exceptions, and de-norm number handling are not important to multimedia applications, IEEE correctness on the single-precision floating-point numbers is sacrificed for performance and simple design. It employs fine-grained clock gating for power saving. The design has 768K transistors in 1.3 mm/sup 2/, fabricated SOI in 90-nm technology. Correct operations have been observed up to 5.6 GHz with 1.4 V and 56/spl deg/C, delivering 44.8 GFlops. Architecture, logic, circuits, and integration are codesigned to meet the performance, power, and area goals.
-
IEEE Symposium on Computer Arithmetic - The vector floating-point unit in a synergistic processor element of a CELL processor
17th IEEE Symposium on Computer Arithmetic (ARITH'05), 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
the vector floating point unit in a synergistic processor element of a cell processor
Symposium on Computer Arithmetic, 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
A fully-pipelined single-precision floating point unit in the synergistic processor element of a CELL processor
Digest of Technical Papers. 2005 Symposium on VLSI Circuits 2005., 2005Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating point unit in the synergistic processor element of a CELL processor is a fully-pipelined 4-way SIMD unit designed to accelerate media and data streaming. It supports 32-bit single-precision floating point and 16-bit integer operands with two different latencies, optimizing the performance of critical single-precision multiply-add operations. It employs fine-grained clock gating for power saving. Architecture, logic, circuits and integration are co-designed to meet the performance, power, and area goals.
Y. Totsuka - One of the best experts on this subject based on the ideXlab platform.
-
A fully pipelined single-precision floating-point unit in the synergistic processor element of a CELL processor
IEEE Journal of Solid-State Circuits, 2006Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit (FPU) in the synergistic processor element (SPE) of a CELL processor is a fully pipelined 4-way single-instruction multiple-data (SIMD) unit designed to accelerate media and data streaming with 128-bit operands. It supports 32-bit single-precision floating-point and 16-bit integer operands with two different latencies, six-cycle and seven-cycle, with 11 FO4 delay per stage. The FPU optimizes the performance of critical single-precision multiply-add operations. Since exact rounding, exceptions, and de-norm number handling are not important to multimedia applications, IEEE correctness on the single-precision floating-point numbers is sacrificed for performance and simple design. It employs fine-grained clock gating for power saving. The design has 768K transistors in 1.3 mm/sup 2/, fabricated SOI in 90-nm technology. Correct operations have been observed up to 5.6 GHz with 1.4 V and 56/spl deg/C, delivering 44.8 GFlops. Architecture, logic, circuits, and integration are codesigned to meet the performance, power, and area goals.
-
IEEE Symposium on Computer Arithmetic - The vector floating-point unit in a synergistic processor element of a CELL processor
17th IEEE Symposium on Computer Arithmetic (ARITH'05), 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
the vector floating point unit in a synergistic processor element of a cell processor
Symposium on Computer Arithmetic, 2005Co-Authors: S.m. Mueller, Hwa-joon Oh, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
-
A fully-pipelined single-precision floating point unit in the synergistic processor element of a CELL processor
Digest of Technical Papers. 2005 Symposium on VLSI Circuits 2005., 2005Co-Authors: Hwa-joon Oh, S.m. Mueller, C. Jacobi, K.d. Tran, S.r. Cottier, B.w. Michael, H. Nishikawa, Y. Totsuka, T. Namatame, N. YanoAbstract:The floating point unit in the synergistic processor element of a CELL processor is a fully-pipelined 4-way SIMD unit designed to accelerate media and data streaming. It supports 32-bit single-precision floating point and 16-bit integer operands with two different latencies, optimizing the performance of critical single-precision multiply-add operations. It employs fine-grained clock gating for power saving. Architecture, logic, circuits and integration are co-designed to meet the performance, power, and area goals.