The Experts below are selected from a list of 297 Experts worldwide ranked by ideXlab platform
Mário Véstias - One of the best experts on this subject based on the ideXlab platform.
-
FPL - Efficient implementation of a single-precision floating-point arithmetic unit on FPGA
2014 24th International Conference on Field Programmable Logic and Applications (FPL), 2014Co-Authors: Wilson José, Ana Rita Silva, Horácio C. Neto, Mário VéstiasAbstract:This paper presents a single precision floating point arithmetic unit with support for multiplication, addition, fused multiply-add, Reciprocal, Square-Root and inverse SquareRoot with high-performance and low resource usage. The design uses a piecewise 2nd order polynomial approximation to implement Reciprocal, Square-Root and inverse Square-Root. The unit can be configured with any number of operations and is capable to calculate any function with a throughput of one operation per cycle. The floatingpoint multiplier of the unit is also used to implement the polynomial approximation and the fused multiply-add operation. We have compared our implementation with other state-of-the-art proposals, including the Xilinx Core-Gen operators, and conclude that the approach has a high relative performance/area efficiency.
-
Efficient implementation of a single-precision floating-point arithmetic unit on FPGA
2014 24th International Conference on Field Programmable Logic and Applications (FPL), 2014Co-Authors: Wilson José, Ana Rita Silva, Horácio Neto, Mário VéstiasAbstract:This paper presents a single precision floating point arithmetic unit with support for multiplication, addition, fused multiply-add, Reciprocal, Square-Root and inverse Square-Root with high-performance and low resource usage. The design uses a piecewise 2nd order polynomial approximation to implement Reciprocal, Square-Root and inverse Square-Root. The unit can be configured with any number of operations and is capable to calculate any function with a throughput of one operation per cycle. The floating-point multiplier of the unit is also used to implement the polynomial approximation and the fused multiply-add operation. We have compared our implementation with other state-of-the-art proposals, including the Xilinx Core-Gen operators, and conclude that the approach has a high relative performance/area efficiency.
Wilson José - One of the best experts on this subject based on the ideXlab platform.
-
FPL - Efficient implementation of a single-precision floating-point arithmetic unit on FPGA
2014 24th International Conference on Field Programmable Logic and Applications (FPL), 2014Co-Authors: Wilson José, Ana Rita Silva, Horácio C. Neto, Mário VéstiasAbstract:This paper presents a single precision floating point arithmetic unit with support for multiplication, addition, fused multiply-add, Reciprocal, Square-Root and inverse SquareRoot with high-performance and low resource usage. The design uses a piecewise 2nd order polynomial approximation to implement Reciprocal, Square-Root and inverse Square-Root. The unit can be configured with any number of operations and is capable to calculate any function with a throughput of one operation per cycle. The floatingpoint multiplier of the unit is also used to implement the polynomial approximation and the fused multiply-add operation. We have compared our implementation with other state-of-the-art proposals, including the Xilinx Core-Gen operators, and conclude that the approach has a high relative performance/area efficiency.
-
Efficient implementation of a single-precision floating-point arithmetic unit on FPGA
2014 24th International Conference on Field Programmable Logic and Applications (FPL), 2014Co-Authors: Wilson José, Ana Rita Silva, Horácio Neto, Mário VéstiasAbstract:This paper presents a single precision floating point arithmetic unit with support for multiplication, addition, fused multiply-add, Reciprocal, Square-Root and inverse Square-Root with high-performance and low resource usage. The design uses a piecewise 2nd order polynomial approximation to implement Reciprocal, Square-Root and inverse Square-Root. The unit can be configured with any number of operations and is capable to calculate any function with a throughput of one operation per cycle. The floating-point multiplier of the unit is also used to implement the polynomial approximation and the fused multiply-add operation. We have compared our implementation with other state-of-the-art proposals, including the Xilinx Core-Gen operators, and conclude that the approach has a high relative performance/area efficiency.
Antonio G. M. Strollo - One of the best experts on this subject based on the ideXlab platform.
-
High-Performance Special Function Unit for Programmable 3-D Graphics Processors
IEEE Transactions on Circuits and Systems I: Regular Papers, 2009Co-Authors: Davide De Caro, Nicola Petra, Antonio G. M. StrolloAbstract:An high-speed special function unit (SFU) is presented in this paper. The system supports the single-precision IEEE-754 floating-point standard and implements faithfully rounded Reciprocal, Square Root, Reciprocal Square Root, logarithm, and exponential functions. The functions are approximated by using a novel constrained piecewise quadratic interpolation technique. In this way, the lookup table size is reduced by 40% with respect to previously proposed techniques, without any loss in accuracy. Error analysis and sizing methodology are presented in the paper. The SFU has been implemented in a 0.18-mum CMOS technology. The circuit is able to operate up to 420-MHz clock frequency, with a power dissipation of 160 mW at 420 MHz. The system can be employed in programmable graphics accelerators and in other applications where high-performance function evaluation is needed.
-
A high performance floating-point special function unit using constrained piecewise quadratic approximation
2008 IEEE International Symposium on Circuits and Systems, 2008Co-Authors: Davide De Caro, Nicola Petra, Antonio G. M. StrolloAbstract:A special function unit, able to compute Square Root, Reciprocal Square Root, logarithm and exponential functions is presented in this paper. The system supports single precision IEEE-754 floating-point standard and uses a novel constrained piecewise quadratic interpolation technique to approximate the implemented functions. The proposed approach allows to reduce look-up table size of 40% with respect to previously proposed techniques. The SFU has been implemented in a test chip in 0.18 mum CMOS. A maximum clock frequency of 420 MHz and a power dissipation of 160 mW@420 MHz have been measured.
-
ISCAS - A high performance floating-point special function unit using constrained piecewise quadratic approximation
2008 IEEE International Symposium on Circuits and Systems, 2008Co-Authors: Davide De Caro, Nicola Petra, Antonio G. M. StrolloAbstract:A special function unit, able to compute Square Root, Reciprocal Square Root, logarithm and exponential functions is presented in this paper. The system supports single precision IEEE-754 floating-point standard and uses a novel constrained piecewise quadratic interpolation technique to approximate the implemented functions. The proposed approach allows to reduce look-up table size of 40% with respect to previously proposed techniques. The SFU has been implemented in a test chip in 0.18 mum CMOS. A maximum clock frequency of 420 MHz and a power dissipation of 160 mW@420 MHz have been measured.
Ana Rita Silva - One of the best experts on this subject based on the ideXlab platform.
-
FPL - Efficient implementation of a single-precision floating-point arithmetic unit on FPGA
2014 24th International Conference on Field Programmable Logic and Applications (FPL), 2014Co-Authors: Wilson José, Ana Rita Silva, Horácio C. Neto, Mário VéstiasAbstract:This paper presents a single precision floating point arithmetic unit with support for multiplication, addition, fused multiply-add, Reciprocal, Square-Root and inverse SquareRoot with high-performance and low resource usage. The design uses a piecewise 2nd order polynomial approximation to implement Reciprocal, Square-Root and inverse Square-Root. The unit can be configured with any number of operations and is capable to calculate any function with a throughput of one operation per cycle. The floatingpoint multiplier of the unit is also used to implement the polynomial approximation and the fused multiply-add operation. We have compared our implementation with other state-of-the-art proposals, including the Xilinx Core-Gen operators, and conclude that the approach has a high relative performance/area efficiency.
-
Efficient implementation of a single-precision floating-point arithmetic unit on FPGA
2014 24th International Conference on Field Programmable Logic and Applications (FPL), 2014Co-Authors: Wilson José, Ana Rita Silva, Horácio Neto, Mário VéstiasAbstract:This paper presents a single precision floating point arithmetic unit with support for multiplication, addition, fused multiply-add, Reciprocal, Square-Root and inverse Square-Root with high-performance and low resource usage. The design uses a piecewise 2nd order polynomial approximation to implement Reciprocal, Square-Root and inverse Square-Root. The unit can be configured with any number of operations and is capable to calculate any function with a throughput of one operation per cycle. The floating-point multiplier of the unit is also used to implement the polynomial approximation and the fused multiply-add operation. We have compared our implementation with other state-of-the-art proposals, including the Xilinx Core-Gen operators, and conclude that the approach has a high relative performance/area efficiency.
Javier Vázouez-castillo - One of the best experts on this subject based on the ideXlab platform.
-
LATINCOM - IEEE-754 Half-Precision Floating-Point Low-Latency Reciprocal Square Root IP-Core
2018 IEEE 10th Latin-American Conference on Communications (LATINCOM), 2018Co-Authors: Cuauhtémoc R. Aguilera-galicia, Omar Longoria-gandara, Oscar A. Guzmán-ramos, Luis Pizano-escalante, Javier Vázouez-castilloAbstract:In different matrix-decomposition techniques for wireless-communication systems, the Reciprocal Square Root (RSR) is a fundamental and recurrent operation, as well in gaming and signal processing systems computation of the RSR is required. Most reported RSR architectures are focused on accelerating high-precision floating-point (FP) units. The IEEE 754–2008 half-precision FP standard offers larger dynamic range than fixed-point systems, fewer hardware resources than single-precision FP and enough precision for some applications. This article reports the FPGA implementation of a low-latency, half-precision floating-point RSR unit. The implementation results show that the proposed design exhibits lower latency and better throughput than Intel and Xilinx RSR IP cores.
-
IEEE-754 Half-Precision Floating-Point Low-Latency Reciprocal Square Root IP-Core
2018 IEEE 10th Latin-American Conference on Communications (LATINCOM), 2018Co-Authors: Cuauhtémoc R. Aguilera-galicia, Omar Longoria-gandara, Oscar A. Guzmán-ramos, Luis Pizano-escalante, Javier Vázouez-castilloAbstract:In different matrix-decomposition techniques for wireless-communication systems, the Reciprocal Square Root (RSR) is a fundamental and recurrent operation, as well in gaming and signal processing systems computation of the RSR is required. Most reported RSR architectures are focused on accelerating high-precision floating-point (FP) units. The IEEE 754-2008 half-precision FP standard offers larger dynamic range than fixed-point systems, fewer hardware resources than single-precision FP and enough precision for some applications. This article reports the FPGA implementation of a low-latency, half-precision floating-point RSR unit. The implementation results show that the proposed design exhibits lower latency and better throughput than Intel and Xilinx RSR IP cores.