The Experts below are selected from a list of 324 Experts worldwide ranked by ideXlab platform

Takahiro Hanyu - One of the best experts on this subject based on the ideXlab platform.

  • ISCAS - High-performance Asynchronous intra-chip communication link based on a multiple-valued current-mode single-track scheme
    2009 IEEE International Symposium on Circuits and Systems, 2009
    Co-Authors: Yo Ohtake, Naoya Onizawa, Takahiro Hanyu
    Abstract:

    This paper presents a high-performance Asynchronous Data-Transfer circuit based on a multiple-valued current-mode single-track scheme for on-chip communication. Since one-bit Data and control information are represented by using a multi-level signal in the proposed single-track scheme, one-bit Data can be transmitted Asynchronously using a single wire between modules. The use of current-mode signaling makes the voltage swing on wires reduced, which achieves high-speed Data Transfer. Moreover, as the number of current sources is reduced by the reduction of wires, it is possible to achieve low power dissipation. Using the proposed circuit, we achieve a throughput of 0.65 Gbps/wires with power consumption of 0.29 mW at 5mm wire length. This presents a 400% increase in throughput, a 57% decrease in power consumption with respect to a conventional Asynchronous circuit, using a 90nm CMOS process.

  • Implementation of a High-Speed Asynchronous Data-Transfer Chip Based on Multiple-Valued Current-Signal Multiplexing
    IEICE Transactions on Electronics, 2006
    Co-Authors: Tomohiro Takahashi, Takahiro Hanyu
    Abstract:

    This paper presents an Asynchronous multiple-valued current-mode Data-Transfer controller chip based on a 1-phase dual-rail encoding technique. The proposed encoding technique enables "one-way delay" Asynchronous Data Transfer because request and acknowledge signals can be transmitted simultaneously and valid states are detected by calculating the sum of dual-rail codewords. Since a key component, a current-to-voltage conversion circuit in a valid-state detector, is tuned so as to obtain a sufficient voltage range to improve switching speed of a comparator, signal detection can be performed quickly in spite of using 6-level signals. It is evaluated using HSPICE simulation with a 0.18-μm CMOS that the throughput of the proposed circuit based on the 1-phase dual-rail scheme attains 435 Mbps/wire which is 2.9 times faster than that of a CMOS circuit based on a conventional 4-phase dual-rail scheme. The test chip is fabricated, and the Asynchronous Data-Transfer behavior of the proposed scheme is confirmed.

  • multiple valued duplex Asynchronous Data Transfer scheme for interleaving in ldpc decoders
    International Symposium on Multiple-Valued Logic, 2005
    Co-Authors: Naoya Onizawa, Akira Mochizuki, Takahiro Hanyu, Vincent Gaudet
    Abstract:

    A novel duplex Asynchronous Data-Transfer scheme based on multiple-valued encoding is proposed for interleaving in low-density parity-check (LDPC) decoders, where high-throughput interleavers between variable and check nodes without clock-distribution problems are highly advantageous. Since control signals and Data from mutual nodes are multiplexed using a multi-level dual-rail codeword, the number of communication steps can be greatly reduced, which results in high-speed communication without any additional wires. The hardware is simply implemented by utilizing a multiple-valued current-mode circuit because all the information can be superposed on the same line. The advantages of the proposed Asynchronous Data-Transfer scheme are discussed in comparison with corresponding synchronous and conventional Asynchronous schemes.

  • Differential Operation Oriented Multiple-Valued Encoding and Circuit Realization for Asynchronous Data Transfer
    IEICE Transactions on Electronics, 2004
    Co-Authors: Tomohiro Takahashi, Naoya Onizawa, Takahiro Hanyu
    Abstract:

    This paper presents an Asynchronous Data Transfer scheme using 2-color 2-phase dual-rail encoding based on a differential operation and its circuit realization. The proposed encoding enables seamless Asynchronous Data Transfer without inserting a spacer, because each logic value is represented by two kinds of codewords with dual-rail, called color Data. Since the difference x - x' between components of a codeword (x, x') becomes constant in every valid state, the Data-arrival state can be detected by calculating the difference x - x'. From the viewpoint of circuit implementation, during the state transition, since the dual-rail x and x' are defined so as to transit differentially, the compatibility with a comparator using a differential amplifier becomes high, which results in reduction of the cycle time. It is evaluated using HSPICE simulation with a 0.18μm CMOS technology that communication speed using the proposed dual-rail encoding becomes 1.4 times faster than that using conventional dual-rail encoding.

  • bidirectional Data Transfer based Asynchronous vlsi system using multiple valued current mode logic
    International Symposium on Multiple-Valued Logic, 2003
    Co-Authors: Takahiro Hanyu, T. Takahashi, Michitaka Kameyama
    Abstract:

    A new Asynchronous Data Transfer scheme using multiple-valued 2-color 1-phase coding, called a bidirectional Data Transfer scheme, is proposed for a high-performance and low-power VLSI system. Valid Data signals of "0" or "1" are represented by binary dual-rail complementary codes, (0,1) and (1,0), and "ODD" and "EVEN" colors are represented by binary dual-rail codes, (0,0) and (1,1), respectively. Control signals from both a transmitter and a receiver are represented by dual-rail multiple-valued coding with superposition of Data and color signals. The use of dual-rail coding makes it easy to detect EVEN and ODD information by calculating the sum of dual-rail codes, even when Data and color information are mixed on the same wires in Asynchronous Data Transfer Since a linear-summation can be implemented by wiring without active devices in multiple-valued bidirectional current-mode circuitry, the proposed circuit for Asynchronous control becomes simple. It is evaluated in a 0.18 /spl mu/m CMOS technology that the switching speed of the proposed Asynchronous Data Transfer scheme is about 1.6-times faster than that of the corresponding binary CMOS implementation under the normalized power dissipation.

Weiguo Liu - One of the best experts on this subject based on the ideXlab platform.

  • ICPP - SWMapper: Scalable Read Mapper on SunWay TaihuLight
    49th International Conference on Parallel Processing - ICPP, 2020
    Co-Authors: Xiaohui Duan, Xiangxu Meng, Bertil Schmidt, Weiguo Liu
    Abstract:

    With the rapid development of next-generation sequencing (NGS) technologies, high throughput sequencing platforms continuously produce large amounts of short read DNA Data at low cost. Read mapping is a performance-critical task, being one of the first stages required for many different types of NGS analysis pipelines. We present SWMapper — a scalable and efficient read mapper for the Sunway TaihuLight supercomputer. A number of optimization techniques are proposed to achieve high performance on its heterogeneous architecture which are centered around a memory-efficient succinct hash index Data structure including seed filtration, duplicate removal, dynamic scheduling, Asynchronous Data Transfer, and overlapping I/O and computation. Furthermore, a vectorized version of the banded Myers algorithm for pairwise alignment with 256-bit vector registers is presented to fully exploit the computational power of the SW26010 processor. Our performance evaluation shows that SWMapper using all 4 compute groups of a single Sunway TaihuLight node outperforms S-Aligner on the same hardware by a factor of 6.2. In addition, compared the state-of-the-art CPU-based mappers RazerS3, BitMapper2, and Hobbes3 running on a 4-core Xeon W-2123v3 CPU, SWMapper achieves speedups of 26.5, 7.8, and 2.6, respectively. Our optimizations achieve an aggregated speedup of 11 compared to the naive implementation on one compute group of an SW26010 processor as well as a strong scaling efficiency of 74% on 128 compute groups.

  • CLUSTER - S-Aligner: Ultrascalable Read Mapping on Sunway Taihu Light
    2017 IEEE International Conference on Cluster Computing (CLUSTER), 2017
    Co-Authors: Xiaohui Duan, Pavan Balaji, Bertil Schmidt, Yuandong Chan, Christian Hundt, Weiguo Liu
    Abstract:

    The availability and amount of sequenced genomes have been rapidly growing in recent years because of the adoption of next-generation sequencing (NGS) technologies that enable high-throughput short-read generation at highly competitive cost. Since this trend is expected to continue in the foreseeable future, the design and implementation of efficient and scalable NGS bioinformatics algorithms are important to research and industrial applications. In this paper, we introduce S-Aligner–a highly scalable read mapper designed for the Sunway Taihu Light supercomputer and its fourth-generationShenWei many-core architecture (SW26010). S-Aligner employs a combination of optimization techniques to overcome both the memory-bound and the compute-bound bottlenecks in the read mapping algorithm. In order to make full use of the compute power of Sunway Taihu Light, our design employs three levels of parallelism: (1) internode parallelism using MPI based on a task-grid pattern, (2) intranode parallelism using multithreading and Asynchronous Data Transfer to fully utilize all 260 cores of the SW26010 many-core processor, and (3) vectorization to exploit the available 256-bit SIMD vector registers. Moreover, we have employed Asynchronous access patterns and Data-sharing strategies during file I/O to overcome bandwidth limitations of the network file system. Our performance evaluation demonstrates that S-Aligner scales almost linearly with approximately 95% efficiency for up to 13,312 nodes (concurrently harnessing more than 3 millioncompute cores). Furthermore, our implementation on a single node outperforms the established RazerS3 mapper running on a platform with eight Intel Xeon E7-8860v3 CPUs while achieving highly competitive alignment accuracy.

Naoya Onizawa - One of the best experts on this subject based on the ideXlab platform.

  • ISCAS - High-performance Asynchronous intra-chip communication link based on a multiple-valued current-mode single-track scheme
    2009 IEEE International Symposium on Circuits and Systems, 2009
    Co-Authors: Yo Ohtake, Naoya Onizawa, Takahiro Hanyu
    Abstract:

    This paper presents a high-performance Asynchronous Data-Transfer circuit based on a multiple-valued current-mode single-track scheme for on-chip communication. Since one-bit Data and control information are represented by using a multi-level signal in the proposed single-track scheme, one-bit Data can be transmitted Asynchronously using a single wire between modules. The use of current-mode signaling makes the voltage swing on wires reduced, which achieves high-speed Data Transfer. Moreover, as the number of current sources is reduced by the reduction of wires, it is possible to achieve low power dissipation. Using the proposed circuit, we achieve a throughput of 0.65 Gbps/wires with power consumption of 0.29 mW at 5mm wire length. This presents a 400% increase in throughput, a 57% decrease in power consumption with respect to a conventional Asynchronous circuit, using a 90nm CMOS process.

  • multiple valued duplex Asynchronous Data Transfer scheme for interleaving in ldpc decoders
    International Symposium on Multiple-Valued Logic, 2005
    Co-Authors: Naoya Onizawa, Akira Mochizuki, Takahiro Hanyu, Vincent Gaudet
    Abstract:

    A novel duplex Asynchronous Data-Transfer scheme based on multiple-valued encoding is proposed for interleaving in low-density parity-check (LDPC) decoders, where high-throughput interleavers between variable and check nodes without clock-distribution problems are highly advantageous. Since control signals and Data from mutual nodes are multiplexed using a multi-level dual-rail codeword, the number of communication steps can be greatly reduced, which results in high-speed communication without any additional wires. The hardware is simply implemented by utilizing a multiple-valued current-mode circuit because all the information can be superposed on the same line. The advantages of the proposed Asynchronous Data-Transfer scheme are discussed in comparison with corresponding synchronous and conventional Asynchronous schemes.

  • Differential Operation Oriented Multiple-Valued Encoding and Circuit Realization for Asynchronous Data Transfer
    IEICE Transactions on Electronics, 2004
    Co-Authors: Tomohiro Takahashi, Naoya Onizawa, Takahiro Hanyu
    Abstract:

    This paper presents an Asynchronous Data Transfer scheme using 2-color 2-phase dual-rail encoding based on a differential operation and its circuit realization. The proposed encoding enables seamless Asynchronous Data Transfer without inserting a spacer, because each logic value is represented by two kinds of codewords with dual-rail, called color Data. Since the difference x - x' between components of a codeword (x, x') becomes constant in every valid state, the Data-arrival state can be detected by calculating the difference x - x'. From the viewpoint of circuit implementation, during the state transition, since the dual-rail x and x' are defined so as to transit differentially, the compatibility with a comparator using a differential amplifier becomes high, which results in reduction of the cycle time. It is evaluated using HSPICE simulation with a 0.18μm CMOS technology that communication speed using the proposed dual-rail encoding becomes 1.4 times faster than that using conventional dual-rail encoding.

  • ISMVL - Multiple-valued duplex Asynchronous Data Transfer scheme for interleaving in LDPC decoders
    35th International Symposium on Multiple-Valued Logic (ISMVL'05), 1
    Co-Authors: Naoya Onizawa, Akira Mochizuki, Takahiro Hanyu, Vincent Gaudet
    Abstract:

    A novel duplex Asynchronous Data-Transfer scheme based on multiple-valued encoding is proposed for interleaving in low-density parity-check (LDPC) decoders, where high-throughput interleavers between variable and check nodes without clock-distribution problems are highly advantageous. Since control signals and Data from mutual nodes are multiplexed using a multi-level dual-rail codeword, the number of communication steps can be greatly reduced, which results in high-speed communication without any additional wires. The hardware is simply implemented by utilizing a multiple-valued current-mode circuit because all the information can be superposed on the same line. The advantages of the proposed Asynchronous Data-Transfer scheme are discussed in comparison with corresponding synchronous and conventional Asynchronous schemes.

Wuchun Feng - One of the best experts on this subject based on the ideXlab platform.

  • a scalable multi engine xpress9 compressor with Asynchronous Data Transfer
    Field-Programmable Custom Computing Machines, 2013
    Co-Authors: Lokendra S Panwar, Ashwin M Aji, Jiayuan Meng, Pavan Balaji, Wuchun Feng
    Abstract:

    Data compression is crucial in large-scale storage servers to save both storage and network bandwidth, but it suffers from high computational cost. In this work, we present a high throughput FPGA based compressor as a PCIe accelerator to achieve CPU resource saving and high power efficiency. The proposed compressor is differentiated from previous hardware compressors by the following features:1) targeting Xpress9 algorithm, whose compression quality is comparable to the best Gzip implementation (level 9), 2) a scalable multi-engine architecture with various IP blocks to handle algorithmic complexity as well as to achieve high throughput, 3) supporting a heavily multi-threaded server environment with an Asynchronous Data Transfer interface between the host and the accelerator. The implemented Xpress9 compressor on Altera Stratix V GS performs 1.6-2.4Gbps throughput with 7 engines on various compression benchmarks, supporting up to 128 thread contexts.

  • ICPADS - A Scalable Multi-engine Xpress9 Compressor with Asynchronous Data Transfer
    2013
    Co-Authors: Lokendra S Panwar, Ashwin M Aji, Jiayuan Meng, Pavan Balaji, Wuchun Feng
    Abstract:

    Data compression is crucial in large-scale storage servers to save both storage and network bandwidth, but it suffers from high computational cost. In this work, we present a high throughput FPGA based compressor as a PCIe accelerator to achieve CPU resource saving and high power efficiency. The proposed compressor is differentiated from previous hardware compressors by the following features:1) targeting Xpress9 algorithm, whose compression quality is comparable to the best Gzip implementation (level 9), 2) a scalable multi-engine architecture with various IP blocks to handle algorithmic complexity as well as to achieve high throughput, 3) supporting a heavily multi-threaded server environment with an Asynchronous Data Transfer interface between the host and the accelerator. The implemented Xpress9 compressor on Altera Stratix V GS performs 1.6-2.4Gbps throughput with 7 engines on various compression benchmarks, supporting up to 128 thread contexts.

Pavan Balaji - One of the best experts on this subject based on the ideXlab platform.

  • CLUSTER - S-Aligner: Ultrascalable Read Mapping on Sunway Taihu Light
    2017 IEEE International Conference on Cluster Computing (CLUSTER), 2017
    Co-Authors: Xiaohui Duan, Pavan Balaji, Bertil Schmidt, Yuandong Chan, Christian Hundt, Weiguo Liu
    Abstract:

    The availability and amount of sequenced genomes have been rapidly growing in recent years because of the adoption of next-generation sequencing (NGS) technologies that enable high-throughput short-read generation at highly competitive cost. Since this trend is expected to continue in the foreseeable future, the design and implementation of efficient and scalable NGS bioinformatics algorithms are important to research and industrial applications. In this paper, we introduce S-Aligner–a highly scalable read mapper designed for the Sunway Taihu Light supercomputer and its fourth-generationShenWei many-core architecture (SW26010). S-Aligner employs a combination of optimization techniques to overcome both the memory-bound and the compute-bound bottlenecks in the read mapping algorithm. In order to make full use of the compute power of Sunway Taihu Light, our design employs three levels of parallelism: (1) internode parallelism using MPI based on a task-grid pattern, (2) intranode parallelism using multithreading and Asynchronous Data Transfer to fully utilize all 260 cores of the SW26010 many-core processor, and (3) vectorization to exploit the available 256-bit SIMD vector registers. Moreover, we have employed Asynchronous access patterns and Data-sharing strategies during file I/O to overcome bandwidth limitations of the network file system. Our performance evaluation demonstrates that S-Aligner scales almost linearly with approximately 95% efficiency for up to 13,312 nodes (concurrently harnessing more than 3 millioncompute cores). Furthermore, our implementation on a single node outperforms the established RazerS3 mapper running on a platform with eight Intel Xeon E7-8860v3 CPUs while achieving highly competitive alignment accuracy.

  • a scalable multi engine xpress9 compressor with Asynchronous Data Transfer
    Field-Programmable Custom Computing Machines, 2013
    Co-Authors: Lokendra S Panwar, Ashwin M Aji, Jiayuan Meng, Pavan Balaji, Wuchun Feng
    Abstract:

    Data compression is crucial in large-scale storage servers to save both storage and network bandwidth, but it suffers from high computational cost. In this work, we present a high throughput FPGA based compressor as a PCIe accelerator to achieve CPU resource saving and high power efficiency. The proposed compressor is differentiated from previous hardware compressors by the following features:1) targeting Xpress9 algorithm, whose compression quality is comparable to the best Gzip implementation (level 9), 2) a scalable multi-engine architecture with various IP blocks to handle algorithmic complexity as well as to achieve high throughput, 3) supporting a heavily multi-threaded server environment with an Asynchronous Data Transfer interface between the host and the accelerator. The implemented Xpress9 compressor on Altera Stratix V GS performs 1.6-2.4Gbps throughput with 7 engines on various compression benchmarks, supporting up to 128 thread contexts.

  • ICPADS - A Scalable Multi-engine Xpress9 Compressor with Asynchronous Data Transfer
    2013
    Co-Authors: Lokendra S Panwar, Ashwin M Aji, Jiayuan Meng, Pavan Balaji, Wuchun Feng
    Abstract:

    Data compression is crucial in large-scale storage servers to save both storage and network bandwidth, but it suffers from high computational cost. In this work, we present a high throughput FPGA based compressor as a PCIe accelerator to achieve CPU resource saving and high power efficiency. The proposed compressor is differentiated from previous hardware compressors by the following features:1) targeting Xpress9 algorithm, whose compression quality is comparable to the best Gzip implementation (level 9), 2) a scalable multi-engine architecture with various IP blocks to handle algorithmic complexity as well as to achieve high throughput, 3) supporting a heavily multi-threaded server environment with an Asynchronous Data Transfer interface between the host and the accelerator. The implemented Xpress9 compressor on Altera Stratix V GS performs 1.6-2.4Gbps throughput with 7 engines on various compression benchmarks, supporting up to 128 thread contexts.