The Experts below are selected from a list of 5424 Experts worldwide ranked by ideXlab platform

Alok Choudhary - One of the best experts on this subject based on the ideXlab platform.

  • Distributed smart disks for I/O-intensive workloads on switched interconnects
    Future Generation Computer Systems, 2006
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary
    Abstract:

    Disk-based smart storage represents processor-embedded active I/O devices that are equipped with a dedicated network interface controller and on-disk storage (memory and disk). Today's I/O-intensive workloads require architectures whose performance scales with increasing storage capacity. The advent of switch-based, low-latency and high-bandwidth network protocols, such as Virtual interface Architecture and InfiniBand, promises to improve both performance and scalability beyond current state-of-the-art I/O interconnect topologies. This paper evaluates a disk-based distributed smart storage architecture composed of fully distributed processor-embedded disks that are interconnected by a high-performance switched network, and investigates the performance of the proposed processing model on different storage interconnects with representative workloads, including TPC-H queries, association rules mining, high-dimensional data clustering, and two-dimensional fast Fourier transform (FFT).

  • Processor-embedded distributed smart disks for I/O-intensive workloads: architectures, performance models and evaluation
    Journal of Parallel and Distributed Computing, 2004
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary, Mahmut Kandemir
    Abstract:

    Processor-embedded disks, or smart disks, with their network interface controller, can in effect be viewed as processing elements with on-disk memory and secondary storage. The data sizes and access patterns of today's large I/O-intensive workloads require architectures whose processing power scales with increased storage capacity. To address this concern, we propose and evaluate disk-based distributed smart storage architectures. Based on analytically derived performance models, our evaluation with representative workloads show that offloading processing and performing point-to-point data communication improve performance over centralized architectures. Our results also demonstrate that distributed smart disk systems exhibit desirable scalability and can efficiently handle I/O-intensive workloads, such as commercial decision support database (TPC-H) queries, association rules mining, data clustering, and two-dimensional fast Fourier transform, among others.

  • International Conference on Computational Science - Design and evaluation of distributed smart disk architecture for I/O-intensive workloads
    Lecture Notes in Computer Science, 2003
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary
    Abstract:

    Smart disks, a type of processor-embedded active I/O devices, with their on-disk memory and network interface controller, can be viewed as processing elements with attached storage. The growing size and access patterns of today's large I/O-intensive applications require architectures whose processing power scales with the storage capacity. We evaluate a distributed smart disk architecture with representative I/O-intensive workloads including TPC-H queries, association rule mining, data clustering, and 2-D fast Fourier transform applications to study the proposed architecture.

Hideharu Amano - One of the best experts on this subject based on the ideXlab platform.

  • Martini: A network interface controller Chip for High Performance Computing with Distributed PCs
    IEEE Transactions on Parallel and Distributed Systems, 2007
    Co-Authors: Konosuke Watanabe, Tomohiro Otsuka, Noboru Tanabe, Tomohiro Kudoh, Junji Yamamoto, Junichiro Tsuchiya, Hiroaki Nishi, Hideharu Amano
    Abstract:

    In this paper, “Martini,” a network interface controller chip for our original network called RHiNET is described. Martini is designed to provide high-bandwidth and low-latency communication with small overhead. To obtain high performance communication, protected user-level zero-copy RDMA communication functions are completely implemented by a hardwired logic. Also, to reduce the communication latency efficiently, we have proposed PIO-based communication mechanisms called “On-the-fly (OTF)” and have implemented them on Martini. The evaluation results show that Martini connected to a 64bit/66MHz PCI-bus achieves 470MByte/s maximum bidirectional bandwidth and 1.74 μsec minimum latency on host-to-host memory copying.

  • PDCAT - Evaluation of network interface controller on DIMMnet-2 Prototype Board
    Sixth International Conference on Parallel and Distributed Computing Applications and Technologies (PDCAT'05), 2005
    Co-Authors: Akira Kitamura, Hideharu Amano, Yasuo Miyabe, Tetsu Izawa, Tomotaka Miyashiro, Konosuke Watanabe, Tomohiro Otsuka, Yoshihiro Hamada, Noboru Tanabe, Hironori Nakajo
    Abstract:

    By recent performance improvement of interconnection networks for a PC cluster, standard I/O bus which connects network interface becomes the performance bottleneck. DIMMnet is a network interface which can solve the problem by using the memory bus instead of PCI bus or other I/O buses. The second generation network interface DIMMnet-2 can be connected with DDR-SDRAM slot by using the indirect accessing to memory and buffers. Although the current board is a prototype using an FPGA, the latency for 8 Bytes data transfer is only 0.441µs.

  • low latency communication on dimmnet 1 network interface plugged into a dimm slot
    Parallel Computing in Electrical Engineering, 2002
    Co-Authors: Noboru Tanabe, Yoshihiro Hamada, Hironori Nakajo, H. Imashiro, Tomohiro Kudoh, Junji Yamamoto, Hideharu Amano
    Abstract:

    DIMMnet-1 is a high performance network interface for PC clusters that can be directly plugged into the DIMM slot of a PC. By using both low latency AOTF (Atomic On-The-Fly) sending and high bandwidth BOTF (Block On-The-Fly) sending, it can overcome the overhead caused by standard I/O such as the PCI bus. Two types of DIMMnet-1 prototypeboards (providing optical and electrical network interfaces) containing a Martini network interface controller chip are currently available. They can be plugged into a 100MHz DIMM slot of a PC with a Pentium-3, Pentium-4 or Athlon processor. The round-trip time for AOTF onthis incompletely tuned DIMMnet-1 is 7.5 times faster than Myrinet2000. The barrier synchronization time for AOTF is 4 times faster than that of an SR8000 supercomputer. Theinter-two-node floating sum operation time is 1903 ns. This shows that DIMMnet-1 holds promise for applications in which scalable performance with traditional approaches is difficult because of frequent data exchange.

  • PARELEC - Low Latency Communication on DIMMnet-1 network interface Plugged into a DIMM Slot
    2002
    Co-Authors: Noboru Tanabe, Yoshihiro Hamada, Hironori Nakajo, H. Imashiro, Tomohiro Kudoh, Junji Yamamoto, Hideharu Amano
    Abstract:

    DIMMnet-1 is a high performance network interface for PC clusters that can be directly plugged into the DIMM slot of a PC. By using both low latency AOTF (Atomic On-The-Fly) sending and high bandwidth BOTF (Block On-The-Fly) sending, it can overcome the overhead caused by standard I/O such as the PCI bus. Two types of DIMMnet-1 prototypeboards (providing optical and electrical network interfaces) containing a Martini network interface controller chip are currently available. They can be plugged into a 100MHz DIMM slot of a PC with a Pentium-3, Pentium-4 or Athlon processor. The round-trip time for AOTF onthis incompletely tuned DIMMnet-1 is 7.5 times faster than Myrinet2000. The barrier synchronization time for AOTF is 4 times faster than that of an SR8000 supercomputer. Theinter-two-node floating sum operation time is 1903 ns. This shows that DIMMnet-1 holds promise for applications in which scalable performance with traditional approaches is difficult because of frequent data exchange.

  • FPL - A General Hardware Design Model for Multicontext FPGAs
    Lecture Notes in Computer Science, 2002
    Co-Authors: Naoto Kaneko, Hideharu Amano
    Abstract:

    We propose a general multicontext hardware design model to promote the use of multicontext FPGAs in wide range of applications. There are two major problems to be overcome in multicontext FPGAs programming. One is the complexity of programming an application in multicontext. The design model provides the context scheduling mechanisms to free the programmers from the low-level context management. The other is the difficulty in scheduling non-algorithmic applications. Our distributed demand-driven dynamicsc heduler facilitates their implementation. We have applied the design model to a network interface controller Martini. The result supports the effectiveness of our design model.

Steve C. Chiu - One of the best experts on this subject based on the ideXlab platform.

  • Distributed smart disks for I/O-intensive workloads on switched interconnects
    Future Generation Computer Systems, 2006
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary
    Abstract:

    Disk-based smart storage represents processor-embedded active I/O devices that are equipped with a dedicated network interface controller and on-disk storage (memory and disk). Today's I/O-intensive workloads require architectures whose performance scales with increasing storage capacity. The advent of switch-based, low-latency and high-bandwidth network protocols, such as Virtual interface Architecture and InfiniBand, promises to improve both performance and scalability beyond current state-of-the-art I/O interconnect topologies. This paper evaluates a disk-based distributed smart storage architecture composed of fully distributed processor-embedded disks that are interconnected by a high-performance switched network, and investigates the performance of the proposed processing model on different storage interconnects with representative workloads, including TPC-H queries, association rules mining, high-dimensional data clustering, and two-dimensional fast Fourier transform (FFT).

  • Processor-embedded distributed smart disks for I/O-intensive workloads: architectures, performance models and evaluation
    Journal of Parallel and Distributed Computing, 2004
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary, Mahmut Kandemir
    Abstract:

    Processor-embedded disks, or smart disks, with their network interface controller, can in effect be viewed as processing elements with on-disk memory and secondary storage. The data sizes and access patterns of today's large I/O-intensive workloads require architectures whose processing power scales with increased storage capacity. To address this concern, we propose and evaluate disk-based distributed smart storage architectures. Based on analytically derived performance models, our evaluation with representative workloads show that offloading processing and performing point-to-point data communication improve performance over centralized architectures. Our results also demonstrate that distributed smart disk systems exhibit desirable scalability and can efficiently handle I/O-intensive workloads, such as commercial decision support database (TPC-H) queries, association rules mining, data clustering, and two-dimensional fast Fourier transform, among others.

  • International Conference on Computational Science - Design and evaluation of distributed smart disk architecture for I/O-intensive workloads
    Lecture Notes in Computer Science, 2003
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary
    Abstract:

    Smart disks, a type of processor-embedded active I/O devices, with their on-disk memory and network interface controller, can be viewed as processing elements with attached storage. The growing size and access patterns of today's large I/O-intensive applications require architectures whose processing power scales with the storage capacity. We evaluate a distributed smart disk architecture with representative I/O-intensive workloads including TPC-H queries, association rule mining, data clustering, and 2-D fast Fourier transform applications to study the proposed architecture.

Wei-keng Liao - One of the best experts on this subject based on the ideXlab platform.

  • Distributed smart disks for I/O-intensive workloads on switched interconnects
    Future Generation Computer Systems, 2006
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary
    Abstract:

    Disk-based smart storage represents processor-embedded active I/O devices that are equipped with a dedicated network interface controller and on-disk storage (memory and disk). Today's I/O-intensive workloads require architectures whose performance scales with increasing storage capacity. The advent of switch-based, low-latency and high-bandwidth network protocols, such as Virtual interface Architecture and InfiniBand, promises to improve both performance and scalability beyond current state-of-the-art I/O interconnect topologies. This paper evaluates a disk-based distributed smart storage architecture composed of fully distributed processor-embedded disks that are interconnected by a high-performance switched network, and investigates the performance of the proposed processing model on different storage interconnects with representative workloads, including TPC-H queries, association rules mining, high-dimensional data clustering, and two-dimensional fast Fourier transform (FFT).

  • Processor-embedded distributed smart disks for I/O-intensive workloads: architectures, performance models and evaluation
    Journal of Parallel and Distributed Computing, 2004
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary, Mahmut Kandemir
    Abstract:

    Processor-embedded disks, or smart disks, with their network interface controller, can in effect be viewed as processing elements with on-disk memory and secondary storage. The data sizes and access patterns of today's large I/O-intensive workloads require architectures whose processing power scales with increased storage capacity. To address this concern, we propose and evaluate disk-based distributed smart storage architectures. Based on analytically derived performance models, our evaluation with representative workloads show that offloading processing and performing point-to-point data communication improve performance over centralized architectures. Our results also demonstrate that distributed smart disk systems exhibit desirable scalability and can efficiently handle I/O-intensive workloads, such as commercial decision support database (TPC-H) queries, association rules mining, data clustering, and two-dimensional fast Fourier transform, among others.

  • International Conference on Computational Science - Design and evaluation of distributed smart disk architecture for I/O-intensive workloads
    Lecture Notes in Computer Science, 2003
    Co-Authors: Steve C. Chiu, Wei-keng Liao, Alok Choudhary
    Abstract:

    Smart disks, a type of processor-embedded active I/O devices, with their on-disk memory and network interface controller, can be viewed as processing elements with attached storage. The growing size and access patterns of today's large I/O-intensive applications require architectures whose processing power scales with the storage capacity. We evaluate a distributed smart disk architecture with representative I/O-intensive workloads including TPC-H queries, association rule mining, data clustering, and 2-D fast Fourier transform applications to study the proposed architecture.

Pei-rong Wang - One of the best experts on this subject based on the ideXlab platform.

  • The design and implementation of a general reduced TCP/IP protocol stack for embedded Web server
    2005 International Conference on Industrial Electronics and Control Applications, 1
    Co-Authors: Zhi-liang Zhu, Xiao-xing Gao, Pei-rong Wang
    Abstract:

    The embedded Web server technology is the combination of embedded device and Internet technology, which provides a flexible remote device monitoring and management function based on Internet browser and it has become an advanced development trend of embedded technology. In this paper, a design and implementation scheme of a general reduced Web server protocol stack which aims at the limited resources characteristic that the embedded system has was proposed based on AT90S8515 single chip that uses RISC technology and RTL-8019 network interface controller hardware platform. The overall design issue was thoroughly discussed and the implementation methods of core protocols in protocol stack such as 802.3, ARP, IP, TCP and HTTP were analyzed in detail