The Experts below are selected from a list of 16959 Experts worldwide ranked by ideXlab platform

Ulrich Bruning - One of the best experts on this subject based on the ideXlab platform.

  • a resource optimized remote memory access architecture for low Latency communication
    International Conference on Parallel Processing, 2009
    Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich Bruning
    Abstract:

    This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.

  • ICPP - A Resource Optimized Remote-Memory-Access Architecture for Low-Latency Communication
    2009 International Conference on Parallel Processing, 2009
    Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich Bruning
    Abstract:

    This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.

Quanlong Guan - One of the best experts on this subject based on the ideXlab platform.

  • TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
    2017
    Co-Authors: Yuxiang Zhang, Quanlong Guan, Lin Cui, Fung Tso, Weijia Jia
    Abstract:

    Nowadays, many web services (e.g., cloud storage) are deployed inside datacenters and may trigger transfers to clients through WAN. TCP congestion control is a vital component for improving the performance (e.g., Latency) of these services. Considering complex networking environment, the default congestion control algorithms on servers may not always be the most efficient, and new advanced algorithms will be proposed. However, adjusting congestion control algorithm usually requires modification of TCP stacks of servers, which is difficult if not impossible, especially considering different operating systems and configurations on servers. In this paper, we propose TCon, a light-weight, flexible and scalable platform that allows administrators (or operators) to deploy any appropriate congestion control algorithms transparently without making any changes to TCP stacks of servers. We have implemented TCon in Open vSwitch (OVS) and conducted extensive test-bed experiments by transparently deploying BBR congestion control algorithm over TCon. Test-bed results show that the BBR over TCon works effectively and the performance stays close to its native implementation on servers, Reducing Latency by 12.76% on average.

  • NPC - TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
    Lecture Notes in Computer Science, 2017
    Co-Authors: Yuxiang Zhang, Quanlong Guan
    Abstract:

    Nowadays, many web services (e.g., cloud storage) are deployed inside datacenters and may trigger transfers to clients through WAN. TCP congestion control is a vital component for improving the performance (e.g., Latency) of these services. Considering complex networking environment, the default congestion control algorithms on servers may not always be the most efficient, and new advanced algorithms will be proposed. However, adjusting congestion control algorithm usually requires modification of TCP stacks of servers, which is difficult if not impossible, especially considering different operating systems and configurations on servers. In this paper, we propose TCon, a light-weight, flexible and scalable platform that allows administrators (or operators) to deploy any appropriate congestion control algorithms transparently without making any changes to TCP stacks of servers. We have implemented TCon in Open vSwitch (OVS) and conducted extensive test-bed experiments by transparently deploying BBR congestion control algorithm over TCon. Test-bed results show that the BBR over TCon works effectively and the performance stays close to its native implementation on servers, Reducing Latency by 12.76% on average.

Mondrian Nussle - One of the best experts on this subject based on the ideXlab platform.

  • a resource optimized remote memory access architecture for low Latency communication
    International Conference on Parallel Processing, 2009
    Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich Bruning
    Abstract:

    This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.

  • ICPP - A Resource Optimized Remote-Memory-Access Architecture for Low-Latency Communication
    2009 International Conference on Parallel Processing, 2009
    Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich Bruning
    Abstract:

    This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.

Yuxiang Zhang - One of the best experts on this subject based on the ideXlab platform.

  • TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
    2017
    Co-Authors: Yuxiang Zhang, Quanlong Guan, Lin Cui, Fung Tso, Weijia Jia
    Abstract:

    Nowadays, many web services (e.g., cloud storage) are deployed inside datacenters and may trigger transfers to clients through WAN. TCP congestion control is a vital component for improving the performance (e.g., Latency) of these services. Considering complex networking environment, the default congestion control algorithms on servers may not always be the most efficient, and new advanced algorithms will be proposed. However, adjusting congestion control algorithm usually requires modification of TCP stacks of servers, which is difficult if not impossible, especially considering different operating systems and configurations on servers. In this paper, we propose TCon, a light-weight, flexible and scalable platform that allows administrators (or operators) to deploy any appropriate congestion control algorithms transparently without making any changes to TCP stacks of servers. We have implemented TCon in Open vSwitch (OVS) and conducted extensive test-bed experiments by transparently deploying BBR congestion control algorithm over TCon. Test-bed results show that the BBR over TCon works effectively and the performance stays close to its native implementation on servers, Reducing Latency by 12.76% on average.

  • NPC - TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
    Lecture Notes in Computer Science, 2017
    Co-Authors: Yuxiang Zhang, Quanlong Guan
    Abstract:

    Nowadays, many web services (e.g., cloud storage) are deployed inside datacenters and may trigger transfers to clients through WAN. TCP congestion control is a vital component for improving the performance (e.g., Latency) of these services. Considering complex networking environment, the default congestion control algorithms on servers may not always be the most efficient, and new advanced algorithms will be proposed. However, adjusting congestion control algorithm usually requires modification of TCP stacks of servers, which is difficult if not impossible, especially considering different operating systems and configurations on servers. In this paper, we propose TCon, a light-weight, flexible and scalable platform that allows administrators (or operators) to deploy any appropriate congestion control algorithms transparently without making any changes to TCP stacks of servers. We have implemented TCon in Open vSwitch (OVS) and conducted extensive test-bed experiments by transparently deploying BBR congestion control algorithm over TCon. Test-bed results show that the BBR over TCon works effectively and the performance stays close to its native implementation on servers, Reducing Latency by 12.76% on average.

Martin Scherer - One of the best experts on this subject based on the ideXlab platform.

  • a resource optimized remote memory access architecture for low Latency communication
    International Conference on Parallel Processing, 2009
    Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich Bruning
    Abstract:

    This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.

  • ICPP - A Resource Optimized Remote-Memory-Access Architecture for Low-Latency Communication
    2009 International Conference on Parallel Processing, 2009
    Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich Bruning
    Abstract:

    This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.