The Experts below are selected from a list of 16959 Experts worldwide ranked by ideXlab platform
Ulrich Bruning - One of the best experts on this subject based on the ideXlab platform.
-
a resource optimized remote memory access architecture for low Latency communication
International Conference on Parallel Processing, 2009Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich BruningAbstract:This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.
-
ICPP - A Resource Optimized Remote-Memory-Access Architecture for Low-Latency Communication
2009 International Conference on Parallel Processing, 2009Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich BruningAbstract:This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.
Quanlong Guan - One of the best experts on this subject based on the ideXlab platform.
-
TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
2017Co-Authors: Yuxiang Zhang, Quanlong Guan, Lin Cui, Fung Tso, Weijia JiaAbstract:Nowadays, many web services (e.g., cloud storage) are deployed inside datacenters and may trigger transfers to clients through WAN. TCP congestion control is a vital component for improving the performance (e.g., Latency) of these services. Considering complex networking environment, the default congestion control algorithms on servers may not always be the most efficient, and new advanced algorithms will be proposed. However, adjusting congestion control algorithm usually requires modification of TCP stacks of servers, which is difficult if not impossible, especially considering different operating systems and configurations on servers. In this paper, we propose TCon, a light-weight, flexible and scalable platform that allows administrators (or operators) to deploy any appropriate congestion control algorithms transparently without making any changes to TCP stacks of servers. We have implemented TCon in Open vSwitch (OVS) and conducted extensive test-bed experiments by transparently deploying BBR congestion control algorithm over TCon. Test-bed results show that the BBR over TCon works effectively and the performance stays close to its native implementation on servers, Reducing Latency by 12.76% on average.
-
NPC - TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
Lecture Notes in Computer Science, 2017Co-Authors: Yuxiang Zhang, Quanlong GuanAbstract:Nowadays, many web services (e.g., cloud storage) are deployed inside datacenters and may trigger transfers to clients through WAN. TCP congestion control is a vital component for improving the performance (e.g., Latency) of these services. Considering complex networking environment, the default congestion control algorithms on servers may not always be the most efficient, and new advanced algorithms will be proposed. However, adjusting congestion control algorithm usually requires modification of TCP stacks of servers, which is difficult if not impossible, especially considering different operating systems and configurations on servers. In this paper, we propose TCon, a light-weight, flexible and scalable platform that allows administrators (or operators) to deploy any appropriate congestion control algorithms transparently without making any changes to TCP stacks of servers. We have implemented TCon in Open vSwitch (OVS) and conducted extensive test-bed experiments by transparently deploying BBR congestion control algorithm over TCon. Test-bed results show that the BBR over TCon works effectively and the performance stays close to its native implementation on servers, Reducing Latency by 12.76% on average.
Mondrian Nussle - One of the best experts on this subject based on the ideXlab platform.
-
a resource optimized remote memory access architecture for low Latency communication
International Conference on Parallel Processing, 2009Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich BruningAbstract:This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.
-
ICPP - A Resource Optimized Remote-Memory-Access Architecture for Low-Latency Communication
2009 International Conference on Parallel Processing, 2009Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich BruningAbstract:This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.
Yuxiang Zhang - One of the best experts on this subject based on the ideXlab platform.
-
TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
2017Co-Authors: Yuxiang Zhang, Quanlong Guan, Lin Cui, Fung Tso, Weijia JiaAbstract:Nowadays, many web services (e.g., cloud storage) are deployed inside datacenters and may trigger transfers to clients through WAN. TCP congestion control is a vital component for improving the performance (e.g., Latency) of these services. Considering complex networking environment, the default congestion control algorithms on servers may not always be the most efficient, and new advanced algorithms will be proposed. However, adjusting congestion control algorithm usually requires modification of TCP stacks of servers, which is difficult if not impossible, especially considering different operating systems and configurations on servers. In this paper, we propose TCon, a light-weight, flexible and scalable platform that allows administrators (or operators) to deploy any appropriate congestion control algorithms transparently without making any changes to TCP stacks of servers. We have implemented TCon in Open vSwitch (OVS) and conducted extensive test-bed experiments by transparently deploying BBR congestion control algorithm over TCon. Test-bed results show that the BBR over TCon works effectively and the performance stays close to its native implementation on servers, Reducing Latency by 12.76% on average.
-
NPC - TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
Lecture Notes in Computer Science, 2017Co-Authors: Yuxiang Zhang, Quanlong GuanAbstract:Nowadays, many web services (e.g., cloud storage) are deployed inside datacenters and may trigger transfers to clients through WAN. TCP congestion control is a vital component for improving the performance (e.g., Latency) of these services. Considering complex networking environment, the default congestion control algorithms on servers may not always be the most efficient, and new advanced algorithms will be proposed. However, adjusting congestion control algorithm usually requires modification of TCP stacks of servers, which is difficult if not impossible, especially considering different operating systems and configurations on servers. In this paper, we propose TCon, a light-weight, flexible and scalable platform that allows administrators (or operators) to deploy any appropriate congestion control algorithms transparently without making any changes to TCP stacks of servers. We have implemented TCon in Open vSwitch (OVS) and conducted extensive test-bed experiments by transparently deploying BBR congestion control algorithm over TCon. Test-bed results show that the BBR over TCon works effectively and the performance stays close to its native implementation on servers, Reducing Latency by 12.76% on average.
Martin Scherer - One of the best experts on this subject based on the ideXlab platform.
-
a resource optimized remote memory access architecture for low Latency communication
International Conference on Parallel Processing, 2009Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich BruningAbstract:This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.
-
ICPP - A Resource Optimized Remote-Memory-Access Architecture for Low-Latency Communication
2009 International Conference on Parallel Processing, 2009Co-Authors: Mondrian Nussle, Martin Scherer, Ulrich BruningAbstract:This paper introduces a new highly optimized architecture for remote memory access (RMA). RMA, using put and get operations, is a one-sided communication function which amongst others is important in current and upcoming Partitioned Global Address Space (PGAS) systems. In this work, a virtualized hardware unit is described which is resource optimized, exhibits high overlap, processor offload and very good Latency characteristics. To start an RMA operation a single HyperTransport packet caused by one CPU instruction is sufficient, thus Reducing Latency to an absolute minimum. In addition to the basic architecture an implementation in FPGA technology is presented together with an evaluation of the target ASIC-implementation. The current system can sustain more than 4.9 million transactions per second on the FPGA and exhibits an end-to-end Latency of 1.2 μs for an 8-byte put operation. Both values are limited by the FPGA technology used for the prototype implementation. An estimation of the performance reachable on ASIC technology suggests that application to application latencies of less than 500 ns are feasible.