The Experts below are selected from a list of 192 Experts worldwide ranked by ideXlab platform
Benghong Lim - One of the best experts on this subject based on the ideXlab platform.
-
virtualizing i o devices on vmware workstation s hosted virtual machine monitor
USENIX Annual Technical Conference, 2001Co-Authors: Jeremy Sugerman, Ganesh Venkitachalam, Benghong LimAbstract:Virtual machines were developed by IBM in the 1960’s to provide concurrent, interactive access to a mainframe computer. Each virtual machine is a replica of the underlying physical machine and users are given the illusion of running directly on the physical machine. Virtual machines also provide benefits like isolation and resource sharing, and the ability to run multiple flavors and configurations of operating systems. VMwareWorkstation brings such mainframe-class virtual machine technology to PC-based desktop and workstation computers. This paper focuses on VMware Workstation’s approach to virtualizing I/O devices. PCs have a staggering variety of hardware, and are usually pre-installed with an operating system. Instead of replacing the pre-installed OS, VMware Workstation uses it to host a user-level application (VMApp) component, as well as to schedule a privileged virtual machine monitor (VMM) component. The VMM directly provides high-performance CPU virtualization while the VMApp uses the host OS to virtualize I/O devices and shield the VMM from the variety of devices. A crucial question is whether virtualizing devices via such a hosted architecture can meet the performance required of high throughput, low latency devices. To this end, this paper studies the virtualization and performance of an Ethernet Adapter on VMware Workstation. Results indicate that with optimizations, VMware Workstation’s hosted virtualization architecture can match native I/O throughput on standard PCs. Although a straightforward hosted implementation is CPU-limited due to virtualization overhead on a 733 MHz Pentium R III system on a 100 Mb/s Ethernet, a series of optimizations targeted at reducing CPU utilization allows the system to match native network throughput. Further optimizations are discussed both within and outside a hosted architecture.
Jeremy Sugerman - One of the best experts on this subject based on the ideXlab platform.
-
virtualizing i o devices on vmware workstation s hosted virtual machine monitor
USENIX Annual Technical Conference, 2001Co-Authors: Jeremy Sugerman, Ganesh Venkitachalam, Benghong LimAbstract:Virtual machines were developed by IBM in the 1960’s to provide concurrent, interactive access to a mainframe computer. Each virtual machine is a replica of the underlying physical machine and users are given the illusion of running directly on the physical machine. Virtual machines also provide benefits like isolation and resource sharing, and the ability to run multiple flavors and configurations of operating systems. VMwareWorkstation brings such mainframe-class virtual machine technology to PC-based desktop and workstation computers. This paper focuses on VMware Workstation’s approach to virtualizing I/O devices. PCs have a staggering variety of hardware, and are usually pre-installed with an operating system. Instead of replacing the pre-installed OS, VMware Workstation uses it to host a user-level application (VMApp) component, as well as to schedule a privileged virtual machine monitor (VMM) component. The VMM directly provides high-performance CPU virtualization while the VMApp uses the host OS to virtualize I/O devices and shield the VMM from the variety of devices. A crucial question is whether virtualizing devices via such a hosted architecture can meet the performance required of high throughput, low latency devices. To this end, this paper studies the virtualization and performance of an Ethernet Adapter on VMware Workstation. Results indicate that with optimizations, VMware Workstation’s hosted virtualization architecture can match native I/O throughput on standard PCs. Although a straightforward hosted implementation is CPU-limited due to virtualization overhead on a 733 MHz Pentium R III system on a 100 Mb/s Ethernet, a series of optimizations targeted at reducing CPU utilization allows the system to match native network throughput. Further optimizations are discussed both within and outside a hosted architecture.
-
USENIX Annual Technical Conference, General Track - Virtualizing I/O Devices on VMware Workstation's Hosted Virtual Machine Monitor
2001Co-Authors: Jeremy Sugerman, Ganesh VenkitachalamAbstract:Virtual machines were developed by IBM in the 1960’s to provide concurrent, interactive access to a mainframe computer. Each virtual machine is a replica of the underlying physical machine and users are given the illusion of running directly on the physical machine. Virtual machines also provide benefits like isolation and resource sharing, and the ability to run multiple flavors and configurations of operating systems. VMwareWorkstation brings such mainframe-class virtual machine technology to PC-based desktop and workstation computers. This paper focuses on VMware Workstation’s approach to virtualizing I/O devices. PCs have a staggering variety of hardware, and are usually pre-installed with an operating system. Instead of replacing the pre-installed OS, VMware Workstation uses it to host a user-level application (VMApp) component, as well as to schedule a privileged virtual machine monitor (VMM) component. The VMM directly provides high-performance CPU virtualization while the VMApp uses the host OS to virtualize I/O devices and shield the VMM from the variety of devices. A crucial question is whether virtualizing devices via such a hosted architecture can meet the performance required of high throughput, low latency devices. To this end, this paper studies the virtualization and performance of an Ethernet Adapter on VMware Workstation. Results indicate that with optimizations, VMware Workstation’s hosted virtualization architecture can match native I/O throughput on standard PCs. Although a straightforward hosted implementation is CPU-limited due to virtualization overhead on a 733 MHz Pentium R III system on a 100 Mb/s Ethernet, a series of optimizations targeted at reducing CPU utilization allows the system to match native network throughput. Further optimizations are discussed both within and outside a hosted architecture.
P Wyckoff - One of the best experts on this subject based on the ideXlab platform.
-
a performance analysis of the ammasso rdma enabled Ethernet Adapter and its iwarp api
International Conference on Cluster Computing, 2005Co-Authors: D Dalessandro, P WyckoffAbstract:Network speeds are increasing well beyond the capabilities of today's CPUs to efficiently handle the traffic. This bottleneck at the CPU causes the processor to spend more of its time handling communication and less time on actual processing. As network speeds reach 10 Gb/s and more, the CPU simply can not keep up with the data. Various methods have been proposed to solve this problem. High performance interconnects, such as Infiniband, have been developed that rely on RDMA and protocol offload in order to achieve higher throughput and lower latency. In this paper we evaluate the feasibility of a similar approach which, unlike existing high performance interconnects, requires no special infrastructure. RDMA over Ethernet, otherwise known as iWARP, facilitates the zero copy exchange of data over ordinary local area networks. Since it is based on TCP, iWARP enables RDMA in the wide area network as well. This paper provides a look into the performance of one of the earliest commodity implementations of this emerging technology, the Ammasso 1100 RNIC
-
CLUSTER - A Performance Analysis of the Ammasso RDMA Enabled Ethernet Adapter and its iWARP API
2005 IEEE International Conference on Cluster Computing, 2005Co-Authors: D Dalessandro, P WyckoffAbstract:Network speeds are increasing well beyond the capabilities of today's CPUs to efficiently handle the traffic. This bottleneck at the CPU causes the processor to spend more of its time handling communication and less time on actual processing. As network speeds reach 10 Gb/s and more, the CPU simply can not keep up with the data. Various methods have been proposed to solve this problem. High performance interconnects, such as Infiniband, have been developed that rely on RDMA and protocol offload in order to achieve higher throughput and lower latency. In this paper we evaluate the feasibility of a similar approach which, unlike existing high performance interconnects, requires no special infrastructure. RDMA over Ethernet, otherwise known as iWARP, facilitates the zero copy exchange of data over ordinary local area networks. Since it is based on TCP, iWARP enables RDMA in the wide area network as well. This paper provides a look into the performance of one of the earliest commodity implementations of this emerging technology, the Ammasso 1100 RNIC
Ganesh Venkitachalam - One of the best experts on this subject based on the ideXlab platform.
-
virtualizing i o devices on vmware workstation s hosted virtual machine monitor
USENIX Annual Technical Conference, 2001Co-Authors: Jeremy Sugerman, Ganesh Venkitachalam, Benghong LimAbstract:Virtual machines were developed by IBM in the 1960’s to provide concurrent, interactive access to a mainframe computer. Each virtual machine is a replica of the underlying physical machine and users are given the illusion of running directly on the physical machine. Virtual machines also provide benefits like isolation and resource sharing, and the ability to run multiple flavors and configurations of operating systems. VMwareWorkstation brings such mainframe-class virtual machine technology to PC-based desktop and workstation computers. This paper focuses on VMware Workstation’s approach to virtualizing I/O devices. PCs have a staggering variety of hardware, and are usually pre-installed with an operating system. Instead of replacing the pre-installed OS, VMware Workstation uses it to host a user-level application (VMApp) component, as well as to schedule a privileged virtual machine monitor (VMM) component. The VMM directly provides high-performance CPU virtualization while the VMApp uses the host OS to virtualize I/O devices and shield the VMM from the variety of devices. A crucial question is whether virtualizing devices via such a hosted architecture can meet the performance required of high throughput, low latency devices. To this end, this paper studies the virtualization and performance of an Ethernet Adapter on VMware Workstation. Results indicate that with optimizations, VMware Workstation’s hosted virtualization architecture can match native I/O throughput on standard PCs. Although a straightforward hosted implementation is CPU-limited due to virtualization overhead on a 733 MHz Pentium R III system on a 100 Mb/s Ethernet, a series of optimizations targeted at reducing CPU utilization allows the system to match native network throughput. Further optimizations are discussed both within and outside a hosted architecture.
-
USENIX Annual Technical Conference, General Track - Virtualizing I/O Devices on VMware Workstation's Hosted Virtual Machine Monitor
2001Co-Authors: Jeremy Sugerman, Ganesh VenkitachalamAbstract:Virtual machines were developed by IBM in the 1960’s to provide concurrent, interactive access to a mainframe computer. Each virtual machine is a replica of the underlying physical machine and users are given the illusion of running directly on the physical machine. Virtual machines also provide benefits like isolation and resource sharing, and the ability to run multiple flavors and configurations of operating systems. VMwareWorkstation brings such mainframe-class virtual machine technology to PC-based desktop and workstation computers. This paper focuses on VMware Workstation’s approach to virtualizing I/O devices. PCs have a staggering variety of hardware, and are usually pre-installed with an operating system. Instead of replacing the pre-installed OS, VMware Workstation uses it to host a user-level application (VMApp) component, as well as to schedule a privileged virtual machine monitor (VMM) component. The VMM directly provides high-performance CPU virtualization while the VMApp uses the host OS to virtualize I/O devices and shield the VMM from the variety of devices. A crucial question is whether virtualizing devices via such a hosted architecture can meet the performance required of high throughput, low latency devices. To this end, this paper studies the virtualization and performance of an Ethernet Adapter on VMware Workstation. Results indicate that with optimizations, VMware Workstation’s hosted virtualization architecture can match native I/O throughput on standard PCs. Although a straightforward hosted implementation is CPU-limited due to virtualization overhead on a 733 MHz Pentium R III system on a 100 Mb/s Ethernet, a series of optimizations targeted at reducing CPU utilization allows the system to match native network throughput. Further optimizations are discussed both within and outside a hosted architecture.
Giuseppe Ciaccio - One of the best experts on this subject based on the ideXlab platform.
-
using a self connected gigabit Ethernet Adapter as a memcpy low overhead engine for mpi
Lecture Notes in Computer Science, 2003Co-Authors: Giuseppe CiaccioAbstract:Memory copies in messaging systems can be a major source of performance degradation in cluster computing. In this paper we discuss a system which can offload a host CPU from most of the overhead of copying data between distinct regions in the host physical memory. The sistem is implemented as a special-purpose Linux device driver operating a generic, non-programmable Gigabit Ethernet Adapter connected to itself. Whenever the descriptor-based DMA engines of the Adapter are instructed to start a data communication, the data are read from the host memory and written to the memory itself thanks to the loopback cable; this is semantically equivalent to a non-blocking memory copy operation performed by the two DMA engines. Suitable completion test/waiting routines are also implemented, in order to provide traditional, blocking semantics in a split-phase fashion. An implementation of MPI using this system in place of traditional memcpy() calls on receive shows a significantly lower receive overhead.
-
PVM/MPI - Using a Self-connected Gigabit Ethernet Adapter as a memcpy() Low-Overhead Engine for MPI
Recent Advances in Parallel Virtual Machine and Message Passing Interface, 2003Co-Authors: Giuseppe CiaccioAbstract:Memory copies in messaging systems can be a major source of performance degradation in cluster computing. In this paper we discuss a system which can offload a host CPU from most of the overhead of copying data between distinct regions in the host physical memory. The sistem is implemented as a special-purpose Linux device driver operating a generic, non-programmable Gigabit Ethernet Adapter connected to itself. Whenever the descriptor-based DMA engines of the Adapter are instructed to start a data communication, the data are read from the host memory and written to the memory itself thanks to the loopback cable; this is semantically equivalent to a non-blocking memory copy operation performed by the two DMA engines. Suitable completion test/waiting routines are also implemented, in order to provide traditional, blocking semantics in a split-phase fashion. An implementation of MPI using this system in place of traditional memcpy() calls on receive shows a significantly lower receive overhead.
-
CLUSTER - Low-cost Gigabit Ethernet at Work
Cluster Computing, 2000Co-Authors: Giuseppe Ciaccio, Giovanni ChiolaAbstract:In this paper we report about the recently completed porting of the Genoa Active Message MAchine (GAMMA) to the Netgear GA620 Gigabit Ethernet Adapter. Such device is a low-cost (less than 300 US dollars at the time of writing) Gigabit Ethernet Adapter with a state-of-art internal architecture based on two programmable on-board processors and substantial amount of on-board RAM, which makes this product a very appealing, cheap alternative to Myrinet. A combination of low end-to-end latency (32 s) and high transmission throughput (103 MByte/s end-to-end) demonstrates the potential for Gigabit Ethernet lightweight protocols to yield messaging performance comparable to the best Myrinet protocols. This result is of interest, given the envisaged drop in cost of Gigabit Ethernet due to the forthcoming transition from fiber optic to UTP cabling and ever increasing mass market production of such standard interconnect.
-
LCN - Exploiting Gigabit Ethernet capacity for cluster applications
27th Annual IEEE Conference on Local Computer Networks 2002. Proceedings. LCN 2002., 1Co-Authors: Giuseppe Ciaccio, M. Ehlert, Bettina SchnorAbstract:In this paper we report about the recently completed porting of GAMMA to the Netgear GA621 Gigabit Ethernet Adapter, and provide a comparison among GAMMA, MPI/GAMMA, TCP/IP and MPICH/TCP, based on the Netgear GA621 and the older Netgear GA620 network Adapters and using different device drivers, in a Gigabit Ethernet cluster of PC running Linux 2.4. GAMMA (the Genoa Active Message Machine) is a lightweight messaging system based on an active message-like paradigm, originally designed for efficient exploitation of Fast Ethernet interconnects. The comparison includes simple latency/bandwidth evaluation of the messaging systems on both Adapters, as well as performance comparisons based on the NAS Parallel Benchmarks and an end-user fluid dynamics application called Modular Ocean Model (MOM). The analysis of results provides useful hints concerning the efficient use of Gigabit Ethernet with clusters of PC. In particular, it emerges that GAMMA on the GA621 Adapter, with a combination of low end-to-end latency (8.5 /spl mu/s) and high throughput (118.4 MByte/s), provides a performing, cost-effective alternative to proprietary high-speed networks, e.g. Myrinet, for a wide range of cluster computing applications.