The Experts below are selected from a list of 5781 Experts worldwide ranked by ideXlab platform

Yao Wang - One of the best experts on this subject based on the ideXlab platform.

  • on complexity modeling of h 264 avc video decoding and its application for energy efficient decoding
    IEEE Transactions on Multimedia, 2011
    Co-Authors: Hao Hu, Yao Wang
    Abstract:

    This paper proposes a new complexity model for H.264/AVC video decoding. The model is derived by decomposing the entire decoder into several decoding modules (DM), and identifying the fundamental operation unit (termed complexity unit or CU) in each DM. The complexity of each DM is modeled by the product of the average complexity of one CU and the number of CUs required. The model is shown to be highly accurate for software video decoding both on Intel Pentium mobile 1.6-GHz and ARM Cortex A8 600-MHz processors, over a variety of video contents at different spatial and temporal resolutions and bit rates. We further show how to use this model to predict the required clock frequency and hence perform dynamic voltage and frequency scaling (DVFS) for energy efficient video decoding. We evaluate achievable power savings on both the Intel and ARM Platforms, by using analytical power models for these two Platforms as well as real experiments with the ARM-based TI OMAP35x EVM board. Our study shows that for the Intel Platform where the dynamic power dominates, a power saving factor of 3.7 is possible. For the ARM processor where the static leakage power is not negligible, a saving factor of 2.22 is still achievable.

  • On Complexity Modeling of H.264/AVC Video Decoding and Its Application for Energy Efficient Decoding
    IEEE Transactions on Multimedia, 2011
    Co-Authors: Yao Wang
    Abstract:

    This paper proposes a new complexity model for H.264/AVC video decoding. The model is derived by decomposing the entire decoder into several decoding modules (DM), and identifying the fundamental operation unit (termed complexity unit or CU) in each DM. The complexity of each DM is modeled by the product of the average complexity of one CU and the number of CUs required. The model is shown to be highly accurate for software video decoding both on Intel Pentium mobile 1.6-GHz and ARM Cortex A8 600-MHz processors, over a variety of video contents at different spatial and temporal resolutions and bit rates. We further show how to use this model to predict the required clock frequency and hence perform dynamic voltage and frequency scaling (DVFS) for energy efficient video decoding. We evaluate achievable power savings on both the Intel and ARM Platforms, by using analytical power models for these two Platforms as well as real experiments with the ARM-based TI OMAP35x EVM board. Our study shows that for the Intel Platform where the dynamic power dominates, a power saving factor of 3.7 is possible. For the ARM processor where the static leakage power is not negligible, a saving factor of 2.22 is still achievable.

Hao Hu - One of the best experts on this subject based on the ideXlab platform.

  • on complexity modeling of h 264 avc video decoding and its application for energy efficient decoding
    IEEE Transactions on Multimedia, 2011
    Co-Authors: Hao Hu, Yao Wang
    Abstract:

    This paper proposes a new complexity model for H.264/AVC video decoding. The model is derived by decomposing the entire decoder into several decoding modules (DM), and identifying the fundamental operation unit (termed complexity unit or CU) in each DM. The complexity of each DM is modeled by the product of the average complexity of one CU and the number of CUs required. The model is shown to be highly accurate for software video decoding both on Intel Pentium mobile 1.6-GHz and ARM Cortex A8 600-MHz processors, over a variety of video contents at different spatial and temporal resolutions and bit rates. We further show how to use this model to predict the required clock frequency and hence perform dynamic voltage and frequency scaling (DVFS) for energy efficient video decoding. We evaluate achievable power savings on both the Intel and ARM Platforms, by using analytical power models for these two Platforms as well as real experiments with the ARM-based TI OMAP35x EVM board. Our study shows that for the Intel Platform where the dynamic power dominates, a power saving factor of 3.7 is possible. For the ARM processor where the static leakage power is not negligible, a saving factor of 2.22 is still achievable.

Dhabaleswar K. Panda - One of the best experts on this subject based on the ideXlab platform.

  • memory scalability evaluation of the next generation Intel bensley Platform with infiniband
    High Performance Interconnects, 2006
    Co-Authors: Matthew J. Koop, Wei Huang, Abhinav Vishnu, Dhabaleswar K. Panda
    Abstract:

    As multi-core systems gain popularity for their increased computing power at low-cost, the rest of the architecture must be kept in balance, such as the memory subsystem. Many existing memory subsystems can suffer from scalability issues and show memory performance degradation with more than one process running. To address these scalability issues, Fully-Buffered DIMMs have recently been introduced. In this paper we present an initial performance evaluation of the next-generation multi-core Intel Platform by evaluating the FB-DIMM-based memory subsystem and the associated InfiniBand performance. To the best of our knowledge this is the first such study of Intel multi-core Platforms with multi-rail InfiniBand DDR configurations. We provide an evaluation of the current-generation Intel Lindenhurst Platform as a reference point. We find that the Intel Bensley Platform can provide memory scalability to support memory accesses by multiple processes on the same machine as well as drastically improved inter-node throughput over InfiniBand. On the Bensley Platform we observe a 1.85 times increase in aggregate write bandwidth over the Lindenhurst Platform. For inter-node MPI-level benchmarks we show bi-directional bandwidth of over 4.55 GB/sec for the Bensley Platform using 2 DDR InfiniBand Host Channel Adapters (HCAs), an improvement of 77% over the current generation Lindenhurst Platform. The Bensley system is also able to achieve a throughput of 3.12 million MPI messages/sec in the above configuration.

  • Hot Interconnects - Memory Scalability Evaluation of the Next-Generation Intel Bensley Platform with InfiniBand
    14th IEEE Symposium on High-Performance Interconnects (HOTI'06), 1
    Co-Authors: Matthew J. Koop, Wei Huang, Abhinav Vishnu, Dhabaleswar K. Panda
    Abstract:

    As multi-core systems gain popularity for their increased computing power at low-cost, the rest of the architecture must be kept in balance, such as the memory subsystem. Many existing memory subsystems can suffer from scalability issues and show memory performance degradation with more than one process running. To address these scalability issues, Fully-Buffered DIMMs have recently been introduced. In this paper we present an initial performance evaluation of the next-generation multi-core Intel Platform by evaluating the FB-DIMM-based memory subsystem and the associated InfiniBand performance. To the best of our knowledge this is the first such study of Intel multi-core Platforms with multi-rail InfiniBand DDR configurations. We provide an evaluation of the current-generation Intel Lindenhurst Platform as a reference point. We find that the Intel Bensley Platform can provide memory scalability to support memory accesses by multiple processes on the same machine as well as drastically improved inter-node throughput over InfiniBand. On the Bensley Platform we observe a 1.85 times increase in aggregate write bandwidth over the Lindenhurst Platform. For inter-node MPI-level benchmarks we show bi-directional bandwidth of over 4.55 GB/sec for the Bensley Platform using 2 DDR InfiniBand Host Channel Adapters (HCAs), an improvement of 77% over the current generation Lindenhurst Platform. The Bensley system is also able to achieve a throughput of 3.12 million MPI messages/sec in the above configuration.

Tom J Adelmeyer - One of the best experts on this subject based on the ideXlab platform.

  • characterization analysis of a server consolidation benchmark
    Virtual Execution Environments, 2008
    Co-Authors: Padma Apparao, Don Newell, Ravi Iyer, Xiaomin Zhang, Tom J Adelmeyer
    Abstract:

    Virtualization is already becoming ubiquitous in data centers for the consolidation of multiple workloads on a single Platform. However, there are very few performance studies of server consolidation workloads in the literature. In this paper, our goal is to analyze the performance characteristics of a representative server consolidation workload. To address this goal, we have carried out extensive measurement and profiling experiments of a newly proposed consolidation workload (vConsolidate). vConsolidate consists of a compute intensive workload, a web server, a mail server and a database application running simultaneously on a single Platform. We start by studying the performance slowdown of each workload due to consolidation on a contemporary multi-core dual-processor Intel Platform. We then look at architectural characteristics such as CPI (cycles per instruction) and L2 MP (L2 misses per instruction) I, and analyze the benefits of larger caches for such a consolidated workload. We estimate the virtualization overheads for events such as context switches, interrupts and page faults and show how these impact the performance of the workload in consolidation. Finally, we also present the execution profile of the server consolidation workload and illustrate the life of each VM in the consolidated environment. We conclude by presenting an approach to developing a preliminary performance model based on the performance.

  • VEE - Characterization & analysis of a server consolidation benchmark
    Proceedings of the fourth ACM SIGPLAN SIGOPS international conference on Virtual execution environments - VEE '08, 2008
    Co-Authors: Padma Apparao, Don Newell, Ravi Iyer, Xiaomin Zhang, Tom J Adelmeyer
    Abstract:

    Virtualization is already becoming ubiquitous in data centers for the consolidation of multiple workloads on a single Platform. However, there are very few performance studies of server consolidation workloads in the literature. In this paper, our goal is to analyze the performance characteristics of a representative server consolidation workload. To address this goal, we have carried out extensive measurement and profiling experiments of a newly proposed consolidation workload (vConsolidate). vConsolidate consists of a compute intensive workload, a web server, a mail server and a database application running simultaneously on a single Platform. We start by studying the performance slowdown of each workload due to consolidation on a contemporary multi-core dual-processor Intel Platform. We then look at architectural characteristics such as CPI (cycles per instruction) and L2 MP (L2 misses per instruction) I, and analyze the benefits of larger caches for such a consolidated workload. We estimate the virtualization overheads for events such as context switches, interrupts and page faults and show how these impact the performance of the workload in consolidation. Finally, we also present the execution profile of the server consolidation workload and illustrate the life of each VM in the consolidated environment. We conclude by presenting an approach to developing a preliminary performance model based on the performance.

Kathryn J Hayes - One of the best experts on this subject based on the ideXlab platform.

  • Handbook of Research on Mobility and Computing: Evolving Technologies and Ubiquitous Impacts - Process Innovation with Ambient Intelligence (AmI) Technologies in Manufacturing SMEs: Absorptive Capacity Limitations
    Handbook of Research on Mobility and Computing, 2011
    Co-Authors: Kathryn J Hayes, Ross L Chapman
    Abstract:

    This chapter considers the potential for absorptive capacity limitations to prevent SME manufacturers benefiting from the implementation of Ambient Intelligence (AmI) technologies. The chapter also examines the role of intermediary organisations in alleviating these absorptive capacity constraints. In order to understand the context of the research, a review of the role of SMEs in the Australian manufacturing industry, plus the impacts of government innovation policy and absorptive capacity constraints in SMEs in Australia is provided. Advances in the development of ICT industry standards, and the proliferation of software and support for the Windows/Intel Platform have brought technology to SMEs without the need for bespoke development. The results from the joint European and Australian AmI-4-SME projects suggest that SMEs can successfully use “external research sub-units” in the form of industry networks, research organisations and technology providers to offset internal absorptive capacity limitations.

  • Process Innovation with Ambient Intelligence (AmI) Technologies in Manufacturing SMEs
    Industrial Engineering, 1
    Co-Authors: Kathryn J Hayes, Ross Chapman
    Abstract:

    This chapter considers the potential for absorptive capacity limitations to prevent SME manufacturers benefiting from the implementation of Ambient Intelligence (AmI) technologies. The chapter also examines the role of intermediary organisations in alleviating these absorptive capacity constraints. In order to understand the context of the research, a review of the role of SMEs in the Australian manufacturing industry, plus the impacts of government innovation policy and absorptive capacity constraints in SMEs in Australia is provided. Advances in the development of ICT industry standards, and the proliferation of software and support for the Windows/Intel Platform have brought technology to SMEs without the need for bespoke development. The results from the joint European and Australian AmI-4-SME projects suggest that SMEs can successfully use “external research sub-units” in the form of industry networks, research organisations and technology providers to offset internal absorptive capacity limitations.