The Experts below are selected from a list of 81 Experts worldwide ranked by ideXlab platform

Xiaojun Yang - One of the best experts on this subject based on the ideXlab platform.

  • prototyping a configurable Cache scratchpad memory with virtualized user level rdma capability
    Trans. High Perform. Embed. Archit. Compil., 2019
    Co-Authors: George Kalokerinos, Vassilis Papaefstathiou, George Nikiforos, Stamatis Kavadias, Xiaojun Yang, Dionisios Pnevmatikatos, Manolis Katevenis
    Abstract:

    We present the hardware design and implementation of a local memory system for individual processors inside future chip multi- processors (CMP). Our memory system supports both implicit commu- nication via Caches, and explicit communication via directly accessible local ("scratchpad") memories and remote DMA (RDMA). We provide run-time configurability of the SRAM blocks that lie near each proces- sor, so that portions of them operate as 2nd level (local) Cache, while the rest operate as scratchpad. We also strive to merge the communi- cation subsystems required by the Cache and scratchpad into one inte- grated Network Interface (NI) and Cache Controller (CC), in order to economize on circuits. The processor interacts with the NI at user-level through virtualized command areas in scratchpad; the NI uses a similar access mechanism to provide efficient support for two hardwaresynchro- nization primitives: counters, and queues. We describe the NI design, the hardware cost, and the latencies of our FPGA-based prototype im- plementation that integrates four MicroBlaze processors, each with 64 KBytes of local SRAM, a crossbar NoC, and a DRAM Controller. One- way, end-to-end, user-level communication completes within about 20 clock cycles for short transfer sizes.

  • fpga implementation of a configurable Cache scratchpad memory with virtualized user level rdma capability
    International Conference on Systems, 2009
    Co-Authors: George Kalokerinos, Vassilis Papaefstathiou, George Nikiforos, Stamatis Kavadias, Dionisios Pnevmatikatos, Manolis Katevenis, Xiaojun Yang
    Abstract:

    We report on the hardware implementation of a local memory system for individual processors inside future chip multiprocessors (CMP). It intends to support both implicit communication, via Caches, and explicit communication, via directly accessible local (“scratchpad”) memories and remote DMA (RDMA). We provide run-time configurability of the SRAM blocks near each processor, so that part of them operates as 2nd level (local) Cache, while the rest operates as scratchpad. We also strive to merge the communication subsystems required by the Cache and scratchpad into one integrated Network Interface (NI) and Cache Controller (CC), in order to economize on circuits. The processor communicates with the NI in user-level, through virtualized command areas in scratchpad; through a similar mechanism, the NI also provides efficient support for synchronization, using two hardware primitives: counters, and queues. We describe the block diagram, the hardware cost, and the latencies of our FPGA-based prototype implementation, which integrates four MicroBlaze processors, each with 64 KBytes of local SRAM, a crossbar NoC, and a DRAM Controller on a Xilinx-5 FPGA. One-way, end-to-end, user-level communication completes within about 30 clock cycles for short transfer sizes.

Manolis Katevenis - One of the best experts on this subject based on the ideXlab platform.

  • prototyping a configurable Cache scratchpad memory with virtualized user level rdma capability
    Trans. High Perform. Embed. Archit. Compil., 2019
    Co-Authors: George Kalokerinos, Vassilis Papaefstathiou, George Nikiforos, Stamatis Kavadias, Xiaojun Yang, Dionisios Pnevmatikatos, Manolis Katevenis
    Abstract:

    We present the hardware design and implementation of a local memory system for individual processors inside future chip multi- processors (CMP). Our memory system supports both implicit commu- nication via Caches, and explicit communication via directly accessible local ("scratchpad") memories and remote DMA (RDMA). We provide run-time configurability of the SRAM blocks that lie near each proces- sor, so that portions of them operate as 2nd level (local) Cache, while the rest operate as scratchpad. We also strive to merge the communi- cation subsystems required by the Cache and scratchpad into one inte- grated Network Interface (NI) and Cache Controller (CC), in order to economize on circuits. The processor interacts with the NI at user-level through virtualized command areas in scratchpad; the NI uses a similar access mechanism to provide efficient support for two hardwaresynchro- nization primitives: counters, and queues. We describe the NI design, the hardware cost, and the latencies of our FPGA-based prototype im- plementation that integrates four MicroBlaze processors, each with 64 KBytes of local SRAM, a crossbar NoC, and a DRAM Controller. One- way, end-to-end, user-level communication completes within about 20 clock cycles for short transfer sizes.

  • fpga implementation of a configurable Cache scratchpad memory with virtualized user level rdma capability
    International Conference on Systems, 2009
    Co-Authors: George Kalokerinos, Vassilis Papaefstathiou, George Nikiforos, Stamatis Kavadias, Dionisios Pnevmatikatos, Manolis Katevenis, Xiaojun Yang
    Abstract:

    We report on the hardware implementation of a local memory system for individual processors inside future chip multiprocessors (CMP). It intends to support both implicit communication, via Caches, and explicit communication, via directly accessible local (“scratchpad”) memories and remote DMA (RDMA). We provide run-time configurability of the SRAM blocks near each processor, so that part of them operates as 2nd level (local) Cache, while the rest operates as scratchpad. We also strive to merge the communication subsystems required by the Cache and scratchpad into one integrated Network Interface (NI) and Cache Controller (CC), in order to economize on circuits. The processor communicates with the NI in user-level, through virtualized command areas in scratchpad; through a similar mechanism, the NI also provides efficient support for synchronization, using two hardware primitives: counters, and queues. We describe the block diagram, the hardware cost, and the latencies of our FPGA-based prototype implementation, which integrates four MicroBlaze processors, each with 64 KBytes of local SRAM, a crossbar NoC, and a DRAM Controller on a Xilinx-5 FPGA. One-way, end-to-end, user-level communication completes within about 30 clock cycles for short transfer sizes.

George Kalokerinos - One of the best experts on this subject based on the ideXlab platform.

  • prototyping a configurable Cache scratchpad memory with virtualized user level rdma capability
    Trans. High Perform. Embed. Archit. Compil., 2019
    Co-Authors: George Kalokerinos, Vassilis Papaefstathiou, George Nikiforos, Stamatis Kavadias, Xiaojun Yang, Dionisios Pnevmatikatos, Manolis Katevenis
    Abstract:

    We present the hardware design and implementation of a local memory system for individual processors inside future chip multi- processors (CMP). Our memory system supports both implicit commu- nication via Caches, and explicit communication via directly accessible local ("scratchpad") memories and remote DMA (RDMA). We provide run-time configurability of the SRAM blocks that lie near each proces- sor, so that portions of them operate as 2nd level (local) Cache, while the rest operate as scratchpad. We also strive to merge the communi- cation subsystems required by the Cache and scratchpad into one inte- grated Network Interface (NI) and Cache Controller (CC), in order to economize on circuits. The processor interacts with the NI at user-level through virtualized command areas in scratchpad; the NI uses a similar access mechanism to provide efficient support for two hardwaresynchro- nization primitives: counters, and queues. We describe the NI design, the hardware cost, and the latencies of our FPGA-based prototype im- plementation that integrates four MicroBlaze processors, each with 64 KBytes of local SRAM, a crossbar NoC, and a DRAM Controller. One- way, end-to-end, user-level communication completes within about 20 clock cycles for short transfer sizes.

  • fpga implementation of a configurable Cache scratchpad memory with virtualized user level rdma capability
    International Conference on Systems, 2009
    Co-Authors: George Kalokerinos, Vassilis Papaefstathiou, George Nikiforos, Stamatis Kavadias, Dionisios Pnevmatikatos, Manolis Katevenis, Xiaojun Yang
    Abstract:

    We report on the hardware implementation of a local memory system for individual processors inside future chip multiprocessors (CMP). It intends to support both implicit communication, via Caches, and explicit communication, via directly accessible local (“scratchpad”) memories and remote DMA (RDMA). We provide run-time configurability of the SRAM blocks near each processor, so that part of them operates as 2nd level (local) Cache, while the rest operates as scratchpad. We also strive to merge the communication subsystems required by the Cache and scratchpad into one integrated Network Interface (NI) and Cache Controller (CC), in order to economize on circuits. The processor communicates with the NI in user-level, through virtualized command areas in scratchpad; through a similar mechanism, the NI also provides efficient support for synchronization, using two hardware primitives: counters, and queues. We describe the block diagram, the hardware cost, and the latencies of our FPGA-based prototype implementation, which integrates four MicroBlaze processors, each with 64 KBytes of local SRAM, a crossbar NoC, and a DRAM Controller on a Xilinx-5 FPGA. One-way, end-to-end, user-level communication completes within about 30 clock cycles for short transfer sizes.

Paul A Reed - One of the best experts on this subject based on the ideXlab platform.

  • a 250 mhz 5 w powerpc microprocessor with on chip l2 Cache Controller
    IEEE Journal of Solid-state Circuits, 1997
    Co-Authors: Gianfranco Gerosa, M Alexander, Jose Alvarez, C Croxton, M Daddeo, A R Kennedy, C Nicoletta, J P Nissen, R Philip, Paul A Reed
    Abstract:

    This RISC microprocessor is a new, high-performance, PowerPC microprocessor designed specifically for the mobile and high volume desktop personal computer markets. It is an advanced superscalar design with six execution units, aggressive upstream branch processing, out-of-order instruction execution, and a tightly integrated "backside" L2 Cache. This dual-issue engine has a four-stage pipeline with dual 32-kB eight-way set-associative L1 Caches and an integrated L2 Controller with on-chip L2 tag supporting up to 1 MB of external SRAM. A thermal assist unit and an instruction Cache throttling mechanism are included for thermal management in mobile applications. A 60X system bus and L2 interface speeds of 100 and 250 MHz are achieved, respectively. This microprocessor achieves workstation class performance (estimated 10 SPECint95 and 9 SPECfp95) while only dissipating 5 W at 250 MHz. The 6.35-million transistor 66.5-mm/sup 2/ die is fabricated in a 2.5-V, 0.3-/spl mu/m, five-layer metal CMOS process.

John H Arends - One of the best experts on this subject based on the ideXlab platform.

  • instruction fetch energy reduction using loop Caches for embedded applications with small tight loops
    International Symposium on Low Power Electronics and Design, 1999
    Co-Authors: Bill Moyer, John H Arends
    Abstract:

    A fair amount of work has been done in recent years on reducing power consumption in Caches by using a small instruction buffer placed between the execution pipe and a larger main Cache. These techniques, however, often degrade the overall system performance. In this paper, we propose using a small instruction buffer, also called a loop Cache, to save power. A loop Cache has no address tag store. It consists of a direct-mapped data array and a loop Cache Controller. The loop Cache Controller knows precisely whether the next instruction request will hit in the loop Cache, well ahead of time. As a result, there is no performance degradation.