The Experts below are selected from a list of 10602 Experts worldwide ranked by ideXlab platform

Marci Sylwestrzak - One of the best experts on this subject based on the ideXlab platform.

  • massively parallel data processing for quantitative total flow imaging with optical coherence microscopy and tomography
    Computer Physics Communications, 2017
    Co-Authors: Marci Sylwestrzak, Daniel Szlag, Paul J Marchand, Ashwin S Kumar, Theo Lasser
    Abstract:

    Abstract We present an application of massively parallel processing of quantitative flow measurements data acquired using spectral optical coherence microscopy (SOCM). The need for massive signal processing of these particular datasets has been a major hurdle for many applications based on SOCM. In view of this difficulty, we implemented and adapted quantitative total flow estimation algorithms on graphics processing units (GPU) and achieved a 150 fold reduction in processing time when compared to a former CPU implementation. As SOCM constitutes the microscopy counterpart to spectral optical coherence tomography (SOCT), the developed processing procedure can be applied to both imaging modalities. We present the developed DLL library integrated in MATLAB (with an example) and have included the source code for adaptations and future improvements. Program summary Program title: CudaOCMproc Catalogue identifier: AFBT_v1_0 Program summary URL: http://cpc.cs.qub.ac.uk/summaries/AFBT_v1_0.html Program obtainable from: CPC Program Library, Queen’s University, Belfast, N. Ireland Licensing provisions: GNU GPLv3 No. of lines in distributed program, including test data, etc.: 913552 No. of bytes in distributed program, including test data, etc.: 270876249 Distribution format: tar.gz Programming language: CUDA/C, MATLAB. Computer: Intel x64 CPU, GPU supporting CUDA technology. Operating system: 64-bit Windows 7 Professional. Has the code been vectorized or parallelized?: Yes, CPU code has been vectorized in MATLAB, CUDA code has been parallelized. RAM: Dependent on users parameters, typically between several gigabytes and several tens of gigabytes Classification: 6.5, 18. Nature of problem: Speed up of data processing in optical coherence microscopy Solution method: Utilization of GPU for massively parallel data processing Additional comments: Compiled DLL library with source code and documentation, example of utilization (MATLAB script with raw data) Running time: 1,8 s for one B-scan (150  × faster in comparison to the CPU data processing time)

  • Real-time massively parallel processing of Spectral Optical Coherence Tomography data on Graphics processing Units
    2016
    Co-Authors: Marci Sylwestrzak, Daniel Szlag, Maciej Szkulmowski, Pio Targowski
    Abstract:

    In this contribution we describe a specialised data processing system for Spectral Optical Coherence Tomography (SOCT) biomedical imaging which utilises massively parallel data processing on a low-cost, Graphics processing Unit (GPU). One of the most significant limitations of SOCT is the data processing time on the main processor of the computer (CPU), which is generally longer than the data acquisition. Therefore, real-time imaging with acceptable quality is limited to a small number of tomogram lines (A-scans). Recent progress in graphics cards technology gives a promising solution of this problem. The newest graphics processing units allow not only for a very high speed three dimensional (3D) rendering, but also for a general purpose parallel numerical calculations with efficiency higher than provided by the CPU. The presented system utilizes CUDA ™ graphic card and allows for a very effective real time SOCT imaging. The total imaging speed for 2D data consisting of 1200 A-scans is higher than refresh rate of a 120 Hz monitor. 3D rendering of the volume data build of 10 000 A-scans is performed with frame rate of about 9 frames per second. These frame rates include data transfer from a frame grabber to GPU, data processing and 3D rendering to the screen. The software description includes data flow, parallel processing and organization of threads. For illustration we show real time high resolution SOCT imaging of human skin and eye

  • real time imaging for spectral optical coherence tomography with massively parallel data processing
    Photonics Letters of Poland, 2010
    Co-Authors: Marci Sylwestrzak, Daniel Szlag, Maciej Szkulmowski, Pio Targowski
    Abstract:

    In this paper the application of massively parallel processing of Spectral Optical Coherence Tomography (SOCT) data with the aid of a low-cost Graphic processing Unit (GPU) is presented. The reported system may be used for real-time imaging of high resolution 2D tomograms or for presenting volume data. The overall imaging speed is over 100 frames/second for 2D tomograms built of 1024 A-scans and 9 frames/second for 3D volume images containing 100 slices of 100 A-scans. This includes t he acquisition of 2048 pixels from a CCD camera per A-scan, data transfer to the processor, all necessary processing and rendering on the screen. As a contribution, the description of a data flow and parallel processing organization in a GPU is given. Full Text: PDF References: M. Wojtkowski et al., Am. J. Ophthalmol. 138, 412 (2004). [CrossRef] M. Wojtkowski et al., Ophthalmology 112, 1734 (2005). [CrossRef] D. Stifter, Applied Physics B Lasers and Optics 88, 337 (2007). [CrossRef] J.E. Stone et al., Journal of Computational Chemistry 28, 2618 (2007). [CrossRef] I.S. Ufimtsev, T.J. Martinez, J. Chem. Theory Comput. 4, 222 (2008). [CrossRef] E. Gutierrez et al., Computational Science 5101, 700 (2008). J. Probst, P. Koch, G. Huttmann, Proc. SPIE 7372, 7372_0Q, (2009). [CrossRef] S. Van der Jeught et al., JBO Letters 15, 30511 (2010). [CrossRef] K. Zhang, J.U. Kang, Opt. Exp. 18, 11772 (2010) [CrossRef] M. Sylwestrzak et al., Proc. SPIE 7391, 73910A (2009) [CrossRef]

Daniel Szlag - One of the best experts on this subject based on the ideXlab platform.

  • massively parallel data processing for quantitative total flow imaging with optical coherence microscopy and tomography
    Computer Physics Communications, 2017
    Co-Authors: Marci Sylwestrzak, Daniel Szlag, Paul J Marchand, Ashwin S Kumar, Theo Lasser
    Abstract:

    Abstract We present an application of massively parallel processing of quantitative flow measurements data acquired using spectral optical coherence microscopy (SOCM). The need for massive signal processing of these particular datasets has been a major hurdle for many applications based on SOCM. In view of this difficulty, we implemented and adapted quantitative total flow estimation algorithms on graphics processing units (GPU) and achieved a 150 fold reduction in processing time when compared to a former CPU implementation. As SOCM constitutes the microscopy counterpart to spectral optical coherence tomography (SOCT), the developed processing procedure can be applied to both imaging modalities. We present the developed DLL library integrated in MATLAB (with an example) and have included the source code for adaptations and future improvements. Program summary Program title: CudaOCMproc Catalogue identifier: AFBT_v1_0 Program summary URL: http://cpc.cs.qub.ac.uk/summaries/AFBT_v1_0.html Program obtainable from: CPC Program Library, Queen’s University, Belfast, N. Ireland Licensing provisions: GNU GPLv3 No. of lines in distributed program, including test data, etc.: 913552 No. of bytes in distributed program, including test data, etc.: 270876249 Distribution format: tar.gz Programming language: CUDA/C, MATLAB. Computer: Intel x64 CPU, GPU supporting CUDA technology. Operating system: 64-bit Windows 7 Professional. Has the code been vectorized or parallelized?: Yes, CPU code has been vectorized in MATLAB, CUDA code has been parallelized. RAM: Dependent on users parameters, typically between several gigabytes and several tens of gigabytes Classification: 6.5, 18. Nature of problem: Speed up of data processing in optical coherence microscopy Solution method: Utilization of GPU for massively parallel data processing Additional comments: Compiled DLL library with source code and documentation, example of utilization (MATLAB script with raw data) Running time: 1,8 s for one B-scan (150  × faster in comparison to the CPU data processing time)

  • Real-time massively parallel processing of Spectral Optical Coherence Tomography data on Graphics processing Units
    2016
    Co-Authors: Marci Sylwestrzak, Daniel Szlag, Maciej Szkulmowski, Pio Targowski
    Abstract:

    In this contribution we describe a specialised data processing system for Spectral Optical Coherence Tomography (SOCT) biomedical imaging which utilises massively parallel data processing on a low-cost, Graphics processing Unit (GPU). One of the most significant limitations of SOCT is the data processing time on the main processor of the computer (CPU), which is generally longer than the data acquisition. Therefore, real-time imaging with acceptable quality is limited to a small number of tomogram lines (A-scans). Recent progress in graphics cards technology gives a promising solution of this problem. The newest graphics processing units allow not only for a very high speed three dimensional (3D) rendering, but also for a general purpose parallel numerical calculations with efficiency higher than provided by the CPU. The presented system utilizes CUDA ™ graphic card and allows for a very effective real time SOCT imaging. The total imaging speed for 2D data consisting of 1200 A-scans is higher than refresh rate of a 120 Hz monitor. 3D rendering of the volume data build of 10 000 A-scans is performed with frame rate of about 9 frames per second. These frame rates include data transfer from a frame grabber to GPU, data processing and 3D rendering to the screen. The software description includes data flow, parallel processing and organization of threads. For illustration we show real time high resolution SOCT imaging of human skin and eye

  • real time imaging for spectral optical coherence tomography with massively parallel data processing
    Photonics Letters of Poland, 2010
    Co-Authors: Marci Sylwestrzak, Daniel Szlag, Maciej Szkulmowski, Pio Targowski
    Abstract:

    In this paper the application of massively parallel processing of Spectral Optical Coherence Tomography (SOCT) data with the aid of a low-cost Graphic processing Unit (GPU) is presented. The reported system may be used for real-time imaging of high resolution 2D tomograms or for presenting volume data. The overall imaging speed is over 100 frames/second for 2D tomograms built of 1024 A-scans and 9 frames/second for 3D volume images containing 100 slices of 100 A-scans. This includes t he acquisition of 2048 pixels from a CCD camera per A-scan, data transfer to the processor, all necessary processing and rendering on the screen. As a contribution, the description of a data flow and parallel processing organization in a GPU is given. Full Text: PDF References: M. Wojtkowski et al., Am. J. Ophthalmol. 138, 412 (2004). [CrossRef] M. Wojtkowski et al., Ophthalmology 112, 1734 (2005). [CrossRef] D. Stifter, Applied Physics B Lasers and Optics 88, 337 (2007). [CrossRef] J.E. Stone et al., Journal of Computational Chemistry 28, 2618 (2007). [CrossRef] I.S. Ufimtsev, T.J. Martinez, J. Chem. Theory Comput. 4, 222 (2008). [CrossRef] E. Gutierrez et al., Computational Science 5101, 700 (2008). J. Probst, P. Koch, G. Huttmann, Proc. SPIE 7372, 7372_0Q, (2009). [CrossRef] S. Van der Jeught et al., JBO Letters 15, 30511 (2010). [CrossRef] K. Zhang, J.U. Kang, Opt. Exp. 18, 11772 (2010) [CrossRef] M. Sylwestrzak et al., Proc. SPIE 7391, 73910A (2009) [CrossRef]

Masatoshi Ishikawa - One of the best experts on this subject based on the ideXlab platform.

  • a cmos vision chip with simd processing element array for 1 ms image processing
    International Solid-State Circuits Conference, 1999
    Co-Authors: Masatoshi Ishikawa, K Ogawa, Takashi Komuro, I Ishii
    Abstract:

    Conventional image processing systems use a video signal as a transmission signal between an image sensor and image processor. The video rate, however, is not fast enough for some applications such as visual feedback for robot control, automobiles, gesture recognition for human interfaces, high speed visual inspection, microscope image processing, and so on. For such applications, the video signal, which is a time-multiplexed signal of pixel data using scanning circuits, is a bottleneck of conventional image processing systems. The key to the realization of high speed and flexible image processing beyond the video signal rate is to remove the bottleneck of data transfer and implement general purpose functionality. In other words, a fully parallel architecture with generality of processing should be adopted instead of scanning circuits. This CMOS vision chip has massively parallel processing elements integrated with photodetectors. These chips have a SIMD massively parallel processing architecture with each processing element (PE) connected to a photodetector (PD) without scanning circuits.

  • 1ms target tracking system using massively parallel processing vision
    Journal of the Robotics Society of Japan, 1997
    Co-Authors: Yoshihiro Nakabo, Idaku Ishii, Masatoshi Ishikawa
    Abstract:

    It is obvious that visual servo systems have various applications for real time robot control. But most conventional vision systems have serious restriction on their performance, since those systems always use CCD cameras to acquire the images which transmitted by serial video signals. Therefore the sampling rate is limited by the video frame rate, and this restriction on the speed is quite insufficient to control of the robot.To solve this problem we have developed a massively parallel processing vision system called SPE-256 in which the photo-detectors and processing elements are directly connected. We have realized high speed visual feedback with 1ms cycle time.In this paper we describe our 1 ms visual feedback system and its performance in the application of high speed target tracking.

  • target tracking algorithm for 1 ms visual feedback system using massively parallel processing
    International Conference on Robotics and Automation, 1996
    Co-Authors: Idaku Ishii, Yoshihiro Nakabo, Masatoshi Ishikawa
    Abstract:

    Most conventional visual feedback systems using CCD camera are restricted by video rates and therefore cannot be adapted to the changing environment sufficiently quickly. To solve this problem we developed a 1 ms visual feedback system using a general purpose massively parallel processing vision system in which photo-detectors and processing elements are directly connected. For high speed visual feedback fast image processing algorithms are also required. In particular, the difference of images between frames is very small in our system because of its high speed frame rate. Using this feature, we can realize several image processing techniques by simpler algorithms. In this paper we propose a simple algorithm for target tracking using the feature of high speed vision, and realize target tracking on the 1 ms visual feedback system.

  • high speed target tracking using massively parallel processing vision
    Intelligent Robots and Systems, 1993
    Co-Authors: Y Yamada, Masatoshi Ishikawa
    Abstract:

    A high speed target tracking system using a two-dimensional vision sensor with cell-parallel processing capabilities is described. A general purpose processing element connected to each sensing element processes various vision tasks in data parallel at a speed far beyond the video rate. The sensor architecture is simple and scalable, and assumes that this vision should be implemented into a VLSI chip. The cycle time of parallel processing vision is 670 /spl mu/s in the case of their experimental target tracking system. By using the fast position signals, the system has attained target tracking at the actuator speed limit.

  • high speed vision system using massively parallel processing
    Intelligent Robots and Systems, 1992
    Co-Authors: Masatoshi Ishikawa, A Morita, N Takayanagi
    Abstract:

    Abstruct-A visual processing system for early vision using massively parallel processing is described. The system named SPE-4k (Sensory processing Element - 4k) has 64x64~4096 processing elements (PE’s) which are directly connected with photo detectors in parallel. Eight PE’s are integrated into an originally developed LSI. An mesh type networks is used for interconnection among PE’s. The system carries out input and processing of visual information at extremely high speed by using massively parallel processing. Edge detection using four neighborhoods, for example, can be done in 3.3~s. Since the PE is designed to be so compact that the system can be integrated, the system is regarded as a scale up model of integrated parallel processing vision system in future. In this paper, architectures of the parallel processing LSI and the SPE-4k system are shown. processing performance of the system is evaluated by carrying out some applications on the system.

Zapata Rodriguez Mireya - One of the best experts on this subject based on the ideXlab platform.

  • Arquitectura escalable SIMD con conectividad jerárquica y reconfigurable para la emulación de SNN
    Universitat Politècnica de Catalunya, 2017
    Co-Authors: Zapata Rodriguez Mireya
    Abstract:

    A biological neural system consists of millions of highly integrated neurons with multiple dynamic functions operating in coordination with each other. Its structural organization is characterized by highly hierarchical assemblies. These assemblies are distinguished by locally dense and globally ispersed connections communicated by spikes traveling through the axon to the target neuron. In the last century, approaching the biological complexity of the cortex by means of hardware architectures has continued to be a challenge still unattainable. This is not only due to the massively parallel processing with support for the communication between neurons in large-scale networks, but also for the need of mechanisms that allow the evolution of the neural network efficiently. In this context, this thesis contributes to the development of an architecture called HEENS (Hardware Emulator of Evolved Neural System), which supports inter-chip connectivity with a ring topology between a Master Chip (MC) controlling one or more Neuromorphic Chips (NCs). The MC is implemented in a PSoC device that integrates a CPU ARM Dual Core together with programmable logic. The ARM is responsible for setting up the communication ring and executing the software application that controls the data configuration transmission from the algorithm and the neural parameters to all NCs in the network. Besides, the MC is in charge of activating the evolution mode of the network, as well as managing the dispatching of reconfiguration data to any of the nodes during the execution. Each NC, in turn, consists of a configurable 2D array of processing Elements (PEs) with a SIMD-like processing scheme implemented on a Kintex7 FPGA. NCs are SNN multiprocessors that support the execution of any neural algorithm based on spikes. A set of custom instructions was designed specifically for this architecture. The NCs support a hierarchical scheme of local and global spikes to mimic the brain structural configuration. Local spikes establish inter-neuronal connectivity within a single chip and the global ones allow inter-modular communication between different chips. The NCs have fixed hub neurons that process local and global spikes, thus allowing inter-modular and intra-modular connectivity. This definition of local and global spikes allows the development of multi-level hierarchical architectures inspired by the brain topologies, and offers excellent scalability. The spike propagation through the multi-chip network is supported by an Aurora / AER-SRT protocol stack. The Aurora protocol encapsulates and de-capsulates the packets transmitted through a high-speed serial link that communicates the platform, while the Synchronous Address Event representation (AER-SRT) protocol manages the data (address events) and controls packets that allow synchronization of the operation of the neural network. Each event encapsulates the address neuron that fires a spike as result of the neural algorithm execution. The definition of local and global synaptic topology is implemented using on-chip RAM blocks, which reduces the combinational logic requirements and, in addition to allowing the dynamic connectivity configuration, permits the development of evolutionary applications by supporting the on-line reconfiguration of both the neural algorithm or the neural and synaptic parameters. HEENS also supports axon programmable delays, which incorporates dynamic features to the network.Un sistema neuronal biológico consiste de millones de neuronas altamente integradas con múltiples funciones dinámicas operando en coordinación entre sí. Su organización estructural se caracteriza por contener agrupaciones altamente jerárquicas. Dichas agrupaciones se distinguen por conexiones localmente densas y globalmente dispersas comunicadas a través de pulsos transitorios (spikes) que viajan por el axón hasta la neurona destino. En el último siglo, aproximarse a la complejidad biológica del cortex mediante arquitecturas de hardware continúa siendo un desafío todavía inalcanzable. Esto se debe, no sólo al masivo procesamiento paralelo con soporte para la comunicación entre neuronas en redes de gran escala, sinó también a la necesidad de mecanismos que permitan la evolución de la red neuronal de forma eficiente. En este marco, esta tesis contribuye al desarrollo de una arquitectura denominada HEENS (Emulador de Hardware para Sistemas Neuronales Evolutivos, Hardware Emulator of Evolved Neural System) que soporta conectividad inter-chip con una topología de anillo entre un chip que actúa de master (MC) y uno o más Chips Neuromórficos (NCs). El MC está implementado en un dispositivo PSoC que integra un CPU ARM Dual Core junto con lógica programable. El ARM se encarga de configurar el anillo de comunicación y de ejecutar la aplicación de software que controla el envío de información de configuración del algoritmo y los parámetros neuronales a todos los NCs de la red. Además, el MC es el encargado de activar el modo de evolución de la red, así como de gestionar el envío de datos de reconfiguración a cualquiera de los nodos durante la ejecución. Cada NC a su vez, está compuesto por un arreglo 2D parametrizable de Elementos de Procesamiento (processing Elements, PEs) con un esquema de procesamiento tipo SIMD implementado sobre una FPGA Kintex7. Los NCs son multiprocesadores SNN que soportan la ejecución de cualquier algoritmo neuronal basado en spikes. Se cuenta con un set de instrucciones personalizadas diseñadas específicamente para esta arquitectura. Imitando la configuración estructural del cerebro los NC soportan un esquema jerárquico con spikes locales y globales. Los spikes locales establecen la conectividad inter-neuronal dentro de un mismo chip, y los globales la comunicación inter-modular entre diferentes chips. Los NC cuentan con neuronas fijas tipo hub que procesan spikes locales y globales que permiten la conectividad inter e intra modulos. La definición de spikes locales y globales permite desarrollar arquitecturas jerárquicas multi-nivel que se inspiran en las topologías del cerebro y ofrecen una escalabilidad excelente. La propagación de spikes a través de la red multi-chip es soportada por una pila de protocolos Aurora/AER-SRT. El protocolo Aurora encapsula y desencapsula los paquetes transmitidos a través del enlace serial de alta velocidad que comunica la plataforma. Mientras que el protocolo Síncrono de Representación de Eventos de Dirección (AER-SRT) gestiona los datos (eventos de dirección) y los paquetes de control que permiten sincronizar la operación de la red neuronal. Cada evento encapsula la dirección de la neurona que genera un spike como resultado del procesamiento del algoritmo neuronal. La definición de topología sináptica local y global es implementada usando bloques de memoria RAM on-chip, lo que reduce los requerimientos de lógica combinacional y, además de facilitar la configuración del conexionado sin modificar el hardware, permite el desarrollo de aplicaciones evolutivas al soportar la reconfiguración on-line tanto del algoritmo neuronal como de los parámetros neuronales y sinápticos. HEENS también admite retardos programables de axón, lo cual incorpora características dinámicas a la red

  • Arquitectura escalable SIMD con conectividad jerárquica y reconfigurable para la emulación de SNN
    Universitat Politècnica de Catalunya, 2017
    Co-Authors: Zapata Rodriguez Mireya
    Abstract:

    A biological neural system consists of millions of highly integrated neurons with multiple dynamic functions operating in coordination with each other. Its structural organization is characterized by highly hierarchical assemblies. These assemblies are distinguished by locally dense and globally ispersed connections communicated by spikes traveling through the axon to the target neuron. In the last century, approaching the biological complexity of the cortex by means of hardware architectures has continued to be a challenge still unattainable. This is not only due to the massively parallel processing with support for the communication between neurons in large-scale networks, but also for the need of mechanisms that allow the evolution of the neural network efficiently. In this context, this thesis contributes to the development of an architecture called HEENS (Hardware Emulator of Evolved Neural System), which supports inter-chip connectivity with a ring topology between a Master Chip (MC) controlling one or more Neuromorphic Chips (NCs). The MC is implemented in a PSoC device that integrates a CPU ARM Dual Core together with programmable logic. The ARM is responsible for setting up the communication ring and executing the software application that controls the data configuration transmission from the algorithm and the neural parameters to all NCs in the network. Besides, the MC is in charge of activating the evolution mode of the network, as well as managing the dispatching of reconfiguration data to any of the nodes during the execution. Each NC, in turn, consists of a configurable 2D array of processing Elements (PEs) with a SIMD-like processing scheme implemented on a Kintex7 FPGA. NCs are SNN multiprocessors that support the execution of any neural algorithm based on spikes. A set of custom instructions was designed specifically for this architecture. The NCs support a hierarchical scheme of local and global spikes to mimic the brain structural configuration. Local spikes establish inter-neuronal connectivity within a single chip and the global ones allow inter-modular communication between different chips. The NCs have fixed hub neurons that process local and global spikes, thus allowing inter-modular and intra-modular connectivity. This definition of local and global spikes allows the development of multi-level hierarchical architectures inspired by the brain topologies, and offers excellent scalability. The spike propagation through the multi-chip network is supported by an Aurora / AER-SRT protocol stack. The Aurora protocol encapsulates and de-capsulates the packets transmitted through a high-speed serial link that communicates the platform, while the Synchronous Address Event representation (AER-SRT) protocol manages the data (address events) and controls packets that allow synchronization of the operation of the neural network. Each event encapsulates the address neuron that fires a spike as result of the neural algorithm execution. The definition of local and global synaptic topology is implemented using on-chip RAM blocks, which reduces the combinational logic requirements and, in addition to allowing the dynamic connectivity configuration, permits the development of evolutionary applications by supporting the on-line reconfiguration of both the neural algorithm or the neural and synaptic parameters. HEENS also supports axon programmable delays, which incorporates dynamic features to the network.Un sistema neuronal biológico consiste de millones de neuronas altamente integradas con múltiples funciones dinámicas operando en coordinación entre sí. Su organización estructural se caracteriza por contener agrupaciones altamente jerárquicas. Dichas agrupaciones se distinguen por conexiones localmente densas y globalmente dispersas comunicadas a través de pulsos transitorios (spikes) que viajan por el axón hasta la neurona destino. En el último siglo, aproximarse a la complejidad biológica del cortex mediante arquitecturas de hardware continúa siendo un desafío todavía inalcanzable. Esto se debe, no sólo al masivo procesamiento paralelo con soporte para la comunicación entre neuronas en redes de gran escala, sinó también a la necesidad de mecanismos que permitan la evolución de la red neuronal de forma eficiente. En este marco, esta tesis contribuye al desarrollo de una arquitectura denominada HEENS (Emulador de Hardware para Sistemas Neuronales Evolutivos, Hardware Emulator of Evolved Neural System) que soporta conectividad inter-chip con una topología de anillo entre un chip que actúa de master (MC) y uno o más Chips Neuromórficos (NCs). El MC está implementado en un dispositivo PSoC que integra un CPU ARM Dual Core junto con lógica programable. El ARM se encarga de configurar el anillo de comunicación y de ejecutar la aplicación de software que controla el envío de información de configuración del algoritmo y los parámetros neuronales a todos los NCs de la red. Además, el MC es el encargado de activar el modo de evolución de la red, así como de gestionar el envío de datos de reconfiguración a cualquiera de los nodos durante la ejecución. Cada NC a su vez, está compuesto por un arreglo 2D parametrizable de Elementos de Procesamiento (processing Elements, PEs) con un esquema de procesamiento tipo SIMD implementado sobre una FPGA Kintex7. Los NCs son multiprocesadores SNN que soportan la ejecución de cualquier algoritmo neuronal basado en spikes. Se cuenta con un set de instrucciones personalizadas diseñadas específicamente para esta arquitectura. Imitando la configuración estructural del cerebro los NC soportan un esquema jerárquico con spikes locales y globales. Los spikes locales establecen la conectividad inter-neuronal dentro de un mismo chip, y los globales la comunicación inter-modular entre diferentes chips. Los NC cuentan con neuronas fijas tipo hub que procesan spikes locales y globales que permiten la conectividad inter e intra modulos. La definición de spikes locales y globales permite desarrollar arquitecturas jerárquicas multi-nivel que se inspiran en las topologías del cerebro y ofrecen una escalabilidad excelente. La propagación de spikes a través de la red multi-chip es soportada por una pila de protocolos Aurora/AER-SRT. El protocolo Aurora encapsula y desencapsula los paquetes transmitidos a través del enlace serial de alta velocidad que comunica la plataforma. Mientras que el protocolo Síncrono de Representación de Eventos de Dirección (AER-SRT) gestiona los datos (eventos de dirección) y los paquetes de control que permiten sincronizar la operación de la red neuronal. Cada evento encapsula la dirección de la neurona que genera un spike como resultado del procesamiento del algoritmo neuronal. La definición de topología sináptica local y global es implementada usando bloques de memoria RAM on-chip, lo que reduce los requerimientos de lógica combinacional y, además de facilitar la configuración del conexionado sin modificar el hardware, permite el desarrollo de aplicaciones evolutivas al soportar la reconfiguración on-line tanto del algoritmo neuronal como de los parámetros neuronales y sinápticos. HEENS también admite retardos programables de axón, lo cual incorpora características dinámicas a la red.Postprint (published version

Hidalgo Costa, Juan Carlos - One of the best experts on this subject based on the ideXlab platform.

  • Um sistema de arquivos distribuido para computadores maciçamente paralelos virtuais
    [s.n.], 2018
    Co-Authors: Hidalgo Costa, Juan Carlos
    Abstract:

    Orientador: Marco Aurelio Amaral HenriquesDissertação (mestrado) - Universidade Estadual de Campinas, Faculdade de Engenharia Eletrica e de ComputaçãoResumo: Os computadores conectados pela Internet oferecem em conjunto um grande poder de cômputo, o qual deve continuar crescendo nos próximos anos. Eles podem ser vistos como um Computador Maciçamente Paralelo Virtual com memória distribuída que pode ser usado na resolução de problemas de grande porte. Existem várias propostas que visam tirar proveito da Internet como um computador virtual, utilizando Java como linguagem independente de plataforma. Entretanto, a maior parte destes projetos não trata, ou trata de forma superficial, a necessidade de se ter um sistema de arquivos que garanta a viabilidade e eficiência do processamento paralelo. Este trabalho propõe um sistema de arquivos distribuído baseado em grupos de servidores e voltado a plataformas para o processamento maciçamente paralelo na Internet. São propostos mecanismos para atender os requisitos fundamentais de sistemas deste tipo, eliminando as principais deficiências dos sistemas de arquivos convencionais. São apresentados e discutidos os resultados obtidos nos testes de uma implementação de referência do sistema de arquivos sobre JOIN, uma plataforma de processamento maciçamente paralelo virtual baseada na Internet. Esta implementação se mostrou confiável, robusta e aumentou a versatilidade da plataforma JOINAbstract: The total computing power offered by all computers connected to Internet is huge and increasing. Since the birth of the World Wide Web, new proposals have been made on how to take advantage of the enormous computing power represented by these computers. The proposals are based on the availability of WWW browsers to access resources distributed in the network and on the proliferation of Java as a platform independent language. The computers are grouped in a kind of massively parallel virtual computer, aimed at solving large problems. File systems are a necessary - and often forgotten - feature of these virtual machines. File systems for worldwide virtual machines should be efficient, highly available and fault tolerant. The virtual machines proposed so far either have no File System or provide very simple and inefficient solutions. This work proposes a distributed file system based on Groups of Servers, which was implemented and tested on top of JOIN, a platform for massively parallel processing on Internet. The test results showed that the file system is reliable, robust and makes the parallel platform more versatileMestradoMestre em Engenharia Elétric