The Experts below are selected from a list of 23910 Experts worldwide ranked by ideXlab platform

Grzegorz Pastuszak - One of the best experts on this subject based on the ideXlab platform.

  • generative multi symbol architecture of the binary arithmetic coder for uhdtv video encoders
    IEEE Transactions on Circuits and Systems I-regular Papers, 2020
    Co-Authors: Grzegorz Pastuszak
    Abstract:

    Binary arithmetic coding is a key part of recent video compression standards. Its throughput is limited by the inherent dependencies existing in the algorithm. As a consequence, a higher bin parallelism leads to lower Clock frequencies. This paper presents an architecture able to exceed limits existing in previous hardware implementations. The architecture exploits less probable symbols as starting points for long series of bins coded in one Clock Cycle. The evaluation for four possible cases of the range value allows its update in the pipeline before the delayed selection based on actual value. The adaptive division into series is proposed to make long series more frequent. To shorten critical paths, rMPS variables computed for symbols coded in the same Clock Cycle are first summed and then added to the low register. Up to 16 bypass-mode symbols can be processed in parallel with context-coded symbols in one Clock Cycle. The architecture is generative, i.e., its throughput can be scaled with resources without strict limits. For example, the binary arithmetic coder synthesized on 90nm TSMC technology which consumes 101.4k gates and operates at the 570 MHz has the average throughput of 13.4 bins per Clock Cycle for the high-quality H.265/HEVC compression.

  • architecture design and efficiency evaluation for the high throughput interpolation in the hevc encoder
    Digital Systems Design, 2013
    Co-Authors: Grzegorz Pastuszak, Maciej Trochimiuk
    Abstract:

    This paper describes the architecture of the high-throughput interpolator used in motion estimation and compensation of the HEVC encoder. The architecture reads eight input samples and produces 64 output samples at each Clock Cycle. Two versions are developed for FGPA and ASIC technologies. Synthesis results show that they can operate at 200 and 400 MHz when implemented in FPGA Aria II and TSMC 130 nm, respectively. This enables encoders to support HD resolutions. The paper also analyzes relation between compression efficiency and hardware complexity of the interpolation in HEVC and H.264/AVC.

  • a novel architecture of arithmetic coder in jpeg2000 based on parallel symbol encoding
    Parallel Computing in Electrical Engineering, 2004
    Co-Authors: Grzegorz Pastuszak
    Abstract:

    This paper presents a high-performance architecture of the context adaptive binary arithmetic coder (CABAC) for the embedded block-coding algorithm in JPEG 2000. The architecture has been developed in two variants to code two or three context-symbol pairs per Clock Cycle. The inverse multiple branch selection (IMBS) method is proposed to minimize critical paths, which originate from causally dependent operations. The designs have been implemented in VHDL and synthesized for FPGA devices. Simulation results show that the two- and three-symbol engines can process about 22 million samples at 77 and 53 MHz working frequency, respectively.

  • a high performance architecture of arithmetic coder in jpeg2000
    International Conference on Multimedia and Expo, 2004
    Co-Authors: Grzegorz Pastuszak
    Abstract:

    This paper presents a high-performance architecture of the arithmetic coder for the embedded block coding algorithm in JPEG2000. The dedicated pipeline architecture, enhanced by the inverse multiple branch selection (IMBS) method, is proposed to code two context-symbol pairs per Clock Cycle. The overall design was implemented in VHDL and synthesized for FPGA devices. Simulation results show that it can process about 17 million samples at 77 MHz working frequency

  • high efficient architectures of the context adaptive binary arithmetic coder for h 264 avc
    International workshopon systems signals and image processing ambient multimedia, 2004
    Co-Authors: Grzegorz Pastuszak
    Abstract:

    This paper presents architecture design of the context adaptive binary arithmetic coding (CABAC) in H.264/AVC. The pipelined architecture has been implemented in two variants targeting Altera FPGA Stratix devices. The first one process one symbol per Clock Cycle at the working frequency of 140 MHz. The second accepts two symbols per Clock Cycle at the frequency of 100 MHz. Evaluation results show that the engines meets throughput requirements of real-time television systems such as: PAL, NTSC, and even HDTV.

Young Ho Kwak - One of the best experts on this subject based on the ideXlab platform.

  • a 120 mhz 1 8 ghz cmos dll based Clock generator for dynamic frequency scaling
    IEEE Journal of Solid-state Circuits, 2006
    Co-Authors: Young Ho Kwak
    Abstract:

    A delay-locked loop (DLL)-based Clock generator for dynamic frequency scaling has been developed in a 0.35-mum CMOS technology. The proposed Clock generator can generate Clock signals ranging from 120 MHz to 1.8 GHz and change the frequency dynamically in a short time. If the Clock generator scales its output frequency dynamically by programming with the same last bit, it takes only one Clock Cycle to lock. In addition, the Clock generator inherits advantages of a DLL. The proposed DLL-based Clock generator occupies 0.07 mm2 and has a peak-to-peak jitter of plusmn6.6 ps at 1.3 GHz

  • a cmos dll based 120mhz to 1 8ghz Clock generator for dynamic frequency scaling
    International Solid-State Circuits Conference, 2005
    Co-Authors: Young Ho Kwak, Seokryung Yoon
    Abstract:

    A DLL-based Clock generator for dynamic frequency scaling is fabricated in a 0.35 /spl mu/m CMOS technology. It generates Clock signals ranging from 120MHz to 1.8GHz. The frequency can be dynamically changed. If the Clock generator scales its output frequency dynamically by programming with the same last bit, it takes only one Clock Cycle to lock. The proposed Clock generator has a jitter of /spl plusmn/6.6ps/sub pp/ at 1.3GHz.

Jun Koyama - One of the best experts on this subject based on the ideXlab platform.

  • Embedded SRAM and Cortex-M0 Core with Backup Circuits using a 60-nm Crystalline Oxide Semiconductor for Power Gating
    IEEE Micro, 2015
    Co-Authors: Hikaru Tamura, Takahiko Ishizu, Wataru Uesugi, Atsuo Isobe, Naoaki Tsutsui, Yutaka Okazaki, Yukio Maehashi, Yasutaka Suzuki, Kiyoshi Kato, Jun Koyama
    Abstract:

    Using data retention circuits that include crystalline oxide semiconductor transistors as backup circuits for power gating, a processor system can reduce standby leakage current significantly. This is effective in the Internet of Things (IoT) applications that require standby power reduction. The crystalline oxide semiconductor transistor can constitute a nonvolatile data retention circuit easily because it exhibits significantly lower off-state current than a silicon transistor and is highly compatible with a CMOS logic circuit. The backup circuit can achieve 2-Clock-Cycle data backup and 4-Clock-Cycle data restore; thus, the processor system can efficiently perform temporally fine-grained power gating and can achieve longer standby times. Furthermore, area overheads due to the backup circuits are kept very small because the crystalline oxide semiconductor transistors are stacked on silicon transistors.

  • Embedded SRAM and Cortex-M0 Core Using a 60-nm Crystalline Oxide Semiconductor
    IEEE Micro, 2014
    Co-Authors: Hikaru Tamura, Takahiko Ishizu, Wataru Uesugi, Atsuo Isobe, Naoaki Tsutsui, Yutaka Okazaki, Yukio Maehashi, Yasutaka Suzuki, Kiyoshi Kato, Jun Koyama
    Abstract:

    Using data retention circuits that include crystalline oxide semiconductor transistors as backup circuits for power gating, a processor system can reduce standby leakage current significantly. This is effective in the Internet of Things (IoT) applications that require standby power reduction. The crystalline oxide semiconductor transistor can constitute a nonvolatile data retention circuit easily because it exhibits significantly lower off-state current than a silicon transistor and is highly compatible with a CMOS logic circuit. The backup circuit can achieve 2-Clock-Cycle data backup and 4-Clock-Cycle data restore; thus, the processor system can efficiently perform temporally fine-grained power gating and can achieve longer standby times. Furthermore, area overheads due to the backup circuits are kept very small because the crystalline oxide semiconductor transistors are stacked on silicon transistors.

Florian Stelzer - One of the best experts on this subject based on the ideXlab platform.

  • performance boost of time delay reservoir computing by non resonant Clock Cycle
    Neural Networks, 2020
    Co-Authors: Florian Stelzer, Andre Rohm, Kathy Ludge, Serhiy Yanchuk
    Abstract:

    Abstract The time-delay-based reservoir computing setup has seen tremendous success in both experiment and simulation. It allows for the construction of large neuromorphic computing systems with only few components. However, until now the interplay of the different timescales has not been investigated thoroughly. In this manuscript, we investigate the effects of a mismatch between the time-delay and the Clock Cycle for a general model. Typically, these two time scales are considered to be equal. Here we show that the case of equal or resonant time-delay and Clock Cycle could be actively detrimental and leads to an increase of the approximation error of the reservoir. In particular, we can show that non-resonant ratios of these time scales have maximal memory capacities. We achieve this by translating the periodically driven delay-dynamical system into an equivalent network. Networks that originate from a system with resonant delay-times and Clock Cycles fail to utilize all of their degrees of freedom, which causes the degradation of their performance.

Veselin N. Ivanovic - One of the best experts on this subject based on the ideXlab platform.

  • superior execution time design of a space spatial frequency optimal filter for highly nonstationary 2d fm signal estimation
    IEEE Transactions on Circuits and Systems, 2018
    Co-Authors: Veselin N. Ivanovic, Nevena R Brnovic
    Abstract:

    Multiple-Clock-Cycle, signal adaptive, and fully pipelined hardware design of the optimal (Wiener) space/spatial-frequency (S/SF) filter is developed in this paper. All implementation and verification details, as well as the extensive comparative analysis, are provided. The developed solution optimizes critical design performances related to the hardware complexity, in line with multiple-Clock-Cycle nature. Variable (signal adaptive) number of Clock Cycles, taken within the execution in different S/SF points, provides this solution to retain the optimized time requirements, as well as high resolution, selectivity, and estimation quality of the corresponding recently proposed signal adaptive filtering solution. However, as the major contribution, the fully pipelined implementation enables the developed design to additionally improve the time required for execution. The achieved improvement corresponds to a Clock Cycle per each S/SF point performed within the estimation that results in the significant comparative improvement in execution time of up to 50% in terms of S/SF points lying outside the local frequency of the estimated 2D frequency-modulated signal. The implementation is tested on a highly nonstationary multicomponent signal and is verified by a field programmable gate array circuit design.

  • signal adaptive hardware implementation of a system for highly nonstationary two dimensional fm signal estimation
    Aeu-international Journal of Electronics and Communications, 2015
    Co-Authors: Veselin N. Ivanovic, Nevena Radovic
    Abstract:

    Abstract Signal adaptive, multiple-Clock-Cycle hardware implementation (MCI) of an optimal (Wiener) filter for highly nonstationary two-dimensional (2D) FM signals estimation is developed here. It uses results of the space/spatial-frequency (S/SF) analysis in real-time processing of nonstationary 2D signals and is based on the correspondence of the filter's region of support to the local frequency (LF) of the filtered 2D signal and on the S/SF analysis-based LF estimation. The MCI approach helps the proposed design to minimize Clock Cycle time and to optimize critical design performances related to the hardware complexity, making it a suitable system for real-time and on-a-chip implementation. However, the major advantage of the proposed design is the ability to take variable (signal adaptive) number of Clock Cycles in different S/SF points within the execution. This property helps the design to optimize the execution time (the main drawback of the classical MCI approaches in comparison to the single-Clock-Cycle ones), but also to provide the highest quality LF estimation, high S/SF resolution, and a very efficient filtering of nonstationary 2D FM signals. In this way, it is qualified as an optimal solution for wide range of practical implementations. The implementation is verified by a field-programmable gate array (FPGA) circuit design.

  • Multiple-Clock-Cycle architecture for the VLSI design of a system for time-frequency analysis
    EURASIP Journal on Advances in Signal Processing, 2006
    Co-Authors: Veselin N. Ivanovic, Radovan Stojanovic, Ljubisa Stankovic
    Abstract:

    Multiple-Clock-Cycle implementation (MCI) of a flexible system for time-frequency (TF) signal analysis is presented. Some very important and frequently used time-frequency distributions (TFDs) can be realized by using the proposed architecture: (i) the spectrogram (SPEC) and the pseudo-Wigner distribution (WD), as the oldest and the most important tools used in TF signal analysis; (ii) the S-method (SM) with various convolution window widths, as intensively used reduced interference TFD. This architecture is based on the short-time Fourier transformation (STFT) realization in the first Clock Cycle. It allows the mentioned TFDs to take different numbers of Clock Cycles and to share functional units within their execution. These abilities represent the major advantages of multiCycle design and they help reduce both hardware complexity and cost. The designed hardware is suitable for a wide range of applications, because it allows sharing in simultaneous realizations of the higher-order TFDs. Also, it can be accommodated for the implementation of the SM with signal-dependent convolution window width. In order to verify the results on real devices, proposed architecture has been implemented with a field programmable gate array (FPGA) chips. Also, at the implementation (silicon) level, it has been compared with the single-Cycle implementation (SCI) architecture.