The Experts below are selected from a list of 291 Experts worldwide ranked by ideXlab platform
Shih-lun Chen - One of the best experts on this subject based on the ideXlab platform.
-
VLSI Implementation of a Cost-Efficient Micro Control Unit With an Asymmetric Encryption for Wireless Body Sensor Networks
IEEE Access, 2017Co-Authors: Shih-lun Chen, Min-chun TuanAbstract:This paper presents a very large-scale integration (VLSI) circuit design of a micro control unit (MCU) for wireless body sensor networks (WBSNs) in cost-intention. The proposed MCU design consists of an asynchronous interface, a multisensor controller, a Register Bank, a hardware-shared filter, a lossless compressor, an encryption encoder, an error correct coding (ECC) circuit, a universal asynchronous receiver/transmitter interface, a power management, and a QRS complex detector. A hardware-sharing technique was added to reduce the silicon area of a hardware-shared filter and provided functions in terms of high-pass, low-pass, and band-pass filters according to the uses of various body signals. The QRS complex detector was designed for calculating QRS information of the ECG signals. In addition, the QRS information is helpful to obtain the heart beats. The lossless compressor consists of an adaptive trending predictor and an extensible hybrid entropy encoder, which provides various methods to compress the different characteristics of body signals adaptively. Furthermore, an encryption encoder based on an asymmetric cryptography technique was designed to protect the private physical information during wireless transmission. The proposed MCU design in this paper contained 7.61k gate counts and consumed 1.33 mW when operating at 200 MHz by using a 90-nm CMOS process. Compared with previous designs, this paper has the benefits of increasing the average compression rate by over 12% in ECG signal, providing body signals analysis, and enhancing security of the WBSNs.
-
An efficient micro control unit VLSI design for wearable electronics and sensor networks
2016 Pan Pacific Microelectronics Symposium (Pan Pacific), 2016Co-Authors: Min-chun Tuan, Shih-lun ChenAbstract:In this paper, an efficient micro control unit (MCU) design is proposed for wearable electronics and wireless sensor networks. It consists of an asynchronous interface, a Register Bank, a reconfigurable filter, a lossless data encoder, an encryption encoder, an error correct coding (ECC) encoder, a power management, a resolution controller and a multi-sensor controller. The asynchronous interface is used to exchange data in different frequency. The reconfigurable filter cooperates with the Register Bank to provide functions of high-pass, low-pass and band-pass filters according to various signals. The lossless data encoder consists of an adaptive predictor and a hybrid entropy encoder, which can use different methods to compress different characteristics of signals adaptively. The encryption and ECC encoders are added to improve the security of data and transmission, respectively. For, long-term usage, the power management is developed for reducing the power consumption of the whole system. The resolution and multi-sensor controllers are designed to adjust the resolution and select different sensors, respectively, according to the characteristic of various signals. In addition, the proposed wearable electronics and wireless sensor network systems include image sensors and processor for the applications of special education, autism children assistance and healthcare. The proposed MCU design was synthesized by a 0.18-μm CMOS process and it can operate at 100-MHz processing rate. This design contains 4.29-K gate counts and its core area is 43k-μm2. Compared with previous designs, this design achieved higher performance, higher security, higher reliability, more functions, more flexibility, higher compatibility and lower cost than previous designs. It is suitable for developing wearable electronics and wireless sensor network systems.
-
VLSI Implementation of a Cost-Efficient Near-Lossless CFA Image Compressor for Wireless Capsule Endoscopy
IEEE Access, 2016Co-Authors: Shih-lun Chen, Chia-wei Shen, Min-chun TuanAbstract:In this paper, a novel near-lossless color filter array (CFA) image compression algorithm based on JPEG-LS is proposed for VLSI implementation. It consists of a pixel restoration, a prediction, a run mode, and entropy coding modules. According to the information of the previous research, a context table and row memory consumed more than 81% hardware cost in a JPEG-LS encoder design. Hence, in this paper, a novel context-free and near-lossless image compression algorithm is presented. Since removing the context model causes decreasing of the compression performance, a novel prediction, run mode, and modified Golomb-Rice coding techniques were used to improve the compression efficiency. The VLSI architecture of the proposed image compressor consists of a Register Bank, a pixel restoration module, a predictor, a run mode module, and an entropy encoder. A pipeline technique was used to improve the performance of this. It contains only 10.9k gate count, and the core area is 30625 μm2 , synthesized by using a 90-nm CMOS process. Compared with the previous JPEG-LS designs, this paper reduces the gate counts by 44.1% and 41.7%, respectively, for five standard and eight endoscopy testing images in CFA format. It also improves the average PSNR values by 0.96 and 0.43 dB, respectively, for the same test images.
-
Ultra-low-cost colour demosaicking VLSI design for real-time video applications
Electronics Letters, 2014Co-Authors: Shih-lun Chen, Huan-rui ChangAbstract:A novel low-complexity and high-quality colour demosaicking algorithm is proposed for very large-scale integration (VLSI) implementation for real-time video applications. It consists of a boundary detector, a boundary mirror model and five green and red-blue colour interpolation models. Two of the five interpolation models can be selected adaptively according to boundary and position information. In addition, a boundary mirror machine and identical direction technique were used to improve the qualities of the reconstructed images. To reduce the hardware cost, memory requirement and power consumption, a hardware-sharing technique and Register Bank design were used to realise the proposed algorithm. The VLSI architecture of this work contains only 2.9 K gate counts and its core area is 35 966 μm2 synthesised by a 0.18 μm CMOS process. The synthesised results show that this design performs an operating frequency of 100 MHz processing rate by consuming only 1.83 mW. Compared with the previous low-complexity designs, this work not only reduces at least 48.2% of gate counts and 96.7% of power consumption but also improves the average CPSNR quality by more than 0.78 dB.
-
A reconfigurable control system design for wireless body sensor network
2014 Asia-Pacific Conference on Computer Aided System Engineering (APCASE), 2014Co-Authors: Shih-lun ChenAbstract:In this paper, a five-level reconfigurable control system is proposed for wireless body sensor network (WBSN). In this system, the bi-direction communication channel is separated into an instruction path and a data path. The instruction path consists of two different kinds of instructions. The first one is command instruction which provides parameters setting for the sensor nodes. The other one is reconfigurable instruction which serves control signals for executing the applications with the characteristics of loops or repetitions. In addition, a low-cost, low-power and high performance micro control unit (MCU) core is designed for this system. It consists of an asynchronous interface, a Register Bank, a reconfigurable filter, a slop-feature forecast, a lossless data encoder, an error correct coding (ECC) encoder, a UART interface, and a power management (PWM). Moreover, a multi-sensor controller and a reconfigurable memory were added to the proposed MCU design for various body sensors controlling and the loops or repetitions executing. This work achieves reductions of 83% instructions by comparing with a multi-sensor system without reconfigurable control system design. It gets both benefits of reducing the loading in software systems and saving power consumptions when transmits instructions. The MCU core for the reconfigurable control system was implemented by 0.18 μm CMOS process technology. The core area is 202,484 μm2 and the power is 2.9 mW operating at 100 MHz frequency.
Philip Sweany - One of the best experts on this subject based on the ideXlab platform.
-
A Code Generation Framework for VLIW Architectures with Partitioned Register Banks
2020Co-Authors: Saurabh Jang, Steve Carr, Philip Sweany, Darla KurasAbstract:Modern computers are taking increasing advantage of the instruction-level parallelism (ILP) available in programs with advances in both machine and compiler design. Unfortunately, large amounts of ILP hardware and aggressive instruction scheduling techniques put large demands on a machine’s Register resources. With large amounts of ILP, it becomes difficult to maintain a single monolithic Register Bank and a high clock rate. The number of ports required for such a Register Bank severely hampers access time [2, 8]. To provide support for large amounts of ILP while retaining a high clock rate, Registers can be partitioned among several different Register Banks. Each Bank is directly accessible by only a subset of the functional units with explicit inter-Bank copies required to move data between Banks. Therefore, a compiler must deal not only with achieving maximal parallelism via aggressive scheduling, but also with data placement to limit inter-Bank copies. This paper describes our approach to partitioning values among available Register Banks. Our method provides flexibility for our retargetable compiler by representing machine dependent features as node and edge weights and by remaining independent of scheduling and Register allocation methods. Preliminary experimentation with our framework has shown a degradation in execution performance of 33% on average when compared to an unrealizable monolithicRegister-Bank VLIW architecture with the same level of ILP. This compares very favorably with other approaches to the same problem. This research was supported by a grant from Texas Instruments.
-
IEEE PACT - Optimizing loop performance for clustered VLIW architectures
Proceedings.International Conference on Parallel Architectures and Compilation Techniques, 2002Co-Authors: Yi Qian, Steve Carr, Philip SweanyAbstract:Modern embedded systems often require high degrees of instruction-level parallelism (ILP) within strict constraints on power consumption and chip cost. Unfortunately, a high-performance embedded processor with high ILP generally puts large demands on Register resources, making it difficult to maintain a single, multi-ported Register Bank. To address this problem, some architectures, e.g. the Texas Instruments TMS320C6x, partition the Register Bank into multiple Banks that are each directly connected only to a sub-set of functional units. These functional unit/Register Bank groups are called clusters.Clustered architectures require that either copy operations or delay slots be inserted when an operation accesses data stored on a different cluster. In order to generate excellent code for such architectures, the compiler must not only spread the computation across clusters to achieve maximum parallelism, but also must limit the effects of intercluster data transfers.Loop unrolling and unroll-and-jam enhance the parallelism in loops to help limit the effects of intercluster data transfers. In this paper, we describe an accurate metric for predicting the intercluster communication cost of a loop and present an integer-optimization problem that can be used to guide the application of unroll-and-jam and loop unrolling considering the effects of both ILP and intercluster data transfers. Our method achieves a harmonic mean speedup of 1.4 - 1.7 on software pipelined loops for both a simulated architecture and the TI TMS320C64x.
-
loop fusion for clustered vliw architectures
Languages Compilers and Tools for Embedded Systems, 2002Co-Authors: Yi Qian, Steve Carr, Philip SweanyAbstract:Embedded systems require maximum performance from a processor within significant constraints in power consumption and chip cost. Using software pipelining, high-performance digital signal processors can often exploit considerable instruction-level parallelism (ILP), and thus significantly improve performance. However, software pipelining, in some instances, hinders the goals of low power consumption and low chip cost. Specifically, the Registers required by a software pipelined loop may exceed the size of the physical Register set.The Register pressure problem incurred by software pipelining makes it difficult to build a high-performance embedded processor with a single, multi-ported Register Bank with enough Registers to support high levels of ILP while maintaining clock speed and limiting power consumption. The large number of ports required to support a single Register Bank severely hampers access time. The port requirement for a Register Bank can be reduced via hardware by partitioning the Register Bank into multiple Banks connected to disjoint subsets of functional units, called clusters. Since a functional unit is not directly connected to all Register Banks, wasted energy and resources can result due to delays incurred when accessing "non-local" Registers.The overhead due to partitioning of the Register set can be ameliorated by using high-level compiler loop optimization techniques such as unrolling, unroll-and-jam and fusion. High-level loop optimizations spread data-independent parallelism across clusters that may not require "non-local" Register accesses and can provide work to hide the latency of any such Register accesses that are needed.In this paper, we examine the effects of loop fusion on DSP loops run on four simulated, clustered VLIW architectures and the Texas Instruments TMS320C64x. Our experiments show a 1.3 -- 2 harmonic mean speedup.
-
LCTES-SCOPES - Loop fusion for clustered VLIW architectures
Proceedings of the joint conference on Languages compilers and tools for embedded systems software and compilers for embedded systems - LCTES SCOPES ', 2002Co-Authors: Yi Qian, Steve Carr, Philip SweanyAbstract:Embedded systems require maximum performance from a processor within significant constraints in power consumption and chip cost. Using software pipelining, high-performance digital signal processors can often exploit considerable instruction-level parallelism (ILP), and thus significantly improve performance. However, software pipelining, in some instances, hinders the goals of low power consumption and low chip cost. Specifically, the Registers required by a software pipelined loop may exceed the size of the physical Register set.The Register pressure problem incurred by software pipelining makes it difficult to build a high-performance embedded processor with a single, multi-ported Register Bank with enough Registers to support high levels of ILP while maintaining clock speed and limiting power consumption. The large number of ports required to support a single Register Bank severely hampers access time. The port requirement for a Register Bank can be reduced via hardware by partitioning the Register Bank into multiple Banks connected to disjoint subsets of functional units, called clusters. Since a functional unit is not directly connected to all Register Banks, wasted energy and resources can result due to delays incurred when accessing "non-local" Registers.The overhead due to partitioning of the Register set can be ameliorated by using high-level compiler loop optimization techniques such as unrolling, unroll-and-jam and fusion. High-level loop optimizations spread data-independent parallelism across clusters that may not require "non-local" Register accesses and can provide work to hide the latency of any such Register accesses that are needed.In this paper, we examine the effects of loop fusion on DSP loops run on four simulated, clustered VLIW architectures and the Texas Instruments TMS320C64x. Our experiments show a 1.3 -- 2 harmonic mean speedup.
-
Optimizing loop performance for clustered VLIW architectures
Proceedings.International Conference on Parallel Architectures and Compilation Techniques, 2002Co-Authors: Yi Qian, Steve Carr, Philip SweanyAbstract:Modem embedded systems often require high degrees of instruction-level parallelism (ILP) within strict constraints on power consumption and chip cost. Unfortunately, a high-performance embedded processor with high ILP generally puts large demands on Register resources, making it difficult to maintain a single, multi-ported Register Bank. To address this problem, some architectures, e.g. the Texas Instruments TMS320C6x, partition the Register Bank into multiple Banks that are each directly connected only to a subset of functional units. These functional unit/Register Bank groups are called clusters. Clustered architectures require that either copy operations or delay slots be inserted when an operation accesses data stored on a different cluster In order to generate excellent code for such architectures, the compiler must not only spread the computation across clusters to achieve maximum parallelism, but also must limit the effects of intercluster data transfers. Loop unrolling and unroll-and-jam enhance the parallelism in loops to help limit the effects of intercluster data transfers. In this paper we describe an accurate metric for predicting the intercluster communication cost of a loop and present an integer-optimization problem that can be used to guide the application of unroll-and-jam and loop unrolling considering the effects of both ILP and intercluster data transfers. Our method achieves a harmonic mean speedup of 1.4-1.7 on software pipelined loops for both a simulated architecture and the TI TMS320C64x.
Santosh Pande - One of the best experts on this subject based on the ideXlab platform.
-
IEEE PACT - Resolving Register Bank conflicts for a network processor
Oceans 2002 Conference and Exhibition. Conference Proceedings (Cat. No.02CH37362), 2003Co-Authors: X. Zhuang, Santosh PandeAbstract:This paper discusses a Register Bank assignment problem for a popular network processor - Intel's IXP. Due to limited data paths, the network processor has a restriction that thesource operands of most ALU instructions must be resident in two different Banks. This results in higher Register pressure and puts additional burden on the Register allocator. The current vendor-provided Register allocator leaves the problem to users, leading to poor compilation interface and low quality code.This paper presents three different approaches for performing Register allocation and Bank assignment. Bank assignment can be performed before Register allocation, can be performed after Register allocation or it could be combined with the Register allocation. We propose a structure called Register conflict graph (RCG) to capture the dual-Bank constraints. To further improve the effectiveness of the algorithm, we also propose some enabling transformations.Our results show the phase ordering of first doing Register allocation and then assigning Banks can reduce the number of spills with affordable costs of additional instructions.
-
resolving Register Bank conflicts for a network processor
International Conference on Parallel Architectures and Compilation Techniques, 2003Co-Authors: X. Zhuang, Santosh PandeAbstract:This paper discusses a Register Bank assignment problem for a popular network processor - Intel's IXP. Due to limited data paths, the network processor has a restriction that thesource operands of most ALU instructions must be resident in two different Banks. This results in higher Register pressure and puts additional burden on the Register allocator. The current vendor-provided Register allocator leaves the problem to users, leading to poor compilation interface and low quality code.This paper presents three different approaches for performing Register allocation and Bank assignment. Bank assignment can be performed before Register allocation, can be performed after Register allocation or it could be combined with the Register allocation. We propose a structure called Register conflict graph (RCG) to capture the dual-Bank constraints. To further improve the effectiveness of the algorithm, we also propose some enabling transformations.Our results show the phase ordering of first doing Register allocation and then assigning Banks can reduce the number of spills with affordable costs of additional instructions.
-
Resolving Register Bank conflicts for a network processor
2003 12th International Conference on Parallel Architectures and Compilation Techniques, 2003Co-Authors: X. Zhuang, Santosh PandeAbstract:We discuss a Register Bank assignment problem for a popular network processor-Intel's IXP. Due to limited data paths, the network processor has a restriction that the source operands of most ALU instructions must be resident in two different Banks. This results in higher Register pressure and puts additional burden on the Register allocator. The current vendor-provided Register allocator leaves the problem to users, leading to poor compilation interface and low quality code. We present three different approaches for performing Register allocation and Bank assignment. Bank assignment can be performed before Register allocation, can be performed after Register allocation or it could be combined with the Register allocation. We propose a structure called Register conflict graph (RCG) to capture the dual-Bank constraints. To further improve the effectiveness of the algorithm, we also propose some enabling transformations. Our results show the phase ordering of first doing Register allocation and then assigning Banks can reduce the number of spills with affordable costs of additional instructions.
Min-chun Tuan - One of the best experts on this subject based on the ideXlab platform.
-
VLSI Implementation of a Cost-Efficient Micro Control Unit With an Asymmetric Encryption for Wireless Body Sensor Networks
IEEE Access, 2017Co-Authors: Shih-lun Chen, Min-chun TuanAbstract:This paper presents a very large-scale integration (VLSI) circuit design of a micro control unit (MCU) for wireless body sensor networks (WBSNs) in cost-intention. The proposed MCU design consists of an asynchronous interface, a multisensor controller, a Register Bank, a hardware-shared filter, a lossless compressor, an encryption encoder, an error correct coding (ECC) circuit, a universal asynchronous receiver/transmitter interface, a power management, and a QRS complex detector. A hardware-sharing technique was added to reduce the silicon area of a hardware-shared filter and provided functions in terms of high-pass, low-pass, and band-pass filters according to the uses of various body signals. The QRS complex detector was designed for calculating QRS information of the ECG signals. In addition, the QRS information is helpful to obtain the heart beats. The lossless compressor consists of an adaptive trending predictor and an extensible hybrid entropy encoder, which provides various methods to compress the different characteristics of body signals adaptively. Furthermore, an encryption encoder based on an asymmetric cryptography technique was designed to protect the private physical information during wireless transmission. The proposed MCU design in this paper contained 7.61k gate counts and consumed 1.33 mW when operating at 200 MHz by using a 90-nm CMOS process. Compared with previous designs, this paper has the benefits of increasing the average compression rate by over 12% in ECG signal, providing body signals analysis, and enhancing security of the WBSNs.
-
An efficient micro control unit VLSI design for wearable electronics and sensor networks
2016 Pan Pacific Microelectronics Symposium (Pan Pacific), 2016Co-Authors: Min-chun Tuan, Shih-lun ChenAbstract:In this paper, an efficient micro control unit (MCU) design is proposed for wearable electronics and wireless sensor networks. It consists of an asynchronous interface, a Register Bank, a reconfigurable filter, a lossless data encoder, an encryption encoder, an error correct coding (ECC) encoder, a power management, a resolution controller and a multi-sensor controller. The asynchronous interface is used to exchange data in different frequency. The reconfigurable filter cooperates with the Register Bank to provide functions of high-pass, low-pass and band-pass filters according to various signals. The lossless data encoder consists of an adaptive predictor and a hybrid entropy encoder, which can use different methods to compress different characteristics of signals adaptively. The encryption and ECC encoders are added to improve the security of data and transmission, respectively. For, long-term usage, the power management is developed for reducing the power consumption of the whole system. The resolution and multi-sensor controllers are designed to adjust the resolution and select different sensors, respectively, according to the characteristic of various signals. In addition, the proposed wearable electronics and wireless sensor network systems include image sensors and processor for the applications of special education, autism children assistance and healthcare. The proposed MCU design was synthesized by a 0.18-μm CMOS process and it can operate at 100-MHz processing rate. This design contains 4.29-K gate counts and its core area is 43k-μm2. Compared with previous designs, this design achieved higher performance, higher security, higher reliability, more functions, more flexibility, higher compatibility and lower cost than previous designs. It is suitable for developing wearable electronics and wireless sensor network systems.
-
VLSI Implementation of a Cost-Efficient Near-Lossless CFA Image Compressor for Wireless Capsule Endoscopy
IEEE Access, 2016Co-Authors: Shih-lun Chen, Chia-wei Shen, Min-chun TuanAbstract:In this paper, a novel near-lossless color filter array (CFA) image compression algorithm based on JPEG-LS is proposed for VLSI implementation. It consists of a pixel restoration, a prediction, a run mode, and entropy coding modules. According to the information of the previous research, a context table and row memory consumed more than 81% hardware cost in a JPEG-LS encoder design. Hence, in this paper, a novel context-free and near-lossless image compression algorithm is presented. Since removing the context model causes decreasing of the compression performance, a novel prediction, run mode, and modified Golomb-Rice coding techniques were used to improve the compression efficiency. The VLSI architecture of the proposed image compressor consists of a Register Bank, a pixel restoration module, a predictor, a run mode module, and an entropy encoder. A pipeline technique was used to improve the performance of this. It contains only 10.9k gate count, and the core area is 30625 μm2 , synthesized by using a 90-nm CMOS process. Compared with the previous JPEG-LS designs, this paper reduces the gate counts by 44.1% and 41.7%, respectively, for five standard and eight endoscopy testing images in CFA format. It also improves the average PSNR values by 0.96 and 0.43 dB, respectively, for the same test images.
X. Zhuang - One of the best experts on this subject based on the ideXlab platform.
-
IEEE PACT - Resolving Register Bank conflicts for a network processor
Oceans 2002 Conference and Exhibition. Conference Proceedings (Cat. No.02CH37362), 2003Co-Authors: X. Zhuang, Santosh PandeAbstract:This paper discusses a Register Bank assignment problem for a popular network processor - Intel's IXP. Due to limited data paths, the network processor has a restriction that thesource operands of most ALU instructions must be resident in two different Banks. This results in higher Register pressure and puts additional burden on the Register allocator. The current vendor-provided Register allocator leaves the problem to users, leading to poor compilation interface and low quality code.This paper presents three different approaches for performing Register allocation and Bank assignment. Bank assignment can be performed before Register allocation, can be performed after Register allocation or it could be combined with the Register allocation. We propose a structure called Register conflict graph (RCG) to capture the dual-Bank constraints. To further improve the effectiveness of the algorithm, we also propose some enabling transformations.Our results show the phase ordering of first doing Register allocation and then assigning Banks can reduce the number of spills with affordable costs of additional instructions.
-
resolving Register Bank conflicts for a network processor
International Conference on Parallel Architectures and Compilation Techniques, 2003Co-Authors: X. Zhuang, Santosh PandeAbstract:This paper discusses a Register Bank assignment problem for a popular network processor - Intel's IXP. Due to limited data paths, the network processor has a restriction that thesource operands of most ALU instructions must be resident in two different Banks. This results in higher Register pressure and puts additional burden on the Register allocator. The current vendor-provided Register allocator leaves the problem to users, leading to poor compilation interface and low quality code.This paper presents three different approaches for performing Register allocation and Bank assignment. Bank assignment can be performed before Register allocation, can be performed after Register allocation or it could be combined with the Register allocation. We propose a structure called Register conflict graph (RCG) to capture the dual-Bank constraints. To further improve the effectiveness of the algorithm, we also propose some enabling transformations.Our results show the phase ordering of first doing Register allocation and then assigning Banks can reduce the number of spills with affordable costs of additional instructions.
-
Resolving Register Bank conflicts for a network processor
2003 12th International Conference on Parallel Architectures and Compilation Techniques, 2003Co-Authors: X. Zhuang, Santosh PandeAbstract:We discuss a Register Bank assignment problem for a popular network processor-Intel's IXP. Due to limited data paths, the network processor has a restriction that the source operands of most ALU instructions must be resident in two different Banks. This results in higher Register pressure and puts additional burden on the Register allocator. The current vendor-provided Register allocator leaves the problem to users, leading to poor compilation interface and low quality code. We present three different approaches for performing Register allocation and Bank assignment. Bank assignment can be performed before Register allocation, can be performed after Register allocation or it could be combined with the Register allocation. We propose a structure called Register conflict graph (RCG) to capture the dual-Bank constraints. To further improve the effectiveness of the algorithm, we also propose some enabling transformations. Our results show the phase ordering of first doing Register allocation and then assigning Banks can reduce the number of spills with affordable costs of additional instructions.