Pub Date : 2026-08-01Epub Date: 2026-03-22DOI: 10.1109/TVLSI.2026.3693724
Liang Wen;Lixun Wang;Jiangong Wang;Yuejun Zhang
This brief proposes an average-6.5T twin cell for a deep sub-micrometer 64 kb SRAM, which utilizes two identical asymmetric single-ended (SE) 6T cells in a column with a shared read assist device to improve read margin and write ability. It enables the SRAM to achieve read-disturb-free, near/sub-threshold operation and compact array layout, resulting in area and energy efficiencies. The average-6.5T SRAM test chip is fabricated using a 65 nm CMOS logic process. Its cell area shows only 5.6% overhead compared to the standard 6T cell, and is smaller than that of other low-voltage SRAMs. Measured full read and write functionality is performed with VDD down to 0.39 V, which is lower than that of standard 6T and 8T SRAMs. In addition, its minimum energy point of 6.3 pJ is obtained at 0.48 V.
{"title":"Average-6.5T Near-Threshold Twin Cell With Shared Read Assist for IoT Applications","authors":"Liang Wen;Lixun Wang;Jiangong Wang;Yuejun Zhang","doi":"10.1109/TVLSI.2026.3693724","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3693724","url":null,"abstract":"This brief proposes an average-6.5T twin cell for a deep sub-micrometer 64 kb SRAM, which utilizes two identical asymmetric single-ended (SE) 6T cells in a column with a shared read assist device to improve read margin and write ability. It enables the SRAM to achieve read-disturb-free, near/sub-threshold operation and compact array layout, resulting in area and energy efficiencies. The average-6.5T SRAM test chip is fabricated using a 65 nm CMOS logic process. Its cell area shows only 5.6% overhead compared to the standard 6T cell, and is smaller than that of other low-voltage SRAMs. Measured full read and write functionality is performed with VDD down to 0.39 V, which is lower than that of standard 6T and 8T SRAMs. In addition, its minimum energy point of 6.3 pJ is obtained at 0.48 V.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2676-2680"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148627504","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-01Epub Date: 2026-03-27DOI: 10.1109/TVLSI.2026.3695877
Hao Ding;Xiangyong Wang;Yang Xu;Junyan Qian
Three-dimensional (3-D) VLSI processor arrays offer high integration density and scalability, but their increasing system and interconnect complexity pose significant challenges to reliable reconfiguration under permanent hardware failures. Most existing reconfiguration approaches primarily target processing element (PE) faults and provide limited support for interconnect-related failures, such as switch and link faults, which can severely constrain feasible reconfiguration solutions and the achievable size of fault-free subarrays at the system level. This article investigates system-level fault-tolerant reconfiguration of reconfigurable 3-D VLSI processor arrays under multicomponent failures, including PE, switch, and link faults. We first analyze the system-level impact of different fault types and show that interconnect-level faults, especially switch failures, impose more pronounced constraints on reconfiguration effectiveness than PE faults. Based on this observation, a general plane-exclusion mechanism is introduced that can be integrated with representative PE-only reconfiguration schemes to enhance tolerance to switch and link failures. Furthermore, a fault-transformation preprocessing mechanism is developed to model switch failures as equivalent link disconnections, enabling unified system-level handling of heterogeneous faults and improving structural robustness. To mitigate excessive interconnect overhead introduced during reconfiguration, a long-interconnect optimization strategy is incorporated. Experimental results demonstrate that the proposed techniques significantly improve reconfiguration capability and scalability. In particular, the fault-transformation mechanism achieves over 90% of the theoretical upper bound in PE utilization under high switch fault densities, nearly doubling PE utilization compared with baseline methods, while the interconnect optimization reduces the number of long interconnects by more than 10% on average across all evaluated cases.
{"title":"System-Level Fault-Tolerant Reconfiguration of 3-D VLSI Processor Arrays Under Multicomponent Failures","authors":"Hao Ding;Xiangyong Wang;Yang Xu;Junyan Qian","doi":"10.1109/TVLSI.2026.3695877","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3695877","url":null,"abstract":"Three-dimensional (3-D) VLSI processor arrays offer high integration density and scalability, but their increasing system and interconnect complexity pose significant challenges to reliable reconfiguration under permanent hardware failures. Most existing reconfiguration approaches primarily target processing element (PE) faults and provide limited support for interconnect-related failures, such as switch and link faults, which can severely constrain feasible reconfiguration solutions and the achievable size of fault-free subarrays at the system level. This article investigates system-level fault-tolerant reconfiguration of reconfigurable 3-D VLSI processor arrays under multicomponent failures, including PE, switch, and link faults. We first analyze the system-level impact of different fault types and show that interconnect-level faults, especially switch failures, impose more pronounced constraints on reconfiguration effectiveness than PE faults. Based on this observation, a general plane-exclusion mechanism is introduced that can be integrated with representative PE-only reconfiguration schemes to enhance tolerance to switch and link failures. Furthermore, a fault-transformation preprocessing mechanism is developed to model switch failures as equivalent link disconnections, enabling unified system-level handling of heterogeneous faults and improving structural robustness. To mitigate excessive interconnect overhead introduced during reconfiguration, a long-interconnect optimization strategy is incorporated. Experimental results demonstrate that the proposed techniques significantly improve reconfiguration capability and scalability. In particular, the fault-transformation mechanism achieves over 90% of the theoretical upper bound in PE utilization under high switch fault densities, nearly doubling PE utilization compared with baseline methods, while the interconnect optimization reduces the number of long interconnects by more than 10% on average across all evaluated cases.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2607-2620"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148628156","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-01Epub Date: 2026-06-24DOI: 10.1109/TVLSI.2026.3694263
Jiangtao Cui;Xinyu Shao;Sheng Zhang
Rapid growth of intelligent vision applications demands high image quality, low latency, and energy efficiency. Traditional pipelines integrating a neural processing unit (NPU) with an image signal processor (ISP) suffer from CPU scheduling delays, memory bandwidth bottlenecks, and limited precision flexibility. To address these challenges, we propose ENLIVEN, a holistically optimized NPU–ISP architecture that codesign system, dataflow, and circuit layers. At the system level, ENLIVEN enables CPU-free, event-driven scheduling to eliminate control overhead and improve processing efficiency. At the dataflow level, a line-to-tile transformation and pipelined multicore execution enhance throughput and memory utilization while ensuring regular computation. At the circuit level, a mixed-precision NPU architecture further improves computational efficiency and image quality. Fabricated in 12-nm technology, ENLIVEN achieves UHD 120-f/s throughput with a peak energy efficiency of 12.08 TOPS/W and a computational density of 2.29 TOPS/mm2 while delivering significantly improved noise suppression compared with conventional pipelines.
{"title":"ENLIVEN: End-to-End NPU--ISP Codesign for Low-Latency and Hardware-Optimized Visual Processing","authors":"Jiangtao Cui;Xinyu Shao;Sheng Zhang","doi":"10.1109/TVLSI.2026.3694263","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3694263","url":null,"abstract":"Rapid growth of intelligent vision applications demands high image quality, low latency, and energy efficiency. Traditional pipelines integrating a neural processing unit (NPU) with an image signal processor (ISP) suffer from CPU scheduling delays, memory bandwidth bottlenecks, and limited precision flexibility. To address these challenges, we propose ENLIVEN, a holistically optimized NPU–ISP architecture that codesign system, dataflow, and circuit layers. At the system level, ENLIVEN enables CPU-free, event-driven scheduling to eliminate control overhead and improve processing efficiency. At the dataflow level, a line-to-tile transformation and pipelined multicore execution enhance throughput and memory utilization while ensuring regular computation. At the circuit level, a mixed-precision NPU architecture further improves computational efficiency and image quality. Fabricated in 12-nm technology, ENLIVEN achieves UHD 120-f/s throughput with a peak energy efficiency of 12.08 TOPS/W and a computational density of 2.29 TOPS/mm<sup>2</sup> while delivering significantly improved noise suppression compared with conventional pipelines.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2537-2544"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148628185","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-01Epub Date: 2026-06-01DOI: 10.1109/TVLSI.2026.3696577
Joel Poncha Lemayian;Ghyslain Gagnon;Kaiwen Zhang;Pascal Giard
Cryptographic wallets play a vital role in securing digital assets within blockchain networks by managing private keys that authorize secure transactions. However, side channel analysis (SCA) attacks have become a serious threat, enabling attackers to extract sensitive information by exploiting algorithmic weaknesses in microcontroller-based wallets, resulting in the loss of millions of dollars in digital assets. In hierarchically deterministic (HD) systems, the compromise of a single primary key can endanger all subsequent child keys, while the use of independent keys for each account introduces complexity and challenges in key management. This work presents HardVault, a field programmable gate array (FPGA)-based cryptocurrency wallet that supports both Bitcoin and Ethereum. HardVault introduces the first hardware wallet architecture that implements both non-deterministic (ND) and HD key generation modes directly in hardware, giving users the flexibility to choose either approach based on their security and usability needs. By leveraging constant-time operations and hardware-enforced private-key isolation, the design significantly improves resilience to SCA attacks. In addition, the architecture prioritizes resource efficiency to minimize area usage without compromising security, making it well-suited for compact, portable hardware wallet applications. Implementation on a ZCU104 FPGA shows that HardVault uses only 27% of available look-up tables (LUTs). Compared to the Trezor One cryptocurrency (crypto) wallet, the proposed implementation achieves $9times $ higher energy efficiency, $8times $ lower latency, and $7times $ higher throughput.
{"title":"HardVault: A Hybrid FPGA-Based Ethereum-Bitcoin Cold Wallet","authors":"Joel Poncha Lemayian;Ghyslain Gagnon;Kaiwen Zhang;Pascal Giard","doi":"10.1109/TVLSI.2026.3696577","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3696577","url":null,"abstract":"Cryptographic wallets play a vital role in securing digital assets within blockchain networks by managing private keys that authorize secure transactions. However, side channel analysis (SCA) attacks have become a serious threat, enabling attackers to extract sensitive information by exploiting algorithmic weaknesses in microcontroller-based wallets, resulting in the loss of millions of dollars in digital assets. In hierarchically deterministic (HD) systems, the compromise of a single primary key can endanger all subsequent child keys, while the use of independent keys for each account introduces complexity and challenges in key management. This work presents HardVault, a field programmable gate array (FPGA)-based cryptocurrency wallet that supports both Bitcoin and Ethereum. HardVault introduces the first hardware wallet architecture that implements both non-deterministic (ND) and HD key generation modes directly in hardware, giving users the flexibility to choose either approach based on their security and usability needs. By leveraging constant-time operations and hardware-enforced private-key isolation, the design significantly improves resilience to SCA attacks. In addition, the architecture prioritizes resource efficiency to minimize area usage without compromising security, making it well-suited for compact, portable hardware wallet applications. Implementation on a ZCU104 FPGA shows that HardVault uses only 27% of available look-up tables (LUTs). Compared to the Trezor One cryptocurrency (crypto) wallet, the proposed implementation achieves <inline-formula> <tex-math>$9times $ </tex-math></inline-formula> higher energy efficiency, <inline-formula> <tex-math>$8times $ </tex-math></inline-formula> lower latency, and <inline-formula> <tex-math>$7times $ </tex-math></inline-formula> higher throughput.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2469-2482"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11540335","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148628295","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-01Epub Date: 2026-03-26DOI: 10.1109/TVLSI.2026.3693050
T. C. Jayasree;S. Shaeen Kalathil;R. Nandakumar;T. Bindima
A novel unified Newton structure (UNS) for fractional delay (FD) filtering is presented in this article, which makes it possible to implement an arbitrary bandwidth filter (ABF) with low complexity and reconfigurability. To control the FD value at a lower computational cost, the Newton structure formulation uses the Hermite interpolation (HI) technique. Multiple variable-bandwidth (BW) responses are generated from a single hardware structure by integrating a fixed filter between two FD filters (FDFs) in the proposed architecture. Detailed performance and complexity analyses on field-programmable gate array (FPGA) platforms are used to assess the hardware efficiency, and the suggested design is further validated by ASIC synthesis results. The proposed structure is suitable for wireless channelizers, since the results show significant savings in hardware utilization along with reduced computational complexity, while maintaining comparable filtering performance.
{"title":"A Novel Unified Newton Structure-Based Fractional Delay Filter for Low-Complexity Reconfigurable Arbitrary Bandwidth Filters","authors":"T. C. Jayasree;S. Shaeen Kalathil;R. Nandakumar;T. Bindima","doi":"10.1109/TVLSI.2026.3693050","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3693050","url":null,"abstract":"A novel unified Newton structure (UNS) for fractional delay (FD) filtering is presented in this article, which makes it possible to implement an arbitrary bandwidth filter (ABF) with low complexity and reconfigurability. To control the FD value at a lower computational cost, the Newton structure formulation uses the Hermite interpolation (HI) technique. Multiple variable-bandwidth (BW) responses are generated from a single hardware structure by integrating a fixed filter between two FD filters (FDFs) in the proposed architecture. Detailed performance and complexity analyses on field-programmable gate array (FPGA) platforms are used to assess the hardware efficiency, and the suggested design is further validated by ASIC synthesis results. The proposed structure is suitable for wireless channelizers, since the results show significant savings in hardware utilization along with reduced computational complexity, while maintaining comparable filtering performance.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2408-2420"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148626407","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Area efficiency is vital for floating-point units (FPUs) in resource-constrained IoT devices. However, existing designs suffer from rigid architectures and costly arithmetic units, limiting performance-area optimization. To this end, this work presents FPUltra, an area-efficient single-precision FPU for cost-sensitive RISC-V cores. FPUltra adopts a novel phase-decoupled control architecture to mitigate timing hazards and improve execution efficiency. A parallel approximate floating-point multiplier (FPM) is designed using combinational logic, based on the Mitchell algorithm with error compensation. A Newton–Raphson-based subinstruction decomposition method is presented to support floating-point division (Fdiv) and square root (Fsqrt). Compared with state-of-the-art FPUs, FPUltra achieves 9%–695% and 101%–14 186% improvements in equivalent slices efficiency (Eq.Slices Eff.) on FPGA and equivalent area efficiency (Eq.Area Eff.) on ASIC, respectively. Our code will be available at https://github.com/LX-IC/FPUltra
{"title":"FPUltra: An Area-Efficient Single-Precision Floating-Point Unit for Cost-Sensitive RISC-V Cores","authors":"Xian Lin;Jiahao Lan;Xin Zheng;Huanxin Zhuang;Huaien Gao;Shuting Cai;Xiaoming Xiong","doi":"10.1109/TVLSI.2026.3690450","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3690450","url":null,"abstract":"Area efficiency is vital for floating-point units (FPUs) in resource-constrained IoT devices. However, existing designs suffer from rigid architectures and costly arithmetic units, limiting performance-area optimization. To this end, this work presents FPUltra, an area-efficient single-precision FPU for cost-sensitive RISC-V cores. FPUltra adopts a novel phase-decoupled control architecture to mitigate timing hazards and improve execution efficiency. A parallel approximate floating-point multiplier (FPM) is designed using combinational logic, based on the Mitchell algorithm with error compensation. A Newton–Raphson-based subinstruction decomposition method is presented to support floating-point division (Fdiv) and square root (Fsqrt). Compared with state-of-the-art FPUs, FPUltra achieves 9%–695% and 101%–14 186% improvements in equivalent slices efficiency (Eq.Slices Eff.) on FPGA and equivalent area efficiency (Eq.Area Eff.) on ASIC, respectively. Our code will be available at <uri>https://github.com/LX-IC/FPUltra</uri>","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2661-2665"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148626604","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-01Epub Date: 2026-03-13DOI: 10.1109/TVLSI.2026.3690726
Wei Chen;Dake Liu
At present, most networks on chips (NoCs) are designed for symmetric multiprocessing architectures for general-purpose computing. While currently designed NoCs provide considerable flexibility, they are accompanied by substantial routing delays and a relatively high cost associated with reorder buffer memory. The significant difference in data transmission length is a characteristic of heterogeneous multi-core systems. We propose an innovative NoC architecture for embedded heterogeneous multi-core digital signal processor (DSP) systems, which includes the design of top-level architecture, router, network interface (NI), and NoC pipeline. Our designed NoC architecture has the advantages of flexibility and low latency. It can not only significantly improve the transmission performance of long data packets but also improve the transmission efficiency of short data packets. Moreover, we propose efficient routing methods, which include an efficient routing algorithm, a hybrid transmission method, a packet-connected circuit transmission method, a short packet transmission method, and multicast and broadcast transmission methods. We also propose fault-tolerant data transmission mechanisms, which include a fault-tolerant data transmission algorithm and a deadlock avoidance method. Our proposed method and mechanism can improve the performance and reliability of NoC, mitigate the impact of router failures, eliminate the need for reorder buffer memory, reduce power consumption, and improve the robustness of NoC. The experiments show the performance advantages of our design in terms of latency, throughput, and silicon overhead in comparison with state-of-the-art NoCs. Our design reduced the latency, area, and power consumption by 47%, 92.6%, and 78.9%, respectively, and improved the throughput by 27%. This verifies its effectiveness in high-end applications.
{"title":"Efficient NoC for Embedded Heterogeneous Multi-Core DSPs","authors":"Wei Chen;Dake Liu","doi":"10.1109/TVLSI.2026.3690726","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3690726","url":null,"abstract":"At present, most networks on chips (NoCs) are designed for symmetric multiprocessing architectures for general-purpose computing. While currently designed NoCs provide considerable flexibility, they are accompanied by substantial routing delays and a relatively high cost associated with reorder buffer memory. The significant difference in data transmission length is a characteristic of heterogeneous multi-core systems. We propose an innovative NoC architecture for embedded heterogeneous multi-core digital signal processor (DSP) systems, which includes the design of top-level architecture, router, network interface (NI), and NoC pipeline. Our designed NoC architecture has the advantages of flexibility and low latency. It can not only significantly improve the transmission performance of long data packets but also improve the transmission efficiency of short data packets. Moreover, we propose efficient routing methods, which include an efficient routing algorithm, a hybrid transmission method, a packet-connected circuit transmission method, a short packet transmission method, and multicast and broadcast transmission methods. We also propose fault-tolerant data transmission mechanisms, which include a fault-tolerant data transmission algorithm and a deadlock avoidance method. Our proposed method and mechanism can improve the performance and reliability of NoC, mitigate the impact of router failures, eliminate the need for reorder buffer memory, reduce power consumption, and improve the robustness of NoC. The experiments show the performance advantages of our design in terms of latency, throughput, and silicon overhead in comparison with state-of-the-art NoCs. Our design reduced the latency, area, and power consumption by 47%, 92.6%, and 78.9%, respectively, and improved the throughput by 27%. This verifies its effectiveness in high-end applications.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2523-2536"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148627028","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-01Epub Date: 2026-07-27DOI: 10.1109/TVLSI.2026.3708948
{"title":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems Publication Information","authors":"","doi":"10.1109/TVLSI.2026.3708948","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3708948","url":null,"abstract":"","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"C2-C2"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11626105","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148627623","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Conventional duty-cycle-based (DCB) maximum power point tracking (MPPT) schemes generally assume a 50% duty cycle for operation at the maximum power point (MPP). However, due to nonideal losses in the piezoelectric transducer (PZT), such as dielectric loss, mechanical damping, and rectifier off-state leakage, the actual optimal duty cycle deviates from this nominal value. This brief develops a theoretical model that incorporates these losses, which is derived, analyzed, and experimentally validated using a piezoelectric energy harvester (PEH) integrated with a bias-flip MPPT regulating rectifier (BMRR). Measurement results show that the optimal duty cycle ranges from 42.3% to 42.8%, under which the rectifier delivers 1.1 times the output power compared to operation at a fixed 50% duty cycle. Furthermore, the proposed rectifier achieves a peak power conversion efficiency (PCE) of 89.6% and demonstrates a 7.4-fold improvement in energy extraction compared to a full-bridge rectifier (FBR).
{"title":"Analysis and Validation of Duty-Cycle-Based MPPT for Piezoelectric Energy Harvesting: Impact of Nonideal Losses on Optimal Duty Cycle","authors":"Xiudeng Wang;Shulin Gao;Libo Qian;Yongyuan Li;Zhangming Zhu","doi":"10.1109/TVLSI.2026.3689112","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3689112","url":null,"abstract":"Conventional duty-cycle-based (DCB) maximum power point tracking (MPPT) schemes generally assume a 50% duty cycle for operation at the maximum power point (MPP). However, due to nonideal losses in the piezoelectric transducer (PZT), such as dielectric loss, mechanical damping, and rectifier <sc>off</small>-state leakage, the actual optimal duty cycle deviates from this nominal value. This brief develops a theoretical model that incorporates these losses, which is derived, analyzed, and experimentally validated using a piezoelectric energy harvester (PEH) integrated with a bias-flip MPPT regulating rectifier (BMRR). Measurement results show that the optimal duty cycle ranges from 42.3% to 42.8%, under which the rectifier delivers 1.1 times the output power compared to operation at a fixed 50% duty cycle. Furthermore, the proposed rectifier achieves a peak power conversion efficiency (PCE) of 89.6% and demonstrates a 7.4-fold improvement in energy extraction compared to a full-bridge rectifier (FBR).","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2641-2645"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148626600","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
This article presents a second-order fully passive noise-shaping (NS) successive approximation register (SAR) analog-to-digital converter (ADC) with $4times $ passive gain for low oversampling ratio (OSR) designs. By adopting the ping-pong residue extraction technique, the proposed architecture suffers no speed penalty compared with the pure SAR ADC. The merged differential residue extraction and parallel integration technique is proposed to provide a $4times $ fully passive gain, and the parasitic capacitance confining high-order implementation is minimized by isolating the integration domain from the sampling domain. Consequently, an aggressive noise transfer function (NTF) prevailing over the ideal second-order is realized. Besides, a highly linear excessive-capacitor-reset (ECR) switching scheme is introduced. Compared with the $V_{text {cm}}$ -based scheme, the ECR scheme reduces the integral nonlinearity (INL) by 50%, merely at the cost of a few more digital logic gates. The prototype achieves 71.3 dB SNDR with a power of $620.2~mu $ W and a bandwidth of 1 MHz, resulting in a FoM of 163.3 dB.
{"title":"A 71.3 dB-SNDR Fully Passive Speed Enhanced 4× OSR Second-Order NS SAR ADC With Excessive-Capacitor-Reset Switching Scheme","authors":"Haoyu Wu;Jialin Hu;Xingshuai Zou;Biao Wang;Xiangyu Mao;Wenbo Luo;Jiaxin Liu;Guanghui Chen;Minhan Zou;Shiheng Yang;Kai Kang;Wanli Zhang;Moufu Kong;Hongshuai Zhang","doi":"10.1109/TVLSI.2026.3686072","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3686072","url":null,"abstract":"This article presents a second-order fully passive noise-shaping (NS) successive approximation register (SAR) analog-to-digital converter (ADC) with <inline-formula> <tex-math>$4times $ </tex-math></inline-formula> passive gain for low oversampling ratio (OSR) designs. By adopting the ping-pong residue extraction technique, the proposed architecture suffers no speed penalty compared with the pure SAR ADC. The merged differential residue extraction and parallel integration technique is proposed to provide a <inline-formula> <tex-math>$4times $ </tex-math></inline-formula> fully passive gain, and the parasitic capacitance confining high-order implementation is minimized by isolating the integration domain from the sampling domain. Consequently, an aggressive noise transfer function (NTF) prevailing over the ideal second-order is realized. Besides, a highly linear excessive-capacitor-reset (ECR) switching scheme is introduced. Compared with the <inline-formula> <tex-math>$V_{text {cm}}$ </tex-math></inline-formula>-based scheme, the ECR scheme reduces the integral nonlinearity (INL) by 50%, merely at the cost of a few more digital logic gates. The prototype achieves 71.3 dB SNDR with a power of <inline-formula> <tex-math>$620.2~mu $ </tex-math></inline-formula>W and a bandwidth of 1 MHz, resulting in a FoM of 163.3 dB.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2331-2343"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148627131","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}