首页 > 最新文献

IEEE Transactions on Very Large Scale Integration (VLSI) Systems最新文献

英文 中文
Average-6.5T Near-Threshold Twin Cell With Shared Read Assist for IoT Applications 平均6.5 t近阈值双电池与共享读取辅助物联网应用
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-03-22 DOI: 10.1109/TVLSI.2026.3693724
Liang Wen;Lixun Wang;Jiangong Wang;Yuejun Zhang
This brief proposes an average-6.5T twin cell for a deep sub-micrometer 64 kb SRAM, which utilizes two identical asymmetric single-ended (SE) 6T cells in a column with a shared read assist device to improve read margin and write ability. It enables the SRAM to achieve read-disturb-free, near/sub-threshold operation and compact array layout, resulting in area and energy efficiencies. The average-6.5T SRAM test chip is fabricated using a 65 nm CMOS logic process. Its cell area shows only 5.6% overhead compared to the standard 6T cell, and is smaller than that of other low-voltage SRAMs. Measured full read and write functionality is performed with VDD down to 0.39 V, which is lower than that of standard 6T and 8T SRAMs. In addition, its minimum energy point of 6.3 pJ is obtained at 0.48 V.
本文提出了一种用于深度亚微米64 kb SRAM的平均6.5 t双单元,该单元利用两个相同的非对称单端(SE) 6T单元在一个列中,具有共享的读辅助装置,以提高读距和写能力。它使SRAM能够实现无读取干扰,近/亚阈值操作和紧凑的阵列布局,从而提高面积和能源效率。平均6.5 t SRAM测试芯片采用65nm CMOS逻辑工艺制造。与标准6T电池相比,其电池面积仅显示5.6%的开销,并且比其他低压sram小。VDD低至0.39 V,比标准的6T和8T sram低,可实现完整的读写功能。此外,在0.48 V时得到了它的最小能量点6.3 pJ。
{"title":"Average-6.5T Near-Threshold Twin Cell With Shared Read Assist for IoT Applications","authors":"Liang Wen;Lixun Wang;Jiangong Wang;Yuejun Zhang","doi":"10.1109/TVLSI.2026.3693724","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3693724","url":null,"abstract":"This brief proposes an average-6.5T twin cell for a deep sub-micrometer 64 kb SRAM, which utilizes two identical asymmetric single-ended (SE) 6T cells in a column with a shared read assist device to improve read margin and write ability. It enables the SRAM to achieve read-disturb-free, near/sub-threshold operation and compact array layout, resulting in area and energy efficiencies. The average-6.5T SRAM test chip is fabricated using a 65 nm CMOS logic process. Its cell area shows only 5.6% overhead compared to the standard 6T cell, and is smaller than that of other low-voltage SRAMs. Measured full read and write functionality is performed with VDD down to 0.39 V, which is lower than that of standard 6T and 8T SRAMs. In addition, its minimum energy point of 6.3 pJ is obtained at 0.48 V.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2676-2680"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148627504","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
System-Level Fault-Tolerant Reconfiguration of 3-D VLSI Processor Arrays Under Multicomponent Failures 多组件故障下三维VLSI处理器阵列的系统级容错重构
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-03-27 DOI: 10.1109/TVLSI.2026.3695877
Hao Ding;Xiangyong Wang;Yang Xu;Junyan Qian
Three-dimensional (3-D) VLSI processor arrays offer high integration density and scalability, but their increasing system and interconnect complexity pose significant challenges to reliable reconfiguration under permanent hardware failures. Most existing reconfiguration approaches primarily target processing element (PE) faults and provide limited support for interconnect-related failures, such as switch and link faults, which can severely constrain feasible reconfiguration solutions and the achievable size of fault-free subarrays at the system level. This article investigates system-level fault-tolerant reconfiguration of reconfigurable 3-D VLSI processor arrays under multicomponent failures, including PE, switch, and link faults. We first analyze the system-level impact of different fault types and show that interconnect-level faults, especially switch failures, impose more pronounced constraints on reconfiguration effectiveness than PE faults. Based on this observation, a general plane-exclusion mechanism is introduced that can be integrated with representative PE-only reconfiguration schemes to enhance tolerance to switch and link failures. Furthermore, a fault-transformation preprocessing mechanism is developed to model switch failures as equivalent link disconnections, enabling unified system-level handling of heterogeneous faults and improving structural robustness. To mitigate excessive interconnect overhead introduced during reconfiguration, a long-interconnect optimization strategy is incorporated. Experimental results demonstrate that the proposed techniques significantly improve reconfiguration capability and scalability. In particular, the fault-transformation mechanism achieves over 90% of the theoretical upper bound in PE utilization under high switch fault densities, nearly doubling PE utilization compared with baseline methods, while the interconnect optimization reduces the number of long interconnects by more than 10% on average across all evaluated cases.
三维(3-D) VLSI处理器阵列具有高集成密度和可扩展性,但其不断增加的系统和互连复杂性给永久性硬件故障下的可靠重构带来了重大挑战。大多数现有的重构方法主要针对处理元件(PE)故障,并对互连相关故障(如交换机和链路故障)提供有限的支持,这严重限制了可行的重构解决方案和系统级无故障子阵列的可实现大小。本文研究了可重构三维VLSI处理器阵列在多组件故障(包括PE、开关和链路故障)下的系统级容错重构。我们首先分析了不同故障类型的系统级影响,并表明互连级故障,特别是交换机故障,比PE故障对重构有效性施加了更明显的约束。在此基础上,引入了一种通用的平面排除机制,该机制可以与典型的pe重构方案相结合,以提高对交换机和链路故障的容忍度。此外,还开发了一种故障转换预处理机制,将交换机故障建模为等效链路断开,实现了异构故障的统一系统级处理,提高了结构的鲁棒性。为了减少重新配置过程中引入的过多互连开销,采用了长互连优化策略。实验结果表明,所提技术显著提高了重构能力和可扩展性。特别是,故障转换机制在高开关故障密度下实现了90%以上的PE利用率理论上限,与基线方法相比,PE利用率几乎翻了一倍,而互连优化在所有评估案例中平均减少了10%以上的长互连数量。
{"title":"System-Level Fault-Tolerant Reconfiguration of 3-D VLSI Processor Arrays Under Multicomponent Failures","authors":"Hao Ding;Xiangyong Wang;Yang Xu;Junyan Qian","doi":"10.1109/TVLSI.2026.3695877","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3695877","url":null,"abstract":"Three-dimensional (3-D) VLSI processor arrays offer high integration density and scalability, but their increasing system and interconnect complexity pose significant challenges to reliable reconfiguration under permanent hardware failures. Most existing reconfiguration approaches primarily target processing element (PE) faults and provide limited support for interconnect-related failures, such as switch and link faults, which can severely constrain feasible reconfiguration solutions and the achievable size of fault-free subarrays at the system level. This article investigates system-level fault-tolerant reconfiguration of reconfigurable 3-D VLSI processor arrays under multicomponent failures, including PE, switch, and link faults. We first analyze the system-level impact of different fault types and show that interconnect-level faults, especially switch failures, impose more pronounced constraints on reconfiguration effectiveness than PE faults. Based on this observation, a general plane-exclusion mechanism is introduced that can be integrated with representative PE-only reconfiguration schemes to enhance tolerance to switch and link failures. Furthermore, a fault-transformation preprocessing mechanism is developed to model switch failures as equivalent link disconnections, enabling unified system-level handling of heterogeneous faults and improving structural robustness. To mitigate excessive interconnect overhead introduced during reconfiguration, a long-interconnect optimization strategy is incorporated. Experimental results demonstrate that the proposed techniques significantly improve reconfiguration capability and scalability. In particular, the fault-transformation mechanism achieves over 90% of the theoretical upper bound in PE utilization under high switch fault densities, nearly doubling PE utilization compared with baseline methods, while the interconnect optimization reduces the number of long interconnects by more than 10% on average across all evaluated cases.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2607-2620"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148628156","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
ENLIVEN: End-to-End NPU--ISP Codesign for Low-Latency and Hardware-Optimized Visual Processing ENLIVEN:端到端NPU- ISP协同设计,用于低延迟和硬件优化的视觉处理
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-06-24 DOI: 10.1109/TVLSI.2026.3694263
Jiangtao Cui;Xinyu Shao;Sheng Zhang
Rapid growth of intelligent vision applications demands high image quality, low latency, and energy efficiency. Traditional pipelines integrating a neural processing unit (NPU) with an image signal processor (ISP) suffer from CPU scheduling delays, memory bandwidth bottlenecks, and limited precision flexibility. To address these challenges, we propose ENLIVEN, a holistically optimized NPU–ISP architecture that codesign system, dataflow, and circuit layers. At the system level, ENLIVEN enables CPU-free, event-driven scheduling to eliminate control overhead and improve processing efficiency. At the dataflow level, a line-to-tile transformation and pipelined multicore execution enhance throughput and memory utilization while ensuring regular computation. At the circuit level, a mixed-precision NPU architecture further improves computational efficiency and image quality. Fabricated in 12-nm technology, ENLIVEN achieves UHD 120-f/s throughput with a peak energy efficiency of 12.08 TOPS/W and a computational density of 2.29 TOPS/mm2 while delivering significantly improved noise suppression compared with conventional pipelines.
智能视觉应用的快速增长要求高图像质量、低延迟和能源效率。传统的集成神经处理单元(NPU)和图像信号处理器(ISP)的管道存在CPU调度延迟、内存带宽瓶颈和精度灵活性有限等问题。为了应对这些挑战,我们提出了ENLIVEN,这是一种整体优化的NPU-ISP架构,可协同设计系统,数据流和电路层。在系统级别,ENLIVEN支持无cpu、事件驱动的调度,以消除控制开销并提高处理效率。在数据流级别,行到块转换和流水线多核执行增强了吞吐量和内存利用率,同时确保了常规计算。在电路层面,混合精度NPU架构进一步提高了计算效率和图像质量。采用12纳米技术制造的ENLIVEN实现了超高清120-f/s的吞吐量,峰值能量效率为12.08 TOPS/W,计算密度为2.29 TOPS/mm2,同时与传统管道相比,显著改善了噪声抑制。
{"title":"ENLIVEN: End-to-End NPU--ISP Codesign for Low-Latency and Hardware-Optimized Visual Processing","authors":"Jiangtao Cui;Xinyu Shao;Sheng Zhang","doi":"10.1109/TVLSI.2026.3694263","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3694263","url":null,"abstract":"Rapid growth of intelligent vision applications demands high image quality, low latency, and energy efficiency. Traditional pipelines integrating a neural processing unit (NPU) with an image signal processor (ISP) suffer from CPU scheduling delays, memory bandwidth bottlenecks, and limited precision flexibility. To address these challenges, we propose ENLIVEN, a holistically optimized NPU–ISP architecture that codesign system, dataflow, and circuit layers. At the system level, ENLIVEN enables CPU-free, event-driven scheduling to eliminate control overhead and improve processing efficiency. At the dataflow level, a line-to-tile transformation and pipelined multicore execution enhance throughput and memory utilization while ensuring regular computation. At the circuit level, a mixed-precision NPU architecture further improves computational efficiency and image quality. Fabricated in 12-nm technology, ENLIVEN achieves UHD 120-f/s throughput with a peak energy efficiency of 12.08 TOPS/W and a computational density of 2.29 TOPS/mm<sup>2</sup> while delivering significantly improved noise suppression compared with conventional pipelines.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2537-2544"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148628185","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
HardVault: A Hybrid FPGA-Based Ethereum-Bitcoin Cold Wallet HardVault:一个基于fpga的混合以太币-比特币冷钱包
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-06-01 DOI: 10.1109/TVLSI.2026.3696577
Joel Poncha Lemayian;Ghyslain Gagnon;Kaiwen Zhang;Pascal Giard
Cryptographic wallets play a vital role in securing digital assets within blockchain networks by managing private keys that authorize secure transactions. However, side channel analysis (SCA) attacks have become a serious threat, enabling attackers to extract sensitive information by exploiting algorithmic weaknesses in microcontroller-based wallets, resulting in the loss of millions of dollars in digital assets. In hierarchically deterministic (HD) systems, the compromise of a single primary key can endanger all subsequent child keys, while the use of independent keys for each account introduces complexity and challenges in key management. This work presents HardVault, a field programmable gate array (FPGA)-based cryptocurrency wallet that supports both Bitcoin and Ethereum. HardVault introduces the first hardware wallet architecture that implements both non-deterministic (ND) and HD key generation modes directly in hardware, giving users the flexibility to choose either approach based on their security and usability needs. By leveraging constant-time operations and hardware-enforced private-key isolation, the design significantly improves resilience to SCA attacks. In addition, the architecture prioritizes resource efficiency to minimize area usage without compromising security, making it well-suited for compact, portable hardware wallet applications. Implementation on a ZCU104 FPGA shows that HardVault uses only 27% of available look-up tables (LUTs). Compared to the Trezor One cryptocurrency (crypto) wallet, the proposed implementation achieves $9times $ higher energy efficiency, $8times $ lower latency, and $7times $ higher throughput.
加密钱包通过管理授权安全交易的私钥,在b区块链网络中保护数字资产方面发挥着至关重要的作用。然而,侧通道分析(SCA)攻击已经成为一种严重的威胁,使攻击者能够通过利用基于微控制器的钱包中的算法弱点来提取敏感信息,导致数百万美元的数字资产损失。在层次确定性(HD)系统中,单个主密钥的泄露可能危及所有后续子密钥,而每个帐户使用独立密钥会给密钥管理带来复杂性和挑战。这项工作提出了HardVault,这是一种基于现场可编程门阵列(FPGA)的加密货币钱包,支持比特币和以太坊。HardVault引入了第一个硬件钱包架构,它直接在硬件中实现了非确定性(ND)和高清密钥生成模式,使用户可以根据他们的安全性和可用性需求灵活地选择任何一种方法。通过利用固定时间操作和硬件强制的私钥隔离,该设计显著提高了对SCA攻击的弹性。此外,该体系结构优先考虑资源效率,在不影响安全性的情况下最大限度地减少面积使用,使其非常适合紧凑、便携的硬件钱包应用程序。在ZCU104 FPGA上的实现表明HardVault仅使用了27%的可用查找表(lut)。与Trezor One加密货币(crypto)钱包相比,提议的实现实现了9倍的能源效率,8倍的延迟和7倍的吞吐量。
{"title":"HardVault: A Hybrid FPGA-Based Ethereum-Bitcoin Cold Wallet","authors":"Joel Poncha Lemayian;Ghyslain Gagnon;Kaiwen Zhang;Pascal Giard","doi":"10.1109/TVLSI.2026.3696577","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3696577","url":null,"abstract":"Cryptographic wallets play a vital role in securing digital assets within blockchain networks by managing private keys that authorize secure transactions. However, side channel analysis (SCA) attacks have become a serious threat, enabling attackers to extract sensitive information by exploiting algorithmic weaknesses in microcontroller-based wallets, resulting in the loss of millions of dollars in digital assets. In hierarchically deterministic (HD) systems, the compromise of a single primary key can endanger all subsequent child keys, while the use of independent keys for each account introduces complexity and challenges in key management. This work presents HardVault, a field programmable gate array (FPGA)-based cryptocurrency wallet that supports both Bitcoin and Ethereum. HardVault introduces the first hardware wallet architecture that implements both non-deterministic (ND) and HD key generation modes directly in hardware, giving users the flexibility to choose either approach based on their security and usability needs. By leveraging constant-time operations and hardware-enforced private-key isolation, the design significantly improves resilience to SCA attacks. In addition, the architecture prioritizes resource efficiency to minimize area usage without compromising security, making it well-suited for compact, portable hardware wallet applications. Implementation on a ZCU104 FPGA shows that HardVault uses only 27% of available look-up tables (LUTs). Compared to the Trezor One cryptocurrency (crypto) wallet, the proposed implementation achieves <inline-formula> <tex-math>$9times $ </tex-math></inline-formula> higher energy efficiency, <inline-formula> <tex-math>$8times $ </tex-math></inline-formula> lower latency, and <inline-formula> <tex-math>$7times $ </tex-math></inline-formula> higher throughput.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2469-2482"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11540335","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148628295","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
A Novel Unified Newton Structure-Based Fractional Delay Filter for Low-Complexity Reconfigurable Arbitrary Bandwidth Filters 一种新的基于统一牛顿结构的分数阶延迟滤波器,用于低复杂度可重构任意带宽滤波器
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-03-26 DOI: 10.1109/TVLSI.2026.3693050
T. C. Jayasree;S. Shaeen Kalathil;R. Nandakumar;T. Bindima
A novel unified Newton structure (UNS) for fractional delay (FD) filtering is presented in this article, which makes it possible to implement an arbitrary bandwidth filter (ABF) with low complexity and reconfigurability. To control the FD value at a lower computational cost, the Newton structure formulation uses the Hermite interpolation (HI) technique. Multiple variable-bandwidth (BW) responses are generated from a single hardware structure by integrating a fixed filter between two FD filters (FDFs) in the proposed architecture. Detailed performance and complexity analyses on field-programmable gate array (FPGA) platforms are used to assess the hardware efficiency, and the suggested design is further validated by ASIC synthesis results. The proposed structure is suitable for wireless channelizers, since the results show significant savings in hardware utilization along with reduced computational complexity, while maintaining comparable filtering performance.
本文提出了一种新的统一牛顿结构(UNS)用于分数延迟(FD)滤波,使得实现低复杂度和可重构的任意带宽滤波器(ABF)成为可能。为了在较低的计算成本下控制FD值,牛顿结构公式使用了Hermite插值(HI)技术。通过在两个FD滤波器(fdf)之间集成一个固定滤波器,可以在单个硬件结构中生成多个可变带宽(BW)响应。通过现场可编程门阵列(FPGA)平台的详细性能和复杂性分析来评估硬件效率,并通过ASIC综合结果进一步验证了所提出的设计。所提出的结构适用于无线信道器,因为结果表明,在保持相当的滤波性能的同时,硬件利用率显著降低,计算复杂性降低。
{"title":"A Novel Unified Newton Structure-Based Fractional Delay Filter for Low-Complexity Reconfigurable Arbitrary Bandwidth Filters","authors":"T. C. Jayasree;S. Shaeen Kalathil;R. Nandakumar;T. Bindima","doi":"10.1109/TVLSI.2026.3693050","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3693050","url":null,"abstract":"A novel unified Newton structure (UNS) for fractional delay (FD) filtering is presented in this article, which makes it possible to implement an arbitrary bandwidth filter (ABF) with low complexity and reconfigurability. To control the FD value at a lower computational cost, the Newton structure formulation uses the Hermite interpolation (HI) technique. Multiple variable-bandwidth (BW) responses are generated from a single hardware structure by integrating a fixed filter between two FD filters (FDFs) in the proposed architecture. Detailed performance and complexity analyses on field-programmable gate array (FPGA) platforms are used to assess the hardware efficiency, and the suggested design is further validated by ASIC synthesis results. The proposed structure is suitable for wireless channelizers, since the results show significant savings in hardware utilization along with reduced computational complexity, while maintaining comparable filtering performance.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2408-2420"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148626407","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
FPUltra: An Area-Efficient Single-Precision Floating-Point Unit for Cost-Sensitive RISC-V Cores FPUltra:用于成本敏感型RISC-V内核的面积高效单精度浮点单元
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-03-12 DOI: 10.1109/TVLSI.2026.3690450
Xian Lin;Jiahao Lan;Xin Zheng;Huanxin Zhuang;Huaien Gao;Shuting Cai;Xiaoming Xiong
Area efficiency is vital for floating-point units (FPUs) in resource-constrained IoT devices. However, existing designs suffer from rigid architectures and costly arithmetic units, limiting performance-area optimization. To this end, this work presents FPUltra, an area-efficient single-precision FPU for cost-sensitive RISC-V cores. FPUltra adopts a novel phase-decoupled control architecture to mitigate timing hazards and improve execution efficiency. A parallel approximate floating-point multiplier (FPM) is designed using combinational logic, based on the Mitchell algorithm with error compensation. A Newton–Raphson-based subinstruction decomposition method is presented to support floating-point division (Fdiv) and square root (Fsqrt). Compared with state-of-the-art FPUs, FPUltra achieves 9%–695% and 101%–14 186% improvements in equivalent slices efficiency (Eq.Slices Eff.) on FPGA and equivalent area efficiency (Eq.Area Eff.) on ASIC, respectively. Our code will be available at https://github.com/LX-IC/FPUltra
在资源受限的物联网设备中,面积效率对于浮点单元(fpu)至关重要。然而,现有的设计受到刚性架构和昂贵的算术单元的影响,限制了性能区域优化。为此,本研究提出了FPUltra,一种面积高效的单精度FPU,用于成本敏感的RISC-V内核。FPUltra采用了一种新颖的相位解耦控制架构,以减少时序风险,提高执行效率。基于带误差补偿的Mitchell算法,采用组合逻辑设计了一种并行近似浮点乘法器。提出了一种基于牛顿-拉夫森的子指令分解方法,以支持浮点除法和平方根。与最先进的fpu相比,FPUltra在FPGA上的等效切片效率(Eq.Slices Eff)和ASIC上的等效面积效率(Eq.Area Eff)分别提高了9%-695%和101% - 14.186%。我们的代码可以在https://github.com/LX-IC/FPUltra上找到
{"title":"FPUltra: An Area-Efficient Single-Precision Floating-Point Unit for Cost-Sensitive RISC-V Cores","authors":"Xian Lin;Jiahao Lan;Xin Zheng;Huanxin Zhuang;Huaien Gao;Shuting Cai;Xiaoming Xiong","doi":"10.1109/TVLSI.2026.3690450","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3690450","url":null,"abstract":"Area efficiency is vital for floating-point units (FPUs) in resource-constrained IoT devices. However, existing designs suffer from rigid architectures and costly arithmetic units, limiting performance-area optimization. To this end, this work presents FPUltra, an area-efficient single-precision FPU for cost-sensitive RISC-V cores. FPUltra adopts a novel phase-decoupled control architecture to mitigate timing hazards and improve execution efficiency. A parallel approximate floating-point multiplier (FPM) is designed using combinational logic, based on the Mitchell algorithm with error compensation. A Newton–Raphson-based subinstruction decomposition method is presented to support floating-point division (Fdiv) and square root (Fsqrt). Compared with state-of-the-art FPUs, FPUltra achieves 9%–695% and 101%–14 186% improvements in equivalent slices efficiency (Eq.Slices Eff.) on FPGA and equivalent area efficiency (Eq.Area Eff.) on ASIC, respectively. Our code will be available at <uri>https://github.com/LX-IC/FPUltra</uri>","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2661-2665"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148626604","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Efficient NoC for Embedded Heterogeneous Multi-Core DSPs 嵌入式异构多核dsp的高效NoC
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-03-13 DOI: 10.1109/TVLSI.2026.3690726
Wei Chen;Dake Liu
At present, most networks on chips (NoCs) are designed for symmetric multiprocessing architectures for general-purpose computing. While currently designed NoCs provide considerable flexibility, they are accompanied by substantial routing delays and a relatively high cost associated with reorder buffer memory. The significant difference in data transmission length is a characteristic of heterogeneous multi-core systems. We propose an innovative NoC architecture for embedded heterogeneous multi-core digital signal processor (DSP) systems, which includes the design of top-level architecture, router, network interface (NI), and NoC pipeline. Our designed NoC architecture has the advantages of flexibility and low latency. It can not only significantly improve the transmission performance of long data packets but also improve the transmission efficiency of short data packets. Moreover, we propose efficient routing methods, which include an efficient routing algorithm, a hybrid transmission method, a packet-connected circuit transmission method, a short packet transmission method, and multicast and broadcast transmission methods. We also propose fault-tolerant data transmission mechanisms, which include a fault-tolerant data transmission algorithm and a deadlock avoidance method. Our proposed method and mechanism can improve the performance and reliability of NoC, mitigate the impact of router failures, eliminate the need for reorder buffer memory, reduce power consumption, and improve the robustness of NoC. The experiments show the performance advantages of our design in terms of latency, throughput, and silicon overhead in comparison with state-of-the-art NoCs. Our design reduced the latency, area, and power consumption by 47%, 92.6%, and 78.9%, respectively, and improved the throughput by 27%. This verifies its effectiveness in high-end applications.
目前,大多数片上网络(noc)都是针对通用计算的对称多处理架构设计的。虽然目前设计的noc提供了相当大的灵活性,但它们伴随着大量的路由延迟和与重排序缓冲存储器相关的相对较高的成本。数据传输长度的显著差异是异构多核系统的一个特征。本文提出了一种新型的嵌入式异构多核数字信号处理器(DSP)系统NoC架构,包括顶层架构、路由器、网络接口(NI)和NoC管道的设计。我们设计的NoC架构具有灵活性和低延迟的优点。它不仅可以显著提高长数据包的传输性能,还可以提高短数据包的传输效率。此外,我们还提出了高效的路由方法,包括高效路由算法、混合传输方法、分组连接电路传输方法、短分组传输方法以及组播和广播传输方法。我们还提出了容错数据传输机制,包括容错数据传输算法和死锁避免方法。我们提出的方法和机制可以提高NoC的性能和可靠性,减轻路由器故障的影响,消除对缓冲存储器的重新排序需求,降低功耗,提高NoC的鲁棒性。实验表明,与最先进的noc相比,我们的设计在延迟、吞吐量和硅开销方面具有性能优势。我们的设计将延迟、面积和功耗分别降低了47%、92.6%和78.9%,并将吞吐量提高了27%。验证了其在高端应用中的有效性。
{"title":"Efficient NoC for Embedded Heterogeneous Multi-Core DSPs","authors":"Wei Chen;Dake Liu","doi":"10.1109/TVLSI.2026.3690726","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3690726","url":null,"abstract":"At present, most networks on chips (NoCs) are designed for symmetric multiprocessing architectures for general-purpose computing. While currently designed NoCs provide considerable flexibility, they are accompanied by substantial routing delays and a relatively high cost associated with reorder buffer memory. The significant difference in data transmission length is a characteristic of heterogeneous multi-core systems. We propose an innovative NoC architecture for embedded heterogeneous multi-core digital signal processor (DSP) systems, which includes the design of top-level architecture, router, network interface (NI), and NoC pipeline. Our designed NoC architecture has the advantages of flexibility and low latency. It can not only significantly improve the transmission performance of long data packets but also improve the transmission efficiency of short data packets. Moreover, we propose efficient routing methods, which include an efficient routing algorithm, a hybrid transmission method, a packet-connected circuit transmission method, a short packet transmission method, and multicast and broadcast transmission methods. We also propose fault-tolerant data transmission mechanisms, which include a fault-tolerant data transmission algorithm and a deadlock avoidance method. Our proposed method and mechanism can improve the performance and reliability of NoC, mitigate the impact of router failures, eliminate the need for reorder buffer memory, reduce power consumption, and improve the robustness of NoC. The experiments show the performance advantages of our design in terms of latency, throughput, and silicon overhead in comparison with state-of-the-art NoCs. Our design reduced the latency, area, and power consumption by 47%, 92.6%, and 78.9%, respectively, and improved the throughput by 27%. This verifies its effectiveness in high-end applications.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2523-2536"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148627028","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
IEEE Transactions on Very Large Scale Integration (VLSI) Systems Publication Information IEEE超大规模集成电路(VLSI)系统学报
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-07-27 DOI: 10.1109/TVLSI.2026.3708948
{"title":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems Publication Information","authors":"","doi":"10.1109/TVLSI.2026.3708948","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3708948","url":null,"abstract":"","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"C2-C2"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11626105","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148627623","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Analysis and Validation of Duty-Cycle-Based MPPT for Piezoelectric Energy Harvesting: Impact of Nonideal Losses on Optimal Duty Cycle 基于占空比的压电能量收集MPPT分析与验证:非理想损耗对最优占空比的影响
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-03-12 DOI: 10.1109/TVLSI.2026.3689112
Xiudeng Wang;Shulin Gao;Libo Qian;Yongyuan Li;Zhangming Zhu
Conventional duty-cycle-based (DCB) maximum power point tracking (MPPT) schemes generally assume a 50% duty cycle for operation at the maximum power point (MPP). However, due to nonideal losses in the piezoelectric transducer (PZT), such as dielectric loss, mechanical damping, and rectifier off-state leakage, the actual optimal duty cycle deviates from this nominal value. This brief develops a theoretical model that incorporates these losses, which is derived, analyzed, and experimentally validated using a piezoelectric energy harvester (PEH) integrated with a bias-flip MPPT regulating rectifier (BMRR). Measurement results show that the optimal duty cycle ranges from 42.3% to 42.8%, under which the rectifier delivers 1.1 times the output power compared to operation at a fixed 50% duty cycle. Furthermore, the proposed rectifier achieves a peak power conversion efficiency (PCE) of 89.6% and demonstrates a 7.4-fold improvement in energy extraction compared to a full-bridge rectifier (FBR).
传统的基于占空比(DCB)的最大功率点跟踪(MPPT)方案通常假设在最大功率点(MPP)运行时占空比为50%。然而,由于压电换能器(PZT)中的非理想损耗,如介电损耗、机械阻尼和整流器的失态泄漏,实际的最佳占空比偏离了这个标称值。本文建立了一个包含这些损耗的理论模型,并使用集成了偏置翻转MPPT调节整流器(BMRR)的压电能量收集器(PEH)推导、分析和实验验证了该模型。测量结果表明,最佳占空比范围为42.3% ~ 42.8%,在此范围内,整流器输出功率是固定占空比50%时输出功率的1.1倍。此外,该整流器的峰值功率转换效率(PCE)为89.6%,与全桥整流器(FBR)相比,能量提取效率提高了7.4倍。
{"title":"Analysis and Validation of Duty-Cycle-Based MPPT for Piezoelectric Energy Harvesting: Impact of Nonideal Losses on Optimal Duty Cycle","authors":"Xiudeng Wang;Shulin Gao;Libo Qian;Yongyuan Li;Zhangming Zhu","doi":"10.1109/TVLSI.2026.3689112","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3689112","url":null,"abstract":"Conventional duty-cycle-based (DCB) maximum power point tracking (MPPT) schemes generally assume a 50% duty cycle for operation at the maximum power point (MPP). However, due to nonideal losses in the piezoelectric transducer (PZT), such as dielectric loss, mechanical damping, and rectifier <sc>off</small>-state leakage, the actual optimal duty cycle deviates from this nominal value. This brief develops a theoretical model that incorporates these losses, which is derived, analyzed, and experimentally validated using a piezoelectric energy harvester (PEH) integrated with a bias-flip MPPT regulating rectifier (BMRR). Measurement results show that the optimal duty cycle ranges from 42.3% to 42.8%, under which the rectifier delivers 1.1 times the output power compared to operation at a fixed 50% duty cycle. Furthermore, the proposed rectifier achieves a peak power conversion efficiency (PCE) of 89.6% and demonstrates a 7.4-fold improvement in energy extraction compared to a full-bridge rectifier (FBR).","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2641-2645"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148626600","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
A 71.3 dB-SNDR Fully Passive Speed Enhanced 4× OSR Second-Order NS SAR ADC With Excessive-Capacitor-Reset Switching Scheme 一种71.3 dB-SNDR全无源速度增强4× OSR二阶NS SAR ADC
IF 3.2 2区 工程技术 Q2 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE Pub Date : 2026-08-01 Epub Date: 2026-06-22 DOI: 10.1109/TVLSI.2026.3686072
Haoyu Wu;Jialin Hu;Xingshuai Zou;Biao Wang;Xiangyu Mao;Wenbo Luo;Jiaxin Liu;Guanghui Chen;Minhan Zou;Shiheng Yang;Kai Kang;Wanli Zhang;Moufu Kong;Hongshuai Zhang
This article presents a second-order fully passive noise-shaping (NS) successive approximation register (SAR) analog-to-digital converter (ADC) with $4times $ passive gain for low oversampling ratio (OSR) designs. By adopting the ping-pong residue extraction technique, the proposed architecture suffers no speed penalty compared with the pure SAR ADC. The merged differential residue extraction and parallel integration technique is proposed to provide a $4times $ fully passive gain, and the parasitic capacitance confining high-order implementation is minimized by isolating the integration domain from the sampling domain. Consequently, an aggressive noise transfer function (NTF) prevailing over the ideal second-order is realized. Besides, a highly linear excessive-capacitor-reset (ECR) switching scheme is introduced. Compared with the $V_{text {cm}}$ -based scheme, the ECR scheme reduces the integral nonlinearity (INL) by 50%, merely at the cost of a few more digital logic gates. The prototype achieves 71.3 dB SNDR with a power of $620.2~mu $ W and a bandwidth of 1 MHz, resulting in a FoM of 163.3 dB.
本文提出了一种二阶全无源噪声整形(NS)逐次逼近寄存器(SAR)模数转换器(ADC),具有4倍无源增益,适用于低过采样比(OSR)设计。通过采用乒乓残馀提取技术,该架构与纯SAR ADC相比没有速度损失。提出了融合微分残差提取和并行积分技术,以提供$4 × $的全无源增益,并通过将积分域与采样域隔离,将限制高阶实现的寄生电容降至最低。因此,一个侵略性的噪声传递函数(NTF)盛行于理想的二阶被实现。此外,还介绍了一种高度线性的过电容复位(ECR)开关方案。与基于$V_{text {cm}}$的方案相比,ECR方案仅以增加几个数字逻辑门为代价,将积分非线性(INL)降低了50%。该样机实现了71.3 dB SNDR,功率为620.2~mu $ W,带宽为1 MHz, FoM为163.3 dB。
{"title":"A 71.3 dB-SNDR Fully Passive Speed Enhanced 4× OSR Second-Order NS SAR ADC With Excessive-Capacitor-Reset Switching Scheme","authors":"Haoyu Wu;Jialin Hu;Xingshuai Zou;Biao Wang;Xiangyu Mao;Wenbo Luo;Jiaxin Liu;Guanghui Chen;Minhan Zou;Shiheng Yang;Kai Kang;Wanli Zhang;Moufu Kong;Hongshuai Zhang","doi":"10.1109/TVLSI.2026.3686072","DOIUrl":"https://doi.org/10.1109/TVLSI.2026.3686072","url":null,"abstract":"This article presents a second-order fully passive noise-shaping (NS) successive approximation register (SAR) analog-to-digital converter (ADC) with <inline-formula> <tex-math>$4times $ </tex-math></inline-formula> passive gain for low oversampling ratio (OSR) designs. By adopting the ping-pong residue extraction technique, the proposed architecture suffers no speed penalty compared with the pure SAR ADC. The merged differential residue extraction and parallel integration technique is proposed to provide a <inline-formula> <tex-math>$4times $ </tex-math></inline-formula> fully passive gain, and the parasitic capacitance confining high-order implementation is minimized by isolating the integration domain from the sampling domain. Consequently, an aggressive noise transfer function (NTF) prevailing over the ideal second-order is realized. Besides, a highly linear excessive-capacitor-reset (ECR) switching scheme is introduced. Compared with the <inline-formula> <tex-math>$V_{text {cm}}$ </tex-math></inline-formula>-based scheme, the ECR scheme reduces the integral nonlinearity (INL) by 50%, merely at the cost of a few more digital logic gates. The prototype achieves 71.3 dB SNDR with a power of <inline-formula> <tex-math>$620.2~mu $ </tex-math></inline-formula>W and a bandwidth of 1 MHz, resulting in a FoM of 163.3 dB.","PeriodicalId":13425,"journal":{"name":"IEEE Transactions on Very Large Scale Integration (VLSI) Systems","volume":"34 8","pages":"2331-2343"},"PeriodicalIF":3.2,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148627131","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
期刊
IEEE Transactions on Very Large Scale Integration (VLSI) Systems
全部 Acc. Chem. Res. ACS Applied Bio Materials ACS Appl. Electron. Mater. ACS Appl. Energy Mater. ACS Appl. Mater. Interfaces ACS Appl. Nano Mater. ACS Appl. Polym. Mater. ACS BIOMATER-SCI ENG ACS Catal. ACS Cent. Sci. ACS Chem. Biol. ACS Chemical Health & Safety ACS Chem. Neurosci. ACS Comb. Sci. ACS Earth Space Chem. ACS Energy Lett. ACS Infect. Dis. ACS Macro Lett. ACS Mater. Lett. ACS Med. Chem. Lett. ACS Nano ACS Omega ACS Photonics ACS Sens. ACS Sustainable Chem. Eng. ACS Synth. Biol. Anal. Chem. BIOCHEMISTRY-US Bioconjugate Chem. BIOMACROMOLECULES Chem. Res. Toxicol. Chem. Rev. Chem. Mater. CRYST GROWTH DES ENERG FUEL Environ. Sci. Technol. Environ. Sci. Technol. Lett. Eur. J. Inorg. Chem. IND ENG CHEM RES Inorg. Chem. J. Agric. Food. Chem. J. Chem. Eng. Data J. Chem. Educ. J. Chem. Inf. Model. J. Chem. Theory Comput. J. Med. Chem. J. Nat. Prod. J PROTEOME RES J. Am. Chem. Soc. LANGMUIR MACROMOLECULES Mol. Pharmaceutics Nano Lett. Org. Lett. ORG PROCESS RES DEV ORGANOMETALLICS J. Org. Chem. J. Phys. Chem. J. Phys. Chem. A J. Phys. Chem. B J. Phys. Chem. C J. Phys. Chem. Lett. Analyst Anal. Methods Biomater. Sci. Catal. Sci. Technol. Chem. Commun. Chem. Soc. Rev. CHEM EDUC RES PRACT CRYSTENGCOMM Dalton Trans. Energy Environ. Sci. ENVIRON SCI-NANO ENVIRON SCI-PROC IMP ENVIRON SCI-WAT RES Faraday Discuss. Food Funct. Green Chem. Inorg. Chem. Front. Integr. Biol. J. Anal. At. Spectrom. J. Mater. Chem. A J. Mater. Chem. B J. Mater. Chem. C Lab Chip Mater. Chem. Front. Mater. Horiz. MEDCHEMCOMM Metallomics Mol. Biosyst. Mol. Syst. Des. Eng. Nanoscale Nanoscale Horiz. Nat. Prod. Rep. New J. Chem. Org. Biomol. Chem. Org. Chem. Front. PHOTOCH PHOTOBIO SCI PCCP Polym. Chem.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1