维普中文期刊产品整合服务
5篇 您的检索式:作者名="Luk Wayne"
    题名 作者 年代 出处 被引量
1基于混合式两阶段的动态部分重构FPGA软硬件划分算法显示文摘动态部分重构的特性大大提高了硬件设计的灵活性,但传统的软硬件划分算法不再适用于针对这类硬件的系统设计。部分研究考虑了动态部分重构的特性,并建立了混合整数线性规划(MILP)模型进行求解。但是由于MILP自身的限制,求解时间特别长,只能处理规模较小的问题。为了能够处理规模较大的问题,并且缩短求解时间,该文对MILP方法进行了详细的分析,并且通过启发式算法确定部分关键任务的状态,从而减小MILP的规模,加快求解速度。实验结果表明:与传统的数学规划方法相比,在求解质量不变的情况下,该算法可以得到最高约200倍的速度提升。马昱春 张超 Luk Wayne 2016清华大学学报(自然科学版)2016,56,3:4
2Towards efficient deep neural network training by FPGA-based batch-level parallelism显示文摘Training deep neural networks(DNNs)requires a significant amount of time and resources to obtain acceptable results,which severely limits its deployment in resource-limited platforms.This paper proposes DarkFPGA,a novel customizable framework to efficiently accelerate the entire DNN training on a single FPGA platform.First,we explore batch-level parallelism to enable efficient FPGA-based DNN training.Second,we devise a novel hardware architecture optimised by a batch-oriented data pattern and tiling techniques to effectively exploit parallelism.Moreover,an analytical model is developed to determine the optimal design parameters for the DarkFPGA accelerator with respect to a specific network specification and FPGA resource constraints.Our results show that the accelerator is able to perform about 10 times faster than CPU training and about a third of the energy consumption than GPU training using 8-bit integers for training VGG-like networks on the CIFAR dataset for the Maxeler MAX5 platform.Cheng Luo Man-Kit Sit Hongxiang Fan Shuanglong Liu Wayne Luk Ce Guo 2020Journal of Semiconductors2020,41,2:2
3A flexible hardware encoder for low-density parity-check codes 显示文摘LEE Dong-U WAYNE Luk WANG Connie Proceedings of the0,12,:1
4Gaussian random number generators显示文摘David B Thomas Wayne Luk Philip H W Leong 2007ACM Computing Surveys2007,39,4:1
5A fully-customized dataflow engine for 3D earthquake simulation with a complex topography显示文摘With HPC(high performance computing)evolving into the exascale era,improvements in computing performance and power efficiency have become increasingly more important.Based on our previous work on enabling earthquake simulations on a large scale on Sunway TaihuLight,we further explore other possibilities to improve the application through a fully-customized hardware design on reconfigurable FPGA(field programmable gate array)devices.We investigate the feasibility and the potential benefits of a complete fixed-point design.We first perform a coarse-resolution-based simulation to analyze the representation range and precision needed to capture both the total energy and the energy distribution of variables over space and time.We then derive a complete fixed-point design that identifies the suitable bitwidth for major categories of variables and dynamically represents the range through a dynamic scaling scheme.Finally,we use the optimized fixed-point design to run a case of the Wenchuan earthquake to demonstrate the potential of supporting large-scale scientific simulations on FPGA devices.The results demonstrate that an 18-bit fixed-point design already provides an almost identical description of the seismic events in the Wenchuan scenario down to a single-precision floating-point version and provides sustainable performance equivalent to 13.1 Intel Xeon Gold 615418-core CPUs or 2.10 Sunway 260-core processors,with performance per watt(power efficiency)improved by 15.3 and 3.72 times compared with the Intel Xeon Gold 615418-core CPUs and the Sunway 260-core processors,respectively.Bingwei CHEN Haohuan FU Wayne LUK Guangwen YANG 2022Science China(Information Sciences)2022,65,5:0
返回顶部 每页显示:
共1页 首页 上一页 第1页 下一页 末页 /1 跳转

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费