Qijun Zhang, Chen Zhang, Zhuoshan Zhou, Haibo Wang, Zhe Zhou, Zhipeng Tu, Guangyu Sun, Zhiyao Xie, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He, Minyi Guo
(2026).
Tackling MoE Communication Bottleneck with Dynamic In-Switch Computing on Multi-GPUs.
ISCA 2026.
Z. He, Y. Wang, Z. Yue, Z. Wu, H. Han, S. Wei, Y. Hu, F. Tu, S. Yin
(2026).
HR-DCIM: High-Reliability Floating-Point Digital CIM Architecture with Unified Low-Cost Iterative Error Correction.
International Symposium on High-Performance Computer Architecture (HPCA), Sydney, Australia.
R. Guo, L. Wang, X. Chen, K. Jiang, Z. Yue, H. Han, Y. Wang, F. Tu, S. Wei, Y. Hu, S. Yin
(2026).
Denim: Heterogeneous Compute-in-Memory Accelerator Exploiting Denoising-similarity for Diffusion Models.
IEEE Journal of Solid-State Circuits (JSSC).
Jinming Ge, Linfeng Du, Likith Anaparty, Shangkun Li, Tingyuan Liang, Afzal Ahmad, Vivek Chaturvedi, Sharad Sinha, Zhiyao Xie, Jiang Xu, Wei Zhang
(2026).
DAPO: Design Structure Aware Pass Ordering in High-Level Synthesis with Graph Contrastive and Reinforcement Learning.
DATE 2026.
J. Chen, Z. Xuan, Y. Li, Z. Yan, Z. Li, Y-F. Yung, Y. Deng, X. Huo, C-Y. Tsui, K-T. Cheng, F. Tu
(2026).
D2CIM: A 28nm 53.3 TFLOPS/W Decoding Digital CIM Macro for Efficient FlashMLA-based LLM Inference.
Custom Integrated Circuits Conference (CICC), Seattle, USA.
Z. Chen, Z. Li, L. Liang, X. Fan, X. Wang, Q. Liu, Z. Gu, Y. Lu, Y. Zhao, P. Qiu, Y. Xie
(2026).
ALOHA: Accelerating Leveled Fully Homomorphic Encryption with Cryptography-Specific Architectures.
ACM Transactions on Architecture and Code Optimization 2026.
C. Li, C. Xue, Y. Ren, X. Dong, Y. Cheng, Y. Hu, F. Bai, Y. Guo, X. Jiang, Q. Wu, Y. Xie
(2026).
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators.
arXiv preprint (2026).
P. Dong, Y. Tan, X. Liu, P. Luo, Y. Liu, D. Pang, S. Ma, X. Huang, S-Y. Liu, D. Zhang, Z. Lu, L. Liang, C-Y. Tsui, F. Tu, L. Zhao, K-T. Cheng
(2026).
A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding.
International Solid-State Circuits Conference (ISSCC), San Francisco, USA.
J. Deng, X. Tang, J. Zhang, Y. Li, L. Zhang, B. Han, H. He, F. Tu, S. Wei, Y. Hu, S. Yin
(2025).
Rethinking Control Flow in Spatial Architectures: Insights into Control Flow Plane Design.
IEEE Transactions on Computers (TC).
J. Chen, T-C. Pang, Y-F. Yung, Y. Deng, A. Yin, X. Huo, L. Liang, Z. Wang, C-Y. Tsui, K-T. Cheng, F. Tu
(2025).
PipeDCIM: 28nm 115.11 TOPS/mm2xTOPS/W@1.24GHz Pipeline Digital CIM Macro with an Auto-Design Tool for Diverse High-Performance AI Scenarios.
European Solid-State Electronics Research Conference (ESSERC), Munich, Germany.
Z. He, Z. Wu, R. Guo, L. Yan, H. Han, Y. Wang, S. Wei, Y. Hu, F. Tu, S. Yin
(2025).
LLM-CIM: A 28nm 126.7 TOPS/W Input-LUT-based Digital CIM Macro with Reconfigurable Matrix Multiplication and Nonlinear Operation Modes for LLMs.
Symposium on VLSI Technology and Circuits (VLSI), Kyoto, Japan.
C. Li, Y. Yin, X. Wu, J. Zhu, Z. Gao, D. Niu, Q. Wu, X. Si, Y. Xie, C. Zhang, G. Sun
(2025).
H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference.
ISCA 2025.
Y. Wu, X. Wang, J. Chen, Z. Zhu, J. He, P. Dong, Y. Tan, X. Zhao, L. Chang, Y. Wang, F. Tu, C-Y. Tsui, K-T. Cheng
(2025).
Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design Optimization.
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD).
S. Ma, J. Wu, Y. Tan, P. Dong, P. Luo, D. Pang, Y. Liu, X. Liu, L. Liang, K-T. Cheng, F. Tu
(2025).
CoXplorer: Multi-staged Co-exploration Framework for AI Model Compression and Accelerator Design.
International Conference on Computer Aided Design (ICCAD), Munich, Germany.
Jian Weng, Boyang Han, Derui Gao, Ruijie Gao, Wanning Zhang, An Zhong, Ceyu Xu, Jihao Xin, Yangzhixin Luo, Lisa Wu Wills, Marco Canini
(2025).
Assassyn: A Unified Abstraction for Architectural Simulation and Implementation.
ISCA 2025.
P. Dong, Y. Tan, X. Liu, P. Luo, Y. Liu, L. Liang, Y. Zhou, D. Pang, M. Yung, D. Zhang, X. Huang, S-Y. Liu, Y. Wu, F. Tian, C-Y. Tsui, F. Tu, K-T. Cheng
(2025).
A 28nm 0.22uJ/Token Memory-Compute-Intensity-Aware CNN-Transformer Accelerator with Hybrid-Attention-Based Layer-Fusion and Cascaded Pruning for Semantic-Segmentation.
International Solid-State Circuits Conference (ISSCC), San Francisco, USA.