publications

Reverse-chronological. * denotes equal contribution.

2026

  1. CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution
    Zihao Ye, Yingyi Huang, Hongyi Jin, and 11 more authors
    arXiv preprint arXiv:2608.12629, 2026
  2. CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents
    Zhongming Yu, Hengjia Yu, Boqin Yuan, and 12 more authors
    arXiv preprint arXiv:2607.25431, 2026
  3. SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution
    Gangda Deng, Zhaoling Chen, Zhongming Yu, and 11 more authors
    In Proceedings of the 43rd International Conference on Machine Learning, 2026
  4. LLM4Cov: Execution-Aware Agentic Learning for High-Coverage Testbench Generation
    Hejia ZhangZhongming Yu, Chia-Tung Ho, and 3 more authors
    In Proceedings of the 43rd International Conference on Machine Learning, 2026
  5. PRO-V-R1: Reasoning Enhanced Programming Agent for RTL Verification
    Yujie Zhao, Zhijing Wu, Boqin Yuan, and 6 more authors
    In Proceedings of the 63rd ACM/IEEE Design Automation Conference, 2026
  6. AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
    Yujie Zhao, Boqin Yuan, Junbo Huang, and 9 more authors
    In Proceedings of the 43rd International Conference on Machine Learning, 2026
  7. Multi-Agent Memory from a Computer Architecture Perspective: Visions and Challenges Ahead
    Zhongming Yu, Naicheng Yu, Hejia Zhang, and 5 more authors
    arXiv preprint arXiv:2603.10062, 2026
    Position paper. Featured on ACM SIGARCH
  8. Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs
    Wentao Ni, Kangqi Zhang, Zhongming Yu, and 7 more authors
    arXiv preprint arXiv:2602.05191, 2026

2025

  1. OrcaLoca: An LLM Agent Framework for Software Issue Localization
    Zhongming YuHejia ZhangYujie Zhao, and 4 more authors
    In Proceedings of the 42nd International Conference on Machine Learning, 2025
  2. MAGE: A Multi-Agent Engine for Automated RTL Code Generation
    Yujie ZhaoHejia Zhang, Hanxian Huang, and 2 more authors
    In Proceedings of the 62nd ACM/IEEE Design Automation Conference, 2025

2024

  1. GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
    Zhongming Yu, Genghan Zhang, Hanxian Huang, and 2 more authors
    arXiv preprint arXiv:2404.03019, 2024

2023

  1. TorchSparse++: Efficient Training and Inference Framework for Sparse Convolution on GPUs
    Haotian Tang, Shang Yang, Zhijian Liu, and 6 more authors
    In Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture, 2023
  2. HyperGef: A Framework Enabling Efficient Fusion for Hypergraph Neural Network on GPUs
    Zhongming YuGuohao Dai, Shang Yang, and 6 more authors
    Proceedings of Machine Learning and Systems, 2023
  3. Exploiting Hardware Utilization and Adaptive Dataflow for Efficient Sparse Convolution in 3D Point Clouds
    Ke Hong*, Zhongming Yu*Guohao Dai, and 4 more authors
    Proceedings of Machine Learning and Systems, 2023
    * stands for equal contribution
  4. CogDL: A Comprehensive Library for Graph Deep Learning
    Yukuo Cen, Zhenyu Hou, Yan Wang, and 8 more authors
    In Proceedings of the ACM Web Conference 2023, 2023
  5. CLAP: Locality Aware and Parallel Triangle Counting with Content Addressable Memory
    Tianyu Fu, Chiyue Wei, Zhenhua Zhu, and 5 more authors
    In 2023 Design, Automation & Test in Europe Conference & Exhibition (DATE), 2023
  6. Sgap: Towards Efficient Sparse Tensor Algebra Compilation for GPU
    Genghan Zhang, Yuetong Zhao, Yanting Tao, and 6 more authors
    CCF Transactions on High Performance Computing, 2023

2022

  1. Understanding GNN Computational Graph: A Coordinated Computation, IO, and Memory Perspective
    Hengrui Zhang*, Zhongming Yu*Guohao Dai, and 4 more authors
    Proceedings of Machine Learning and Systems, 2022
    * stands for equal contribution
  2. Heuristic Adaptability to Input Dynamics for SpMM on GPUs
    Guohao Dai, Guyue Huang, Shang Yang, and 6 more authors
    In Proceedings of the 59th ACM/IEEE Design Automation Conference, 2022
  3. FastFold: Reducing AlphaFold Training Time from 11 Days to 67 Hours
    Shenggan Cheng, Xuanlei Zhao, Guangyang Lu, and 7 more authors
    arXiv preprint arXiv:2203.00854, 2022
  4. Benchmarking GNN-Based Recommender Systems on Intel Optane Persistent Memory
    Yuwei Hu, Jiajie Li, Zhongming Yu, and 1 more author
    arXiv preprint arXiv:2207.11918, 2022

2021

  1. Exploiting Online Locality and Reduction Parallelism for Sampled Dense Matrix Multiplication on GPUs
    Zhongming YuGuohao Dai, Guyue Huang, and 2 more authors
    In 2021 IEEE 39th International Conference on Computer Design (ICCD), 2021