About

I am a Ph.D. student at the University of Washington, advised by Prof. Ang Li in the Pᴺ Computer Engineering Lab. I research energy-efficient computer architectures for AI and general-purpose computing, with a focus on reconfigurable accelerators and GPU-like parallel systems. My work spans architecture and compiler design, performance and power modeling, RTL implementation, and silicon tape-out.

My research has been published in or accepted to leading conferences in computer architecture and reconfigurable computing, including ISCA, MICRO, ICCD, and FCCM.

I received my B.S. in Electrical Engineering from Shanghai Jiao Tong University and my M.S. from the University of Washington. I have interned at Intel and Qualcomm, and am currently a Compute Performance Intern at NVIDIA.

Summer 2027: I am seeking research internships in GPU architecture, AI accelerators, and performance modeling.

News

Earlier updates

Publications

My name is highlighted in bold. Full CV (PDF) · Code on GitHub

First-Author Publications & Workshops

  1. DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays

    Jiayi Wang, Ang Da Lu, Zhichen Zeng, and Ang Li

    ISCA 2026 Int'l Symposium on Computer Architecture, Raleigh, NC · Jun–Jul 2026 · pp. 2648–2663 · 19.1% acceptance

  2. MFSA: A Multi-Format Systolic Array with Native Micro Scaling Support

    Jiayi Wang, Kearnan Bishop, and Ang Li

    ICCD 2026 44th IEEE Int'l Conference on Computer Design, Hong Kong · Nov 2026 · 26% acceptance

  3. TransDot: An Area-Efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines

    Jiayi Wang, Maohua Nie, Sin-Chen Lin, C.-J. Richard Shi, and Ang Li

    FCCM 2026 IEEE Int'l Symp. on Field-Programmable Custom Computing Machines, Atlanta, GA · May 2026 · 24.6% acceptance

  4. DICE: Efficient Thread-Pipelined SIMT Execution on a Statically Scheduled Reconfigurable Spatial Fabric

    Jiayi Wang, Ang Da Lu, and Ang Li

    GPGPU 2026 18th Workshop on General Purpose Processing Using GPU, w/ ASPLOS 2026, Pittsburgh, PA · Mar 2026

  5. piPE-SA: Enabling Deeply Pipelined Processing Elements in Systolic Arrays

    Jiayi Wang, Chenyi Wang, and Ang Li

    EMC²-11 Energy Efficient ML and Cognitive Computing Workshop, w/ ASPLOS 2026, Pittsburgh, PA · Mar 2026

  6. Hybrid Aerial-Aquatic Vehicle for Large-Scale High-Spatial-Resolution Marine Observation

    Jiayi Wang, Yiwei Yang, Jiajin Wu, Zheng Zeng, Di Lu, and Lian Lian

    OCEANS 2019 Marseille · pp. 1–7

Collaborative Publications

  1. Step-dLLM: Adaptive Step-aware Sparse Attention for Efficient Diffusion LLM Inference

    Zhichen Zeng, Xichong Zhang, Junpan Wu, Yifei Zuo, Chi-Chih Chang, Jiayi Wang, Maohua Nie, Ji Liu, Ang Li, and Banghua Zhu

    NeurIPS 2026 Conference on Neural Information Processing Systems · Accepted

  2. HierSVA: A Data Synthesis Pipeline, Dataset, and Benchmark for LLM-Driven Hierarchical Hardware Formal Verification

    Maohua Nie, Jiang Zhu, Jingqun Zhang, Zhichen Zeng, Jiayi Wang, Sibo Zhang, Jialin Wang, and C.-J. Richard Shi

    NeurIPS 2026 Conference on Neural Information Processing Systems · Accepted

  3. CacheFlex: Direct Software-Managed Access to Higher-Level Cache for Scalable Vector Support

    Jingqun Zhang, Maohua Nie, Jiayi Wang, Weihang Li, Rishi Sappidi, Shwet Chitnis, and Ang Li

    MICRO 2026 59th IEEE/ACM Int'l Symposium on Microarchitecture · ~25% acceptance

  4. All-Digital Bluetooth Low Energy (BLE) Backscatter ASIC Using Standard I/O Pad Drivers in 180nm CMOS

    Ryan H. Lee, Kate L. Tseng, Te Min Yu, Andrew Pan, Jiayi Wang, James Rosenthal, Kevin J. Ho, Ang Li, and Matthew Reynolds

    RFID 2026 IEEE Int'l Conference on RFID

  5. CacheFlex: Explicitly Control What You Need in Your Cache

    Jingqun Zhang, Weihang Li, Maohua Nie, Yung-Jen Cheng, Jiayi Wang, Shwet Chitnis, and Ang Li

    EMC²-11 w/ ASPLOS 2026, Pittsburgh, PA · Mar 2026

  6. STEP: Spatially Threaded Execution Pipeline

    Ang Da Lu, Jiayi Wang, and Ang Li

    LATTE 2026 6th Workshop on Languages, Tools, and Techniques for Accelerator Design, w/ ASPLOS 2026

  7. Precision-Aware Communication in CGRAs

    Shwet Chitnis, Fergus Xu, Ayush Kulkarni, Jiayi Wang, Jingqun Zhang, Arjun Raje, and Ang Li

    FCCM 2026 IEEE Int'l Symp. on Field-Programmable Custom Computing Machines

  8. DORA: Open-Source Infrastructure for Prototyping Reconfigurable Fabrics

    Shwet Chitnis, Jiayi Wang, Fergus Xu, Ayush Kulkarni, Rampranav Navendran, Juwon Jun, Jingqun Zhang, and Ang Li

    OSCAR 2026 Workshop on Open-Source Computer Architecture Research, w/ ISCA 2026, Raleigh, NC · Jun 2026 · poster

Patents

  1. DICE: Enabling Efficient General-Purpose SIMT Execution with Coarse-Grained Reconfigurable Arrays

    Ang Li and Jiayi Wang

    Patent U.S. Provisional Application 64/012,051 · filed Mar 20, 2026

  2. An Aerial-Aquatic Vehicle Combining a Fixed-Wing Aircraft and a Glider

    Zheng Zeng, Jiayi Wang, Yiwei Yang, Jiajin Wu, Lian Lian, and Di Lu

    Patent Chinese Invention Patent CN 109204812 B · granted Nov 17, 2020

Selected Research

Reconfigurable AI Hardware & Accelerator Microarchitecture Research

UW Pᴺ Computer Engineering Lab · PI Prof. Ang Li · Aug 2025 – Present

I design reconfigurable arithmetic units and systolic arrays for AI workloads with diverse numerical formats and matrix shapes. My work spans shared datapaths for mixed-precision computation, native block-scaling support, and flexible arrays for small, non-square GEMMs in LLM serving.

  • TransDot — reconfigurable arithmetic: developed a unified FPU supporting FP4/FP8/FP16/FP32 fused multiply-add, subword SIMD, and dot products with FP16/FP32 accumulation, replacing separate per-format FPUs in FPGA AI engines.
  • MFSA — multi-format systolic arrays: designed a weight-stationary array with a shared dot-product datapath supporting formats from FP32 to FP4/INT4. Introduced online block-scale rebasing to incorporate MX/NVFP4 scales directly into partial sums and reduce block-scaling overhead.
  • piPE-SA — pipelined processing elements: developed timing support for multi-stage PEs in weight- and output-stationary arrays, enabling exploration of PE pipeline depth alongside array size and dataflow.
  • Flexible arrays (ongoing): designing a twisted-torus fabric with an orthogonal wavefront to eliminate diagonal data skew and support runtime partitioning into independently scheduled strips, targeting higher utilization on small, non-square GEMMs.
Related publications: MFSA ICCD 2026 · TransDot FCCM 2026 · piPE-SA EMC²-11 · 1 paper under review

High-Efficiency CGRA–GPU/CPU Integration Research

UW Pᴺ Computer Engineering Lab · PI Prof. Ang Li · Oct 2024 – Present

I combine statically scheduled CGRAs with SIMT execution to reduce register-file traffic and instruction-control overhead while preserving the CUDA programming model. This work connects architecture design, compiler development, performance and power modeling, and RTL/FPGA prototyping.

  • DICE — GPU architecture: designed a CGRA-based replacement for the SM's SIMD backend, with a Temporal Memory Coalescing Unit for pipelined memory access and swizzled thread-bank mapping to reduce register-file bank conflicts.
  • Performance & energy: achieved 1.90× energy efficiency, 42% lower power, and 1.16× performance relative to a modeled NVIDIA Turing SM baseline in GPU-Rodinia evaluation using GPGPU-Sim/GPUWattch. DICE is the subject of a U.S. provisional patent application.
  • Compiler & implementation: built a p-graph compiler, contributed to STEP for mapping SIMT kernels onto the spatial fabric, and extended GPGPU-Sim/GPUWattch for evaluation. Led an 8-member RTL team developing the RTL implementation and FPGA prototype for architectural validation.
  • CPU integration (ongoing): developing a core-private SIMT engine connected to each CPU core's cache hierarchy, with an in-memory thread-launch mechanism to support fine-grained irregular kernels without driver launches or explicit data transfers.
Related publications: DICE ISCA 2026 · DICE (thread-pipelined SIMT) GPGPU 2026 · STEP LATTE 2026 · 1 paper under review

Coarse-Grained Reconfigurable Architecture Framework Research

UW Pᴺ Computer Engineering Lab · PI Prof. Ang Li · Mar 2024 – Present

I develop reusable infrastructure for generating and prototyping CGRAs, enabling architecture exploration through configurable processing elements, interconnects, and control logic.

  • RTL generation: built a parameterized Python framework for customizing processing elements, routing networks, and control flow; used across the lab's CGRA projects for architecture exploration and FPGA prototyping.
  • Precision-aware communication: investigated how operand precision informs data movement and interconnect design in CGRAs.
Related publications: DORA OSCAR 2026 · Precision-Aware Communication in CGRAs FCCM 2026

GPU Multi-Tenant Power Management Research

UW Pᴺ Computer Engineering Lab · PI Prof. Ang Li · 2026 – Present

I investigate how shared GPU power limits create interference between co-located LLM workloads, and how per-tenant power management can protect latency-sensitive inference.

  • Hardware characterization: measured power-induced interference on H100/B200-class GPUs, examining how a co-located workload increases a prioritized tenant's time-to-first-token under a shared power cap.
  • Architecture proposal: proposed per-partition power metering and independent per-GPC voltage/frequency control, coordinated by a power arbiter that reallocates budgets at sub-millisecond intervals based on tenant priority, execution phase, and unused capacity.
  • Simulation results: evaluated the design using real LLM traces and a serving simulator with power models calibrated on H100 and B200. In simulation, the design restored the prioritized tenant's latency to near its standalone value.

AI Agent Development

SABOL Vision Lab, Cultivate Learning · PI Prof. Gail E. Joseph · Feb 2026 – Present

I build multimodal AI agents that analyze classroom video to assess teacher–child interactions and support real-time teacher coaching. VISTA (Video Interaction Scoring & Teaching Assistant) combines video-based assessment with live audio/video analysis.

  • VISTA-QEI: developed a video-analysis agent that scores teacher–child interactions using the QEI framework and generates coaching reports grounded in video evidence.
  • VISTA-RT: designed a real-time coaching agent integrating live audio/video analysis, rolling context memory, and prioritized prompts delivered through wearable interfaces.
  • Prototyping & evaluation: implemented working demos and evaluation tools, and presented prototypes to collaborators in AI and early-childhood education.

High-Speed Interface Integration Research

UW Silicon Systems Research Lab · PI Prof. C.-J. Richard Shi · Mar 2022 – Dec 2023

I led RTL design and verification for CPHY, a DDR/UCIe-compatible PHY and LPDDR4/4x memory controller, and carried out physical design and SoC integration for a TSMC N28 tape-out.

  • Design & verification: led RTL development, mixed-signal integration, and verification of the memory controller and PHY.
  • Physical design: performed TSMC N28 physical implementation, including analog/mixed-signal integration, through tape-out.
  • SoC integration: integrated the controller and PHY with a RISC-V CPU and a deep-learning accelerator.

Research Interests

AI & General-Purpose Architecture

Computer architecture for AI and general-purpose workloads across CPUs, GPGPUs, and accelerators.

Reconfigurable & Spatial Computing

Reconfigurable and spatial/dataflow architectures — FPGA and CGRA — for efficient, flexible execution.

Memory & Cross-Stack Co-Design

Memory systems, cross-stack co-design, and multimodal AI systems for real-world applications.

Education

Sep 2023 – Jun 2028 (expected)
Ph.D. in Computer Engineering — University of Washington, Seattle

GPA 4.00 · Computer architecture, reconfigurable AI hardware architecture, AI systems, digital VLSI · Advisor: Prof. Ang Li

Sep 2020 – Aug 2022
M.S. in Electrical Engineering — University of Washington, Seattle

GPA 4.00 · Computer architecture, digital VLSI

Sep 2016 – Jun 2020
B.S. in Electrical Engineering — Shanghai Jiao Tong University

GPA 3.70 · Power electronics, embedded systems

Summer 2019
Summer School, Computer Science — UC Los Angeles

GPA 4.00

Industry Experience

Summer 2026
Compute Performance Intern — NVIDIA Corporation

Working on GPU acceleration of Cadence EDA tools.

Jun 2024 – Sep 2024
Research Intern, GPGPU — Qualcomm Technologies

Investigated architectural optimizations in commercial GPGPU/NPU platforms for sparse matmul; explored sparse flows in TensorRT, PyTorch, and CUTLASS using NVIDIA 2:4 Sparse Tensor Cores; benchmarked across ResNet, MobileNet, and BERT.

Oct 2022 – Dec 2023
Technical Staff — Orora Design Technologies

Developed C++ automation tools for automatic SystemVerilog model generation of analog circuits.

Mar 2021 – Sep 2021
Technical Enablement Engineer Intern — Intel Corporation

Debugged and verified DDR and PCIe high-speed interfaces on IA-32 platform products.

Aug 2019 – Feb 2020
Hardware Engineer Intern — Signify (Philips Lighting)

Worked on a dual-mode self-adaptive LED/HID driver circuit for the EU market.

Silicon & Chip Photos

A selection of die shots from tape-outs and prototypes I've contributed to.

Teaching, Mentoring & Service

Graduate Teaching Assistant — University of Washington (Jan 2024 – Present)

Autumn 2026 Current
EE538: Post-Tapeout Chip Test
Spring 2026
EE478/526: VLSI Capstone
Autumn 2025
EE478/526 (cont.): VLSI Chip Test

Designed customized PCB test boards for taped-out chips.

Spring 2025
EE478/526: VLSI Capstone

Developed a complete TSMC 180nm tape-out flow with the Hammer flow and custom TCL plugins — Genus synthesis, Innovus P&R, Tempus STA, Calibre DRC/LVS, padframe generation, and hierarchical flows with extracted timing models. Mentored 5 student tape-out projects.

Autumn 2024
EE476: VLSI-I
Spring 2024
EE271: Digital Circuits and Systems
Winter 2024
EE477/525: VLSI-II

Service & Mentoring

Dec 2025 – Present
Technical Member — UW Semiconductor Society

Contributed to student community activities in semiconductor design and systems research.

Winter 2025
M.S. Admissions Triage Committee — University of Washington

Reviewed master's applicant materials and provided evaluation feedback.

Expertise & Skills

Architecture & Modeling

AI hardware; CPU/GPGPU/FPGA/CGRA architecture; GPGPU-Sim/GPUWattch (Accel-Sim/AccelWattch); gem5.

ASIC Design

End-to-end design, tape-out & mentoring in TSMC 180nm and N28 — RTL, verification, synthesis, P&R, STA, signoff (DRC/LVS/PEX).

EDA Tools

Cadence (Xcelium, Genus, Innovus, Tempus), Synopsys, Mentor/Calibre.

Programming

CUDA, Python, C++, SystemVerilog; strong debugging on complex codebases.

Contact

For research collaborations or Summer 2027 internship opportunities, please contact me by email.