🏫 I’m a senior student at Zhejiang University, majoring in Computer Science and Technology.
🔭 I’m interested in Efficient AI through algorithm–system co-design, focusing on hardware-friendly sparse and quantized module design as well as efficient inference strategies.
🚀 I’m currently a leader of ZJUSCT, a super computing team at Zhejiang University which has won several international super computing competitions.
FlashQLA v0.1.1 has been released, adding intra-card sequence parallelism for the backward pass and SM100 support. Check out the GitHub repository.
TriAttention has been accepted to ICML 2026. It is an efficient KV cache compression method for long reasoning in large language models. By leveraging Q/K concentration in pre-RoPE space and a trigonometric series to score key importance, TriAttention matches Full Attention reasoning accuracy on AIME25 while achieving 2.5x higher throughput or 10.7x KV memory reduction. Check out the project page and code!
FlashQLA v0.1.0 has been released as an open-source high-performance linear attention kernel library for Qwen. Check out the GitHub repository.
A collection of my research publications and academic papers.
Selected systems and infrastructure projects around efficient AI, LLM inference, and high-performance computing.
Research intern in Qwen Infra, working on efficient inference infrastructure for Qwen.
Recipient
First Prize
Application Innovation Award
Third Prize
First Prize
Second Prize
4th Place
Second Prize
Third Prize
Visit my personal blog to read about my thoughts, experiences, and technical insights.