Projects
Selected systems and infrastructure projects around efficient AI, LLM inference, and high-performance computing.
Project • April 2026
A high-performance linear attention kernel library for Qwen, built on TileLang. FlashQLA optimizes GDN Chunked Prefill with fused kernels, gate-driven context parallelism, and hardware-friendly reformulation.