Projects

Selected systems and infrastructure projects around efficient AI, LLM inference, and high-performance computing.

FlashQLA: Flash Qwen Linear Attention

FlashQLA: Flash Qwen Linear Attention

ProjectApril 2026

A high-performance linear attention kernel library for Qwen, built on TileLang. FlashQLA optimizes GDN Chunked Prefill with fused kernels, gate-driven context parallelism, and hardware-friendly reformulation.
Nifty tech tag lists from Wouter Beeftink