TLX Block Attention: A Warp-Specialized Blackwell Kernel for Fixed-Block Sparse Self-Attention

2026-05-26 04:26 GMT · 2 months ago aimagpro.com

Code available at: https://github.com/facebookresearch/ads_model_kernel_library  In this post, we present the design of TLX Block Attention — a Triton kernel targeting NVIDIA Blackwell GPUs that exploits compile-time knowledge of a block-diagonal…