2026  14

August  2

One Formula, Two Jobs: How RoPE and Timestep Embedding Are Built

August 6, 2026 · 15 min · 2989 words · Yunsheng Ni

Absorbed MLA vs. Naive MLA: Same Linear FLOPs, 3.4x Core Attention

August 2, 2026 · 18 min · 3665 words · Yunsheng Ni

March  4

Progressive CUDA GEMM Optimization: From Memory-Bound to Swizzling

March 22, 2026 · 11 min · 2155 words · Yunsheng Ni

Loss Reduction in Distributed Training

March 21, 2026 · 4 min · 809 words · Yunsheng Ni

Computing Global Gradient Norm in Distributed Training: TP, DP_Shard, DP_Replicate, EP, and PP

March 21, 2026 · 5 min · 900 words · Yunsheng Ni

Demystifying FlashAttention: Forward, Backward, and Triton Implementation

March 15, 2026 · 11 min · 2338 words · Yunsheng Ni

February  5

The Devil in the Details: Engineering Tricks for SOTA Video Models

February 25, 2026 · 6 min · 1211 words · Yunsheng Ni

Deep Dive into Triton GEMM Optimization: From Naive Tiling to Hopper TMA

February 11, 2026 · 15 min · 3019 words · Yunsheng Ni

Roofline Analysis of LLMs on H200: Performance Modeling and Recomputation Strategies

February 9, 2026 · 6 min · 1217 words · Yunsheng Ni

From DDPM to Flow Matching: The Evolution of Generative Trajectories

From DiT to Hunyuan: The Evolution of adaLN-Zero in Generative Models

January  3

Beyond Theoretical FLOPs: Analyzing MFU, HFU, and Attention Overhead in Transformers

January 31, 2026 · 7 min · 1285 words · Yunsheng Ni

Visualizing 3D Attention: Bridging the Gap Between 1D Sequences and 3D Space

January 26, 2026 · 2 min · 403 words · Yunsheng Ni

GPU & Network Constants

January 25, 2026 · 2 min · 238 words · Yunsheng Ni