From DiT to Hunyuan: The Evolution of adaLN-Zero in Generative Models

From DiT to Hunyuan Video, adaLN-Zero remains the gold standard for conditioning. Here’s how this zero-initialized module works and why it persists in the era of Flow Matching.

Beyond Theoretical FLOPs: Analyzing MFU, HFU, and Attention Overhead in Transformers

A rigorous breakdown of FLOPs in Llama-style architectures, deriving the relationship between linear projections and quadratic attention overhead, with insights into sample packing efficiency.

Visualizing 3D Attention: Bridging the Gap Between 1D Sequences and 3D Space

An interactive tool to visualize the mapping between 1D token sequences and 3D (T, H, W) sliding windows.

GPU & Network Constants

A quick reference of Dense FLOPS and Unidirectional Bandwidth for A100, H100, H200, and Blackwell.