KV caching, FlashAttention, and continuous batching - how a trained model actually generates tokens fast.
MLA: Compressing the KV Cache Without Decompressing It
llm-systems
inference
attention
Speculative Decoding: Exact Sampling Is Only Half the Problem
llm-systems
inference
sampling
The Weight Error Is Not the Error That Matters
quantization
numerics
llm-systems
One Outlier Sets Everyone’s Precision
quantization
numerics
llm-systems
Continuous Batching: The GPU Was Underfilled, Not Underpowered
llm-systems
inference
serving
What Ordinary Autograd Saves: Tiling the Forward Pass Is Only Half the Algorithm
llm-systems
training
attention
autograd
FlashAttention: The Matrix You Never Have to Write Down
llm-systems
inference
attention
No matching items


