AI Infra
Systems, not recipes. If it is a cluster, a kernel, a serving engine, or a cost/SLO problem, it lives here.
The World Model Platform is the end-to-end case study.
- Performance — rooflines, accelerators, profiling
- Training systems — parallelism, checkpointing, elasticity
- Inference systems — batching, KV cache, disaggregation
- Cluster & platform — Kubernetes, schedulers, GitOps
- Data plane — storage, pipelines, lineage
- Kernels & compilers — Triton, CUDA, LLVM
- Deployment — rollout, serving paths