source: PyTorch Blog: PyTorch 2.14 Release Blog

level: technical

pytorch 2.14 is now available, bringing a new gpu math backend called nvgemm that uses cutlass kernels with epilogue fusion and low-precision support. the release also adds an nccl2 backend for distributed training, making fault tolerance a first-class concept with in-place process-group reconfiguration and one-sided rma windows. apple silicon gains native linear algebra routines like svd, qr, and cholesky, along with a migration of more operators to hand-written metal kernels.

the release includes 2,995 commits from 487 contributors. nvgemm autotunes alongside triton and aten, and supports scaled and nvfp4 gemm. the nccl2 backend implements the full collective contract with nonblocking communicators. on apple silicon, a five-part reduction rewrite and metal kernel migration reduce per-op compilation cost and kernel launch latency. experimental support for complex-valued tensors in torch.compile is opt-in, decomposing operations into real and imaginary parts.

pytorch 2.14 builds on recent releases: 2.12 added a device-agnostic graph api, and 2.13 introduced flexattention on apple silicon and a cudedsl code path. the new release also adds torch.switch for multi-way branching and torch.while_loop capture in cuda graphs. declarative dynamic shapes via @dynamic_spec are shared across torch.compile, torch.export, and make_fx. platform support expands to rocm 7.14, intel xpu native graph capture, and nvidia rubin.

why it matters: these updates can speed up training and inference on gpus and apple silicon, and make distributed training more resilient to node failures.


source: PyTorch Blog: PyTorch 2.14 Release Blog