torchtitan
A PyTorch-native platform for pretraining and post-training generative AI models at scale
Quick start · Models · Features · RL · Docs · Contributing
torchtitan is a clean-room, PyTorch-native implementation of large-scale
training. One readable codebase takes a model from pretraining through
supervised fine-tuning to reinforcement learning, composing FSDP, tensor,
pipeline, context, and expert parallelism without rewriting the model.
- Composable parallelism. Models declare how they shard; FSDP2, TP, PP, CP,
and EP stack into N-D layouts through
DeviceMeshand DTensor. - Built for research. Minimal, hackable code; every run is a Python recipe you can read, copy, and modify.
- Pretraining to RL in one stack. TitanRL runs the same model definitions and kernels in the trainer and in vLLM generation, with a bitwise-reproducible batch-invariant mode.
- Frontier models and techniques. DeepSeek, Qwen, Kimi, gpt-oss, Llama, and FLUX, with MXFP8/NVFP4 low precision, Dist-MoE, CUDA graphs, and more.
[!NOTE]
torchtitanis under active development. To use the latest features, use a recent PyTorch nightly.
Latest news
- [2026/08] TitanRL is a hackable RL stack for scaling and debugging. It reuses TorchTitan model definitions and kernels across training and vLLM generation and supports batch-invariant mode.
- [2025/11] AMD released an optimized fork of
torchtitanfor AMD GPUs. - [2025/10] We released
torchtitanv0.2.0. - [2025/10] SkyPilot now supports
torchtitan! See the tutorial.
Older news
- [2025/07] We published instructions on how to add a model to
torchtitan. - [2025/04] Our paper was accepted by ICLR 2025.
- [2024/12] GPU MODE lecture on torchtitan.
- [2024/07] Presentation at PyTorch Conference 2024.
Quick start
1. Install with a PyTorch nightly (Python 3.11+). Replace cu132 with your CUDA or ROCm build.
git clone https://github.com/pytorch/torchtitan && cd torchtitan
pip install --pre torch --index-url https://download.pytorch.org/whl/nightly/cu132
pip install -r requirements.txt
2. Train a debug model. It uses the tokenizer bundled with the tests, so no download is needed.
NGPU=1 ./run_train.sh # Llama 3 debug model, 10 steps
3. Train a real model. Download the tokenizer (Llama weights require access on Hugging Face), then select a recipe by module and function name:
python scripts/download_hf_assets.py --repo_id meta-llama/Llama-3.1-8B --assets tokenizer --hf_token=...
MODULE=torchtitan_recipes.models.llama3 CONFIG=llama3_8b ./run_train.sh # 8 GPUs
4. Make it yours. A recipe is a Python function returning a complete config, so changing a run means writing a function, not passing flags:
# my_recipes.py
from torchtitan_recipes.models.llama3 import llama3_8b
def llama3_8b_tp2():
config = llama3_8b()
config.parallelism.tensor_parallel_degree = 2
return config
MODULE=my_recipes CONFIG=llama3_8b_tp2 ./run_train.sh
See the configuration guide for the full model,
and torchtitan_recipes for the verified recipes.
Supported models
| Model | Sizes | Modality | Verified recipes | RL |
|---|---|---|---|---|
| Llama 3 | 1B – 405B | Text | llama3_8b, llama3_70b |
✅ |
| Qwen3 | 0.6B – 32B dense; 30B-A3B, 235B-A22B MoE | Text | qwen3_14b |
✅ |
| Qwen3.5 / 3.6 / 3.8 | 0.8B – 397B-A17B; 2.4T-A95B (3.8) | Text, vision | ✅ text | |
| DeepSeek V3 | 16B, 236B, 671B | Text | deepseek_v3_671b (+ Dist-MoE, NVFP4) |
|
| DeepSeek V4 | Flash, Pro | Text | Soon | |
| gpt-oss | 20B, 120B | Text | ✅ | |
| Kimi K2.7 | Moonlight-16B-A3B, Kimi-VL-A3B, K2.5 | Text, vision | ||
| Kimi K3 | K3 | Text, vision | Soon | |
| Muse Glimmer | 30B | Text, vision | muse_glimmer_30b |
✅ |
| FLUX | dev, schnell | Text-to-image | flux_dev, flux_schnell |
Every model has debug configurations exercised in CI. Verified recipes have been run in training, convergence, or performance workflows. To add a model, see the model guide.
Features
Parallelism
- FSDP2, HSDP, and DDP
- Tensor parallel, including async TP
- Pipeline parallel: 1F1B, Interleaved 1F1B, GPipe, and other
torch.distributed.pipeliningschedules - Context parallel: all-gather KV and Ulysses
- Expert parallel with DeepEP / HybridEP dispatch
- Dist-MoE: fused dispatch, expert compute, and combine on Blackwell
Memory and performance
- Activation checkpointing: full, selective, and regional
- Regional
torch.compileand CUDA graphs - MXFP8 and NVFP4 training
- FlashAttention 3/4, FlexAttention, and varlen attention
- BF16 optimizer states and FSDP CPU offload
Training
- Grain data pipeline with packing, mixing, and checkpointable state
- Pretraining, SFT on chat datasets, and LoRA
- Muon (DistMuon), EMA, multi-token prediction
- Warmup-stable-decay LR schedule and validation
Post-training with TitanRL
- One model definition shared by the trainer and vLLM
- Bitwise trainer/generator parity in batch-invariant mode
- Async rollouts orchestrated with Monarch; weight sync over TorchStore
- GRPO and DAPO losses; DAPO Math, Search-R1, and Verifiers examples
Checkpointing
- Distributed checkpointing, including async saves
- Load and save Hugging Face safetensors directly
- Conversion scripts between HF and DCP
Debugging and observability
- Metrics: loss, memory, throughput, TFLOPs, and MFU in TensorBoard or W&B
- Profiling, memory snapshots, and Flight Recorder
- Fake-backend dry runs of large layouts on one GPU
- Deterministic mode and loss comparison across commits
- SDC replay and structured logging
Extensibility
- Python recipes selected with
--moduleand--config - Component overrides for custom kernels and modules
- Experiments: GraphTrainer, TorchFT, Transformers modeling backend
We report performance on up to 512 GPUs and verify loss convergence across parallelisms and techniques.
Test status
| Hardware | Integration tests | Unit tests |
|---|---|---|
| CPU | ||
| NVIDIA GPU | ||
| AMD GPU (ROCm) | ||
| RL (NVIDIA GPU) |
Installation
The quick start installs from source, which is the recommended way to get the latest features. Other options:
Nightly builds
pip install --pre torch --index-url https://download.pytorch.org/whl/nightly/cu132 --force-reinstall
pip install --pre torchtitan --index-url https://download.pytorch.org/whl/nightly/cu132
Replace cu132 with another CUDA version or a ROCm build (e.g. rocm10.0).
Stable releases
pip install torchtitan
# or
conda install conda-forge::torchtitan
Each stable release pins the nightly versions of torch and torchao it was
validated against; see RELEASE.md.
Optional dependencies
- Low precision (MXFP8, NVFP4): install a
torchaonightly matching yourtorchbuild. It is not inrequirements.txtso it cannot be resolved against a differenttorch:USE_CPP=0 pip install --pre --upgrade torchao --index-url https://download.pytorch.org/whl/nightly/cu132 - RL:
scripts/install_rl_env.shinstalls the vLLM, Monarch, and TorchStore nightlies and the FlashAttention build for your GPU. See the RL quick start. - Extras for vision, FLUX, RL, and development are also available as
pip install -e ".[vlm,flux,rl,dev]". - Importing from elsewhere: the source tree runs as is; to import
torchtitanfrom another directory, install it in editable mode without re-resolving dependencies:pip install -e . --no-deps.
Multi-node training with Slurm
Submit scripts/multinode_trainer.slurm with
sbatch from the repository root. Set the node count in both the #SBATCH
header and the torchrun --nnodes argument:
#SBATCH --ntasks=2
#SBATCH --nodes=2
srun torchrun --nnodes 2 ...
If a node does not have 8 GPUs, also adjust --nproc_per_node and
#SBATCH --gpus-per-task.
Documentation
| Topic | Where |
|---|---|
| Configuring runs and writing recipes | torchtitan/config |
| Adding a model | torchtitan/models |
| Parallelisms and distributed techniques | torchtitan/distributed |
| Datasets and data loading | torchtitan/components/data |
| Checkpointing and HF conversion | torchtitan/components/checkpointer |
| Reinforcement learning | torchtitan/rl |
| Debugging, reproducibility, and validating numerics | docs/debugging.md |
| Validation and evaluation | docs/evaluation.md |
| Benchmarks | docs/benchmarks |
| Tests and CI | tests |
Contributing
We look forward to your contributions!
- New ideas should start in the
experimentsfolder; follow the experiments guidelines. - For fixes and contributions to core, follow the contributing guidelines.
Citation
Our paper gives a detailed look at the parallelisms and optimizations in
torchtitan, with advice on when to use each technique:
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.
@inproceedings{
liang2025torchtitan,
title={TorchTitan: One-stop PyTorch native solution for production ready {LLM} pretraining},
author={Wanchao Liang and Tianyu Liu and Less Wright and Will Constable and Andrew Gu and Chien-Chin Huang and Iris Zhang and Wei Feng and Howard Huang and Junjie Wang and Sanket Purandare and Gokul Nadathur and Stratos Idreos},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=SFN6Wm7YBI}
}
License
Source code is made available under a BSD 3 license. You may have other legal obligations that govern your use of other content linked in this repository, such as the license or terms of service for third-party data and models.