← 开源
pytorch

torchtitan

A PyTorch native platform for training generative AI models

Model DevelopmentLLM trainingPython
在 GitHub 打开
增长势头
+324 小时新增 Star+0.1%
5.80k
Star
1.03k
Fork
+19
本周
100
贡献者
创建于 2023-12-13 · 更新于 2026-10-11 · 今日第 2275 名
主要开发者
README

torchtitan

A PyTorch-native platform for pretraining and post-training generative AI models at scale

arXiv ICLR devlogs license pip conda

Quick start · Models · Features · RL · Docs · Contributing

torchtitan is a clean-room, PyTorch-native implementation of large-scale training. One readable codebase takes a model from pretraining through supervised fine-tuning to reinforcement learning, composing FSDP, tensor, pipeline, context, and expert parallelism without rewriting the model.

  • Composable parallelism. Models declare how they shard; FSDP2, TP, PP, CP, and EP stack into N-D layouts through DeviceMesh and DTensor.
  • Built for research. Minimal, hackable code; every run is a Python recipe you can read, copy, and modify.
  • Pretraining to RL in one stack. TitanRL runs the same model definitions and kernels in the trainer and in vLLM generation, with a bitwise-reproducible batch-invariant mode.
  • Frontier models and techniques. DeepSeek, Qwen, Kimi, gpt-oss, Llama, and FLUX, with MXFP8/NVFP4 low precision, Dist-MoE, CUDA graphs, and more.

[!NOTE] torchtitan is under active development. To use the latest features, use a recent PyTorch nightly.

Latest news

  • [2026/08] TitanRL is a hackable RL stack for scaling and debugging. It reuses TorchTitan model definitions and kernels across training and vLLM generation and supports batch-invariant mode.
  • [2025/11] AMD released an optimized fork of torchtitan for AMD GPUs.
  • [2025/10] We released torchtitan v0.2.0.
  • [2025/10] SkyPilot now supports torchtitan! See the tutorial.

Older news

  • [2025/07] We published instructions on how to add a model to torchtitan.
  • [2025/04] Our paper was accepted by ICLR 2025.
  • [2024/12] GPU MODE lecture on torchtitan.
  • [2024/07] Presentation at PyTorch Conference 2024.

Quick start

1. Install with a PyTorch nightly (Python 3.11+). Replace cu132 with your CUDA or ROCm build.

git clone https://github.com/pytorch/torchtitan && cd torchtitan
pip install --pre torch --index-url https://download.pytorch.org/whl/nightly/cu132
pip install -r requirements.txt

2. Train a debug model. It uses the tokenizer bundled with the tests, so no download is needed.

NGPU=1 ./run_train.sh   # Llama 3 debug model, 10 steps

3. Train a real model. Download the tokenizer (Llama weights require access on Hugging Face), then select a recipe by module and function name:

python scripts/download_hf_assets.py --repo_id meta-llama/Llama-3.1-8B --assets tokenizer --hf_token=...
MODULE=torchtitan_recipes.models.llama3 CONFIG=llama3_8b ./run_train.sh   # 8 GPUs

4. Make it yours. A recipe is a Python function returning a complete config, so changing a run means writing a function, not passing flags:

# my_recipes.py
from torchtitan_recipes.models.llama3 import llama3_8b

def llama3_8b_tp2():
    config = llama3_8b()
    config.parallelism.tensor_parallel_degree = 2
    return config
MODULE=my_recipes CONFIG=llama3_8b_tp2 ./run_train.sh

See the configuration guide for the full model, and torchtitan_recipes for the verified recipes.

Supported models

Model Sizes Modality Verified recipes RL
Llama 3 1B – 405B Text llama3_8b, llama3_70b ✅
Qwen3 0.6B – 32B dense; 30B-A3B, 235B-A22B MoE Text qwen3_14b ✅
Qwen3.5 / 3.6 / 3.8 0.8B – 397B-A17B; 2.4T-A95B (3.8) Text, vision ✅ text
DeepSeek V3 16B, 236B, 671B Text deepseek_v3_671b (+ Dist-MoE, NVFP4)
DeepSeek V4 Flash, Pro Text Soon
gpt-oss 20B, 120B Text ✅
Kimi K2.7 Moonlight-16B-A3B, Kimi-VL-A3B, K2.5 Text, vision
Kimi K3 K3 Text, vision Soon
Muse Glimmer 30B Text, vision muse_glimmer_30b ✅
FLUX dev, schnell Text-to-image flux_dev, flux_schnell

Every model has debug configurations exercised in CI. Verified recipes have been run in training, convergence, or performance workflows. To add a model, see the model guide.

Features

Parallelism

  • FSDP2, HSDP, and DDP
  • Tensor parallel, including async TP
  • Pipeline parallel: 1F1B, Interleaved 1F1B, GPipe, and other torch.distributed.pipelining schedules
  • Context parallel: all-gather KV and Ulysses
  • Expert parallel with DeepEP / HybridEP dispatch
  • Dist-MoE: fused dispatch, expert compute, and combine on Blackwell

Memory and performance

Training

  • Grain data pipeline with packing, mixing, and checkpointable state
  • Pretraining, SFT on chat datasets, and LoRA
  • Muon (DistMuon), EMA, multi-token prediction
  • Warmup-stable-decay LR schedule and validation

Post-training with TitanRL

Checkpointing

Debugging and observability

Extensibility

We report performance on up to 512 GPUs and verify loss convergence across parallelisms and techniques.

Test status

Hardware Integration tests Unit tests
CPU CPU Unit Test
NVIDIA GPU Integration Tests H100 Tests B200 Tests GPU Unit Tests
AMD GPU (ROCm) Integration Tests (ROCm)
RL (NVIDIA GPU) RL Integration Tests RL B200 Tests RL GPU Unit Tests

Installation

The quick start installs from source, which is the recommended way to get the latest features. Other options:

Nightly builds

pip install --pre torch --index-url https://download.pytorch.org/whl/nightly/cu132 --force-reinstall
pip install --pre torchtitan --index-url https://download.pytorch.org/whl/nightly/cu132

Replace cu132 with another CUDA version or a ROCm build (e.g. rocm10.0).

Stable releases

pip install torchtitan
# or
conda install conda-forge::torchtitan

Each stable release pins the nightly versions of torch and torchao it was validated against; see RELEASE.md.

Optional dependencies

  • Low precision (MXFP8, NVFP4): install a torchao nightly matching your torch build. It is not in requirements.txt so it cannot be resolved against a different torch:
    USE_CPP=0 pip install --pre --upgrade torchao --index-url https://download.pytorch.org/whl/nightly/cu132
    
  • RL: scripts/install_rl_env.sh installs the vLLM, Monarch, and TorchStore nightlies and the FlashAttention build for your GPU. See the RL quick start.
  • Extras for vision, FLUX, RL, and development are also available as pip install -e ".[vlm,flux,rl,dev]".
  • Importing from elsewhere: the source tree runs as is; to import torchtitan from another directory, install it in editable mode without re-resolving dependencies: pip install -e . --no-deps.

Multi-node training with Slurm

Submit scripts/multinode_trainer.slurm with sbatch from the repository root. Set the node count in both the #SBATCH header and the torchrun --nnodes argument:

#SBATCH --ntasks=2
#SBATCH --nodes=2

srun torchrun --nnodes 2 ...

If a node does not have 8 GPUs, also adjust --nproc_per_node and #SBATCH --gpus-per-task.

Documentation

Topic Where
Configuring runs and writing recipes torchtitan/config
Adding a model torchtitan/models
Parallelisms and distributed techniques torchtitan/distributed
Datasets and data loading torchtitan/components/data
Checkpointing and HF conversion torchtitan/components/checkpointer
Reinforcement learning torchtitan/rl
Debugging, reproducibility, and validating numerics docs/debugging.md
Validation and evaluation docs/evaluation.md
Benchmarks docs/benchmarks
Tests and CI tests

Contributing

We look forward to your contributions!

Citation

Our paper gives a detailed look at the parallelisms and optimizations in torchtitan, with advice on when to use each technique: TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

@inproceedings{
   liang2025torchtitan,
   title={TorchTitan: One-stop PyTorch native solution for production ready {LLM} pretraining},
   author={Wanchao Liang and Tianyu Liu and Less Wright and Will Constable and Andrew Gu and Chien-Chin Huang and Iris Zhang and Wei Feng and Howard Huang and Junjie Wang and Sanket Purandare and Gokul Nadathur and Stratos Idreos},
   booktitle={The Thirteenth International Conference on Learning Representations},
   year={2025},
   url={https://openreview.net/forum?id=SFN6Wm7YBI}
}

License

Source code is made available under a BSD 3 license. You may have other legal obligations that govern your use of other content linked in this repository, such as the license or terms of service for third-party data and models.