AutoEP (Automatic Expert Parallelism)
AutoEP automatically detects MoE layers in Hugging Face models and replaces them with EP-enabled versions, requiring zero model code changes. It follows the pattern of AutoTP (Automatic Tensor Parallelism).
This API is separate from the explicit deepspeed.moe.layer.MoE layer API.
For the explicit DeepSpeed MoE layer API, see Mixture of Experts (DeepSpeed MoE).
Built-in AutoEP presets: mixtral (Mixtral), qwen3_moe (Qwen3-MoE),
qwen3_5_moe (Qwen3.5-MoE), deepseek_v2 (DeepSeek-V2),
deepseek_v3 (DeepSeek-V3), and minimax_m3 (MiniMax-M3).
The preset name means AutoEP knows the router, expert, and weight naming
patterns for that model family. Running a Hugging Face model also requires a
Transformers build that exposes the matching config/model classes,
model.config.model_type value, and fused expert layout.
Preset |
Minimum Transformers version |
Notes |
|---|---|---|
|
|
|
|
|
Also covers Qwen2-MoE when the installed Transformers build uses the
validated fused expert layout. Qwen3-MoE classes appear in |
|
|
Requires the Qwen3.5 text-backbone |
|
|
|
|
|
|
|
|
Requires the MiniMax-M3 text-backbone |
ZeRO compatibility: Stages 0, 1, and 2, plus constrained Stage 3
support. Stage 3 requires AutoEP-managed MoE layers and does not support native
DeepSpeed MoE layers, AutoTP, tensor model parallelism from mpu, sequence
parallelism, hpZeRO secondary tensor groups, non-1 expert tensor
parallelism, or quantized gradients. Stage 3 AutoEP checkpoints are saved
partition-natively in the zero_pp_rank_* shard files and support
same-topology load, module-only loads (load_module_only),
optimizer-state-free loads (load_optimizer_states=False), and Universal
Checkpoint conversion. Optimizer-including Universal Checkpoint loads can
resume with a different data-parallel world size, a different autoep_size,
or both, when the target autoep_size divides the model’s expert count.
Weights-only/module-only Universal Checkpoint loads use the converted
fp32.pt parameter files and support the same data-parallel and
autoep_size topology changes.
Usage:
{
"expert_parallel": {
"enabled": true,
"autoep_size": 4,
"preset_model": "mixtral"
}
}
Experimental regional torch.compile
AutoEP can keep its router, token movement, expert computation, and collectives
in eager mode while compiling the surrounding decoder blocks with vanilla
torch.compile. This targets fragmented attention, normalization, residual,
and dense backward work without capturing AutoEP communication in the graph.
The path is opt-in. Enable it in the DeepSpeed configuration, then call
engine.compile() after initialization:
{
"compile": {
"autoep_non_moe": true
}
}
engine, optimizer, _, _ = deepspeed.initialize(
model=model,
model_parameters=model.parameters(),
config=ds_config,
)
engine.compile()
The call must happen after deepspeed.initialize() so AutoEP replacement is
complete. DeepSpeed discovers each AutoEPMoELayer and regionally compiles
its direct callable parent with fullgraph=False and dynamic=False.
This is usually a decoder block; a callable model root with a direct AutoEP
child is also supported. The AutoEP layer is an explicit compiler-disabled
graph break, so routing, AllToAll dispatch/combine, and expert execution remain
eager. An AutoEP layer used as the model root has no surrounding region and
is rejected.
The decoder blocks contain repeated attention, normalization, residual, and dense work targeted by this optimization. Compiling these regions bounds the traced code and lets structurally identical blocks reuse compiled graphs. Modules outside the selected regions, such as top-level embeddings and the language-model head, remain eager. They are not intrinsically incompatible with compilation, but extending the region would need separate validation and performance measurements. If the selected region is the model root, its non-MoE operations are included.
compile.autoep_non_moe defaults to false, preserving the existing
full-model behavior of engine.compile(). Setting the option alone does
not compile the model; execution stays eager until engine.compile() is
called.
The initial experimental path supports vanilla torch.compile with the
standard comm backend, sequence and pipeline parallel sizes of one, and
ZeRO stages 0, 1, and 2. Distributed performance and parity validation currently
target ZeRO stage 1. It rejects DeepEP, DeepCompile, AutoEP+AutoTP folding,
sequence or pipeline parallelism, ZeRO stage 3, optimizer or parameter offload,
compiled autograd, DeepCompile schedules, and any fullgraph or dynamic
value other than False instead of silently changing the requested behavior.
How it works:
During
deepspeed.initialize(), AutoEP scans the model for MoE layers using preset-defined patterns (router name, expert name, weight shapes).Detected MoE blocks are replaced with
AutoEPMoELayer, which uses TorchTitan’s grouped GEMM kernels and AllToAll token dispatch.EP/EDP process groups are created automatically based on
autoep_size.Expert parameters are marked for expert-data-parallel gradient reduction; router and shared-expert parameters use standard data-parallel reduction.
Router outputs and activation checkpointing:
Models using Hugging Face’s model-level router-logit recording capture the existing gate output; AutoEP does not compute a second projection just to populate an unused cache. Models whose MoE blocks return router logits directly compute them locally when constructing the return value, preserving that return contract and its gradients. No router-logit tensor is stored on the layer, so checkpoint replay early-stop and exceptions cannot leave a router-logit cache keeping the autograd graph alive between training steps.
Communication backend (optional):
The expert AllToAll can be carried by DeepEP instead of the default collectives. This is opt-in and off by default; jobs that set nothing keep the existing path unchanged.
{
"expert_parallel": {
"enabled": true,
"autoep_size": 8,
"comm_backend": "deepep",
"comm_num_sm": 12,
"comm_qp_margin": 4,
"comm_max_tokens_per_rank": 4096
}
}
comm_backend:"comm"(default) usesdeepspeed.commcollectives;"deepep"uses DeepEP’s dispatch and combine kernels.comm_num_sm: SMs given to communication. Default 12.comm_qp_margin: RDMA queue pairs reserved beyond one per SM. Default 4.comm_max_tokens_per_rank: largest per-rank token count the job will produce, which ismicro_batch_size * seq_lenwhen sequences are padded to a fixed length. Required whencomm_backendis"deepep"because the DeepEP buffer is sized statically and must use the same capacity on every rank. A batch that exceeds it is an error.
For autoep_size > 1, DeepEP receives the router output directly, bypassing
the collective backend’s sorting, token expansion, and split-count exchange.
Shared experts and router-logit outputs retain the same behavior. The EP
communicator is initialized once before each layer’s first DeepEP buffer is
constructed, including when the caller supplied a lazily initialized process
group. This initialization does not run on subsequent forwards. The standard
comm and autoep_size=1 paths are unchanged.
On 16 H100s across two nodes, replaying routing captured from real training, DeepEP reduced payload AllToAll time from roughly 100 ms to 48 ms per step. A full SFT step on Qwen3.5-MoE went from roughly 325 ms to 266 ms, a 1.2x speedup that removes about 18% of the step, reproduced across two independent jobs (1.21x and 1.24x). Both backends are measured in the same job, on the same pods and alternating, since the same measurement varied by a quarter between jobs; the figures are medians rather than single observations, and DeepEP’s own median moved by 0.3% between the two jobs while the collective baseline moved by 2.5%. The advantage grows with routing imbalance: at the most skewed step measured, the collective path degraded to 116 ms while DeepEP stayed flat.
Within one model, all MoE layers that agree on EP group, expert count, top-k,
hidden size, capacity, comm_num_sm and comm_qp_margin share a single
DeepEP buffer, which is every layer of a normal model. A buffer reserves fabric
resources that are not reported as device memory and that run out: measured on
32 H100s across four nodes, the twenty-eighth buffer per rank fails inside
ncclDevCommCreate, so one buffer per layer put a 27-layer ceiling on the
backend there. Buffers are also slow to build, about 15 seconds each on 16
H100s and 22 on 32, so sharing removes minutes of startup as well. The buffer is
released once the last layer holding it is torn down.
Sharing stops at the model. Two models converted separately get their own buffers even on the same EP group with identical geometry, because they are driven independently: an actor and a frozen reference model in a reinforcement learning loop need not reach their MoE layers in any fixed order relative to each other, and a shared DeepEP communication context would make that order matter. The cost is one extra buffer per model against a ceiling of 27.
comm_num_sm matters because communication competes with the expert GEMM for
SMs. The default of 12 was chosen by measuring whole steps: 8 SMs gave a median
297.9 ms against 265.4 ms at 12, and larger budgets were slower again.
comm_qp_margin exists because DeepEP’s automatic queue-pair count assumes
it is alone on the fabric, which exhausts the queue pairs ZeRO and the
data-parallel groups have already claimed in a training step.
DeepEP row weighting implementation (experimental):
DeepEP dispatch returns one received row per routed assignment and one FP32
weight per row. row_weighting_impl selects how AutoEP multiplies those rows
by their weights at the existing score_apply boundary:
{
"expert_parallel": {
"enabled": true,
"autoep_size": 8,
"comm_backend": "deepep",
"comm_max_tokens_per_rank": 4096,
"row_weighting_impl": "fused"
}
}
"auto" (default) resolves to "eager", preserving the existing eager
expression exactly. "fused" runs a separate Triton pointwise operator for
(rows.float() * weights).to(rows.dtype). It does not reduce over top-k, does
not change where BF16/FP16 rounding occurs, and does not replace DeepEP’s
combine; the output remains one weighted row per received row in the same row
order.
The forward product and row gradient match eager’s rounding. The FP32 gradient
of the routing weight sums the same products in a different order, so it need
not be bitwise equal to eager’s; neither summation is consistently closer to
an FP64 reference. Comparisons should use gradient errors relative to the
gradient norm after backward and before optimizer clipping in engine.step().
Adam’s first update can differ on the scale of the learning rate when a
near-zero gradient changes sign, even if the overall gradients agree closely.
"fused" is rejected, rather than silently ignored, when AutoEP cannot honor
it:
comm_backendis not"deepep"orautoep_size=1, because the call sites exist only inside the DeepEP route;Triton is unavailable, the device is not CUDA, or the build is ROCm;
rows are not bfloat16 or float16;
weights are not FP32
[N, 1]tensors on the same CUDA device;rows or weights are not contiguous, or rows are not shaped
[N, H].
The operator also accepts FP16 rows, but the current DeepEP dispatch supports
BF16 rows only. Correct backward replay through DeepEP additionally requires
preserving the cached dispatch layout; that correction is independent of row
weighting. The separate MoE gradient-norm correction affects the
FP16_Optimizer wrapper, which is also used by some BF16 configurations
(for example, BF16 with BF16 gradient accumulation without ZeRO).
The model-level gradient comparison samples gradients before the wrapper
computes the norm and clips them in engine.step(). GPU validation applies
both independent corrections; neither is part of this opt-in change.
Requirements and limits:
The
deep_eppackage must be installed. It is imported only when this backend is selected, so installations without it are unaffected.DeepEP v2 requires NCCL 2.30.4 or newer, built with GIN support. Below that version the transport is unavailable regardless of the network.
DeepEP v1 (the legacy
BufferAPI, using NVSHMEM and IBGDA) is not supported.bfloat16 only. DeepEP’s dispatch kernel takes bfloat16 rows, so selecting this backend for an fp16 or fp32 run is rejected rather than silently downgraded.
Not compatible with folded tensor parallelism (
expert_tensor_parallel_size > 1), which is rejected at setup.
Python cyclic GC (experimental):
Large Python model graphs can accumulate cyclic objects during training, and a
generation-2 collection pauses one rank’s Python thread, which is then exposed
as collective wait time on every expert-parallel rank. The top-level
disable_python_gc option addresses this. It is process-wide rather than
AutoEP-specific, so it is documented with the general configuration options.
Fused weighted restore (experimental):
After the combine all-to-all, AutoEP holds one row per routed assignment and has
to turn it back into one row per token. combine_impl selects how:
{
"expert_parallel": {
"enabled": true,
"autoep_size": 16,
"preset_model": "qwen3_moe",
"combine_impl": "fused_weighted_sum"
}
}
"auto" (default) resolves to "weighted_sum", which scatters the rows into
a zero-filled [tokens * top_k, hidden] buffer, widens it to FP32 to apply the
routing weights, and reduces over top-k. "fused_weighted_sum" computes the
weighted reduction in one kernel launch: each program owns a whole token and
walks the hidden dimension in chunks, multiplying each routed row by its score
in FP32 and accumulating the products in slot order in FP32. FP contraction is
disabled so each product rounds before it is added; the result is cast only
once to the output dtype. Neither the scattered buffer nor the FP32
intermediate is allocated. At the canonical shape the FP32 intermediate alone
is 64 MiB per layer.
The top-k summation order can differ from eager, so results are within the existing numerical tolerances, not bitwise unchanged. Only the forward reduction changes: the backward, collectives, router, grouped GEMM and expert-major reorder are untouched.
"fused_weighted_sum" is rejected, rather than quietly ignored, when it would
have nothing to replace or would change semantics:
tensor_parallel.autotp_sizegreater than 1, which uses folded tensor parallelism and restores combined tokens from assignment metadata instead;expert_tensor_parallel_sizegreater than 1;comm_backend="deepep"with expert parallelism, because DeepEP already restores and reduces its routed rows;a resolved
score_applyother than"post";activations that are not bfloat16, float16, or float32, a non-CUDA device, or a build without Triton.
Failing fast matters for measurement: a run that asked for the fused reduction and silently got the eager one would report the difference between an implementation and itself.
Fused rotary position embedding (experimental):
RoPE belongs to the attention layers rather than to AutoEP, so it is not
configured under expert_parallel. DeepSpeed provides an opt-in
installer that runs Hugging Face’s apply_rotary_pos_emb with a fused Triton
kernel:
from deepspeed.ops.triton_ops.fused_rotary_pos_emb import replace_rotary_pos_emb
patched = replace_rotary_pos_emb(model)
The eager function computes q * cos + rotate_half(q) * sin in several
elementwise kernels, each reading and writing the whole query or key tensor.
The fused kernel reads each tensor once and writes it once, in the forward and in
the backward, and applies each position’s cos and sin to all of its
heads. Every product and sum is rounded to the input dtype where the eager
expression rounds it, so the outputs, and the gradients for the queries and
keys, equal eager’s element for element as compared by torch.equal, which
does not distinguish +0.0 from -0.0.
Attention modules look apply_rotary_pos_emb up in their modeling module, so
the installer replaces it there: the replacement applies to every model of that
architecture in the process, not only to model, and
restore_rotary_pos_emb() undoes it. A modeling module is patched only if one
of the model’s submodules is defined in it, it is listed in
SUPPORTED_ROTARY_MODULES (Llama, Mistral, Mixtral, Qwen2, Qwen2-MoE, Qwen3,
Qwen3-MoE and DeepSeek-V3), and its apply_rotary_pos_emb and
rotate_half still have the code of the split-half expression: the same
bytecode, names, constants and defaults, whatever their docstrings. Other
architectures define functions of the same name that rotate only part of the
head dimension or add casts, so a module whose function has changed is left
alone with a warning. The return value is the number of modeling modules
patched: 1 for Qwen3-30B-A3B.
The replacement runs the kernel when the queries, keys, cos and sin are
bfloat16 or float16 CUDA tensors of one dtype, the head dimension is even, at
most 512 and has unit stride, and cos and sin do not require grad. For
any other input it runs the original function, so the result is exactly eager’s,
and logs a warning once. fused_apply_rotary_pos_emb in the same module
applies the kernel directly and raises on unsupported inputs instead.
Requirements and limits:
The kernel needs CUDA with Triton. On ROCm or without Triton, the replaced function runs eager.
All four tensors must be on one device. Launches use that device’s current stream even if another device is current, and restore the caller’s current device afterwards.
unsqueeze_dim1 (heads before the sequence) or 2 (sequence before the heads), withcosandsinof shape[batch, sequence, head_dim]or[1, sequence, head_dim].For dense layouts, such as attention’s transposed projection, outputs keep the strides of the queries and keys, as eager’s do; other non-contiguous inputs can give outputs laid out differently from eager’s. The gradients for the queries and keys take the layout of the outputs, whereas eager’s follow the incoming gradient, so weight gradients computed from them further back, such as the query projection’s, can differ from eager’s in the last bits.
Transformers releases whose
apply_rotary_pos_embstill takesposition_ids, such as 4.51, are left unpatched.First-order gradients for the queries and keys;
cosandsinare constants. Differentiating those gradients again (double backward) raises.torch.compileandtorch.functransforms are not covered.
Constraints:
autoep_sizemust dividenum_expertsfor all detected MoE layers.autoep_size=1is valid: all experts remain local (no AllToAll), useful for functional testing on a single GPU.AutoEP currently cannot be combined with AutoTP (
tensor_parallel.autotp_size > 1) or tensor model parallelism frommpu; support is planned as follow-up work.AutoEP with ZeRO Stage 3 is supported only without sequence parallelism, hpZeRO secondary tensor groups, non-1 expert tensor parallelism, or quantized gradients.
Regular checkpoint save/load requires matching
autoep_size. To changeautoep_sizeor data-parallel world size across runs for the same AutoEP-detected model topology, convert the checkpoint to Universal Checkpoint format and load it withcheckpoint.load_universal; see the Universal Checkpointing tutorial for the detailed flow and constraints.DeepSeek-V2 and DeepSeek-V3 AutoEP do not support load-balance expert bias yet. The built-in DeepSeek presets disable it by default; explicit non-null values fail.