TL;DR

Padding is a significant source of wasted compute when training LLMs. This remains a problem for hybrid Transformer-SSM models, where methods like sequence packing are not easily applicable. We managed to eliminate ~90% of padding-related overhead by applying a model-agnostic approach that utilizes micro-batch-level truncation and padding-aware micro-batching. This approach provides a simple and broadly applicable alternative across existing and future model architectures that dramatically improves training efficiency.

Addressing inefficiencies in online-RL training

Training efficiency is a key concern in large-scale model development, affecting both compute usage and wall-clock time, and ultimately cost. These tradeoffs become especially pronounced in online-RL training, where training runs are long-lived and compute-intensive, making inefficiencies compound quickly.

Over the past few months, our Labs in Front series has explored several directions for improving the efficiency of our online-RL training pipelines. This includes work from our colleagues on improving data efficiency through dynamic data snoozing, which primarily optimizes the rollout generation phase, as well as efforts to scale multi-node deployments of judge LLMs efficiently, which target the reward computation phase. Together, these efforts reflect a broader push to reduce waste and improve utilization across different stages of the training process.

In this post, we focus on the policy update phase of the training process, which encompasses the reference forward pass and the training forward–backward pass. Our focus on this phase complements prior work in this series on inference and reward computation. We identify padding-related waste within the policy update phase as a significant source of inefficiency, and show how to mitigate it in a model-agnostic way.

Training systems built for transformers meet other architectures

Much of the modern LLM training stack is built around transformer models, and many high-impact efficiency optimizations are implemented specifically for transformer workloads. At the same time, non-transformer and hybrid architectures are increasingly common in practice, often motivated by the need to address known inefficiencies and limitations of standard transformers. Examples include: Qwen3-next (DeltaNet), Nemotron (Mamba2), Granite (Mamba2), and our own Jamba models (Mamba). As these architectures gain adoption, gaps emerge where transformer-specific optimizations no longer apply and previously hidden sources of inefficiency become visible.

We illustrate this setting by training Jamba2-3B, our hybrid model architecture that interleaves attention layers with Mamba state-space layers, using VeRL, a popular online-RL training framework for LLMs. VeRL inherits many of these transformer-centric design choices, which surface clearly when training hybrid architectures like Jamba.

One of the most prominent missing optimizations in this setting is efficient handling of padding, the addition of dummy tokens so variable-length sequences fit a fixed tensor shape, which can become a major source of wasted compute.

Handling padding waste

Training batches typically contain sequences with highly variable lengths, which is handled by padding all sequences in a batch to the same length through the addition of dummy tokens. While padding makes batching possible, these tokens carry no information and are still processed by the model, consuming compute and memory without contributing useful work. Prior work analyzing common NLP datasets shows that under standard fixed-length batching, padding can account for up to about 50% of all processed tokens, and in some realistic cases even more (Krell et al., 2021).

In RL training, this inefficiency is further amplified by the variability in model-generated response lengths, which can differ substantially even for similar prompts. As a result, padding-related waste affects multiple phases of the policy update step, most notably the reference forward pass and the training forward–backward pass. Since these phases account for a significant fraction of overall step time, excessive padding can become a bottleneck for training efficiency.

Padding is wasted compute
Padding waste in a fixed-length micro-batch.
A micro-batch padded to a fixed maximum sequence length (e.g., 8 tokens, as shown here) contains real tokens followed by padding tokens. In this example, 16 out of 40 tokens (40%) are padding, meaning a substantial fraction of the computation and memory is spent on tokens that carry no information.

Eliminating padding in transformer training: sequence packing

For transformer training, the canonical fix is sequence packing, where multiple variable-length sequences are concatenated into a single longer sequence, completely eliminating the need for padding. Transformers interleave token-wise layers, such as MLPs and normalization layers, with attention layers. Token-wise layers operate independently on each token and therefore handle packed sequences without any special treatment.

Attention layers, however, mix information across positions in the sequence. In many pretraining and SFT pipelines, naive packing without explicitly enforcing sequence boundaries often works in practice, as models learn to treat special sequence start and end tokens as soft delimiters. Online-RL is more sensitive to inference–training mismatch: a difference in rollout-time and training-time conditioning can introduce instability. For this reason, it is crucial to enforce sequence boundaries so that tokens from one sequence do not attend to tokens from another.

Modern attention implementations support this by respecting sequence boundaries and avoiding attention across sequences. Crucially, this does not use naive masking, which would still incur quadratic computation over the entire concatenated sequence. Instead, efficient implementations take sequence boundaries into account directly to restrict attention computation so that tokens only attend within their original sequence.

How transformers avoid it_ Sequence packing
Sequence packing eliminates padding by concatenation.
Multiple variable-length sequences are concatenated into a single dense sequence, avoiding padding entirely. Efficient attention implementations respect sequence boundaries so that tokens attend only within their original sequence, eliminating padding while avoiding unnecessary attention computation.

Sequence packing beyond transformers

Sequence packing is an optimization that requires explicit, architecture-specific support. Transformers benefit from mature implementations that handle packed sequences efficiently, but this is the result of substantial, dedicated engineering, not a property of sequence models in general.

The same requirement applies to other classes of sequence models, including state-space and hybrid architectures. Supporting multiple independent sequences concatenated along the sequence dimension requires each architecture to implement its own correct and efficient handling of sequence boundaries. Absent such support, sequence packing cannot be applied directly.

Mamba is one concrete example of this broader pattern. While there are research efforts exploring how to add packing support to Mamba (e.g., PackMamba: Efficient Processing of Variable-Length Sequences in Mamba Training; Xu et al., 2024), doing so requires non-trivial changes and careful validation. From a training infrastructure perspective, this increases implementation complexity, expands the correctness surface, and tightly couples the training stack to a specific architecture.

This led us to ask whether we could eliminate most padding-related inefficiency without relying on architecture-specific changes to the model’s computation.

An alternative angle: reducing padding outside the model

Rather than asking how to make the model support packed sequences, we asked this question:

How much padding can we eliminate before tensors ever reach the model?

This is not a replacement for sequence packing, nor a claim that kernel-level optimizations are unimportant. Instead, it represents another point in the design space, one that emphasizes minimal changes, low risk, fast iteration, and broad applicability.

In practice, this meant making a small number of targeted changes. At a high level, we:

  • Truncate excess padding in each micro-batch.
  • Rearrange micro-batches to enable further padding truncation.

These changes affect only padding tokens and how sequences are grouped into micro-batches. Because micro-batches are used purely for memory management and gradients are accumulated across them to form the effective training batch, we are free to change this grouping. Outputs on all non-padding tokens remain identical (verified by direct output comparison), and models trained with the proposed method reach the same downstream quality as with standard padding.

The following sections describe how this works and what it buys us in practice.

Excess padding truncation

By default, VeRL pads sequences to a fixed maximum sequence length determined by the training configuration. This maximum is composed of the maximum prompt length and the maximum response length: the former effectively controls which prompts are allowed into training (with longer prompts filtered out), while the latter caps the number of tokens the model is allowed to generate. Since most prompts and responses are significantly shorter than these configured maxima, this padding scheme often results in large amounts of unnecessary padding. This effect becomes especially pronounced when training on mixed datasets with heterogeneous prompt and response lengths. We start by considering a given micro-batch and eliminating as much of this padding as possible.

We address this with micro-batch-level truncation. We identify contiguous blocks of padding at the boundaries of the tensors and truncate the tensors to remove them.

To make truncation as effective as possible, we first handle a practical detail: VeRL’s two-sided padding. Prompts are left-padded and completions are right-padded, which splits padding across both ends of the sequence. When padding is split this way, truncation removes padding from each side separately, limiting how much total padding can be eliminated.

To fix this, before truncating, we move left-padding to the right, grouping all padding contiguously at the end of the sequence (implemented as a tensor roll, with attention masks and position IDs adjusted accordingly). This change is purely a tensor reordering step, but it allows truncation to remove more padding.

Micro-batch-level padding truncation
Micro-batch-level padding truncation.
Sequences are initially padded to a fixed maximum length. Left-padding is shifted to the right to make padding contiguous, and then padding is removed by truncating the micro-batch tensors to the maximum sequence length within the micro-batch.

Intermediate results

Experimental setting: We evaluated this approach across a range of training setups, including more complex multi-dataset regimes. To keep the presentation focused, we report results from a standard single-dataset online-RL setup on GSM8K, run on a single H100 node (8 GPUs). We use a batch size of 32, a maximum prompt length of 4k tokens, and a maximum response length of 8k tokens. Reported numbers are average policy update step times over 10 consecutive steps (rounded to the nearest second), shown for both Qwen2.5-7B-Base and Jamba2-3B.

Results: After applying padding truncation, the policy update step time, which includes the reference forward pass and the training forward–backward pass, drops significantly:

  • Jamba2-3B: 45s → 20s (~56% reduction)
  • Qwen2.5-7B: 63s → 22s (~65% reduction)
training step time with padding truncation
Effect of micro-batch-level padding truncation on policy update step time.
Policy update step time before and after applying micro-batch-level padding truncation.

Analysis: The magnitude of this speedup is due to a long-tailed length distribution and the need to accommodate rare long examples in realistic multi-dataset training. While this specific dataset could use a smaller configured maximum, realistic mixed-dataset regimes require accommodating such long-tail examples. In this GSM8K setup, prompts are typically a few hundred tokens (often ~200), but are padded to a 4k maximum, while responses are usually ~500–1000 tokens with rare outliers up to 6–8k. As a result, many micro-batches contain large amounts of padding, and truncation removes a substantial fraction of padded tokens, especially after consolidating left-padding, leading to the large observed gains.

Nevertheless, even after these large gains, a clear gap remains relative to sequence packing (at 16s for Qwen and 14s for Jamba). The reason is length variance within the micro-batch itself: shorter sequences are still forced to match the longest sequence they were grouped with, leaving non-trivial padding that truncation alone could not eliminate.

This raised a natural question: could we narrow this gap further, in an architecture-agnostic way, by being smarter about how sequences are grouped into micro-batches?

Padding-aware dynamic micro-batching

Even with aggressive truncation, padding waste depends heavily on how sequences are grouped. In a micro-batch, all sequences are padded to match the longest sequence, forcing shorter sequences to perform unnecessary work.

To address this, we introduce padding-aware dynamic micro-batching: we sort sequences by length and construct micro-batches incrementally in that order. Before adding a sequence to the micro-batch being constructed, we compute the micro-batch token count after truncation, where all sequences are padded to the length of the longest sequence currently in the micro-batch, and start a new micro-batch when adding the sequence would exceed the token budget.

This approach achieves two goals. First, by grouping sequences of similar length, it reduces the amount of necessary padding and enables more aggressive truncation. Second, it keeps compute and memory usage predictable across micro-batches by enforcing a fixed token budget. As a result, micro-batches may contain different numbers of sequences, but have similar total token counts, which is a much better proxy for compute and memory usage than sequence count.

Micro-batching techniques: Naive vs Padding Aware
Padding-aware dynamic micro-batching.
Standard micro-batching can group very short and very long sequences together, forcing all sequences in the micro-batch to match the longest one. Padding-aware dynamic micro-batching groups sequences of similar length under a fixed token budget, reducing intra-batch variance and enabling more aggressive padding truncation.

Impact of padding-aware dynamic micro-batching

Adding padding-aware dynamic micro-batching on top of truncation further reduces policy update step time to:

  • Jamba2-3B: 20s → 14s (additional ~30% reduction)
  • Qwen2.5-7B: 22s → 16s (additional ~27% reduction)
Training Step Time with Micro-Batch Rearrangement
Effect of padding-aware micro-batching on policy update step time.
Policy update step time with micro-batch-level padding truncation alone, and with truncation plus padding-aware micro-batching.

We can now zoom out and compare the final configuration, external padding minimization, against sequence packing.

For non-transformer and hybrid models, such as Jamba, correct and efficient processing of packed sequences is typically not supported. Feeding packed sequences into these models without architecture-specific support leads to information leakage between sequences and would invalidate training. Nevertheless, it is possible to run such configurations and measure the resulting step time. We therefore report sequence packing results only as a timing reference, indicating the step time that would be achievable if padding were eliminated entirely, not as a usable or correct training configuration.

Training Step Time_ Sequence Packing vs External Padding Minimization
Comparison to sequence packing.
Policy update step time for default batching, external padding minimization, and sequence packing.

Across both architectures, the picture is consistent:

  • For Qwen2.5-7B, external padding minimization reduces policy update step time by ~75% relative to default VeRL (63s → 16s), achieving ~93% of the improvement obtained when padding is eliminated via sequence packing (12s).
  • For Jamba2-3B, it reduces policy update step time by ~69% (45s → 14s), achieving ~91% of the improvement relative to a padding-free reference (11s), reported here using sequence packing as a timing proxy.

In other words, with an architecture-agnostic approach, we eliminate almost all of the padding-related inefficiency, and close most of the performance gap to sequence packing.

Closing thoughts

Modern training stacks are heavily optimized for transformers. When training another model architecture, those optimizations are often lost, and seemingly small issues like padding can become a dominant cost.

What we found encouraging is that sometimes there isn’t a need for architecture-specific changes to get most of the performance back. In our case, a small set of architecture-agnostic mechanisms in the training infrastructure, including excess padding truncation and padding-aware dynamic micro-batching, eliminated most of the padding-related inefficiency and achieved policy update step times close to those obtained when padding is removed entirely.

Crucially, this applies not just to Jamba, but also to standard transformer models like Qwen. That generality matters: approaches like this can accelerate the adoption of novel architectures by reducing the need to reimplement complex, model-specific internal features (such as sequence packing support) before achieving reasonable training efficiency. 

The broader lesson is not that these exact modifications apply everywhere, nor that kernel- or model-level optimizations are unimportant. Rather, it’s often worth first looking for model-agnostic fixes that eliminate wasted work before committing to riskier internal changes, especially when training non-transformer architectures inside a transformer-centric training stack.