Packing¶
Packing operators combine multiple short tokenized records into a single sequence for training. The goal of packing is to ensure that the model spends less time processing padding tokens, which do not contain data that we want the model to learn but are simply there so that the structure of our training sequence matches what the model expects. For example, if a model is trained for an input sequence length of 4096 tokens, but many of the individual samples are only a few hundred tokens long, then much of the space allocated for the input sequence would be wasted on padding, which makes our overall training pipeline much less efficient than it could be. Instead of this, many training stacks support a mechanism to indicate that tokens from multiple documents are included in the same input sequence.
In the general case, optimal packing is an NP-hard problem, but there are several useful heuristics that seem to work well in practice. That said, packing is not always necessary for a training pipeline; if the input records are generally all the same size and are close to a fixed target length, or if the model training pipeline cannot correctly handle document boundaries within a single training sequence, or if keeping a single record per example is important for the workload, then packing can be excluded from your pipeline configuration.
Choosing A Packing Operator. Zephon ships with two built-in packing operators that
support a common set of packing algorithms and configuration options: pack_flat and
pack_sequences. Both packing operators take a max_length argument that defines the
maximum token capacity of each packed sequence that will be created. This is typically
the sequence length of the model plus one, because we need a final label for the last
token in each sequence. For example, if you use
SampleBatch.to_training(return_labels=True), you will find that if you pack with a
maximum length of 4097, that there will be 4096 input tokens and 4096 label tokens,
shifted one to the right.
The pack_flat operator is for the simple and common case in which you want to
concatenate the tokenization fields in the record into fixed-length sequences:
"""Concatenate tokenized records into fixed-length sequences."""
from zephon import Pipeline
# max_length is the model's context length plus one, so to_training can shift
# the labels by a position and still fill the context.
pipeline = (
Pipeline(work_source)
.tokenize(tokenizer_id="gpt2", field="text")
.pack_flat(max_length=4097, algorithm="wrap", pad_token_id=0)
)
for record in pipeline:
# input_ids, attention_mask, and positions, which resets to zero at each
# document the sequence contains.
print(sorted(record.payload))
The output payload of pack_flat is designed to work directly with the
SampleBatch.to_training() conversion methods and includes a positions array that marks
the boundaries of the tokens from each input sample in the output payload.
The pack_sequences operator should be used when you have code downstream of packing that
needs to be able to access the individual records that make up each packed result, which
is common in pipelines for multimodal data:
"""Pack records while keeping each document reachable."""
from zephon import Pipeline
pipeline = (
Pipeline(work_source)
.tokenize(tokenizer_id="gpt2", field="text")
.pack_sequences(max_length=4097, num_bins=8, algorithm="first_fit")
)
for record in pipeline:
# packed_samples, not an assembled sequence: to_training cannot convert
# this, which is the trade for reaching the individual documents.
print(len(record.payload["packed_samples"]))
Here, the output payload contains a packed_samples list rather than an assembled
training sequence. With the bin-packing algorithms (first_fit and best_fit, described
below), this list preserves the individual input payloads, but with wrap and
best_fit_wrap, it contains slices of the token field and other length-aligned fields,
so those algorithms should not be used when you need to retain arbitrary fields from the
original records.
Choosing A Packing Algorithm. While the packing operator determines the shape of the
packed output records, the packing algorithm determines how Zephon chooses to fill those
packed outputs. Each of the algorithms below can be used with either pack_flat or
pack_sequences; the main choice that you need to make is whether you need an algorithm
that will maximize the number of useful tokens in each packed example (even if it
requires splitting up the tokens from a single record across multiple packed outputs) or
if you want to be sure that the input records stay intact in the packed outputs (even if
doing this imposes a cost on the efficiency of the packing).
The wrap algorithm treats the input data as a first-in, first-out stream of tokens and
fills each packed sequence to max_length, splitting input records across packed
sequences when necessary. At each flush, any remaining tokens that cannot fill a complete
sequence are dropped, and Zephon logs a warning with the number of tokens that were
discarded.
"""Fill every sequence completely, splitting records where needed."""
from zephon import Pipeline
# A FIFO stream of tokens cut at max_length. Tokens left over at a flush that
# cannot fill a sequence are dropped, and Zephon logs how many.
pipeline = (
Pipeline(work_source)
.tokenize(tokenizer_id="gpt2", field="text")
.pack_flat(max_length=4097, algorithm="wrap", pad_token_id=0)
.batch(microbatch_size=8)
)
If you need to prioritize keeping input records intact, the first_fit and best_fit
bin-packing based algorithms are your go-to options because they will never split an input
record across different packed outputs. Instead, they maintain a configurable number of
bins and place each incoming record into one of the bins that has enough space to fit the
sample. By convention, the wrap algorithm is commonly used in pretraining, whereas bin
packing algorithms are used in SFT.
As their names suggest, the first_fit algorithm will choose the first bin it encounters
that has enough capacity to fit the sample, while the best_fit algorithm selects the bin
that would have the least remaining capacity once the input record was placed inside of
it. Both of these algorithms require a num_bins argument that controls how many
partially filled sequences are kept open for each packing group. Because first_fit and
best_fit do not split records, any input records that are longer than the max_length
value cannot be packed and will be dropped. At each flush, every open bin is emitted
regardless of whether or not it is full. The pack_flat operator will pad these partial
sequences to max_length, and pack_sequences will emit them as-is, so you can expect up
to num_bins partially filled sequences per flush.
"""Pack without ever splitting a record."""
from zephon import Pipeline
# num_bins partially filled sequences stay open; first_fit takes the first bin
# with room, best_fit the one left with the least. Records longer than
# max_length fit nowhere and are dropped, and each flush emits every open bin.
pipeline = (
Pipeline(work_source)
.tokenize(tokenizer_id="gpt2", field="text")
.pack_flat(max_length=4097, algorithm="first_fit", num_bins=8, pad_token_id=0)
.batch(microbatch_size=8)
)
Finally, best_fit_wrap is a hybrid packing algorithm that maintains a buffer of
candidate records and selects records that fit together efficiently, but also has the
option to split up input records across packed outputs when it is necessary in order to
finish a packed sequence.
When a partially filled sequence needs to be emitted, pack_flat pads it to max_length
using the required pad_token_id argument. This also applies to the partial sequence
produced by best_fit_wrap at each flush. You don’t need to mask this padding yourself:
pack_flat records how much padding it added, and SampleBatch.to_training excludes
those positions from the loss (see
Building Training Batches). The pack_sequences operator
simply returns the constituent payloads without doing any additional padding.
Most users should opt for the wrap algorithm when they are trying to maximize token
utilization and the first_fit algorithm when they must keep records intact. The
pack_flat and pack_sequences entries in the API reference
cover the best_fit, best_fit_wrap, and other packing options.
Packing Homogeneity. By default, Zephon can pack samples from different upstream
Dataset components together into a single training sequence. For many training
pipelines this is perfectly fine, but there are cases where you might want to enforce
some rules about the structure of the samples that the model will encounter in the
training sequence. For example, you might not want to allow general web text to share a
context window with samples that include detailed instructions or reasoning traces, or
where your downstream processing code needs to handle the transformation of different
kinds of data formats in different ways.
To support these use cases, Zephon’s packing operators support a homogeneity setting
which will enforce rules about how the packing algorithms are allowed to combine samples
from different mixture components into a single packed record. Setting
homogeneity="full" requires that each packed output is comprised only of samples that
come from a single mixture component, while setting homogeneity="group" and adding in a
dictionary of groups that organizes the input Dataset names into groups that are
allowed to be packed together will allow samples from the same group to be packed
together:
"""Restrict which mixture components may share a packed sequence."""
from zephon import Pipeline
from zephon.io import Dataset
from zephon.ops import DomainGroups
from zephon.work import StaticMixtureWorkSource
mixture = {"web": 0.4, "books": 0.2, "math": 0.2, "code": 0.2}
datasets = [
Dataset.from_path(name, f"s3://my-bucket/corpora/{name}") for name in mixture
]
work_source = StaticMixtureWorkSource(datasets, mixture, seed=42)
# homogeneity="full" would keep every sequence to a single component. Each
# group keeps its own bins and is flushed separately, so a flush can leave a
# partial sequence per group.
pipeline = (
Pipeline(work_source)
.tokenize(tokenizer_id="gpt2", field="text")
.pack_flat(
max_length=4097,
algorithm="first_fit",
num_bins=8,
pad_token_id=0,
homogeneity="group",
groups=DomainGroups({"prose": ["web", "books"], "formal": ["math", "code"]}),
)
)
The downside of adding the homogeneity constraints is utilization; by default, any sample can help fill any open packed sequence. With a homogeneity constraint in place, the packing operators need to maintain separate pools for each permitted group of samples, which adds overhead and reduces the efficiency of packing for components that are relatively rare in the training data. Each group is also flushed separately, so every flush can produce a partial sequence for each group.