Basic Concepts¶
Zephon is a data loader: it’s the component of your model training stack that is responsible for turning your raw input data (think strings or jpegs) into the form that your model training loop wants to consume (tensors) as quickly and reliably as possible. To do this work, Zephon needs you to do three things that we will be covering in the rest of this user guide:
Describe the input data that you have as one or more
Datasetobjects,Decide on your training curriculum via a
WorkSourcethat references theDatasetdefinitions you provided,Define the processing steps required to turn your input data into the format your model training loop expects via a
Pipelinethat operates on the curriculum defined by theWorkSource.
The training framework you are using is responsible for defining the model itself, configuring the optimizer, and running the steps that consume the prepared data that Zephon provides. Let’s take a (slightly) closer look at how the abstractions Zephon provides fit together.
A Dataset describes a collection of input samples, like documents, conversations, or
text-image pairs. Most of the datasets we work with in model training are made up of a
collection of files that we refer to as shards. You create a Dataset by giving Zephon a
name and the path (cloud or local) to your data. Zephon discovers the shards and their
sample counts, providing the metadata needed to plan a training run without keeping the
full collection of samples in memory. You can dive deeper into the data formats that
Zephon supports out of the box and their various pros/cons for different model training
use cases in Working with Datasets.
A WorkSource defines the training curriculum: which samples to select from the available
Datasets, how to mix them, how to shuffle them, and how often they should be repeated
during training. It uses the metadata that the Dataset objects provide to plan out the
work, but leaves the actual fetching and parsing of the samples that the shards contain to
the downstream Pipeline. This separation allows you to re-use the same underlying data
for different training curricula without needing to reorganize the stored data, and to
change the data processing pipeline without impacting the curriculum.
WorkSources: The Training Curriculum does a deep dive into the
StaticMixtureWorkSource that is the primary WorkSource implementation that Zephon
currently ships with and is the workhorse (see what I did there?) for mixing and combining
samples into an overall curriculum.
Finally, the Pipeline object starts from the curriculum generated by the WorkSource,
fetches the actual sample payloads, and then applies any processing steps required to
those samples to prepare them for training. We refer to these processing steps as
operators before the training loop consumes the resulting data through the Pipeline as
a regular Python iterable. Depending on how your data has been prepared, these operators
may be as simple as tokenization or batching or they may involve a sequence of complex
built-in or user-defined transformations.
Pipelines: From a Curriculum to Training Batches goes into detail on
the different kinds of operators that Zephon supports and how you can adapt them to fit
the needs of your model training loop.