Basic Concepts

Zephon is a data loader: it’s the component of your model training stack that is responsible for turning your raw input data (think strings or jpegs) into the form that your model training loop wants to consume (tensors) as quickly and reliably as possible. To do this work, Zephon needs you to do three things that we will be covering in the rest of this user guide:

  1. Describe the input data that you have as one or more Dataset objects,

  2. Decide on your training curriculum via a WorkSource that references the Dataset definitions you provided,

  3. Define the processing steps required to turn your input data into the format your model training loop expects via a Pipeline that operates on the curriculum defined by the WorkSource.

The training framework you are using is responsible for defining the model itself, configuring the optimizer, and running the steps that consume the prepared data that Zephon provides. Let’s take a (slightly) closer look at how the abstractions Zephon provides fit together.

A Dataset describes a collection of input samples, like documents, conversations, or text-image pairs. Most of the datasets we work with in model training are made up of a collection of files that we refer to as shards. You create a Dataset by giving Zephon a name and the path (cloud or local) to your data. Zephon discovers the shards and their sample counts, providing the metadata needed to plan a training run without keeping the full collection of samples in memory. You can dive deeper into the data formats that Zephon supports out of the box and their various pros/cons for different model training use cases in Working with Datasets.

A WorkSource defines the training curriculum: which samples to select from the available Datasets, how to mix them, how to shuffle them, and how often they should be repeated during training. It uses the metadata that the Dataset objects provide to plan out the work, but leaves the actual fetching and parsing of the samples that the shards contain to the downstream Pipeline. This separation allows you to re-use the same underlying data for different training curricula without needing to reorganize the stored data, and to change the data processing pipeline without impacting the curriculum. WorkSources: The Training Curriculum does a deep dive into the StaticMixtureWorkSource that is the primary WorkSource implementation that Zephon currently ships with and is the workhorse (see what I did there?) for mixing and combining samples into an overall curriculum.

Finally, the Pipeline object starts from the curriculum generated by the WorkSource, fetches the actual sample payloads, and then applies any processing steps required to those samples to prepare them for training. We refer to these processing steps as operators before the training loop consumes the resulting data through the Pipeline as a regular Python iterable. Depending on how your data has been prepared, these operators may be as simple as tokenization or batching or they may involve a sequence of complex built-in or user-defined transformations. Pipelines: From a Curriculum to Training Batches goes into detail on the different kinds of operators that Zephon supports and how you can adapt them to fit the needs of your model training loop.