# Why Zephon? **Your data loader is a hidden variable.** If you resume a training run on 8 GPUs instead of the 16 GPUs it started on, you would expect to be training on the same data in the same order. That is not the case. The data loader quietly changes what each GPU sees. [Mosaic Streaming](https://docs.mosaicml.com/projects/streaming/en/latest/distributed_training/elastic_determinism.html) calls the property you might have been expecting *elastic determinism*: yielding the same global batches no matter how many GPUs you train on. Most data loaders do not even offer it. Isolating variables is the basis of experimental research, as [early systems work](https://multicians.org/InstrumentationPaper.html) stressed. If two runs differ in one parameter and everything else behaves deterministically, the difference in outcome belongs to that parameter. This is not the case when the loader silently changes the data order, making it a hidden variable in your experiments.
Resuming a run on a different GPU count. If we continue training after a fault or cluster resize event, without an elastically deterministic data loader, the loss curve depends on the number of GPUs we train with. Drag the divider to compare the two loaders; each is shown against its own uninterrupted reference.
*Any* variation in the input data sequence may confound the results of ablations. The comparison above shows how the loss curve of a training run behaves when continuing training on fewer nodes after an interruption. The change in GPU topology causes a standard data loader to change the data order. However, with Zephon, every topology receives the same global batches, resulting in overlapping loss curves. This is even measurable in downstream evaluations.
In the downstream evaluation above, the swings in loss correspond to a variation of eval scores across GPU counts that is above relevant thresholds of public evaluation suites like DCLM Core v1 or FineWeb, as well as our internal evaluation suites here at Datology. We observe a maximum swing of ~1pp for Core v1 and ~0.56pp for FineWeb. **Why this is hard.** Preserving the data order would be straightforward if the pipeline hands samples through one at a time. Modern pipelines, however, often don't do this. Splitting documents into subsequences, packing subsequences into bins, and reordering samples are common things for a stateful n-to-m pipeline. Such a pipeline has no sample index. Progress is not a number you can just write down in this case; it is spread across in-flight operator state, partially filled bins, and records that have been reordered. So how does Zephon manage writing checkpoints, and what does it do on resumption? Zephon handles this by checkpointing how far the run has progressed through the data rather than the internal state of the pipeline, and by periodically bringing its operators back to a clean point that a resumed run can restart from. This is what lets Zephon bound the state of an otherwise potentially infinitely unbounded pipeline. [How Pipelines Run](pipelines/how_pipelines_run.md) describes this mechanism in detail, and [Distributed Training](pipelines/distributed_training.md) covers what you need to configure for a multi-rank run. **Why we released it.** When we went looking for something that already implemented elastic determinism, came with useful operations, and was decoupled from underlying file formats, we could not find anything. So we built Zephon for ourselves and for the teams we work with. Once we had it running, it was clear it could help anyone training a model, so we decided to release it to everyone. We are doing this because of where we think the AI landscape is headed. We want to help create a world where any team can build its own foundation model. Building models requires reliable experiments, and non-deterministic data loading makes experiments unreliable. If everyone has to rebuild their data loading infrastructure from scratch before they can even run a reliable experiment, this is not very efficient.