litData

litData is a data processing and loading library developed by Lightning AI, the team behind PyTorch Lightning. Here, we are interested in the binary format that litData reads and writes as a collection of records with a consistent schema that are written into individual files called chunks along with an index.json file that describes the contents of the chunks. (The chunks correspond to what Zephon refers to as shards.)

The litData format is the one that we use for our large pretraining runs at DatologyAI for both our text and multimodal models; we adopted the format because we liked the flexibility that it provided for representing both tensors/numPy arrays and complex nested Python dictionaries depending on the data schema that we needed for the file, and its performance as it relies on a fast PyTree serializer.

Preparing litData Shards. To use the litData format with Zephon, install the project with the litdata extra, which includes all of the dependencies needed for reading litData shards:

examples/guide/datasets/litdata_install.sh
# Install Zephon with the litdata extra.

pip install "zephon[litdata]"

Let’s look at an example of writing out a small litData dataset using the litData library:

examples/guide/datasets/litdata_write_shards.py
"""Write a small litData dataset, index.json included."""

from litdata import optimize


def samples(count: int) -> list[dict[str, object]]:
    """Return the records to serialize, one dict per sample."""
    return [{"text": f"sample {i}", "label": i % 2} for i in range(count)]


def serialize(record: dict[str, object]) -> dict[str, object]:
    """Write each record through unchanged; litdata pickles this to its workers."""
    return record


optimize(
    fn=serialize,
    inputs=samples(1000),
    output_dir="data/litdata_demo",
    chunk_bytes="64MB",
)

Note that this code generates both the shards and the index.json file that describes the shard contents and structure; be sure to keep the index and the shards together if you move the dataset to or from object storage. More examples of tools for generating and preparing samples for litData shards are available in litData’s data preparation guide.

Reading litData Shards. Once the litData files have been written, you can define a Dataset by passing the dataset location to the Dataset.from_path method with the fmt="litdata" argument. Zephon will infer the format if fmt is not passed. You can then use a DatasetInspector to examine the reconstructed sample payloads to confirm that the samples are in the form that you expect:

examples/guide/datasets/litdata_dataset.py
"""Define and inspect a litData Dataset."""

from zephon.debug import DatasetInspector
from zephon.io import Dataset

# fmt is optional: Zephon infers the format from the index.json beside the
# chunks. Keep the two together when you move the dataset.
dataset = Dataset.from_path("litdata_demo", "data/litdata_demo", fmt="litdata")

with DatasetInspector(dataset) as inspector:
    print(inspector.read(shard_id=0, sample_index=0))

Zephon also supports litData shards that are compressed using Zstandard. Depending on the contents of the shards, this may significantly reduce the size of the shards and thus the storage/networking cost of moving them around. Keep in mind that the compressed shards will be fully decompressed before they can be read, so plan for this when configuring the size of your local cache if you are reading the compressed shards from remote storage.