litData¶
litData is a data processing and loading
library developed by Lightning AI, the team behind
PyTorch Lightning. Here, we are interested in
the binary format that litData reads and writes as a collection of records with a
consistent schema that are written into individual files called chunks along with an
index.json file that describes the contents of the chunks. (The chunks correspond to
what Zephon refers to as shards.)
The litData format is the one that we use for our large pretraining runs at DatologyAI
for both our text and multimodal models; we adopted the format because we liked the
flexibility that it provided for representing both tensors/numPy arrays and complex nested
Python dictionaries depending on the data schema that we needed for the file, and its
performance as it relies on a fast PyTree serializer.
Preparing litData Shards. To use the litData format with Zephon, install the project
with the litdata extra, which includes all of the dependencies needed for reading litData
shards:
# Install Zephon with the litdata extra.
pip install "zephon[litdata]"
Let’s look at an example of writing out a small litData dataset using the litData
library:
"""Write a small litData dataset, index.json included."""
from litdata import optimize
def samples(count: int) -> list[dict[str, object]]:
"""Return the records to serialize, one dict per sample."""
return [{"text": f"sample {i}", "label": i % 2} for i in range(count)]
def serialize(record: dict[str, object]) -> dict[str, object]:
"""Write each record through unchanged; litdata pickles this to its workers."""
return record
optimize(
fn=serialize,
inputs=samples(1000),
output_dir="data/litdata_demo",
chunk_bytes="64MB",
)
Note that this code generates both the shards and the index.json file that describes the
shard contents and structure; be sure to keep the index and the shards together if you
move the dataset to or from object storage. More examples of tools for generating and
preparing samples for litData shards are available in
litData’s data preparation guide.
Reading litData Shards. Once the litData files have been written, you can define a
Dataset by passing the dataset location to the Dataset.from_path method with the
fmt="litdata" argument. Zephon will infer the format if fmt is not passed. You can then
use a DatasetInspector to examine the reconstructed sample payloads to confirm that the
samples are in the form that you expect:
"""Define and inspect a litData Dataset."""
from zephon.debug import DatasetInspector
from zephon.io import Dataset
# fmt is optional: Zephon infers the format from the index.json beside the
# chunks. Keep the two together when you move the dataset.
dataset = Dataset.from_path("litdata_demo", "data/litdata_demo", fmt="litdata")
with DatasetInspector(dataset) as inspector:
print(inspector.read(shard_id=0, sample_index=0))
Zephon also supports litData shards that are compressed using Zstandard. Depending on the contents of the shards, this may significantly reduce the size of the shards and thus the storage/networking cost of moving them around. Keep in mind that the compressed shards will be fully decompressed before they can be read, so plan for this when configuring the size of your local cache if you are reading the compressed shards from remote storage.