Mosaic (MDS)¶
Mosaic Data Shards, or MDS, is
the binary format used by the MosaicML Streaming data loader. An MDS dataset consists of a
collection of shard files along with an index.json file that describes their contents.
Each MDS record is a dictionary of named fields whose encodings are specified when the
dataset is written. These encodings support different kinds of training data, including
text, numerical arrays, and images. If you already have data preparation tooling that
writes MDS shards for MosaicML Streaming, you can use those shards directly in Zephon
without needing to convert them to another file format.
Preparing MDS Shards. To use the MDS format with Zephon, install the project with the
mds extra, which includes the MosaicML streaming library and the dependencies needed
for reading both uncompressed shards and shards compressed using Zstandard:
# Install Zephon with the mds extra.
pip install "zephon[mds]"
Let’s look at an example of writing out a small MDS dataset using the MDSWriter class
from the streaming library:
"""Write a small MDS dataset with the streaming library."""
import numpy as np
from streaming import MDSWriter
# One encoding per field; every record written must match what it declares.
columns = {"text": "str", "tokens": "ndarray"}
with MDSWriter(out="data/mds_demo", columns=columns, compression="zstd") as writer:
for i in range(1000):
writer.write(
{"text": f"sample {i}", "tokens": np.arange(i % 32 + 1, dtype=np.int32)}
)
The columns argument maps each field name to the encoding that should be used to write
its values. For example, a text field can use the str encoding while an array of tokens
can use the ndarray encoding. Each record that you write must provide values that are
compatible with its declared encoding. More details and examples are available in the
MosaicML Streaming data preparation guide.
Reading MDS Shards. Once the MDS files have been written, you can define a Dataset
by passing in the dataset location to the Dataset.from_path function with the
fmt="mds" argument. Zephon will infer the format if fmt is not passed. You can then use
a DatasetInspector to examine the decoded sample payloads:
"""Define and inspect an MDS Dataset."""
from zephon.debug import DatasetInspector
from zephon.io import Dataset
dataset = Dataset.from_path("mds_demo", "data/mds_demo", fmt="mds")
with DatasetInspector(dataset) as inspector:
print(inspector.read(shard_id=0, sample_index=0))
As is the case for other binary shard formats, keep the binary shard files and the
corresponding index.json file together when moving the dataset so that Zephon can read
them. Zstandard compression can make the shards cheaper to store and transport, but they
will be fully decompressed before any records are read, so account for their decompressed
sizes when planning your
local cache settings when processing datasets in
remote storage.