Parquet¶
What Is It? Apache Parquet is an open-source and widely-adopted columnar file format for data storage and retrieval. It is ubiquitous in the data warehouse community as a way of efficiently storing and querying datasets, and it can be read and written by most databases and analytical data processing tools in every major programming language. It strikes a reasonable balance between the easy inspection and editing of JSONL and the training efficiency of binary shard formats.
What Is It Good For? Parquet is the most common choice for training datasets that have a natural tabular structure, or when the training is being done at a company that has standardized on Parquet for other data workloads and has lots of tooling to support working with it.
What Is It Bad For? Parquet’s columnar layout is often less efficient when your training curriculum requires frequent random access to individual records, which is typically the case. Zephon needs to read and decode the row groups containing the requested samples, which often requires processing more data than the samples themselves. This makes the organization of your shards and your data access patterns substantially more important for you to think about when you are training against Parquet datasets. Unless there is an organizational reason to use Parquet, we prefer training-optimized binary shard formats over it.
How Does Zephon Collect Dataset Metadata? Each *.parquet file under the Dataset path
is a shard. Parquet embeds useful metadata in each file, including its row count, so
Zephon can discover a Dataset without a separate index file and without having to parse
each record in the file, as it does for JSONL files. Still, the footer of each file needs
to be read, which at large scale can take a significant amount of time.
For larger Parquet datasets that are comprised of a large number of shards, Zephon
includes an indexing tool that can be used to create an index.json file that the Dataset
can use to quickly collect shard metadata information without needing to parse each of the
individual shards to collect their metadata for faster discovery at the cost of a one-time
index creation run:
# Build an index.json so Zephon can plan without reading every footer.
python -m zephon.build_index parquet s3://my-bucket/corpora/wikipedia
If you have existing infrastructure for writing out Parquet files using distributed data
processing frameworks like Spark or Ray, you might want to write out an index.json file
while creating the dataset.
Reading Parquet in Zephon
"""Define and inspect a Parquet Dataset."""
from zephon.debug import DatasetInspector
from zephon.io import Dataset
# Parquet carries its row count in each file footer, so a Dataset can be
# discovered without an index.json -- an index just makes it faster.
dataset = Dataset.from_path("wikipedia", "s3://my-bucket/corpora/wikipedia")
with DatasetInspector(dataset) as inspector:
print(inspector.read(shard_id=0, sample_index=0))