Vortex¶
Vortex is a columnar file format designed to support both efficient
random access and efficient compressed storage. Like Parquet, it integrates into the Arrow
data ecosystem, which makes it worth considering if your existing data preparation tools
work well with Arrow tables. Unlike litData and MDS, Vortex files are self-describing and
do not ship with a separate index.json file to describe how their contents are structured
and stored. You can think of Vortex as a modern Parquet with fast random access. You can
read more about the Vortex format’s structure in the
project’s documentation.
Preparing Vortex Shards. To use the Vortex format with Zephon, install the project
with the vortex extra which includes the vortex-data dependency (note that Vortex
requires Python >= 3.11):
# Install Zephon with the vortex extra (needs Python 3.11 or newer).
pip install "zephon[vortex]"
Let’s look at an example of writing a small collection of records into Vortex files:
"""Write records into Vortex shards."""
import vortex as vx
records = [{"text": f"sample {i}", "label": i % 2} for i in range(1000)]
# One file per shard; 64-256 MB each is the range we recommend.
for shard, start in enumerate(range(0, len(records), 250)):
array = vx.array(records[start : start + 250])
vx.io.write(array, f"data/vortex_demo/shard{shard}.vortex")
The Vortex python library can also write data from Arrow tables, which provides a path for converting existing tabular datasets into this format:
"""Convert an Arrow table into a Vortex shard."""
import pyarrow as pa
import vortex as vx
table = pa.table({"text": ["first", "second"], "label": [0, 1]})
vx.io.write(vx.array(table), "data/vortex_demo/shard0.vortex")
Although Vortex does not require a separate index file, Zephon provides an indexing tool
that collects and packages the shard names, file sizes, and record counts into an
index.json file that helps for efficiently planning the work to be done during training
without requiring Zephon to load each Vortex shard individually:
# Build an index.json so Zephon need not open every shard to plan.
python -m zephon.build_index vortex data/vortex_demo
Reading Vortex Shards in Zephon. Once the Vortex files have been written, you can
define a Dataset object by passing the dataset location and the fmt="vortex" argument
to the Dataset.from_path function. Zephon will infer the format if fmt is not passed.
You can then use the DatasetInspector to examine the decoded sample payloads:
"""Define and inspect a Vortex Dataset."""
from zephon.debug import DatasetInspector
from zephon.io import Dataset
# Vortex files are self-describing, so the index is an optimization, not a
# requirement. Compression is handled inside the format: no extra local copy.
dataset = Dataset.from_path("vortex_demo", "data/vortex_demo", fmt="vortex")
with DatasetInspector(dataset) as inspector:
print(inspector.read(shard_id=0, sample_index=0))
Note that Vortex handles compression operations within the file format itself, so Zephon
reads the .vortex files directly without requiring an extra, decompressed copy of the
shard to be created.