Vortex

Vortex is a columnar file format designed to support both efficient random access and efficient compressed storage. Like Parquet, it integrates into the Arrow data ecosystem, which makes it worth considering if your existing data preparation tools work well with Arrow tables. Unlike litData and MDS, Vortex files are self-describing and do not ship with a separate index.json file to describe how their contents are structured and stored. You can think of Vortex as a modern Parquet with fast random access. You can read more about the Vortex format’s structure in the project’s documentation.

Preparing Vortex Shards. To use the Vortex format with Zephon, install the project with the vortex extra which includes the vortex-data dependency (note that Vortex requires Python >= 3.11):

examples/guide/datasets/vortex_install.sh
# Install Zephon with the vortex extra (needs Python 3.11 or newer).

pip install "zephon[vortex]"

Let’s look at an example of writing a small collection of records into Vortex files:

examples/guide/datasets/vortex_write_shards.py
"""Write records into Vortex shards."""

import vortex as vx

records = [{"text": f"sample {i}", "label": i % 2} for i in range(1000)]

# One file per shard; 64-256 MB each is the range we recommend.
for shard, start in enumerate(range(0, len(records), 250)):
    array = vx.array(records[start : start + 250])
    vx.io.write(array, f"data/vortex_demo/shard{shard}.vortex")

The Vortex python library can also write data from Arrow tables, which provides a path for converting existing tabular datasets into this format:

examples/guide/datasets/vortex_from_arrow.py
"""Convert an Arrow table into a Vortex shard."""

import pyarrow as pa
import vortex as vx

table = pa.table({"text": ["first", "second"], "label": [0, 1]})

vx.io.write(vx.array(table), "data/vortex_demo/shard0.vortex")

Although Vortex does not require a separate index file, Zephon provides an indexing tool that collects and packages the shard names, file sizes, and record counts into an index.json file that helps for efficiently planning the work to be done during training without requiring Zephon to load each Vortex shard individually:

examples/guide/datasets/vortex_index.sh
# Build an index.json so Zephon need not open every shard to plan.

python -m zephon.build_index vortex data/vortex_demo

Reading Vortex Shards in Zephon. Once the Vortex files have been written, you can define a Dataset object by passing the dataset location and the fmt="vortex" argument to the Dataset.from_path function. Zephon will infer the format if fmt is not passed. You can then use the DatasetInspector to examine the decoded sample payloads:

examples/guide/datasets/vortex_dataset.py
"""Define and inspect a Vortex Dataset."""

from zephon.debug import DatasetInspector
from zephon.io import Dataset

# Vortex files are self-describing, so the index is an optimization, not a
# requirement. Compression is handled inside the format: no extra local copy.
dataset = Dataset.from_path("vortex_demo", "data/vortex_demo", fmt="vortex")

with DatasetInspector(dataset) as inspector:
    print(inspector.read(shard_id=0, sample_index=0))

Note that Vortex handles compression operations within the file format itself, so Zephon reads the .vortex files directly without requiring an extra, decompressed copy of the shard to be created.