zephon.io¶
Public IO API: datasets, the in-memory shard primitive, and store options.
Configure store behavior via Pipeline.options(io_options=StoreOptions(...)).
- class CacheOptions(enabled=_NotSupplied.TOKEN, root=_NotSupplied.TOKEN, limit_bytes=_NotSupplied.TOKEN, rg_cache_bytes=_NotSupplied.TOKEN, keep_zip=_NotSupplied.TOKEN, validate_hash=_NotSupplied.TOKEN, download_retry=_NotSupplied.TOKEN, download_timeout=_NotSupplied.TOKEN, open_retry_attempts=_NotSupplied.TOKEN, open_retry_initial_backoff=_NotSupplied.TOKEN, open_retry_max_backoff=_NotSupplied.TOKEN, min_slack_bytes=_NotSupplied.TOKEN, max_slack_bytes=_NotSupplied.TOKEN)[source]¶
Bases:
objectUser-configurable knobs that control cache behaviour.
- class Dataset(name, backend, path=None, _ids=None, _counts=None)[source]¶
Bases:
objectDescriptor for a dataset with index and backend info.
Instances are created via
from_path(file-backed) orfrom_dict(in-memory/testing). They always provide:name: a user-facing identifier used in mixturesbackend: opaque metadata that lets FetchOp build a reader laterpath: original filesystem path if file-backed, otherwiseNone
File-backed datasets also carry a few-KB handle to the node-local shard catalog (set by
from_path); this is what travels in ctx, notshard_meta.Shard counts are not stored as a mapping; read them via
ids()/counts()/total()/max_count()/shard_count().Backend kinds used by the internal store builder:
litdata/mds/jsonl/parquet/vortex:kindandpath(no per-shardshardsgraph — that lives in the catalog now)inmem:kindandshards(dict[int, InMemoryShard])
Note: this class does not expose any method to fetch rows; IO is delegated to an internal shard store owned by the FetchOp.
- raw_bytes()[source]¶
Return per-shard byte sizes (
int64), aligned withids().File-backed datasets read the catalog’s on-disk shard sizes; in-memory shards size their resident payloads lazily (
InMemoryShard.raw_bytes).- Return type:
- classmethod from_path(name, path, *, fmt=None)[source]¶
Construct a file-backed dataset descriptor.
Performs a count-only discovery: it obtains
shard_id+num_rowsas small numpy arrays (the work source’s input) without materializing the per-shardshard_metagraph. The full columnar catalog is built once per node later, by the Engine’sfinalize()(or lazily by the store builder for Engine-less use).- Parameters:
- Return type:
Supported formats:
litdatadirectories containingindex.jsonstructured withconfigandchunksmdsdirectories containingindex.jsonstructured withshardsjsonldirectories where*.jsonlfiles act as shards
Special URI schemes:
hf://org/name[@rev]/[config/]splitis served through the HuggingFace backend: the split’s uploaded files when their format is readable, else HuggingFace’s Parquet conversion when it is complete and built from the requested commit.fmtpicks the source in that format; appending~originalor~parquetto the revision (hf://org/name@~parquet/split) forces one. A partial conversion is used only withZEPHON_HF_ALLOW_PARTIAL=1. Uploaded files are read as stored: the dataset card’s reader options andfeaturescasting are not applied. The dataset’spathbecomes a URI pinned to the resolved commit, source and config.
- Returns:
a descriptor populated with shard counts and a catalog handle for later IO.
- Return type:
- Raises:
FileNotFoundError – if
pathdoes not exist.ValueError – if
pathis not a directory or format unsupported.
- class InMemoryShard(rows)[source]¶
Bases:
objectList-backed shard useful for tests and quickstarts.
- class ParquetRGCacheOptions(enabled=_NotSupplied.TOKEN, root=_NotSupplied.TOKEN, limit_bytes=_NotSupplied.TOKEN, min_free_bytes=_NotSupplied.TOKEN)[source]¶
Bases:
objectNode-shared decoded Parquet row-group cache configuration.
enabled=Noneenables the cache with the shard cache or an explicitroot. An omitted root uses the shard cache’s reserved.parquet-rg-cachechild. An omitted limit uses the deprecated shard option when present, otherwise the larger of 4 GiB and 10% of the shard cache limit. These defaults resolve when the store is built so option merges retain the distinction between omitted and explicit values. ExplicitNoneresets a previously supplied field to its automatic default.
- class StoreOptions(cache=<factory>, parquet_rg_cache=<factory>)[source]¶
Bases:
objectTop-level IO store options passed to FetchOp.
- cache: CacheOptions¶
- parquet_rg_cache: ParquetRGCacheOptions¶
- resolved_parquet_rg_cache()[source]¶
Resolve decoded-cache defaults against the shard-cache options.
- Return type: