zephon.types¶
Canonical data model shared across the core data-loading pipeline.
- class ContributorRef(cursor, is_last_child=True)[source]¶
Bases:
objectReference to a contributing child derived from a base sample.
cursorpinpoints the base offset and lineage;is_last_child=Truedenotes the sole contributor (or tombstone) that closes that base offset for eviction purposes.- cursor: SampleCursor¶
- class SampleBatch(records)[source]¶
Bases:
objectA batch of SampleRecord.
- records: tuple[SampleRecord, ...]¶
- to_training(*, tokens_field='auto', return_labels=False, dtype='auto', extra_fields=(), ignore_index=-100, rename_fields=None, flatten=False, exclude_fields=(), return_num_valid_tokens=False, return_loss_mask=False, return_cu_seqlens=False, return_padding_mask=False, eos_mask_loss=False, eos_token_id=None, position_mode='preserve')[source]¶
Convert batch to training-ready format with optional LM label generation.
A
"loss_mask"payload field (a per-token supervised-vs-not mask, e.g. from chat-template tokenization) is recognized automatically, like"positions". Withreturn_labels=Trueit is consumed: labels whose mask entry (shifted into label alignment,mask[:, 1:]) is 0 are set toignore_index. Setreturn_loss_mask=Trueto emit the final label-aligned mask, derived aslabels != ignore_indexafter supervision, padding, and optional EOS masking. Listing"loss_mask"inextra_fieldstogether withreturn_labelsis an error: extra fields are sliced input-aligned ([:, :-1]), the wrong alignment for a label mask. Withoutreturn_labelsit is surfaced stacked and unshifted, like any data field.- Parameters:
tokens_field (str) – Field name containing token IDs. “auto” detects common names (input_ids, tokens, token_ids, ids), in that order. An explicit field takes precedence. Independent of EOS masking; output uses “input_ids” before rename_fields.
return_labels (bool) – If True, generates next-token prediction labels by shifting tokens. input_ids becomes tokens[:, :-1] and labels becomes tokens[:, 1:]. Extra fields are sliced like input_ids (drop the last token), so they stay aligned with it.
dtype (Any) – Tensor dtype for stacking. Use “auto” to detect (prefers torch.long if available, else np.int64, else returns lists). Use None to explicitly return lists instead of tensors. A requested loss mask always uses float32 (Python floats for lists), independently of this dtype. Requested cumulative sequence lengths and per-row maxima use int32 (Python ints for lists). With flatten=True, the maximum is always a Python int.
extra_fields (Sequence[str]) – Additional fields to include and stack (e.g., [“attention_mask”]). When return_labels=True they are sliced like input_ids (drop the last token) to stay aligned with it. A
"positions"field (emitted bypack_flat) is surfaced automatically and need not be listed here. Extra fields cannot overwrite generated output fields, such as input_ids or labels.ignore_index (int) – Loss-ignore sentinel (default
-100). Withreturn_labels, each record’s trailing pad labels (meta.padding_length) are set to this value.rename_fields (Mapping[str, str] | None) – Optional
{emitted_key: new_key}mapping applied to the output dict before exclusion (e.g.{"input_ids": "input"}). Source keys must be present in the output, except for empty batches, where missing sources are ignored. All emitted fields are renamed, including for empty batches, and their final names must be unique. Swaps are allowed.flatten (bool) – Flatten stacked scalar token fields from
[B, S]to[B * S]after per-sequence shifting and masking. The ids/texts fields remain per-record lists.exclude_fields (Sequence[str]) – Output names to omit after renaming, e.g.
("ids", "texts"). Absent names are ignored.return_num_valid_tokens (bool) – Include a Python int counting labels unequal to ignore_index after masking (zero for empty batches). Requires return_labels=True.
return_loss_mask (bool) – Include a label-aligned “loss_mask” with 1.0 for non-ignored labels and 0.0 otherwise. Requires return_labels=True. Uses float32 for tensors/arrays, Python floats for lists, and an empty list for empty batches. Follows flatten, rename_fields, and exclude_fields like other sequence fields.
return_padding_mask (bool) – Include a boolean input-aligned “padding_mask”, True only for trailing padding described by each record’s
meta.padding_length(missing metadata means no padding). Uses bool tensors/arrays or Python bools for lists. With labels, drops the last token like input_ids; then follows flatten, rename_fields, and exclude_fields. Independent of loss_mask: real unsupervised tokens are not padding. Requires 2-D scalar token rows for tensor/array output. Empty batches return []. Does not infer padding from token IDs or ignore_index.return_cu_seqlens (bool) – Include “cu_seqlens” and “max_seqlen” derived from zeros in input-aligned “positions”. Requires a positions sequence matching the tokens in every payload, starting at zero in every nonempty input row. With flatten=False, boundaries have shape [B, K], where K is the largest boundary count in the batch, padded with each row’s input length. Maxima have shape [B]. With flatten=True, boundaries are one compact cumulative sequence across the batch, ending at the total input length, and the maximum is a Python int. Boundaries and per-row maxima use int32 for tensors/arrays, Python ints for lists. Empty batches return []/[] for boundaries/maxima, or [0]/0 when flattened. Renaming and exclusion apply normally. Padding segments are preserved; this metadata does not itself enable attention masking or change labels. Always uses the payload positions, independently of position_mode. Works with or without return_labels.
eos_mask_loss (bool) –
If True, set labels to ignore_index wherever the corresponding input token equals eos_token_id. Requires return_labels=True and a known EOS ID (see eos_token_id). For packed
[BOS, A, EOS, BOS, B, EOS], this masks the EOS -> BOS prediction, not the A -> EOS or B -> EOS predictions. Requested loss masks and valid-token counts reflect this masking, together with padding and supervision masks.Defaults to False to preserve existing label generation. For documents bracketed with BOS and EOS, this keeps EOS -> BOS transitions supervised. Enable it when the training recipe excludes predictions made from EOS tokens.
CAREFUL: matching is by token ID, not document boundaries. If BOS and EOS share an ID, this also masks predictions after BOS, including the first content token.
eos_token_id (int | None) – Nonnegative EOS ID, required when eos_mask_loss=True. Supplying an ID alone does not enable masking and warns when eos_mask_loss=False. Automatic EOS propagation is in progress; for now, account for Zephon EOS overrides and tokenizer defaults when choosing the ID.
position_mode (Literal['preserve', 'sequence']) –
How to emit model positions. “preserve” (default) forwards payload positions, sliced like input_ids, and omits them when absent. With pack_flat(emit_positions=True), these restart at zero for each packed segment: two three-token documents have positions
[0, 1, 2, 0, 1, 2]. Preserving them keeps the packer’s document structure and existing caller behavior, and is a natural default for independent packed examples. It does not itself prevent attention across documents.”sequence” generates
[0, 1, ..., S-1]in every input row, including when payload positions are absent. The same example becomes[0, 1, 2, 3, 4, 5]. Use it for recipes that number a packed row continuously. Numbering restarts per row even with flatten=True. Generated positions use the input dtype/device; empty batches return an empty list.For reference, TorchTitan’s packed-text loaders preserve segment-reset positions by default, whereas Megatron gives the user control via –reset-position-ids and defaults to continuous positions.
Resetting is not required for standard RoPE correctness: with fixed rotary frequencies and isolated documents, a common position offset cancels from within-document attention scores. These choices can change training with cross-document attention or absolute position embeddings. They never change loss masking. Requested cu_seqlens/max_seqlen still use the original payload positions, so document boundaries survive continuous numbering; those original positions are still required for this metadata.
- Returns:
“ids”: List of sample IDs (always list, not stacked)
”texts”: List of text strings (always list, not stacked)
”input_ids”: Stacked token tensor (shifted if return_labels=True)
”labels”: Shifted labels tensor (only if return_labels=True), with each record’s trailing pad labels and loss-masked positions set to ignore_index
”positions”: Payload positions sliced to match input_ids, or generated per-row positions when position_mode=”sequence”
Any extra_fields as stacked tensors (shifted if return_labels=True)
”num_valid_tokens”: Number of non-ignored labels (only if return_num_valid_tokens=True)
”loss_mask”: Float mask of non-ignored labels if return_loss_mask=True; otherwise the raw payload mask if present and return_labels=False
”cu_seqlens”, “max_seqlen”: Cumulative segment boundaries and maximum segment lengths (only if return_cu_seqlens=True)
”padding_mask”: Boolean input padding mask (only if return_padding_mask=True)
- Return type:
Dictionary with
- Raises:
TypeError – If payloads are neither dicts nor arrays, or exclude_fields is a string instead of a sequence of field names.
ValueError – If tokens_field cannot be auto-detected or is missing, if a loss_mask is present in only some payloads or misaligned with the token field, if “loss_mask” is listed in extra_fields with return_labels=True, or if rename_fields references a missing source key in a nonempty batch or produces a key collision. Also if return_num_valid_tokens or return_loss_mask is used without return_labels, or an extra field would overwrite a generated output field. Also if requested padding lengths fall outside their token rows or tokens are not 2-D for a tensor/array padding mask. Also if requested boundaries cannot be derived from aligned, zero-start positions or represented as int32, position_mode is unknown, eos_token_id is not a nonnegative integer, or eos_mask_loss is used without return_labels=True or a known EOS ID.
- Warns:
UserWarning – If eos_token_id is supplied with eos_mask_loss=False.
- class SampleCursor(chunk_id, chunk_offset, sample_id, lineage=<factory>)[source]¶
Bases:
objectStable ordering key that survives fan-out across the pipeline.
- class SampleMeta(sample_id, lane_id, chunk_id, chunk_offset=0, component_sample_counts=<factory>, component_token_counts=None, lineage=<factory>, tags=<factory>)[source]¶
Bases:
objectLightweight metadata that uniquely identifies a sample in a shard.
lineagetracks the deterministic position of this record after any fan-out. Operators that split inputs must callchild()in the order elements are emitted so downstream consumers observe an ordering identical to the single-threaded execution semantics enforced by the runner.Contributor and tombstone flags are stored inside
tagsunder"_contributors"and"_tombstone"because in simple pipelines (notify_monotone path) that increases the IPC overhead since the public schema grows.Component Contribution Tracking¶
Samples track which mixture components they contain via two fields:
component_sample_counts: Maps component_id -> number of original samples from that component. For a regular sample this is{cid: 1}. For a packed sample combining 3 from component 0 and 2 from component 1:{0: 3, 1: 2}.component_token_counts: Maps component_id -> token count from that component. None until packing occurs after tokenization. When packing tokenized samples, captures tokens per component, e.g.,{0: 300, 1: 200}.This design supports ensure_mixture tracking per-component contributions:
weight="samples": usecomponent_sample_countsdirectlyweight="tokens": usecomponent_token_countsif set, else distribute total token count proportionally bycomponent_sample_counts
The separation allows packing before or after tokenization:
Pack before tokenize: only sample counts known at pack time
Pack after tokenize: both sample and token counts computed at pack time
- with_lineage(path)[source]¶
Return a new
SampleMetawhere the lineage is replaced bypath.- Return type:
- property cursor: SampleCursor¶
Return a
SampleCursorordering key for this metadata.
- contribution_refs()[source]¶
Return contributor references used for eviction/progress.
Operators that don’t set
contributors(1:1 outputs) are treated as emitting a single contributor withis_last_child=Trueusing their own cursor.- Return type:
tuple[ContributorRef, …]
- contribution_cursors()[source]¶
Return the cursors for all contributors.
- Return type:
tuple[SampleCursor, …]
- property contributors: tuple[ContributorRef, ...]¶
- property is_flush_sentinel: bool¶
True if this record is a flush sentinel (triggers accumulator flush).
- property padding_length: int | None¶
Trailing pad-token count of a flat
pack_flatrecord, elseNone.
- with_contributors(value)[source]¶
Return a new
SampleMetawith contributors set/cleared in tags.- Return type:
- class SampleRecord(meta, payload)[source]¶
Bases:
objectSample payload bundled with its metadata for transport through stages.
- meta: SampleMeta¶