Skip to content

Catalogue sample

cutana.catalogue_sample

Bounded catalogue reads for discovery and previews, never a full-file scan.

read_catalogue_sample(path, sample_size=100, seed=42)

Read a bounded discovery sample and report the population estimate.

CSV samples come from the first 10,000 rows. Parquet samples come from bounded windows in up to four randomly selected row groups. Neither is a uniform whole-catalogue sample. FITS binary tables support mapped row access.

The returned frame carries a plain positional index in every format, and that is part of the contract. The three readers have no common row numbering that can be produced within the read budget -- a Parquet row's absolute position needs the row counts of every group before it, which is a metadata scan proportional to the catalogue. So the index means "n-th row of this sample" and nothing more; callers naming a row to the user should quote its SourceID, which is unique and findable in their file.

Parameters:

Name Type Description Default
path

CSV, Parquet or FITS catalogue path.

required
sample_size

Maximum returned rows (at most 10,000).

100
seed

Reproducible local random seed.

42

Returns:

Type Description

Tuple of dataframe, population count, count-is-estimated, sampling scope.

Raises:

Type Description
ValueError

For unsupported formats or invalid sample sizes.

CatalogueValidationError

If a FITS catalogue carries no table extension.