Catalogue sample
cutana.catalogue_sample
¶
Bounded catalogue reads for discovery and previews, never a full-file scan.
read_catalogue_sample(path, sample_size=100, seed=42)
¶
Read a bounded discovery sample and report the population estimate.
CSV samples come from the first 10,000 rows. Parquet samples come from bounded windows in up to four randomly selected row groups. Neither is a uniform whole-catalogue sample. FITS binary tables support mapped row access.
The returned frame carries a plain positional index in every format, and that is part
of the contract. The three readers have no common row numbering that can be produced
within the read budget -- a Parquet row's absolute position needs the row counts of
every group before it, which is a metadata scan proportional to the catalogue. So the
index means "n-th row of this sample" and nothing more; callers naming a row to the
user should quote its SourceID, which is unique and findable in their file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
CSV, Parquet or FITS catalogue path. |
required | |
sample_size
|
Maximum returned rows (at most 10,000). |
100
|
|
seed
|
Reproducible local random seed. |
42
|
Returns:
| Type | Description |
|---|---|
|
Tuple of dataframe, population count, count-is-estimated, sampling scope. |
Raises:
| Type | Description |
|---|---|
ValueError
|
For unsupported formats or invalid sample sizes. |
CatalogueValidationError
|
If a FITS catalogue carries no table extension. |