Skip to content

Validation sampling

cutana.validation_sampling

Choosing which rows a validation check looks at.

Three checks had grown their own copy of "if the catalogue is large, take a deterministic sample", each with its own threshold and its own bare random_state=42. One copy, one seed.

The seed matters more than it looks: validation runs twice in a normal session -- once when the catalogue is loaded and once before the run -- and a report that named different rows each time would read as a catalogue that keeps changing.

sample_for_validation(catalogue_df, sample_size, what=None)

The rows to check: all of them, or a deterministic sample of a large catalogue.

Parameters:

Name Type Description Default
catalogue_df DataFrame

The catalogue.

required
sample_size int

Most rows to return.

required
what Optional[str]

What is being checked, for the log line. A sample that is not mentioned reads as a full pass, and a clean report over 0.1% of a catalogue is worth saying out loud.

None

Returns:

Type Description
DataFrame

The whole catalogue, or sample_size rows of it chosen the same way every time.