Skip to content

Fits dataset

cutana.fits_dataset

FITS Dataset management for Cutana - handles process-level FITS file caching.

This module provides the FITSDataset class that manages FITS file loading and caching at the process level to avoid reloading the same files across sub-batches.

FITSDataset(config, profiler=None, job_tracker=None, process_name=None)

Manages process-level FITS file caching to avoid reloading same files across sub-batches.

This class handles: - Process-level FITS caching - Loading only missing FITS files - Smart memory management to free unused files - Cleanup on completion

initialize_from_sources(source_batch)

Initialize the dataset by preparing FITS sets for all sources.

Parameters:

Name Type Description Default
source_batch List[Dict[str, Any]]

List of all source dictionaries for the process

required

prepare_sub_batch(sub_batch)

Prepare FITS data for a sub-batch, loading only missing files.

Parameters:

Name Type Description Default
sub_batch List[Dict[str, Any]]

List of source dictionaries for this sub-batch

required

Returns:

Type Description
Dict[str, Tuple[HDUList, Dict[str, WCS]]]

Dictionary of FITS data needed for this sub-batch

free_unused_after_sub_batch(current_sub_batch, remaining_sub_batches)

Free FITS files that won't be needed in remaining sub-batches.

Parameters:

Name Type Description Default
current_sub_batch List[Dict[str, Any]]

Current sub-batch that was just processed

required
remaining_sub_batches List[List[Dict[str, Any]]]

List of remaining sub-batches

required

cleanup()

Clean up all remaining FITS files in the cache.

load_fits_sets(fits_sets, fits_extensions, config=None, profiler=None)

Load FITS files for given FITS file sets.

Parameters:

Name Type Description Default
fits_sets List[tuple]

List of FITS file set tuples

required
fits_extensions List[str]

List of FITS extensions to load

required
config DotMap

Configuration DotMap (unused, kept for compatibility)

None
profiler Optional[PerformanceProfiler]

Optional performance profiler

None

Returns:

Type Description
Dict[str, Tuple[HDUList, Dict[str, WCS]]]

Dictionary mapping fits_path -> (hdul, wcs_dict)

prepare_fits_sets_and_sources(source_batch)

Parse FITS paths and group sources by their FITS file sets.

Now uses the extract_fits_sets function from catalogue_preprocessor for consistency.

Parameters:

Name Type Description Default
source_batch List[Dict[str, Any]]

List of source dictionaries

required

Returns:

Type Description
Dict[tuple, List[Dict[str, Any]]]

Dictionary mapping FITS file sets (as tuples) to lists of sources

get_selected_band_names(config)

Band names named by selected_extensions, or None to load every file.

Shared by the orchestrator's dataset loading and by create_cutouts_direct: both have to narrow a FITS set the same way, and when they disagree the tensor ends up with more channels than channel_weights has entries, which combine_channels then applies positionally.

A selection naming no band extract_filter_name can produce is not band selection, so narrowing is off. In practice that is the UI on non-Euclid data: it fills selected_extensions from analyse_source_catalogue, whose name is extract_filter_name's own output, so a file the recogniser cannot classify arrives as UNKNOWN — a label with no band behind it to narrow on. A Python caller can put anything there, HDU names included, and it is read the same way.

Decided from the selection rather than from the match result, which would conflate "this names no band" with "this names a band the set does not have"; the second case would then load whatever the set happens to contain.

Parameters:

Name Type Description Default
config DotMap

Cutana configuration DotMap.

required

Returns:

Type Description
Optional[Set[str]]

Set of band names (e.g. {"VIS"}), or None when no filtering applies.

select_fits_set_bands(fits_set, band_names)

Narrow a FITS set to the requested bands, preserving the catalogue order.

Order is preserved because channel_weights is applied positionally against the loaded channels; reordering here would silently pair weights with the wrong bands.

Whether the selection is band selection at all is decided once, by get_selected_band_names; band_names is None when it is not.

Parameters:

Name Type Description Default
fits_set tuple

FITS file paths forming one set (one tile's channels).

required
band_names Optional[Set[str]]

Bands to keep, from get_selected_band_names, or None to keep the set unchanged.

required

Returns:

Type Description
tuple

The filtered set, in catalogue order.

Raises:

Type Description
ValueError

If the selection names bands but matches no file in the set. The configuration asks for data this set does not contain, and loading the set whole would hand combine_channels the wrong bands under the requested names. What a miss costs is the caller's decision: the direct path lets it out, while FITSDataset._load_missing_fits_files skips that set, having already refused a selection that misses every set.