Fits dataset
cutana.fits_dataset
¶
FITS Dataset management for Cutana - handles process-level FITS file caching.
This module provides the FITSDataset class that manages FITS file loading and caching at the process level to avoid reloading the same files across sub-batches.
FITSDataset(config, profiler=None, job_tracker=None, process_name=None)
¶
Manages process-level FITS file caching to avoid reloading same files across sub-batches.
This class handles: - Process-level FITS caching - Loading only missing FITS files - Smart memory management to free unused files - Cleanup on completion
initialize_from_sources(source_batch)
¶
Initialize the dataset by preparing FITS sets for all sources.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_batch
|
List[Dict[str, Any]]
|
List of all source dictionaries for the process |
required |
prepare_sub_batch(sub_batch)
¶
Prepare FITS data for a sub-batch, loading only missing files.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sub_batch
|
List[Dict[str, Any]]
|
List of source dictionaries for this sub-batch |
required |
Returns:
| Type | Description |
|---|---|
Dict[str, Tuple[HDUList, Dict[str, WCS]]]
|
Dictionary of FITS data needed for this sub-batch |
free_unused_after_sub_batch(current_sub_batch, remaining_sub_batches)
¶
Free FITS files that won't be needed in remaining sub-batches.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
current_sub_batch
|
List[Dict[str, Any]]
|
Current sub-batch that was just processed |
required |
remaining_sub_batches
|
List[List[Dict[str, Any]]]
|
List of remaining sub-batches |
required |
cleanup()
¶
Clean up all remaining FITS files in the cache.
load_fits_sets(fits_sets, fits_extensions, config=None, profiler=None)
¶
Load FITS files for given FITS file sets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fits_sets
|
List[tuple]
|
List of FITS file set tuples |
required |
fits_extensions
|
List[str]
|
List of FITS extensions to load |
required |
config
|
DotMap
|
Configuration DotMap (unused, kept for compatibility) |
None
|
profiler
|
Optional[PerformanceProfiler]
|
Optional performance profiler |
None
|
Returns:
| Type | Description |
|---|---|
Dict[str, Tuple[HDUList, Dict[str, WCS]]]
|
Dictionary mapping fits_path -> (hdul, wcs_dict) |
prepare_fits_sets_and_sources(source_batch)
¶
Parse FITS paths and group sources by their FITS file sets.
Now uses the extract_fits_sets function from catalogue_preprocessor for consistency.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_batch
|
List[Dict[str, Any]]
|
List of source dictionaries |
required |
Returns:
| Type | Description |
|---|---|
Dict[tuple, List[Dict[str, Any]]]
|
Dictionary mapping FITS file sets (as tuples) to lists of sources |
get_selected_band_names(config)
¶
Band names named by selected_extensions, or None to load every file.
Shared by the orchestrator's dataset loading and by create_cutouts_direct:
both have to narrow a FITS set the same way, and when they disagree the tensor
ends up with more channels than channel_weights has entries, which
combine_channels then applies positionally.
A selection naming no band extract_filter_name can produce is not band
selection, so narrowing is off. In practice that is the UI on non-Euclid data:
it fills selected_extensions from analyse_source_catalogue, whose name
is extract_filter_name's own output, so a file the recogniser cannot classify
arrives as UNKNOWN — a label with no band behind it to narrow on. A Python
caller can put anything there, HDU names included, and it is read the same way.
Decided from the selection rather than from the match result, which would conflate "this names no band" with "this names a band the set does not have"; the second case would then load whatever the set happens to contain.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DotMap
|
Cutana configuration DotMap. |
required |
Returns:
| Type | Description |
|---|---|
Optional[Set[str]]
|
Set of band names (e.g. |
select_fits_set_bands(fits_set, band_names)
¶
Narrow a FITS set to the requested bands, preserving the catalogue order.
Order is preserved because channel_weights is applied positionally against
the loaded channels; reordering here would silently pair weights with the
wrong bands.
Whether the selection is band selection at all is decided once, by
get_selected_band_names; band_names is None when it is not.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fits_set
|
tuple
|
FITS file paths forming one set (one tile's channels). |
required |
band_names
|
Optional[Set[str]]
|
Bands to keep, from |
required |
Returns:
| Type | Description |
|---|---|
tuple
|
The filtered set, in catalogue order. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the selection names bands but matches no file in the set.
The configuration asks for data this set does not contain, and loading
the set whole would hand |