Skip to content

Direct cutout

cutana.direct_cutout

Fast in-process cutout generation for small batches.

Provides create_cutouts_direct() for rapid cutout creation without subprocess overhead. Reuses the same vectorized processing pipeline as the full orchestrator but runs entirely in-process with mmap FITS loading.

Typical use cases: - Preview generation (< 200 cutouts) - Quick-look analysis (< 500 cutouts) - Interactive exploration in notebooks

Requests usually carry few cutouts spread across many tiles, so the cost is dominated by opening/reading tiles (I/O bound). Distinct tiles are processed on a thread pool: FITS I/O releases the GIL, while a process pool's per-task pickling of cutout arrays/WCS plus interpreter startup measured far slower for this shape.

For large catalogues (> 1000 sources), use StreamingOrchestrator instead.

create_cutouts_direct(catalogue_df, config, max_workers=None, *, log_set_progress=True)

Create cutouts directly in-process without subprocess overhead.

This is the fast path for small batches (< ~1000 sources). It runs the same vectorized processing pipeline as the full orchestrator but avoids subprocess spawning, JSON/TOML serialization, shared memory IPC, and job tracking.

FITS files are loaded with memory mapping for fast access. Each distinct FITS set (tile) is loaded and closed in isolation, so peak memory scales with the number of concurrently processed tiles rather than the whole request.

Parameters:

Name Type Description Default
catalogue_df DataFrame

DataFrame with columns: SourceID, RA, Dec, diameter_pixel (or diameter_arcsec), fits_file_paths.

required
config DotMap

Cutana configuration DotMap (from get_default_config()).

required
max_workers Optional[int]

Number of worker threads for processing tiles concurrently. None (default) auto-selects min(n_tiles, effective_cpus, 8) using the k8s/cgroup-aware CPU count. Pass 1 to force serial processing; a single-tile request always runs serially regardless.

None
log_set_progress bool

When True (default) and the request spans more than one FITS set, log a per-set completion heartbeat as tiles finish. Callers that invoke this many times in a tight loop (e.g. once per catalogue) should pass False — the heartbeat is then pure noise and the caller's own per-batch logging is the right place for progress.

True

Returns:

Type Description
List[Dict[str, Any]]

List of batch result dicts (one per FITS set, in catalogue grouping order),

List[Dict[str, Any]]

each containing: - "cutouts": ndarray of shape (N_sources, H, W, N_channels) - "metadata": list of per-source metadata dicts - "wcs": list of WCS dicts per source - "channel_names": list of channel name strings

Raises:

Type Description
ValueError

If catalogue_df is empty, max_workers < 1, or selected_extensions names bands that match no file in one of the catalogue's FITS sets. The last is the one a misconfigured run actually hits: the orchestrator and streaming backends log it and skip that set, but here it reaches the caller.

KeyError

If required columns are missing.

RuntimeError

If no valid cutouts were generated.