Direct cutout
cutana.direct_cutout
¶
Fast in-process cutout generation for small batches.
Provides create_cutouts_direct() for rapid cutout creation without subprocess overhead. Reuses the same vectorized processing pipeline as the full orchestrator but runs entirely in-process with mmap FITS loading.
Typical use cases: - Preview generation (< 200 cutouts) - Quick-look analysis (< 500 cutouts) - Interactive exploration in notebooks
Requests usually carry few cutouts spread across many tiles, so the cost is dominated by opening/reading tiles (I/O bound). Distinct tiles are processed on a thread pool: FITS I/O releases the GIL, while a process pool's per-task pickling of cutout arrays/WCS plus interpreter startup measured far slower for this shape.
For large catalogues (> 1000 sources), use StreamingOrchestrator instead.
create_cutouts_direct(catalogue_df, config, max_workers=None, *, log_set_progress=True)
¶
Create cutouts directly in-process without subprocess overhead.
This is the fast path for small batches (< ~1000 sources). It runs the same vectorized processing pipeline as the full orchestrator but avoids subprocess spawning, JSON/TOML serialization, shared memory IPC, and job tracking.
FITS files are loaded with memory mapping for fast access. Each distinct FITS set (tile) is loaded and closed in isolation, so peak memory scales with the number of concurrently processed tiles rather than the whole request.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
catalogue_df
|
DataFrame
|
DataFrame with columns: SourceID, RA, Dec, diameter_pixel (or diameter_arcsec), fits_file_paths. |
required |
config
|
DotMap
|
Cutana configuration DotMap (from get_default_config()). |
required |
max_workers
|
Optional[int]
|
Number of worker threads for processing tiles concurrently. None (default) auto-selects min(n_tiles, effective_cpus, 8) using the k8s/cgroup-aware CPU count. Pass 1 to force serial processing; a single-tile request always runs serially regardless. |
None
|
log_set_progress
|
bool
|
When True (default) and the request spans more than one FITS set, log a per-set completion heartbeat as tiles finish. Callers that invoke this many times in a tight loop (e.g. once per catalogue) should pass False — the heartbeat is then pure noise and the caller's own per-batch logging is the right place for progress. |
True
|
Returns:
| Type | Description |
|---|---|
List[Dict[str, Any]]
|
List of batch result dicts (one per FITS set, in catalogue grouping order), |
List[Dict[str, Any]]
|
each containing: - "cutouts": ndarray of shape (N_sources, H, W, N_channels) - "metadata": list of per-source metadata dicts - "wcs": list of WCS dicts per source - "channel_names": list of channel name strings |
Raises:
| Type | Description |
|---|---|
ValueError
|
If catalogue_df is empty, max_workers < 1, or |
KeyError
|
If required columns are missing. |
RuntimeError
|
If no valid cutouts were generated. |