Skip to content

Production Dataset

This page documents the maintained production dataset path. The active pixel workflow reads the self-contained Hugging Face-style dataset folder at /work/data/OceanVariableReconstruction through ArgoGeoTIFFGriddedPatchDataset, the maintained production dataloader path.

Use Data Sources for native product properties, Depth Alignment for ARGO-to-GLORYS vertical resampling, and Data Export for rebuilding the packaged folder.

Dataset Root

Expected local structure:

/work/data/OceanVariableReconstruction/
  manifest.yaml
  rasters/
  argo/argo_profiles_on_grid.zarr/
  data/argo_glors_ostia_ssh.zarr/
  masks/world_land_mask_glorys_0p1.tif
  metadata/
  indices/

The active config is src/depth_recon/configs/px_space/training_super_config.yaml. It sets data.dataset.core.geotiff_root_dir to the package root and uses dataset-root-relative mask paths such as masks/world_land_mask_glorys_0p1.tif.

Dataset Assembly

At runtime, the loader:

  1. Reads manifest.yaml from the package root.
  2. Resolves dense raster paths under root-level rasters/.
  3. Opens compact ARGO profiles from argo/argo_profiles_on_grid.zarr.
  4. Builds a deterministic land-mask-derived patch grid from masks/.
  5. Assigns train/val from split.val_year when overlapping patches are enabled.
  6. Reads three ordered surface rasters, target or fixed depth support, and compact ARGO profiles lazily per sample.

Only metadata caches are written under dataset.core.metadata_cache_dir. Model-facing tensors are produced on demand.

Patch Grid Concept

A patch is a fixed-size window on the 0.1 degree GLORYS grid. The production configuration uses tile_size: 128, so each patch covers 128 by 128 grid cells. Instead of placing every patch once in a non-overlapping grid, the loader moves the window by patch_stride. The default GeoTIFF stride is 32 cells, which means nearby patches overlap by 75% of their width and height.

The same world grid is reused every time the dataset is instantiated. That makes the patch locations deterministic: changing the date range changes which timesteps are available, but it does not move the patch boundaries.

Patch Filtering

The grid is built from the packaged GLORYS-aligned land-mask GeoTIFF: /work/data/OceanVariableReconstruction/masks/world_land_mask_glorys_0p1.tif. In that mask, 1 means land and 0 means ocean. For each candidate patch, the dataset computes the fraction of land pixels and keeps the patch when land_fraction <= dataset.grid.max_land_fraction.

The default cap is 0.30, so normal retained patches are at least 70% ocean. Patches centered on the Mediterranean, Baltic, Red Sea, and Hudson Bay are force-included with a relaxed land cap through dataset.grid.force_include_regions, because those water bodies are narrow or coastline-heavy and would otherwise lose useful ocean context around coastlines.

Hard-area sampling can temporarily extend these relaxed grid regions through dataset.finetune_sampling.hard_regions when dataset.finetune_sampling.enabled=true and dataset.finetune_sampling.relax_land_filter=true. That run-specific extension lets coast-heavy patches enter whichever splits are named by apply_to_splits. The current local preset applies a 50/50 hard/easy mix to both train and val; HPC presets disable this row filter.

Overlapping patches mean an ARGO profile is not tied to only one spatial context. If a profile falls inside several retained patch bounds, it can contribute support to each matching (patch, date) row. The stride-32 GeoTIFF preset increases these contexts compared with the earlier half-overlap grid, giving each profile more local visual neighborhoods during training.

Patch Registry Storage

During dataset instantiation, the loader expands the retained patch table across the available OSTIA dates, then filters those dates to the GLORYS and sea-level coverage already present on disk. The resulting registry is a table of (patch_id, date) rows and becomes dataset.rows.

Each patch row stores the grid indices, latitude/longitude bounds, center coordinates, land_fraction, ocean_fraction, invalid_fraction, and any force-include metadata. Each date row stores the timestep, split assignment, and optional ARGO-support count. Cache filenames include the grid source, stride, tile size, land threshold, temporal window, split policy, and mask metadata so a changed configuration creates a new cache instead of reusing stale rows.

The cache is metadata only. It records where patches are and which dates are valid; it does not store precomputed GLORYS, OSTIA, sea-level, or ARGO tensors.

Sample Read Path

When training asks for an item, __getitem__ reads one registry row, converts the stored patch bounds back into source-file slices, and lazily loads the matching GLORYS, OSTIA, and sea-level data for that date. ARGO profiles are selected by the patch bounds and the configured temporal window, projected onto the GLORYS depth axis, then rasterized into the sample tensors and validity masks.

Spatial And Temporal Semantics

  • dataset.grid.tile_size controls patch height/width.
  • dataset.grid.resolution_deg controls patch pixel spacing.
  • dataset.grid.patch_stride controls patch overlap; values below tile_size require split.val_year.
  • dataset.grid.max_land_fraction filters land-heavy patches from the committed GLORYS-aligned world mask.
  • dataset.grid.force_include_regions keeps patches centered on the Mediterranean, Baltic, Red Sea, and Hudson Bay up to a relaxed land fraction so the training registry retains those water bodies.
  • dataset.finetune_sampling.* can filter configured splits to named hard-region patch centers and add those boxes as run-specific relaxed land-fraction regions.
  • dataset.sampling.temporal_window_days controls the centered ARGO profile search window for each patch date.
  • dataset.selection.require_argo_for_train is false in maintained presets so all otherwise eligible training locations are retained; dataset.selection.require_argo_for_val remains true.
  • dataset.selection.require_argo_for_all defaults to false so global inference can cover rows without ARGO observations.
  • split.val_year defaults to 2016, assigning that year to validation and all other years to training.

Depth Semantics

  • Normal dense training uses GLORYS thetao for y; optional Stage 1 uses an online deterministic surface-offset target with per-depth confidence.
  • ARGO POTM_CORRECTED is projected from DEPH_CORRECTED samples onto the GLORYS depth axis before rasterization, matching GLORYS potential-temperature thetao. Pass --temperature-source in-situ only to reproduce older exports that used measured TEMP.
  • dataset.depth_axis_m exposes the physical GLORYS depth levels to inference and export code.

Output Contract

Each sample returns three-channel [SST, SSS, ADT] eo, x, y, x_valid_mask, y_valid_mask, x_valid_mask_1d, land_mask, date, optional Stage 1 y_supervision_weight, and optional coords/info. x_valid_mask is ARGO observation support. land_mask uses dated GLORYS support in the normal path and the fixed-reference depth mask in Stage 1, with surface/disk-mask fallbacks. The common on-disk mask is loaded only by prediction/export paths when final cleanup is needed.

Stage 1 and Stage 2 keep identical [sst, sss, adt] channel order. See Vertical-Offset Pretraining for leakage and scientific-evaluation requirements.

See Data Contract for the full tensor contract.

Finetuning

Hard-area row sampling is enabled in the local/standard presets and disabled in the HPC presets. The local preset applies it to both train and val, keeps all rows whose patch centers fall inside configured hard regions, and adds a deterministic sample of easy rows to target hard_fraction: 0.5.

When data.dataset.finetune_sampling.relax_land_filter=true, the same hard-region boxes are also added as run-specific relaxed land-fraction regions before the patch registry is built. The configured polygons are provisional diagnostic regions, not literature-backed scientific boundaries.