Python API
The primary public entry point is ctdcast.report().
Main entry point
- ctdcast.reports._index.report(nc_dir: Path, out_dir: Path, *, profiles_path: Path | None = None, groupings_yaml: Path | None = None, section_yaml: Path | None = None, ladcp_dir: Path | None = None, ladcp_profiles_path: Path | None = None, ladcp_pattern: str | None = None, ship_track_nc: Path | None = None, generate: dict[str, bool] | None = None, force: bool = False, skip_existing: bool = False, section_style: str = 'pcolormesh', timeseries_style: str = 'pcolormesh', vmin_override: dict[str, float] | None = None, vmax_override: dict[str, float] | None = None, cruise_info: dict[str, Any] | None = None, config: ReportConfig | None = None, cast_filter: int | list[int] | None = None, sal_range: tuple[float, float] | None = None, trim_soak: bool = False, dbar_step: int = 1, drop_stub: bool = False) int[source]
Generate the full ctdcast HTML report suite.
- Parameters:
nc_dir – Directory containing per-cast
.ncfiles.out_dir – Root output directory.
profiles_path – Path to compiled
profiles.nc. Required for section and time series pages.groupings_yaml – Cast groupings —
sections:andtimeseries:. Conventionallyctd_groupings.yaml.section_yaml – Superseded spelling of groupings_yaml; accepted, and used only when groupings_yaml is absent.
ladcp_dir – Directory containing processed LADCP
.matfiles namedNNN.matorNNNb.mat(letter-suffix variants supported).ladcp_profiles_path – Path to compiled
ladcp_profiles.nc. When present, the summary page gains a Velocity section (cruise-wide U/V overview panels) and a data-inventory pill for the file.ladcp_pattern – Optional filename glob for non-standard LADCP naming, e.g.
"msm_142_1_*.mat". The*is replaced with the zero-padded cast number. Seefind_ladcp_file().ship_track_nc – Path to a ship-track netCDF for the Leaflet map background line.
generate – Dict of booleans controlling which page types to build. Keys:
"stations","sections","timeseries","index","map". Missing keys default toTrue.force – Regenerate all pages regardless of file modification times.
skip_existing – If True, skip any page whose output HTML already exists, regardless of whether the source files are newer. Use this to fill in only missing pages without touching anything already generated. Takes precedence over the mtime check but is overridden by
force.section_style –
"pcolormesh"or"contourf"for section figures.timeseries_style –
"pcolormesh"or"contourf"for time series figures.vmin_override – Per-variable colormap limit overrides (e.g.
{"SA": 34.5}).vmax_override – Per-variable colormap limit overrides (e.g.
{"SA": 34.5}).cruise_info – Cruise metadata dict (from the
cruise_info:block inconfig.yaml).config – Frozen display configuration (GEBCO path, figsizes, map bounds, colormaps). Defaults to DEFAULT_REPORT_CONFIG.
cast_filter – If set, rebuild only the station page for this cast number (implies
generate={"stations": True, rest False}).sal_range –
(sal_min, sal_max)— records withctd_salinity_1outside this range are excluded from all station page plots. The NC files are not modified. Excluded record count is shown in each page header.trim_soak – If True, apply pre-soak detection on each cast: cut the first 60 s (pump activation) and any records up to the last near-surface record after pump-on. Passed through to
generate_station_page().dbar_step – Subsample the pressure axis by this step for section and timeseries plots (default 1, full 1-dbar resolution).
build_profiles()always stores 1-dbar data; this controls plot-time resolution only.drop_stub – If True, a cast-page section that applies but whose figures all failed to render is dropped and the survivors renumber over it, instead of keeping the heading with an “unavailable” placeholder. Passed through to
generate_station_page().
- Returns:
The number of pages that failed to build (0 on full success), so the CLI can exit non-zero when a requested page could not be generated.
- Return type:
int
Best-available file selection
The supported entry point for consumers outside ctdcast (e.g. caldip) to read the
best available ctdcast file per cast — stage 3 if present, else stage 2, else stage 1 —
without reimplementing the precedence ladder. Import it as ctdcast.select_best_available.
- ctdcast.select_best_available(root: Path | str) list[tuple[tuple[int, str], Path, int]][source]
Return
(cast_id, best_path, source_stage)per cast, sorted by identity.The one place the best-available precedence lives for every consumer (the report,
build_profiles, the LADCP compile).best_pathis the highest-stage file present (3 > 2 > 1);source_stageis that stage, or 0 when the best file is a flat/suffix-less file that does not state its own stage (the compatibility shim only assumes stage 1, so 0 = “unknown, guessed stage 1” stays distinct from a stated stage 1).- Parameters:
root (Path or str) – The instrument stage root.
- Returns:
One
((cast_num, cast_suffix), path, source_stage)per cast, sorted by cast identity.- Return type:
list of (tuple(int, str), Path, int)
Cast identity
Cast identity: cast-number parsing, expansion, and compact formatting.
Single home for the cast-identity operations shared across the package:
parse a cast number and letter suffix from a filename (cast_id_from_name()),
expand a cast_numbers config spec to order-preserving pairs or ints
(expand_cast_ids() / expand_cast_numbers()), format a single cast id as
a zero-padded string (format_cast_id()), and format a list of cast numbers
compactly with collapsed ranges (compact_cast_list()).
- ctdcast.identity.cast_id_from_name(name: str) tuple[int, str] | None[source]
Extract
(cast_num, cast_suffix)from a cast filename stem.Uses the last 3+-digit group in the stem as the cast number, so cruise or leg numbers earlier in the name (e.g. the
142inmsm_142_1_001_1sec) are not mistaken for the cast number. Letter suffixes are recognised whether directly appended (mixsed2_004b) or underscore-separated (mixsed2_004_b). ReturnsNonewhen no 3+-digit group is present.A per-cast file may carry a stage suffix (
_stage1…_stage3; seectdcast.processors.stage_layout). It is stripped first, so a lettered cast likemixsed2_029_b_stage1parses to(29, "b")— otherwise the trailing_stage1defeats the end-anchored letter-suffix match below and the cast collapses to(29, ""), colliding with plain cast 29.
- ctdcast.identity.compact_cast_list(nums: list[int]) str[source]
Format a cast number list compactly, collapsing consecutive runs into ranges.
Example: [131, 133, 134, 136, 163] → “131, 133–134, 136, 163”.
- ctdcast.identity.expand_cast_ids(cast_numbers: list) list[tuple[int, str]][source]
Expand a
cast_numbersspec to(number, suffix)pairs, preserving order.A plain cast NNN and its lettered sibling NNNb are distinct events, so identity is the pair, not the bare number. A bare int or range names plain events only (suffix
""); a “NNNb” string names the lettered event. Input order is preserved and duplicates are kept — section ordering relies on the author’s given order, so sorting and de-duplication are not applied.Raises
ValueErroron a malformed entry, so a bad config fails loudly rather than silently dropping or coercing casts.
- ctdcast.identity.expand_cast_numbers(cast_numbers: list) list[int][source]
Expand a
cast_numbersspec to a flat list of station numbers, in order.The integer view of
expand_cast_ids()— the letter suffix is dropped, so a “10b” entry contributes station10. Use this where callers key on the integer station (LADCP files, map positions, membership); useexpand_cast_ids()where a plain cast must be distinguished from its lettered sibling (section and timeseries profile selection).
- ctdcast.identity.format_cast_id(cast_num: int, cast_suffix: str = '') str[source]
Format a cast identity as a zero-padded id string, e.g.
(10, "b") -> "010b".The single formatter for the
NNN/NNNbconvention used in output filenames (cast_010b.html), page labels, and cast pills. Changing the pad width or suffix rule here changes it everywhere.
Figure builders
draw_*_fig functions build and return a matplotlib Figure (or None when the
dataset lacks the required variables).
Figure builders: each draw_*_fig returns a matplotlib Figure or None.
Encoding to base64 PNG for embedding in a page is done separately by the
_make_*_b64 wrappers in ctdcast.reports._plots.
- ctdcast.plotters.plots.draw_all_sections_map_fig(sections_data: list[dict[str, Any]], all_lats: list[float], all_lons: list[float], legend_outside: bool = False, *, target_h: float = 4.5, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a Figure showing all section tracks coloured by section.
- ctdcast.plotters.plots.draw_aux_profiles_fig(ds: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a O₂ sat, fluorescence, turbidity profiles Figure (downcast + pale upcast).
- ctdcast.plotters.plots.draw_clock_offset_fig(series: list[CastClock], verdict: ClockVerdict, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) plt.Figure | None[source]
Return a Figure of per-cast clock offset vs cast number, with segment means and changepoints.
Points, not a line — the offsets are whole-second-quantised, so a line would imply a precision that is not there. Each segment mean is drawn as a horizontal rule over its cast range and each changepoint as a vertical rule. Returns
Nonewhen there is nothing to plot (no segments — ano_clock_pairorinsufficientverdict).
- ctdcast.plotters.plots.draw_cruise_map_fig(all_meta: list[dict], *, target_h: float = 4.0, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a Figure of all cast positions (no single-cast highlight).
- ctdcast.plotters.plots.draw_ct_sa_sigma0_fig(ds: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a CT, SA, σ₀ profiles side-by-side Figure (downcast + grey upcast).
- ctdcast.plotters.plots.draw_ladcp_bottomtrack_fig(ladcp_path: Path | None, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a LADCP bottom-track U and V vs depth Figure.
- ctdcast.plotters.plots.draw_overview_panel_fig(ds_prof: Dataset, var: str, label: str, bathy_depths: ndarray | None = None, style: str = 'pcolormesh', vmin: float | None = None, vmax: float | None = None, cast_groups: dict[str, list[int]] | None = None, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a Figure of var vs pressure × cast number (cruise overview panel).
- ctdcast.plotters.plots.draw_pressure_time_fig(ds: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a pressure vs elapsed time Figure (cast trajectory + bottle stops).
- ctdcast.plotters.plots.draw_qc_histogram_fig(nc_path: Path, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a per-variable data-value distribution Figure, or None.
One histogram panel per science variable that carries a
{var}_qccompanion. Grey bars are all finite data; coloured bars are the good data (QARTOD pass, flag 1 only — so suspect, fail and missing all show as grey above the colour), on shared bin edges so the two are directly comparable. The gross-range suspect and fail thresholds recorded on the_qccompanion are drawn as dashed/dotted lines when they fall within the plotted range. Reads nc_path directly, so the distribution is the file’s own — untrimmed — data and flags. Ports the logic of oceanarray’sdraw_data_histogram.- Parameters:
nc_path – Path to a per-cast stage file (stage 2 or 3, i.e. one carrying
_qc).cfg – The frozen report configuration;
clean_spinesdrives the despine.
- ctdcast.plotters.plots.draw_section_fig(ds_prof: Dataset, var: str, label: str, x_vals: ndarray, x_label: str, title: str = '', style: str = 'pcolormesh', bathy_depths: ndarray | None = None, bathy_x: ndarray | None = None, cast_labels: list | None = None, vmin: float | None = None, vmax: float | None = None, figsize: tuple[float, float] | None = None, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a Figure of var vs pressure × x_vals.
- ctdcast.plotters.plots.draw_section_map_fig(lats: list[float], lons: list[float], cast_nums: list[int], title: str = '', min_margin: float = 0.03, min_margin_lon: float | None = None, *, fig_w: float = 4.5, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a GEBCO map Figure with the section track.
- ctdcast.plotters.plots.draw_section_ts_histogram_fig(ds_prof: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a CT–SA 2-D count histogram Figure (log₁₀ colour) for section profiles.
- ctdcast.plotters.plots.draw_section_ts_o2_fig(ds_prof: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a CT–SA histogram Figure coloured by median O₂ saturation per bin.
- ctdcast.plotters.plots.draw_section_ts_profiles_fig(ds_prof: Dataset, x_vals: ndarray, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a Figure of per-cast CT–SA profiles coloured by along-track distance.
- ctdcast.plotters.plots.draw_sensor_diff_fig(ds: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a primary minus secondary sensor difference profiles Figure.
- ctdcast.plotters.plots.draw_stability_fig(ds: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a N² and Turner angle (2-panel) Figure.
- ctdcast.plotters.plots.draw_station_map_fig(lat: float, lon: float, all_meta: list[dict], target_h: float = 4.5, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a GEBCO map Figure with all casts and this cast highlighted.
- ctdcast.plotters.plots.draw_timeseries_fig(ds_prof: Dataset, var: str, label: str, style: str = 'pcolormesh', vmin: float | None = None, vmax: float | None = None, figw: float | None = None, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a Figure of var vs cast time × pressure, both down and upcast.
- ctdcast.plotters.plots.draw_ts_density_fig(ds: Dataset, ladcp_path: Path | None = None, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a CT/SA/σ₀ profiles Figure, optionally alongside LADCP U/V.
- ctdcast.plotters.plots.draw_ts_diagram_fig(ds: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a T-S diagram Figure colored by O₂ saturation.
- ctdcast.plotters.plots.draw_ts_diagram_timeseries_fig(ds_ts: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a CT–SA diagram Figure for all timeseries profiles, coloured by time.
- ctdcast.plotters.plots.draw_ts_updown_fig(ds: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a CT–SA scatter Figure: downcast in blue, upcast in red, σ₀ contours.
- ctdcast.plotters.plots.draw_updown_diff_fig(ds: Dataset, *, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Figure | None[source]
Return a downcast minus upcast profiles Figure: ΔCT, ΔSA, Δσ₀.
- ctdcast.plotters.plots.map_panel(ax: Axes, cax: Axes, xl0: float, xl1: float, yl0: float, yl1: float, *, cfg: ReportConfig) None[source]
Draw the GEBCO depth field into ax and its ‘Depth (m)’ colorbar into cax.
cax is a pre-placed colorbar axes from
_map_layout(); when there is no GEBCO to draw it is hidden so no empty box remains.
- ctdcast.plotters.plots.section_figsize_and_slot(p_max_dbar: float, dist_km: float) tuple[tuple[float, float], str][source]
Return figure size and CSS slot class for a section pcolormesh plot.
The figure width is always a canonical slot width (one of
SLOTS), so the rendered PNG fills its slot at the same oversample as every other figure and the browser never rescales it — otherwise a section rendered at a between-slots width is squeezed into the nearest slot box, shrinking its baked-in fonts.Aspect is calibrated (
_SECTION_STRETCH) so KTout (416 dbar, 94 km) is short at full width. The widest slot whose resulting height stays within_MAX_SECTION_His chosen; height then follows the data aspect (floored at_MIN_SECTION_H). A section too deep for even the narrowest slot keeps that slot’s width and accepts the height cap.- Parameters:
p_max_dbar – Maximum pressure in the section (dbar).
dist_km – Total along-track distance of the section (km).
- Returns:
tuple of
((fig_w, fig_h), css_slot)wherefig_wequals the slot’scanonical inch width and
css_slotis one of"slot-full","slot-twothirds","slot-half", or"slot-third".
Layer-1 primitives
ax-taking primitives that draw into a caller-supplied axes and create no Figure, so
a panel shared by more than one page type has a single implementation.
Layer-1 ax-taking primitives that draw into a provided axes and create no Figure.
- ctdcast.plotters.primitives.mesh_field(ax: Axes, fig: Figure, x: np.ndarray, y: np.ndarray, data2d: np.ndarray, *, cmap: Colormap, norm: Normalize, cmap_name: str, bounds: np.ndarray, style: str, cbar_label: str = '') Colorbar[source]
Draw a pcolormesh/contourf field with a matched discrete colorbar into ax; return the colorbar.
The colorbar has a fixed inch width (not a fraction of the host axes), so its thickness and the resulting right margin are identical on every field figure regardless of slot width, and its labels are ~6 round values (
nice_colorbar_ticks()) rather than one per discretisation boundary, with cbar_label written as a title on top (unit_colorbar()) so it does not widen the figure.make_axes_locatable(viaunit_colorbar(reserve=True)) is used only because this axes is free-aspect. It attaches the colorbar to the divider of the axes’ box at layout time; if something resizes that box afterwards —set_aspect("equal", adjustable="box"), or a hand-placed map layout — the cax tracks the pre-resize box and ends up the wrong size. The rule:make_axes_locatablefor free-aspect axes; hand-reserved inches whenever the aspect is locked or the axes are hand-placed. Do not addset_aspect("equal")to a figure that colorbars through here without switching to the reserved-inches path.
- ctdcast.plotters.primitives.nice_colorbar_ticks(vmin: float, vmax: float, *, max_ticks: int = 6) ndarray[source]
Return at most max_ticks nicely-rounded tick positions in
[vmin, vmax].Decoupled from the colorbar’s colour discretisation: a 20-level
BoundaryNormbar can still show ~6 round labels (e.g. 34.8, 34.9, … 35.2) instead of one label per boundary. UsesMaxNLocatorwith round step multiples so labels land on clean values.- Parameters:
vmin (float) – Data range of the colorbar.
vmax (float) – Data range of the colorbar.
max_ticks (int) – Maximum number of ticks (approximate; the locator may return a few fewer).
- Returns:
Tick positions, clipped to
[vmin, vmax].- Return type:
numpy.ndarray
- ctdcast.plotters.primitives.sigma0_isopycnals(ax: Axes, x: np.ndarray, y: np.ndarray, data2d: np.ndarray) None[source]
Overlay the 27.7 and 27.8 σ₀ isopycnal contours (labelled) on ax, swallowing contour failures.
- ctdcast.plotters.primitives.unit_colorbar(target: Axes, mappable: ScalarMappable, *, unit: str = '', ticks: np.ndarray | None = None, extend: str = 'neither', reserve: bool = False, title_loc: str = 'center') Colorbar[source]
Draw the report-standard colorbar with the unit as a title on top.
One entry point, two placement strategies so the appearance (bar width, gap, tick choice, unit-on-top) is set in a single place regardless of how the axes was laid out.
- Parameters:
target (matplotlib Axes) – When reserve is False, a colorbar axes already reserved by a hand layout. When reserve is True, the host plot axes, into which a fixed-inch cax is appended — allowed only for free-aspect axes (see
mesh_field()).mappable (matplotlib ScalarMappable) – The artist to map (
pcolormesh,contourfset, …).unit (str) – Text placed above the bar (
cax.set_title) rather than as a rotated side label — reads cleanly and, unlike a side label, does not widen the figure. A unit ("m s⁻¹") or a full label ("CT (°C)"); empty renders no title.ticks (numpy.ndarray, optional) – Explicit tick positions (e.g. from
nice_colorbar_ticks()).extend (str) –
"neither"/"both"/"min"/"max"— pointed ends for out-of-range.reserve (bool) – Append a fixed-inch cax to target instead of treating it as the cax.
title_loc (str) – Horizontal anchor for the on-top title —
"center"(default),"left"or"right"."left"anchors the title at the thin bar’s left edge so it extends right into the margin, clear of a figure’s top-left annotations (e.g. the cast-marker strip on field figures); centering it over the thin bar would instead overhang the plot.
- Return type:
matplotlib.colorbar.Colorbar
Base64 encoders
_make_*_b64 wrappers render a figure builder’s Figure to an embedded base64 PNG,
returning None on any exception so a missing figure never prevents a page from
being written.
Base64 PNG encoders — thin wrappers that render a Figure for a page.
Each _make_*_b64 builds a figure via a draw_*_fig in
ctdcast.plotters.plots and encodes it with render_b64(). Two use a
custom wrapper instead of render_b64(): _make_all_sections_map_b64 still
delegates drawing to its draw_*_fig but needs a post-tight_layout
adjustment, and _make_ladcp_section_b64 does its own plotting because it
returns a list of RenderedPanel rather than a single figure.
- class ctdcast.reports._plots.RenderedPanel(b64: str | None, title: str = '', short: str = '', figsize: tuple[float, float] | None = None, slot: str | None = None)[source]
A rendered figure plus the layout metadata the HTML template needs.
- b64: str | None
Base64-encoded PNG string, or
Nonewhen the figure could not be rendered.
- figsize: tuple[float, float] | None = None
Figure dimensions
(width, height)in inches, orNonewhen not recorded.
- short: str = ''
Short label used in
<figcaption>elements (e.g."CT","U").
- slot: str | None = None
CSS slot class (e.g.
"slot-full") matching the PNG aspect ratio, orNone.
- title: str = ''
Long descriptive title (used in
altattributes and headings).
Plotting parameters
Package-wide constants: plotting parameters, CNV aliases, variable metadata, CCHDO conventions.
Compile-time constants only — the per-run display settings (GEBCO path,
clean_spines, figsizes, map bounds, colormap overrides) live in the frozen
ctdcast.config.report_config.ReportConfig, built once and threaded down.
Section headers mark what kind of constant each block holds, because that determines who may change it and what breaks when they do:
- Contract
Changing it makes output wrong or non-conformant. Requires a code review and a version bump.
- Science
Changing it gives a different but equally valid answer. Per-cruise overrides go in
display.variables:in the cruiseconfig.yaml; usectdcast.config.loader.load_display_config().- Derived
Computed from another constant; must live here to avoid drift.
- Deferred
Belongs in
oceanvisonce that package exists.
- ctdcast.config.parameters.SBE_PREFIX: str = 'sbe_'
Prefix marking a SeaBird-derived channel — quantities ctdcast recomputes (density → sigma0) or does not use. Stage 1 keeps them under this prefix (faithful translation); stage 2 drops them by default; build_profiles never bins them; CCHDO export excludes them. The contract lives here, not as a bare
"sbe_"literal scattered across those sites.
- ctdcast.config.parameters.SECTION_BIOGEO_VARS: tuple[str, ...] = ('ctd_oxygen_1', 'ctd_fluor', 'ctd_turbidity')
Biogeochemical variables drawn on section, overview, and timeseries pages, in order.
- ctdcast.config.parameters.SECTION_PHYSICS_VARS: tuple[str, ...] = ('conservative_temperature', 'absolute_salinity', 'sigma0')
Physics variables drawn on section, overview, and timeseries pages, in order. Use
vlabel(var)for the axis/panel label andVARIABLES[var]["label"]for the short caption. Defined here so the three report modules share one source of truth and cannot silently diverge from each other or from VARIABLES.
- ctdcast.config.parameters.is_sbe_channel(name: str) bool[source]
Return whether name is a SeaBird-derived channel (carries
SBE_PREFIX).
- ctdcast.config.parameters.resolve_sensor_var(ds: xr.Dataset, var: str) str[source]
Return the name to use for var in ds, applying the single/dual-sensor rule.
A variable may be stored plain (single sensor, e.g.
ctd_oxygen) or suffixed (dual sensor, e.g.ctd_oxygen_1). This resolves whichever form ds actually holds: a suffixed var falls back to its plain form, and a plain var falls back to the_1form. Returns var unchanged when neither is present, so the caller’s draw function then returnsNonefor the missing variable.- Parameters:
ds – The dataset whose variables to resolve against.
var – The requested variable name (plain or
_1/_2suffixed).
- Returns:
The name present in ds, or var unchanged if neither form is found.
- Return type:
str
- ctdcast.config.parameters.vlabel(var: str, prefix: str = '') str[source]
Return a matplotlib axis label for var using the VARIABLES registry.
Format is
"Label (units)"whenlabel_unitsis non-empty, or just"Label"when there are no units (e.g. dimensionless quantities). Units are always in the Unicode display form fromlabel_units— never the ASCII-safeunitsstring used for netCDF attributes.- Parameters:
var – VARIABLES key (e.g.
"conservative_temperature").prefix – Optional prefix prepended to the label component only, not the units. Use
"Δ"to produce difference labels such as"ΔCT (°C)".
- Returns:
Ready-to-use axis label string. Falls back to var itself when var is not in VARIABLES.
- Return type:
str
- ctdcast.config.parameters.vlabel_html(var: str, prefix: str = '') str[source]
Return
vlabel()as HTML-ready text — mathtext subscripts as Unicode.Use this wherever a variable label is written into HTML (a figure caption, a table cell): matplotlib needs the mathtext form, but HTML must show
σ₀, not the literal$\sigma_0$. One helper so the label-form choice lives in one place.
- ctdcast.config.parameters.vunit(var: str) str[source]
Return the Unicode display unit for var, or
""when dimensionless.The unit half of
vlabel(), for field colorbars that place the unit as a title on top of the bar rather than a"Label (units)"side label. Uses thelabel_unitsdisplay form (never the ASCIIunitsnetCDF string). Works for both canonical and single-/dual-sensor-resolved names (both carry the samelabel_unitsinVARIABLES).
File-level metadata
Cruise-level metadata written into the compiled files as ACDD-1.3 global
attributes. global_attrs composes the derived coverage
bounds, provenance, embargo licence, people, and platform attributes;
platforms resolves the vessel registry and derives the
EXPOCODE; people turns the structured contributors
list into the semicolon-delimited ACDD strings.
Assemble a compiled file’s global attributes from config + derived bounds.
Three layers, kept apart on purpose (see the file-level-metadata design note):
derived —
geospatial_*andtime_coverage_*bounds,date_created. Computed from the data at write time, never authored and never copied up from a per-cast file. Copying the first cast’s latitude up into the cruise file states a bounding box containing one station; the fix is to compute.authored —
title,project,acknowledgement, people, embargo. Taken fromcruise_info:in the cruise config, once, at the level it is true.identity —
cruise,platform_*andexpocode. Constant for a ctdcast file, so all three are globals. (CCHDO’s exchange format storesexpocodeper profile, because a file there may span cruises; a ctdcast file never does, so that per-profile projection belongs with a CCHDO exporter, not here.)
The one rule that decides where a fact goes: a global attribute must be true of the entire file. Anything that varies within the file (cast lat/lon, station name, sensor serial) is a variable, not an attribute.
Conformance target is ACDD-1.3 (with CF). We take OG1’s vocabularies (C89 roles, EDMO institutions, L06 platform, L08 access policy) without adopting OG1’s glider-mission entity model.
- ctdcast.config.global_attrs.ATTR_GROUPS: tuple[tuple[str, tuple[str, ...]], ...] = (('Identity & discovery', ('title', 'cruise', 'cast_id', 'platform', 'platform_name', 'platform_ices_code', 'platform_vocabulary', 'expocode', 'id', 'naming_authority', 'project', 'internal_mission_identifier', 'summary', 'program', 'references')), ('Spatiotemporal coverage', ('geospatial_lat_min', 'geospatial_lat_max', 'geospatial_lat_units', 'geospatial_lon_min', 'geospatial_lon_max', 'geospatial_lon_units', 'geospatial_vertical_min', 'geospatial_vertical_max', 'geospatial_vertical_units', 'geospatial_vertical_positive', 'time_coverage_start', 'time_coverage_end', 'time_coverage_duration', 'time_coverage_resolution')), ('People & institutions', ('contributor_name', 'contributor_role', 'contributor_role_vocabulary', 'contributor_id', 'contributor_email', 'contributing_institutions', 'contributing_institutions_id', 'contributing_institutions_role', 'contributing_institutions_role_vocabulary', 'institution', 'publisher_name', 'publisher_email', 'publisher_url')), ('Rights & access', ('acknowledgement', 'license', 'access_constraint', 'date_available', 'license_after_embargo', 'citation', 'doi')), ('Provenance & processing', ('date_created', 'date_modified', 'tracking_id', 'source_tracking_id', 'source_cnv', 'source_mat', 'processing_stage', 'history', 'creator_name', 'creator_type', 'creator_id', 'source', 'processing_level', 'data_mode', 'data_mode_meaning', 'Conventions', 'featureType', 'cdm_data_type', 'pressure_units', 'pressure_spacing_dbar', 'cchdo_software_version', 'cchdo_parameters_version')))
Canonical grouping + order of the compiled-file global attributes. Single source of truth for both the on-disk write order (
order_attrs()) and the inventory page’s grouped tables (group_attrs()). The order within each group is the canonical order; the group titles are the HTML section headings.
- ctdcast.config.global_attrs.CREATION_NOTE = 'file created by ctdcast (https://github.com/ocean-uhh/ctdcast)'
The note recording what wrote this file, appended by each builder through
ctdcast.processors.history.append_history()– which already stamps the time and the ctdcast version, so this carries only the part that helper does not: where to find the software.
- ctdcast.config.global_attrs.DATA_MODES: dict[str, str] = {'D': 'delayed-mode', 'M': 'mixed', 'P': 'provisional', 'R': 'real-time'}
OceanSITES Reference Table 4 data modes. Single letters rather than OG1’s spelled-out forms because OG1’s own examples are inconsistent — it shows both
sea008_..._delayedandsp032_..._R— and oceanarray already uses the OceanSITES letters, so one convention serves both packages.
- ctdcast.config.global_attrs.EXPOCODE_PLACEHOLDER_ICES = '{ICES}'
Stand-ins for the two halves of an EXPOCODE when config cannot supply them. Braces are not legal in an EXPOCODE (it is ICES code +
YYYYMMDD, both alphanumeric), so a placeholder can never be mistaken for a real code by a reader or a regex — the same reasoning asUNKCRUISEinctdcast.config.parameters: conspicuous beats plausible. [Contract — a file carrying one of these is provisional. Do not publish it.]
- ctdcast.config.global_attrs.OTHER_GROUP = 'Other'
Title for attributes not named in
ATTR_GROUPS(never dropped).
- ctdcast.config.global_attrs.aggregate_identity(per_cast: list[dict[str, str]], cruise_info: dict[str, Any] | None = None) dict[str, str][source]
Lift identity attributes from the per-cast files being compiled.
The compiled product describes one cruise, so every per-cast file should agree about which cruise it is. Rather than trusting config and hoping, the inputs are compared:
constant across every cast that states it → lifted;
a cruise-defining attribute (
cruise,expocode) varying →ValueError. Two cruises in one directory is a mistake, not a merge. The previous behaviour (combine_attrs="drop_conflicts") silently discarded the attribute, so a stray cast from another cruise madecruisevanish rather than fail;a ship-describing attribute (the
platform_*block) varying → warn, and take cruise_info’s value if it has one, else omit the attribute. Disagreement here is usually registry drift — aplatform_vocabularyURI edited between two stage-1 runs, a vessel renamed — which says nothing about whether these casts are one cruise. Failing the whole compile over it, with a message about legs sharing a root, misdirects the reader towards a problem they do not have;absent from every cast → fall back to cruise_info and warn, which is the pre-stage-2 path and the case for files written before identity was recorded at stage 1.
The key set is
_IDENTITY_KEYS, whichidentity_attrs()filters its own output through, so the two cannot drift apart.- Parameters:
per_cast (list of dict) – Each per-cast file’s global attributes.
cruise_info (dict or None) – Fallback for attributes no cast states.
- Return type:
dict of str to str
- Raises:
ValueError – If a cruise-defining attribute (
cruise,expocode) takes more than one value across the inputs.
- ctdcast.config.global_attrs.canonical_attr_order() list[str][source]
Return the flat canonical order of all named global attributes.
- ctdcast.config.global_attrs.coverage_attrs(*, lats: _Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | complex | bytes | str | _NestedSequence[complex | bytes | str], lons: _Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | complex | bytes | str | _NestedSequence[complex | bytes | str], vertical_min: float | None = None, vertical_max: float | None = None, vertical_units: str = 'dbar', times: _Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | complex | bytes | str | _NestedSequence[complex | bytes | str] | None = None) dict[str, str][source]
Return the derived ACDD coverage attributes.
Every value here is computed from the data, so a multi-station file states a bounding box that brackets all its stations — the property the accompanying test asserts.
- Parameters:
lats (array-like) – Per-profile latitudes/longitudes (NaNs ignored).
lons (array-like) – Per-profile latitudes/longitudes (NaNs ignored).
vertical_min (float, optional) – Shallowest and deepest levels, in vertical_units. Omitted when None.
vertical_max (float, optional) – Shallowest and deepest levels, in vertical_units. Omitted when None.
vertical_units (str) – Units of the vertical bounds.
"dbar"for a pressure grid (the honest label — the profiles are gridded on pressure, not converted to metres).times (array-like of datetime64, optional) – Per-profile times; drives
time_coverage_start/end/duration.
- Returns:
ACDD
geospatial_*andtime_coverage_*attributes.- Return type:
dict of str to str
- ctdcast.config.global_attrs.cruise_expocode(cruise_info: dict[str, Any] | None) str | None[source]
Return the cruise EXPOCODE, with placeholders for what config omits.
A present but unusable slug (ambiguous, no ICES code, or a forbidden code) is a config error that
ctdcast.config.platforms.derive_expocode()raises on — but at build time a wrong EXPOCODE is worse than none and a hard crash mid-pipeline is worse still, so here it is downgraded to a loud warning andNone. The resolver keeps its strict raising contract for direct callers; this is the lenient wrapper the file builders use.A missing
platform/ship_slugorstart_dateis different: the cruise is mid-processing and has simply not settled that value yet. ReturningNonethere meant the compiled file carried no cruise identifier at all, and — because the departure date appears nowhere else (see the CCHDO exemplar, which encodes it only inside the EXPOCODE) — the fact was lost rather than merely deferred. So the shape is kept and the missing half is filled with a conspicuous placeholder:29OD{YYYYMMDD} departure date not yet set {ICES}20260709 ship not yet resolved to an ICES code
With neither half given there is nothing to defer — no cruise is being described — so that returns
Noneas before rather than a placeholder on every profile.The file can therefore be generated and read, and states plainly which fact is absent. Callers that must not see a placeholder — the filename builder above all — test with
is_placeholder_expocode().- Parameters:
cruise_info (dict or None) – The
cruise_info:mapping.- Returns:
The EXPOCODE, possibly containing placeholders;
Noneonly when the slug is present but unusable.- Return type:
str or None
- ctdcast.config.global_attrs.cruise_global_attrs(cruise_info: dict[str, Any] | None, *, lats: _Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | complex | bytes | str | _NestedSequence[complex | bytes | str] | None = None, lons: _Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | complex | bytes | str | _NestedSequence[complex | bytes | str] | None = None, vertical_min: float | None = None, vertical_max: float | None = None, vertical_units: str = 'dbar', times: _Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | complex | bytes | str | _NestedSequence[complex | bytes | str] | None = None, source: str | None = None, config: dict[str, Any] | None = None, now: datetime | None = None) dict[str, str][source]
Compose the full global-attribute set for a compiled cruise file.
Merges, in order: authored discovery fields (
title/summary/project/program/cruise_id/acknowledgement), derived coverage bounds, provenance/conformance tags, licence/embargo, people (creator + contributors), and platform attributes. Theexpocodeis not here — the caller emits it as anN_PROFcoordinate.- Parameters:
cruise_info (dict or None) – The
cruise_info:mapping. None or empty yields only derived + provenance attributes.lats – Passed to
coverage_attrs().lons – Passed to
coverage_attrs().vertical_min – Passed to
coverage_attrs().vertical_max – Passed to
coverage_attrs().vertical_units – Passed to
coverage_attrs().times – Passed to
coverage_attrs().source (str, optional) – Product key (
"ctd","ladcp"). Passed toctdcast.config.people.contributor_attrs(), which writes only the roles each contributor scoped to this product (theirroles: {all, ctd, ladcp}mapping) plus their unscoped roles — so a product’s own processing personnel are credited on that file alone.config (dict, optional) – The whole cruise config. Only used to read
processing.profiles_dbarfor the grid token indataset_identity(); without it a CTD file is identified at the 1 dbar default.now (datetime, optional) – Creation timestamp (injectable for tests).
- Returns:
Ready to merge into the dataset’s
attrs.- Return type:
dict of str to str
- ctdcast.config.global_attrs.cruise_name(cruise_info: dict[str, Any] | None) str | None[source]
Return the cruise identifier from config, accepting either spelling.
cruiseis the attribute written to the file, and is the preferred config key: for every other authored field –title,summary,project,program– the key and the attribute are the same word, and making this one an exception is what let acruise_idattribute leak into every file.cruise_idstays accepted, because configs are written by hand and both spellings are already in use.One resolver rather than the precedence repeated at each of the six call sites, so a config cannot be read one way by the compiler and another by the report.
- Parameters:
cruise_info (dict or None) – The
cruise_info:mapping.- Returns:
The identifier, or
Nonewhen neither key is set.- Return type:
str or None
- ctdcast.config.global_attrs.data_mode_with_meaning(mode: str | None, *, warn: bool = True) tuple[str, str][source]
Return a validated OceanSITES
data_modeand itsdata_mode_meaning.The single place the mode-to-meaning pairing is made, so the two attributes can never drift. An unset or out-of-vocabulary mode falls back to
"P"(provisional) — the honest default for a file nobody has declared finished — optionally warning.- Parameters:
mode – The candidate mode (e.g. from
cruise_info.data_modeor an existing attribute).warn – Emit a warning when mode is present but not an OceanSITES data mode.
- Returns:
The validated mode and its meaning, both drawn from
DATA_MODES.- Return type:
tuple of str
- ctdcast.config.global_attrs.dataset_filename(cruise_info: dict[str, Any] | None, product: str, *, config: dict[str, Any] | None = None, grid: str | None = None) str | None[source]
Return
<id>.ncfor a compiled product, or None when noidexists.ACDD notes that the
idmay be the file name without its suffix, which is what this enforces, so the file on disk and the identifier inside it cannot drift apart.
- ctdcast.config.global_attrs.dataset_identity(cruise_info: dict[str, Any] | None, product: str, *, config: dict[str, Any] | None = None, grid: str | None = None) dict[str, str][source]
Build the ACDD identity attributes for one compiled product.
Two identifiers, sharing a tail:
id 29OD20260709_R_ctd_2dbar internal_mission_identifier mixsed2_20260709_R_ctd_2dbar
idfollows OG1’s<platform_serial>_<start_date>_<data_mode>shape, with the EXPOCODE standing in for platform-plus-date — it already is the ICES ship code followed by the departure date, so repeating the date would be redundant.internal_mission_identifieris the institution’s own name for the cruise and carries the date explicitly, since a project short name likemixsed2has none. The product and grid then distinguish the files within one cruise.The leading
<expocode>_ctdis exactly CCHDO’s own file name (740H20200119_ctd.nc), so a ctdcast identifier is a strict extension of the convention ctdcast already reads._separates fields and must not occur inside one — the OceanSITES rule, adopted so the identifier can be split back apart.- Parameters:
cruise_info (dict or None) – The
cruise_info:mapping. Readsdata_mode,naming_authorityandinternal_id.product (str) –
"ctd"or"ladcp".config (dict, optional) – The whole cruise config, for the grid lookup.
grid (str, optional) – Overrides the derived grid token.
- Returns:
id,naming_authority,internal_mission_identifier,data_modeanddata_mode_meaning— only those that can be built.- Return type:
dict of str to str
- ctdcast.config.global_attrs.grid_token(product: str, config: dict[str, Any] | None = None) str | None[source]
Return the vertical-grid token for a product, e.g.
"2dbar"or"10m".The grid belongs in the identifier because it is part of what the file is, not a parameter of how it was made: regridding the same casts at 1 dbar and at 3 dbar yields two products that a user may legitimately want to keep side by side, and nothing else in the name would tell them apart.
- Parameters:
product (str) –
"ctd"or"ladcp".config (dict, optional) – The whole cruise config, read for
processing.profiles_dbar.
- Returns:
A token safe for an identifier field (no
_), or None if unknown.- Return type:
str or None
- ctdcast.config.global_attrs.group_attrs(attrs: dict[str, Any]) list[dict[str, Any]][source]
Split attrs into the canonical groups for display.
Groups appear in
ATTR_GROUPSorder; within each group the attributes keep the order they appear in attrs (i.e. the file order), so the rendered page can be checked visually against the canonical spec. Unnamed attributes fall into a trailingOTHER_GROUP. Empty groups are omitted.- Returns:
[{"title": str, "rows": [(key, value), ...]}, ...]. The key isrows(notitems) so a Jinja template’sgroup.rowsdoes not collide with the dict.items()method.- Return type:
list of dict
- ctdcast.config.global_attrs.identity_attrs(cruise_info: dict[str, Any] | None, *, include_expocode: bool = True) dict[str, str][source]
Return the attributes that identify which cruise and ship this is.
cruise, theplatform_*block, andexpocode. These are the facts fixed the moment a cast is taken, so they are written at stage 1 onto each per-cast file and lifted unchanged into the compiled products — unlike coverage (computed per file) or people and rights (authored, and revisable for years afterwards).- Parameters:
cruise_info (dict or None) – The
cruise_info:mapping.Noneor empty returns{}, which is what keeps a call path supplying no config unchanged.include_expocode (bool, default True) – Emit
expocode. This switch exists for a caller that wants the rest of the identity without it.
- Returns:
Only the attributes that could be resolved.
- Return type:
dict of str to str
- ctdcast.config.global_attrs.is_placeholder_expocode(expocode: str | None) bool[source]
Return True when expocode carries a placeholder for a missing config value.
- ctdcast.config.global_attrs.license_attrs(cruise_info: dict[str, Any]) dict[str, str][source]
Return the
license(and embargo companions) for a cruise.An embargo is not a licence: a CC BY grant is irrevocable and takes effect the instant it is written, so an embargoed file must carry a self-describing restriction statement, not
CC-BY-4.0. This function writes:a moratorium free-text
licenseciting NERC L08 and COAR whencruise_info.embargois present. The release date isembargo.untilwhen set, otherwiseend_date + 2 years. Machine readers also getdate_availableandaccess_constraint;a bare CC BY statement when
cruise_info.licensenames it and there is no embargo;nothing when neither is configured — an absent
licenseasserts nothing, which is safer mid-cruise than a wrong claim.
- Parameters:
cruise_info (dict) – The
cruise_info:mapping.- Returns:
licenseand, under embargo,date_available/access_constraint.- Return type:
dict of str to str
- ctdcast.config.global_attrs.order_attrs(attrs: dict[str, Any]) dict[str, Any][source]
Return attrs in canonical order.
Named attributes come first in
canonical_attr_order(); any remaining ones keep their original relative order and follow (so nothing is dropped and an unrecognised attribute stays visible rather than being silently reordered away).
- ctdcast.config.global_attrs.provenance_attrs(now: datetime | None = None) dict[str, str][source]
Return
date_created/date_modifiedand CF/ACDD conformance tags.Deliberately not
history. Every other writer appends to that attribute throughctdcast.processors.history.append_history(), so a layer that returns it in a dict merged with.update()silently replaces whatever the caller had already recorded – the ordering of the merge becomes load-bearing. The software-provenance note lives inCREATION_NOTEand is appended, like every other line.- Parameters:
now (datetime, optional) – Creation timestamp; defaults to the current UTC time. Injectable so a test can assert an exact value.
- Return type:
dict of str to str
Research-vessel registry: config slug → identity, ICES code, and EXPOCODE.
A cruise’s EXPOCODE is derived, not allocated:
EXPOCODE = <4-character ICES ship code> + <departure date YYYYMMDD>
where the departure date is the day the ship left port (cruise_info.start_date),
which is not necessarily the first cast. The only piece that must be looked up
is the ICES ship code, and looking it up by vessel name is unsafe: names carry
accents, punctuation, and generational ambiguity (four hulls have been called
Meteor; two live C17 entries are called Odon de Buen). A wrong code yields a
well-formed EXPOCODE filed against somebody else’s cruise — a silent, permanent
error.
So the registry is keyed on a short ASCII slug (odb, msm, meteor3),
and this module refuses to guess:
an unknown slug raises;
a slug listed in
ambiguous_slugs(e.g. baremeteor) raises with the message naming the alternatives;a slug whose
ices_codeisnull(a vessel with no code yet) raises rather than emitting a malformed EXPOCODE;a derived code appearing in
forbidden_codesraises — each is a real trap (deprecated codes, same-name different-ship) that a name lookup would fall into.
Source registry: ctdcast.config.platforms platforms.yaml.
- exception ctdcast.config.platforms.PlatformError[source]
A platform slug could not be resolved, or its ICES code is unusable.
Raised rather than returning a sentinel because every caller of this module is about to write an EXPOCODE into a file, and a wrong or missing code there is worse than a hard failure at build time.
- ctdcast.config.platforms.ambiguous_slugs() dict[str, str][source]
Return the
ambiguous_slugstraps (slug → why it is refused), forctdcast list.
- ctdcast.config.platforms.derive_expocode(platform: str | dict[str, Any], start_date: str | date) str[source]
Derive the CCHDO/GO-SHIP EXPOCODE for a cruise.
EXPOCODE = <ICES ship code> + <departure date YYYYMMDD>.- Parameters:
platform (str or dict) – A
platforms.yamlslug, or an inline platform record carrying at leastices_code(seeresolve_platform_spec()).start_date (str or datetime.date) – The departure date from port (
cruise_info.start_date). Not the first cast; the two can differ (MSM142 departs 2026-03-27, first cast 2026-03-29).
- Returns:
e.g.
"29OD20260709"forodbdeparting 2026-07-09.- Return type:
str
- Raises:
PlatformError – If a slug is unknown/ambiguous, the platform has no ICES code, the derived code is in
forbidden_codes, or start_date is unparseable.
- ctdcast.config.platforms.expocode_from_cruise_info(cruise_info: dict[str, Any]) str | None[source]
Derive the EXPOCODE from a cruise config, or
Nonewhen not derivable.Returns
None(rather than raising) when the config supplies noship_slugor nostart_date— a cruise mid-processing may not have settled its departure date, and that is not an error. A present slug that is ambiguous, unknown, or lacks an ICES code still raises viaderive_expocode(), because those are authoring mistakes.- Parameters:
cruise_info (dict) – The
cruise_info:mapping from the cruise config.- Returns:
The EXPOCODE, or
Nonewhen the slug/start_dateis absent.- Return type:
str or None
Notes
The slug is read from
cruise_info["platform"], falling back tocruise_info["ship_slug"].ship(the free-text display name) is never used as a slug — name lookup is exactly the ambiguity this registry avoids.
- ctdcast.config.platforms.forbidden_codes() dict[str, str][source]
Return the
forbidden_codestraps (ICES code → why it is refused), forctdcast list.
- ctdcast.config.platforms.load_platforms() dict[str, dict[str, Any]][source]
Return the platform registry keyed by slug.
Returns an empty mapping when the registry file is absent, so a config that supplies no
ship_slugstill works (EXPOCODE is simply not derived).
- ctdcast.config.platforms.parse_config_date(value: Any) date | None[source]
Coerce a config date value to a
datetime.date, orNone.Accepts a
date/datetime(YAML parses a bare2026-03-27as adate), an ISO"YYYY-MM-DD"string, or a compact"YYYYMMDD"string. ReturnsNoneforNoneor any value it cannot parse, leaving the caller to decide whether that is an error. Shared by the EXPOCODE derivation and the embargo-date logic so both accept exactly the same date forms.
- ctdcast.config.platforms.platform_attrs(platform: str | dict[str, Any]) dict[str, str][source]
Return ACDD platform global attributes for a slug or inline platform.
Emits
platform,platform_vocabulary,platform_nameandplatform_ices_codefrom the resolved record, skipping any it omits. Returns an empty mapping when a slug does not resolve to a usable record (the caller decides whether that is fatal).- Parameters:
platform (str or dict) – A
platforms.yamlslug, or an inline platform record (seeresolve_platform_spec()).- Returns:
Platform attributes ready to merge into a dataset.
- Return type:
dict of str to str
- ctdcast.config.platforms.platform_display_name(platform: str | dict[str, Any]) str | None[source]
Return the registry display name for a slug or inline platform, or
None.Uses the plain
name(notnative_name), for the report masthead and for defaulting the free-textshipinctdcast initfrom the chosen slug — deriving the name from the ICES code (unambiguous) rather than resolving a typed name into a code (ambiguous). ReturnsNonewhen a slug does not resolve.
- ctdcast.config.platforms.resolve_platform(slug: str) dict[str, Any][source]
Return the registry record for slug.
- Parameters:
slug (str) – The
cruise_info.ship_slugvalue — a key inplatforms.yaml.- Returns:
The platform’s fields (
name,ices_code,platform, …).- Return type:
dict
- Raises:
PlatformError – If slug is listed in
ambiguous_slugs(quoting its message) or is not a key in the registry.
- ctdcast.config.platforms.resolve_platform_spec(platform: str | dict[str, Any]) dict[str, Any][source]
Resolve a config platform value (a registry slug or an inline dict).
A string is a slug looked up in
platforms.yaml. A mapping is used inline — for a vessel not (yet) in the shared registry — and should carry at leastices_code(for the EXPOCODE) andname, plus optionalplatform(L06 category) andplatform_vocabulary. This mirrors the inline institution form, so a user can name a new vessel in their own config without first editingplatforms.yaml.- Parameters:
platform (str or dict) – A
platforms.yamlslug, or an inline platform record.- Returns:
The platform record (registry entry, or the inline mapping as given).
- Return type:
dict
Contributors and institutions: config → validated records → netCDF attributes.
The conventions store people as parallel delimited strings aligned by
position — contributor_name, contributor_email, contributor_role
and contributor_id are separate attributes whose n-th elements describe the
same person. That representation has two silent failure modes, and this module
exists to make both impossible:
A delimiter inside a value. OG1 specifies comma-separated, OceanSITES semicolon-separated. Comma is unsafe —
"A. Sanchez Franks, NOC"is a real string from a CCHDO header, and splitting it yields two people. ctdcast writes semicolon-separated (per OceanSITES) and refuses any person-level value containing;or,, so the output survives a reader that assumes either.Institution names are the exception and are checked for
;only: EDMO’s official names contain commas unavoidably — “University of Hamburg, Institute of Oceanography” — which is itself the demonstration that comma cannot serve as the delimiter forcontributing_institutions.Lists of unequal length. Four names and three emails parses cleanly and attributes the wrong address to the wrong person. The strings are generated from one structured list here, so they cannot drift.
Note the direction of travel. amocatlas.contributors parses delimited
strings out of other people’s files and is deliberately lenient — it splits on
both delimiters and warns heuristically. ctdcast writes its own files and
controls the delimiter, so it can be strict instead, which is the stronger
guarantee.
Roles come from NERC C89 (BODC data/cruise roles) — the vocabulary that fits a
research-cruise dataset (Cruise principal scientist, Project principal
investigator, Project collaborator, …). It is the only NVS collection
with cruise-scoped terms, and so the only one able to record a chief scientist —
PS, whose definition explicitly allows the role to be shared. A role may be given in config as its full
prefLabel or its short C89 code (PS, PI, CO); it is normalised to the
prefLabel on write, and contributor_role_vocabulary names C89.
- ctdcast.config.people.ALL_SCOPE = 'all'
Scope keys accepted inside a contributor’s
roles:mapping.allmeans every compiled file; the others name one product, so a role listed there is written only when that file is built.
- ctdcast.config.people.C89_ROLES: dict[str, str] = {'CI': 'Project co-investigator', 'CO': 'Project collaborator', 'CP': 'Cruise participant', 'DC': 'Dataset supply contact', 'DI': 'Cruise dataset principal investigator', 'DM': 'Data manager', 'DP': 'Project dataset principal investigator', 'MC': 'Cruise data manager', 'MP': 'Project data manager', 'PA': 'Project administrator', 'PD': 'Project research assistant', 'PG': 'Project research student', 'PI': 'Project principal investigator', 'PM': 'Project manager', 'PS': 'Cruise principal scientist', 'RA': 'Project technician', 'TC': 'Cruise technical contact', 'TS': 'Cruise technician'}
Convenience alias for the default vocabulary’s terms (code → prefLabel).
- ctdcast.config.people.CREATOR_TYPES = ('person', 'group', 'institution', 'position')
ACDD 1.3
creator_typevalues.
- ctdcast.config.people.DEFAULT_INSTITUTION_ROLE = 'CONMEM'
Applied to an institution entry that names no role.
- ctdcast.config.people.DEFAULT_INSTITUTION_ROLE_VOCABULARY = 'C59'
Used when
cruise_info.institution_role_vocabularyis absent.
- ctdcast.config.people.DEFAULT_ROLE_VOCABULARY = 'C89'
Used when
cruise_info.role_vocabularyis absent.
- ctdcast.config.people.FORBIDDEN_IN_VALUE = (';', ',')
Characters that may never appear inside a field value, because a reader assuming either convention’s delimiter would mis-split the string.
- ctdcast.config.people.INSTITUTION_ROLE_VOCABULARIES: dict[str, dict[str, Any]] = {'C59': {'terms': {'CONLEAD': 'Project Lead', 'CONMEM': 'Project Member', 'DATM': 'Project Data Management', 'FUND': 'Project Funding body'}, 'title': 'BODC organisation roles within activities and projects', 'uri': 'https://vocab.nerc.ac.uk/collection/C59/current/'}, 'W08': {'terms': {'CONT0001': 'Manufacturer', 'CONT0002': 'Owner', 'CONT0003': 'Operator', 'CONT0004': 'PI', 'CONT0005': 'Technical Coordinator', 'CONT0006': 'Data scientist', 'CONT0007': 'Service Provider'}, 'title': 'SensorML Contact Section Terms', 'uri': 'https://vocab.nerc.ac.uk/collection/W08/current/'}}
Vocabularies for INSTITUTION roles — a different question from a person’s role, and a different list. OG1 names W08 for this (
contributing_institutions_role_vocabulary), and W08’s organisational terms (Operator, Owner, Manufacturer, Service Provider) genuinely are organisation roles, which is why W08 belongs here even though it was the wrong list for people. C59 is the default because it is project-shaped — lead, member, data management, funding body — which is what an institution is to a funded cruise, and because it can name a funder, which W08 cannot.
- ctdcast.config.people.ROLE_VOCABULARIES: dict[str, dict[str, Any]] = {'C89': {'terms': {'CI': 'Project co-investigator', 'CO': 'Project collaborator', 'CP': 'Cruise participant', 'DC': 'Dataset supply contact', 'DI': 'Cruise dataset principal investigator', 'DM': 'Data manager', 'DP': 'Project dataset principal investigator', 'MC': 'Cruise data manager', 'MP': 'Project data manager', 'PA': 'Project administrator', 'PD': 'Project research assistant', 'PG': 'Project research student', 'PI': 'Project principal investigator', 'PM': 'Project manager', 'PS': 'Cruise principal scientist', 'RA': 'Project technician', 'TC': 'Cruise technical contact', 'TS': 'Cruise technician'}, 'title': 'BODC dataset roles', 'uri': 'https://vocab.nerc.ac.uk/collection/C89/current/'}, 'G04': {'terms': {'001': 'resourceProvider', '002': 'custodian', '003': 'owner', '004': 'user', '005': 'distributor', '006': 'originator', '007': 'pointOfContact', '008': 'principalInvestigator', '009': 'processor', '010': 'publisher', '011': 'author', '012': 'sponsor', '013': 'coAuthor', '014': 'collaborator', '015': 'editor', '016': 'mediator', '017': 'rightsHolder', '018': 'contributor', '019': 'funder', '020': 'stakeholder'}, 'title': 'ISO 19115 CI_RoleCode', 'uri': 'https://vocab.nerc.ac.uk/collection/G04/current/'}}
The role vocabularies ctdcast can author in, keyed by short name. Each maps concept id → prefLabel, and each
uriis what goes intocontributor_role_vocabulary.These are alternatives to CHOOSE BETWEEN, never things to convert between. A mapping from one to another is necessarily lossy — C89 distinguishes the chief scientist from the funded PI, G04 cannot — and a converted file has no way to record that the detail was dropped. So ctdcast validates and writes in whichever vocabulary the config declares, and does not translate. A consumer needing a different vocabulary can map it themselves, knowing their own tolerance for the loss; that judgement is not ours to make.
- ctdcast.config.people.SEPARATOR = '; '
Separator written between entries in the parallel attribute strings. OceanSITES mandates
;; OG1 says,. Semicolon is the safe choice because personal names and “Name, Institution” strings contain commas.
- ctdcast.config.people.check_contributors(cruise_info: dict[str, Any]) tuple[list[str], list[str]][source]
Validate the
contributors/creatorblock of a cruise config.Returns
(errors, warnings). Errors are conditions that would produce a misleading file — a value containing a delimiter, an unknown role, an unresolvable institution. Warnings are omissions that are legitimate while a cruise is in progress but should be settled before publication.- Parameters:
cruise_info (dict) – The
cruise_info:mapping from the cruise config.- Returns:
errors (list of str)
warnings (list of str)
- ctdcast.config.people.contributor_attrs(cruise_info: dict[str, Any], source: str | None = None) dict[str, str][source]
Build the netCDF global attributes describing people and institutions.
Generates the parallel strings from one structured list, so they cannot fall out of alignment. An attribute whose every entry would be empty is omitted entirely rather than written as
"; ; ; "— an empty delimited string asserts four empty values, whereas an absent attribute asserts nothing was recorded.contributing_institutionscomes fromcruise_info.institutions, a list independent of the people: one person may sit at several institutions and one institution may send several people, so the two must never be zipped together.- Parameters:
cruise_info (dict) – The
cruise_info:mapping from the cruise config.source (str or None, optional) – Provenance hint forwarded to
entry_roles()when resolving each entry’s roles;Noneuses the default resolution.
- Returns:
Global attributes, ready to merge into the dataset.
- Return type:
dict of str to str
- ctdcast.config.people.entry_roles(entry: dict[str, Any], source: str | None = None) list[str][source]
Return the roles a contributor holds on one product file, in order.
A contributor states their roles once, in any of three forms:
role: PI # one role, every file roles: [PS, PI] # several roles, every file roles: # scoped: some roles only on one product all: [PI] ctd: [DC] ladcp: [DI]
The scoped form is what lets one person be a project PI everywhere but the supply contact for the CTD file only, without repeating their name and ORCID and without a second contributor list.
- Parameters:
entry (dict) – One contributor mapping.
source (str or None) – Product key being written (
"ctd","ladcp").Nonereturns the union across every scope, which is what validation wants.
- Returns:
Role tokens as written in config (codes or prefLabels),
allfirst then the product’s own, with duplicates removed and order preserved.- Return type:
list of str
- ctdcast.config.people.institution_registry_paths(extra: Path | str | None = None) list[Path][source]
Return the registry files to merge, lowest precedence first.
Institutions are looked up in three places so that a user who installed ctdcast from PyPI never has to edit package data — which lives in site-packages and is overwritten on upgrade:
the packaged
ctdcast/config/institutions.yaml(the shipped defaults);~/.config/ctdcast/institutions.yaml(or$CTDCAST_CONFIG_DIR), for institutions an individual or lab uses across cruises;a path given by
cruise_info.institutions_file, for a registry kept beside the cruise config and version-controlled with it.
Later files win key-by-key. A fourth route needs no registry at all: an entry in
cruise_info.institutionsmay be written inline with its ownnameandid— seeresolve_institution(). Inline is the most reproducible option, since the config then carries everything the file asserts.- Parameters:
extra (Path or str, optional) – The
institutions_filefrom the cruise config.- Returns:
Existing files only, in merge order.
- Return type:
list of Path
- ctdcast.config.people.institution_role_vocabulary(cruise_info: dict[str, Any]) tuple[str, dict[str, Any]][source]
Return
(name, spec)for the institution-role vocabulary in use.- Parameters:
cruise_info (dict) – The
cruise_info:mapping;institution_role_vocabularyselects by short name and defaults toDEFAULT_INSTITUTION_ROLE_VOCABULARY.- Returns:
name (str)
spec (dict)
- Raises:
KeyError – If the named vocabulary is unknown.
- ctdcast.config.people.institutions_with_source(extra: Path | str | None = None, inline: list[Any] | None = None) list[dict[str, str]][source]
Return every resolvable institution with the source that supplies it, for
ctdcast list.Follows
load_institutions()’ merge order (packaged < user directory < config file, later winning slug by slug) but keeps the winning source per slug instead of discarding it, then appends the inlinecruise_info.institutionsentries that carry their ownname/idand bypass the registry entirely. Keeping the merge here rather than re-walking the paths in the CLI means the precedence lives in one place.Each item is
{slug, name, id, source};sourceis"packaged","user", the config file’s path, or"inline". Inline entries have an emptyslug— they are named, not keyed.
- ctdcast.config.people.load_institutions(extra: Path | str | None = None) dict[str, dict[str, Any]][source]
Return the merged institution registry, keyed by slug.
Merges every file from
institution_registry_paths(), later files overriding earlier ones slug by slug. Returns an empty mapping when no registry exists, so a config that defines its institutions inline still works.- Parameters:
extra (Path or str, optional) – The
institutions_filefrom the cruise config.- Returns:
Slug → institution entry.
- Return type:
dict
- ctdcast.config.people.orcid_uri(orcid: str) str[source]
Return the resolvable URI for an ORCID given in either accepted form.
Accepts a bare identifier (
0000-0001-8773-7838) or a URL (https://orcid.org/0000-0001-8773-7838) and always returns the URL, so the value written tocontributor_idis canonical regardless of how it was typed.- Parameters:
orcid (str) – Bare identifier or orcid.org URL.
- Returns:
https://orcid.org/<identifier>.- Return type:
str
- Raises:
ValueError – If orcid matches neither form.
- ctdcast.config.people.resolve_institution(item: Any, registry: dict[str, dict[str, Any]]) dict[str, Any] | None[source]
Resolve one
institutions:entry to{name, id, role}.An entry may be written three ways, so that a user who installed ctdcast from PyPI — and therefore cannot edit the packaged registry — can still name an institution the registry does not know:
institutions: - uhh # a registry slug - slug: ulpgc # a slug, with a role role: CONLEAD - name: "Agencia Estatal de Investigación" # inline, no registry id: "https://ror.org/003x0zc53" role: FUND
- Parameters:
item (str or dict) – One entry from
cruise_info.institutions.registry (dict) – The loaded institution registry, keyed by slug.
- Returns:
{"name": str, "id": str, "role": str}, orNonewhen a slug does not resolve.- Return type:
dict or None
- ctdcast.config.people.role_choices(vocabulary: str | None = None) list[str][source]
Return the prefLabels of a role vocabulary, in definition order.
Intended for interactive prompting, where a numbered list of human-readable roles is wanted rather than concept codes.
- Parameters:
vocabulary (str, optional) – Short name from
ROLE_VOCABULARIES; defaults toDEFAULT_ROLE_VOCABULARY.- Returns:
prefLabels, ordered as the vocabulary defines them (cruise-scoped terms first, for C89).
- Return type:
list of str
- ctdcast.config.people.role_vocabulary(cruise_info: dict[str, Any]) tuple[str, dict[str, Any]][source]
Return
(name, spec)for the vocabulary this config authors in.- Parameters:
cruise_info (dict) – The
cruise_info:mapping;role_vocabularyselects the vocabulary by short name and defaults toDEFAULT_ROLE_VOCABULARY.- Returns:
name (str)
spec (dict) –
uri,titleandtermsfor the selected vocabulary.
- Raises:
KeyError – If the named vocabulary is not one of
ROLE_VOCABULARIES.
cnv_header is the counterpart that reads inwards rather
than composing outwards: pure functions over a CNV header’s text, recovering what
the deck unit and SBE Data Processing did to a cast before ctdcast saw it. It opens
no files — the header travels on every stage-1 file in raw_metadata.
Parse the acquisition (*) and start-time (#) header of an SBE CNV/HEX file.
These are pure functions over header text — no file I/O, no xarray. They recover
what a Sea-Bird deck unit and Seasave did to a cast before ctdcast ever read it:
the deck-unit conductivity advance (applied in hardware at acquisition and recorded
nowhere ctdcast previously looked), the system/NMEA clock pair, and which clock the
start_time coordinate is anchored to.
The two verbatim blocks are ground truth; every typed field is a best-effort read on
top of them. A line the classifier does not recognise is kept in verbatim and
skipped — never dropped, never fatal. Absence of a deck-unit line means a different
instrument class, not “unaligned”: it yields an empty DeckUnit, not an error.
Header lines are matched with str.splitlines(), which handles the CRLF endings
these Windows-written files carry, and whitespace inside a value is collapsed before
a timestamp is parsed (some writers emit NMEA UTC with a double space).
The header text is not read from a CNV file: it is already carried on every stage-1
file in the raw_metadata global attribute. header_from_raw_metadata()
extracts it from that JSON envelope, so nothing here opens a file.
- class ctdcast.config.cnv_header.Acquisition(deck_unit: DeckUnit, clocks: Clocks, seasave_version: str | None, verbatim: str)[source]
Everything recovered from the
*acquisition block of a CNV/HEX header.- seasave_version: str | None
- verbatim: str
- ctdcast.config.cnv_header.CONFORMANCE_MATCH = 'match'
Conformance-tick states for the “Matches reference” column. Three states, never two: a dash (no reference) must never read as a cross (differs).
- class ctdcast.config.cnv_header.Clocks(system_upload: str | None = None, system_utc: str | None = None, nmea_utc: str | None = None, offset_seconds: float | None = None, system_dt: datetime | None = None, nmea_dt: datetime | None = None)[source]
The two acquisition clocks and their offset.
offset_secondsisnmea_utc - system_utc(positive when the system clock is slow relative to GPS), orNoneif either timestamp is absent or unparseable.system_dt/nmea_dtare the same two clocks already parsed todatetimewhile computing the offset, exposed so callers need not re-parse the verbatim strings.- nmea_dt: datetime | None = None
- nmea_utc: str | None = None
- offset_seconds: float | None = None
- system_dt: datetime | None = None
- system_upload: str | None = None
- system_utc: str | None = None
- class ctdcast.config.cnv_header.ConformanceTick(key: str, label: str, state: str, reference: str, source: str, detail: str, variables: str = '')[source]
One correction row’s agreement with a documented reference value.
stateisCONFORMANCE_MATCH/CONFORMANCE_DIFFER/CONFORMANCE_NO_REFERENCE, rendered ✓ / ✗ / — and never conflated: a dash means “no reference to check against”, not “differs”. A match means the value equals a documented typical value, NOT that it is “correct”; a differ is a deviation to judge, NOT “wrong” — every reference is configuration-dependent, sosourcenames where the reference came from and the tick is a match indicator, not a verdict.keyaligns the tick to theCorrectionit belongs to ("celltm","align", a suffixed"wildedit_2"); several ticks share one key when a correction expands per sensor/channel (labeldisambiguates the sub-rows).detailis the parsed-vs-reference text;variablesnames the variable(s) the step modified, for a display column (conductivity_1for a cell-thermal-mass row, the smoothed channels for a filter row,all channels: raw→physicalfor datcnv). The comparison is producer-agnostic: it checks the parameters, not who applied them, so the same references serve a future stage-3 correction ledger.- detail: str
- key: str
- label: str
- reference: str
- source: str
- state: str
- variables: str = ''
- class ctdcast.config.cnv_header.Correction(label: str, key: str, producer: str, version: str, parameters: str)[source]
One correction in the ledger, split into producer, version and parameters.
The deck-unit alignment and each Sea-Bird Data Processing module become a record.
labelis the display name ("align (deck)","celltm","celltm (2)"for a repeat);keyis the attribute suffix ("align","celltm","celltm_2").parametersis the salient parameter string, possibly empty.- key: str
- label: str
- parameters: str
- producer: str
- version: str
- class ctdcast.config.cnv_header.DeckUnit(model: str | None = None, firmware: str | None = None, advance: dict[str, float]=<factory>, scans_averaged: int | None = None)[source]
The SBE 11plus deck unit’s acquisition-time configuration.
advancemaps a channel name exactly as the header spells it ("primary conductivity","voltage 0") to its advance in seconds. It is a per-channel map, not a boolean or single value: a V 5.0 unit advances the two conductivity channels by different amounts, and that asymmetry is the whole reason this parse exists. An empty map withmodelNonemeans no deck-unit line was found.- advance: dict[str, float]
- firmware: str | None = None
- model: str | None = None
- scans_averaged: int | None = None
- class ctdcast.config.cnv_header.ProcessingChain(steps: list[ProcessingStep] = <factory>, verbatim: str = '')[source]
The SBE Data Processing module chain, in file order.
stepsis one entry per contiguous run of a module’s lines, so a module that runs twice (SBE sanctions running Wild Edit more than once) yields two steps and is never deduplicated.verbatimis the#processing region exactly as read, so an unrecognised line is preserved there even when it contributes no step.- steps: list[ProcessingStep]
- verbatim: str = ''
- class ctdcast.config.cnv_header.ProcessingStep(module: str, params: dict[str, str]=<factory>)[source]
One SBE Data Processing module in the
#chain.moduleis the name exactly as the header spells it (datcnv,wildedit,Derive).paramsmaps each field to its verbatim value in file order — keys keep the SBE channel names the header uses (low_pass_A_vars,action t090C); values are never split, so a comma-collection such asno, min = 0.000, ...and a date’s trailing[datcnv_vars = 15]stay intact.- module: str
- params: dict[str, str]
- class ctdcast.config.cnv_header.SbeHistoryNote(timestamp: str, version: str, stage: str, note: str)[source]
One SBE Data Processing module rendered as a
historyline’s components.timestampis the module’s verbatim SBE stamp (noZ);stageis the module name as spelled;noteis its salient-parameter body. The deck-unit align has no timestamp and is not a history note (it iscorrection_align).- note: str
- stage: str
- timestamp: str
- version: str
- class ctdcast.config.cnv_header.SensorCalibration(kind: str, label: str, serial: str, slope: str, offset: str, slope_nondefault: bool, offset_nondefault: bool)[source]
A frequency sensor’s drift Slope/Offset, read from the CNV
<Sensors>config.Answers “was a calibration correction already applied before ctdcast read the cast?” Sea-Bird applies
corrected = slope * computed + offsetatdatcnvto the derived physical value (not the raw frequency), so a non-identity slope/offset means a drift or span correction is already baked into the data.kindis"temperature"/"conductivity"/"pressure";labeldisambiguates dual sensors ("temperature_1","temperature_2"), or is the bare kind when a sensor is single.slopeandoffsetare kept verbatim as the config spells them, so no precision is lost.slope_nondefault/offset_nondefaultflag each value that departs from its identity (slope 1.0, offset 0.0) independently, so the display can mark exactly which one carries a correction;drift_appliedis True when either does. Only the frequency sensors are read: a voltage sensor’s Slope/Offset (pH, transmissometer) is a native calibration, not a drift knob.- property drift_applied: bool
True when a drift or span correction is baked into either value.
- kind: str
- label: str
- offset: str
- offset_nondefault: bool
- serial: str
- slope: str
- slope_nondefault: bool
- class ctdcast.config.cnv_header.StartTime(value: str | None, clock: str, anchor: str, source: str = '')[source]
The
# start_timevalue and the provenance named in its bracket.sourceis the verbatim bracket text ("System UTC, first data scan.");clockandanchorare its resolved halves.sourceis""when the line carries no bracket.- anchor: str
- clock: str
- source: str = ''
- value: str | None
- ctdcast.config.cnv_header.build_correction_ledger(header_text: str) dict[str, str | float][source]
Build the stage-1 correction ledger from an SBE header.
Records what was done to the cast before ctdcast read it, so a later stage does not re-apply a correction already made. Two verbatim blocks (
sbe_acquisition,sbe_processing) are the ground truth; the structured attributes are convenience summaries over them: the module order, the deck-unit alignment, onecorrection_<module>per step in the chain (a repeat suffixed_2,_3…), and the time coordinate’s source and offset. Every module gets a summary — deciding which “count” as corrections would require predicting them. An attribute is written only when the header supports it; an absentcorrection_<name>means “not recorded”, never “not done”. Returns an empty mapping for an empty or missing header (e.g. a LADCP file), so the caller can update attributes unconditionally.
- ctdcast.config.cnv_header.cal_value_nondefault(value: str, default: float) bool[source]
Return whether a calibration value departs from its default beyond tolerance.
Decides whether a stored slope (default 1.0) or offset (default 0.0) carries a drift/span correction, so the cast page can amber-tint and bold it — used both on the
sensor_calibrationsheader parse and, since the sensor catalog, on the per-castSENSOR_*attributes.- Parameters:
value (str) – The stored slope or offset, verbatim as read from the SBE header or a
SENSOR_*catalog attribute.default (float) – The identity value it is compared against —
1.0for a slope,0.0for an offset.
- Returns:
True when value parses and lies more than the identity tolerance (
_CAL_IDENTITY_TOL, 5e-7) from default; False when it is within tolerance or cannot be parsed — a drift that cannot be read is not claimed.- Return type:
bool
- ctdcast.config.cnv_header.conformance_advisories(header_text: str) list[str][source]
Where a cast’s processing deviates from a documented reference, as hedged prose.
A sibling of
provenance_advisories()(which reports structural facts); this reports comparisons to a recommendation. Two sources, one list: every parameter tick thatCONFORMANCE_DIFFER`s becomes a sentence, plus three structural checks — oxygen present but not aligned (SBE p.86), a temperature/conductivity channel low-pass smoothed when SBE smooths pressure only (p.100), and Bin Average without a preceding Loop Edit. Every string cites its source and hedges: a deviation is a value to judge, never asserted "wrong", because each reference is configuration-dependent. Empty header, or an instrument outside the SBE 9 family (see :func:`conformance_supported) →[].
- ctdcast.config.cnv_header.conformance_supported(header_text: str) bool[source]
Return True when documented references exist for this cast’s instrument (the SBE 9 family).
The values in
_CONFORMANCE_REFS(celltm α/τ, filter-pressure tc, deck conductivity advance) are the SBE 9 / 11plus defaults. Other instruments (SBE 19plus, 25, …) have different documented defaults that are not yet encoded, so conformance is reported only for the SBE 9 family; every other instrument yields no ticks and, on the report, a note naming it. An unrecognised header is unsupported.
- ctdcast.config.cnv_header.conformance_ticks(header_text: str, rename_map: dict[str, str] | None = None) list[ConformanceTick][source]
Compare a cast’s processing parameters against documented references, one tick per row.
A sibling of
correction_records(): the ticks align to those records bykey, and a correction whose parameters are per-sensor/per-channel lists (deck align, cell thermal mass) expands into one tick per channel/cell. Modules with no documented reference (datcnv, binavg, loop edit, wfilter, Derive, …) get a singleCONFORMANCE_NO_REFERENCEtick — a dash, never a cross. The comparison reads the structuredProcessingStepparameters and is producer-agnostic, so the same references serve a future stage-3 ledger. rename_map is the cast’s{sbe_column_name: variable_name}(from the reader’scnv_original_name), used to canonicalise the Variables column to the names the reader actually applied — see_canonical_var(). Returns an empty list for an empty header, or for an instrument outside the SBE 9 family (seeconformance_supported()).
- ctdcast.config.cnv_header.correction_records(header_text: str) list[Correction][source]
Return the structured corrections (deck-unit align + each module) in file order.
The display counterpart of the flat
correction_<key>ledger attributes: same records, split into producer / version / parameters instead of one string.
- ctdcast.config.cnv_header.detect_instrument(header_text: str) str | None[source]
Return the CTD model from the header’s first
* Sea-Bird …line, or None.The spelling’s spacing is inconsistent across firmware (
SBE 9vsSBE19plus), so it is normalised to a single space afterSBE(SBE19plus→SBE 19plus). Returns None for a header with no such line (a bare hex, a LADCP file, an empty header).
- ctdcast.config.cnv_header.header_from_raw_metadata(raw_metadata: str | None) str | None[source]
Extract the verbatim SBE header text from a stage-1 file’s
raw_metadataattr.raw_metadatais the JSON envelope seasenselib writes and stage 1 carries through:{"schema": ..., "raw_format": "sbe-cnv", "blocks": {"header": "<* and # lines>", ...}}. Returns theblocks.headerstring, orNoneif the attribute is absent, is not JSON, or does not carry a header string. The schema string is not required to match, so a schema bump degrades gracefully rather than raising — only a change to the envelope shape yieldsNone.
- ctdcast.config.cnv_header.parse_processing_chain(header_text: str) ProcessingChain[source]
Parse the
#SBE Data Processing module chain into ordered steps.Module identity is the token before the first
_(datcnv_date->datcnv), which is correct for every SBE#-block module — they are all single tokens.A new step starts when the module changes or when a field reappears within the current run. The second rule catches a module run twice as adjacent blocks — the case SBE’s manual explicitly sanctions for Wild Edit — where the module name never changes: the repeat of
wildedit_date(or any field already seen) opens a fresh step, so the first run’s parameters are never overwritten. This works even for a dialect that emits no_dateline at all, unlike splitting on_date. Repeats are therefore kept, never deduplicated.Values are split on the first
=only, leaving comma-collections and bracketed second fields verbatim. A line that is notkey = valueis preserved inProcessingChain.verbatimbut adds no step, so an unrecognised dialect degrades to verbatim rather than to an empty chain.
- ctdcast.config.cnv_header.parse_sensor_block(header_text: str) list[dict][source]
Parse the CNV
<Sensors>config block into one raw record per channel.The single reader of the block. Each record has
channel(int),element(the type tag, e.g."TemperatureSensor"),sensor_id,serial,calibration_date(verbatim),slopeandoffset(defaulting"1.0"/"0.0"when absent),comment(the sensor-comment body, e.g."Temperature, 2"or"Free") for the caller to resolve to a role, andraw_block(the whole<sensor>element verbatim, for provenance). Returns[]for an empty header or one with no<Sensors>block. Role interpretation and calibration-date normalisation are the reader’s concern (seectdcast.readers.metadata.parse_sensor_channels()), so they stay out of here to keep the parser free of the sensor-role vocabulary.
- ctdcast.config.cnv_header.parse_star_block(header_text: str) Acquisition[source]
Parse the
*acquisition block of an SBE CNV/HEX header.Recognised lines populate the typed fields; every
*line (bar the sensor XML) is preserved inAcquisition.verbatim. Unrecognised lines are not fatal. Absence of the deck-unit line yields an emptyDeckUnit.
- ctdcast.config.cnv_header.parse_start_time(header_text: str) StartTime[source]
Parse
# start_timeand resolve its bracket to a(clock, anchor)pair.The provenance bracket is optional: a
start_timeline without one still yields its timestamp value, withclock/anchor"unknown". A file with nostart_timeline at all (e.g. a raw HEX, which carries no#block) returnsvalue=None.
- ctdcast.config.cnv_header.provenance_advisories(header_text: str) list[str][source]
Structural implications of the SBE ledger — what the file is and what was done.
Each is a fact about the cast and its consequence, not a comparison to any recommendation. The single source for both a stage-1 warning and the note under the cast-page provenance table, so the two never drift. (The SBE-conformance check — where a cast deviates from Sea-Bird’s published recommendation — is a separate, later concern and deliberately not here.)
- ctdcast.config.cnv_header.sbe_history_notes(header_text: str) list[SbeHistoryNote][source]
Return one history note per timestamped SBE Data Processing module, file order.
Attributed to Sea-Bird so a stage-1 file’s
historyshows what SBE did before ctdcast, oldest-first. A module with no_dateline — the third-party dialect that emits none — is skipped, because a history line needs a stamp and it has none: such a module still appears incorrection_<module>andsbe_processing_orderbut not inhistory. The deck-unit align has no timestamp either and is recorded only ascorrection_align. The note body is the same summary the ledger uses.
- ctdcast.config.cnv_header.sensor_calibrations(header_text: str) list[SensorCalibration][source]
Read each frequency sensor’s drift Slope/Offset from the CNV
<Sensors>config.Surfaces whether a drift or span correction is already baked into the data before ctdcast read it (see
SensorCalibration). A projection overparse_sensor_block(): only temperature, conductivity and pressure sensors are returned, in config order; voltage sensors are skipped because their Slope/Offset is a native calibration rather than a datcnv drift knob. Dual sensors are labelled<kind>_1/<kind>_2and a lone sensor keeps the bare kind. Returns an empty list for an empty header or a header with no<Sensors>block.
- ctdcast.config.cnv_header.start_time_clock(bracket: str) str[source]
Classify a
time_coordinate_sourcebracket to its clock:system/nmea/unknown.The public single point of truth for “which clock is the time coordinate anchored to”, used by the stage-2 clock applier’s gate so it does not re-parse the bracket with a fragile substring. An empty or unrecognised bracket resolves to
"unknown".
Acquisition-clock diagnostic
clock reads each cast’s System/NMEA clock pair off the
header and classifies the cruise’s clock error (constant, step, drift, or too
little to tell), recursively segmenting the whole-second offsets into levels and
claiming a drift only when a rate positively fits. It is read-only — it finds the
error and generates a paste-ready processing.clock block; applying the
correction is a stage-2 step.
Cruise-scale acquisition-clock diagnostic: per-cast NMEA-vs-System offsets, changepoints, verdict.
Reads the acquisition (System) and GPS (NMEA) clock pair off each cast’s raw Sea-Bird header —
already on every stage-1 file in raw_metadata — and classifies the cruise’s clock error as a
constant offset, a drift, or one or more step changes. Read-only: this finds the error; whether
a correction is applied is a separate stage-2 decision, and turns not only on the kind here but
on whether the time coordinate is already on GPS (start_time sourced from NMEA) — the two are
orthogonal.
Sign convention, stated once because a sign error here is silent and catastrophic:
offset = NMEA - System — the seconds added to acquisition time to obtain corrected (GPS)
time. Positive means the acquisition clock ran behind GPS. This matches oceanarray’s
_apply_clock_offset, so the value transfers without a flip.
The offsets are whole seconds, so structure is read as a histogram of levels, not a fitted
line: three contiguous casts at one value are a level, not scatter. classify_offsets()
recurses, splitting at each detected step until every segment is flat (sd below the quantisation).
A segment that will not flatten is a drift only if a rate positively fits it — otherwise it is a
noisy level, reported as a constant/step with a resolution caveat rather than a drift.
classify_offsets() is pure; clock_offsets() reads the cast files.
- class ctdcast.analysis.clock.CastClock(cast_id: str, system_utc: datetime, nmea_utc: datetime, offset_seconds: float)[source]
One cast’s acquisition/GPS clock pair, with both raw times so the sign is checkable.
Carries both times and the cast id (not just the derived offset): the per-cast table shows both so a reader can confirm the sign, and segments are labelled by cast range.
- cast_id: str
- nmea_utc: datetime
- offset_seconds: float
- system_utc: datetime
- class ctdcast.analysis.clock.ClockSegment(first_cast: str, last_cast: str, start: datetime, end: datetime, offset_seconds: float, sd_seconds: float, n_casts: int)[source]
A run of casts sharing one clock offset, between changepoints — one row of the segment table.
- property cast_range: str
Cast-id span as
"001-003"(or just"001"for a single-cast segment).
- end: datetime
- first_cast: str
- last_cast: str
- n_casts: int
- offset_seconds: float
- sd_seconds: float
- start: datetime
- class ctdcast.analysis.clock.ClockVerdict(kind: str, segments: list[ClockSegment], note: str, resolution_caveat: str = '')[source]
The cruise’s clock classification: kind, level segments and a plain-language note.
kindis"constant","drift","step","no_clock_pair"or"insufficient"."drift"is claimed only when a rate is measured (the note quotes slope, R² and residual sd) — failing to resolve into flat segments is not evidence of a rate, so a merely noisy cruise is a"constant"(or"step") with a resolution caveat, never a drift."no_clock_pair"(files scanned, none carrying a pair) is distinct from"insufficient"(too few casts, or no files) so a renderer selects onkindalone, never on the note prose.resolution_caveatis non-empty when a segment’s sd exceeds the quantisation: a statement about resolution (per-cast comparison noise, or possibly-hidden structure), carried beside the verdict rather than as a verdict of its own. Changepoints are exposed viachangepoints.- property changepoints: list[datetime]
Acquisition times of the segment boundaries (empty unless the verdict is a step).
- kind: str
- note: str
- resolution_caveat: str = ''
- segments: list[ClockSegment]
- ctdcast.analysis.clock.cast_number(cast_id: str) int[source]
Integer cast number from a formatted cast id (
"001"-> 1,"029b"-> 29).
- ctdcast.analysis.clock.classify_offsets(series: list[CastClock], *, n_scanned: int | None = None) ClockVerdict[source]
Classify a cruise’s offset series (sorted by acquisition time) into a
ClockVerdict.Recursively segments the series and reads the result: one segment is a constant, several are a step. A segment that will not flatten triggers a rate test — only a rate that positively fits is a drift; otherwise the segments stand and the non-flat sd becomes a resolution caveat (per-cast comparison noise, or possibly-hidden structure). n_scanned, when given, is reported in the too-few message so a reader can tell “no clocks in these headers” from “few casts present”; the cause of an empty result is not asserted here.
- ctdcast.analysis.clock.clock_offsets(root: Path | str) tuple[list[CastClock], int, dict[str, int]][source]
Return the casts carrying both clocks, the number of casts scanned, and a coordinate census.
Discovers casts with
group_by_cast(), the same best-available machinery the report uses, so it works across the nestedstageN/layout, a flatnc_dir, or a tree processed only to stage 2/3 — and skips compiled products and unnumbered files. The header (raw_metadata) rides every stage, so the raw clock pair is read from each cast’s stage-1 file when present, else its lowest available stage. A cast whose clock is missing/unparseable is skipped and counted in a single warning. The census counts which clock each cast’sstart_timebracket anchors the coordinate to (systemvsnmea) — the fact the stage-2 gate keys on, orthogonal to the offset structure. Only determined sources are counted; a bracketless header ("unknown") is left out rather than tallied as a source it never named.
- ctdcast.analysis.clock.coordinate_summary(coordinate_counts: dict[str, int], scanned: int) str[source]
One line stating which clock the time coordinate is anchored to, and the correction consequence.
Orthogonal to the offset verdict: a cruise can show a real offset yet need no correction because its coordinate was already on GPS (OdB). Reports the count, not a single answer — mixed values across a cruise are themselves a finding.
nmeaneeds no correction;systemis where a non-zero offset would be applied.
- ctdcast.analysis.clock.correction_status(coordinate_counts: dict[str, int]) str[source]
Whether a clock correction applies, from the coordinate census:
apply/on_gps/undetermined.The single source of truth for the gate both the CLI and the report read, so they cannot disagree. Any cast on the System clock means a correction is relevant (
apply); a coordinate only ever on GPS needs none (on_gps); an empty census means nostart_timebracket was parsed, so whether a correction applies isundetermined— which must NOT be reported as “already on GPS”, a source that was never measured.
- ctdcast.analysis.clock.suggested_config_yaml(verdict: ClockVerdict) str | None[source]
Paste-ready
processing.clockYAML for a step/constant verdict, orNoneif inapplicable.One
segmentsentry per segment, keyed on inclusive cast-number ranges (thectd_groupings.yamlidiom). ReturnsNonefordrift(no applier),no_clock_pairandinsufficient. The sign is emitted so it is copied, never hand-computed — that is the one thing here that is silently catastrophic to get backwards. A[[lo, hi]]range expands to plain casts only, so a segment holding lettered casts (029b) or gaps gets a warning comment to verify by hand rather than silently mis-scoping the correction.Each segment also carries
n_castsandclock_offset_sd_seconds— the evidence that justified the offset. The applier records these with the correction so the number’s uncertainty travels with it; sourcing them here (not from a re-measurement at apply time) keeps the recorded statistics describing the population that actually produced the number.
Section manifest
Each report page is described by a Profile of
Section and
Panel entries, resolved by
resolve() into a numbered, rendered report. The
model is package-neutral; each page’s concrete registry lives in its own generator
module. _anchors maps the old hand-authored #s-* anchors onto the new
section ids for one release.
Section-manifest model and resolver for report pages.
A report page is described by a Profile: an ordered sequence of
Section (and Expand) entries, each naming Panel ids
that render figures or tables. resolve() walks a profile once against a
render context and returns a ResolvedReport whose sections are numbered
over the rendered subset — absent sections leave no gap, and identity is the
section id (a stable slug), not the integer.
This module holds only the model and the resolution algorithm — it is
package-neutral (names no variable, page, or science) and is vendored
byte-identical to the sister repos. Each page’s concrete registry (its
Ctx builder, panels, sections, and profiles) lives in that page’s own
module — grid’s in reports/_grid.py, and so on — never in a plural
_manifests.py companion.
- class ctdcast.reports._manifest.Expand(over: Callable[[Any], Sequence[Any]], section: Callable[[Any], Section])[source]
A single manifest entry that becomes N sections at resolution time.
- Parameters:
over (Callable) – Returns the sequence of items to expand over (e.g. the variables present on the page, in display order).
section (Callable) – Builds one
Sectionper item yielded byover.
- over: Callable[[Any], Sequence[Any]]
- class ctdcast.reports._manifest.Panel(id: str, render: ~collections.abc.Callable[[~typing.Any], str | None], kind: ~typing.Literal['figure', 'html', 'table'] = 'figure', slot: str | ~collections.abc.Callable[[~typing.Any], str] | None = 'full', caption: str | None = None, applies_to: ~collections.abc.Callable[[~typing.Any], bool] = <function _always>, unavailable_if: ~collections.abc.Callable[[~typing.Any], str | None] = <function _unavailable_never>)[source]
One figure or table, addressable by id.
- Parameters:
id (str) – Unique panel identifier, e.g.
"temperature_field". Also the anchor a caption or cross-reference can point at.render (Callable) – Adapter that returns the panel’s payload — a base64 PNG string for a
"figure", or ready-to-emit markup for"html"/"table"— orNonewhen the panel applies but no output could be produced (a plot raised, or the variable is absent), which the resolver turns into a.warnstub. Figure adapters wrap the existing_make_*_b64functions unchanged.kind ({"figure", "html", "table"}, optional) – Content discriminator. The template macro branches on it so that only
"html"/"table"payloads are emitted|safe; a"figure"payload is always escaped into an imagesrc.slot (str or Callable or None, optional) –
SLOTSkey giving the panel’s display width, or a callablectx -> slotthat computes it (e.g. from a section’s aspect ratio). Belongs to the panel, not the section, so a panel is the same width wherever it is placed. The resolver calls it when callable.Nonemarks a figure that is not rendered through the slot system: it is emitted as a bare.figthat fills the content column (noslot-*class, so no rendered-width contract).caption (str, optional) – Caption text rendered beneath the panel.
applies_to (Callable, optional) – Predicate deciding whether the panel is attempted at all.
Falseomits the panel silently (it is excluded, not merely unavailable).unavailable_if (Callable, optional) – Precondition checked before
render: given the context, return a reason string when the panel applies but cannot be produced (e.g. a required metadata field is missing), orNoneto proceed. A returned reason becomes a.warnstub with that reason andrenderis not called. This is the channel for defects knowable from context; arenderthat still returnsNonegets the generic stub reason, because that is the “should have worked and didn’t” case.
- applies_to() bool
Return True for any context (the default
applies_topredicate).
- caption: str | None = None
- id: str
- kind: Literal['figure', 'html', 'table'] = 'figure'
- render: Callable[[Any], str | None]
- slot: str | Callable[[Any], str] | None = 'full'
Return None for any context (the default
unavailable_ifpredicate).
- class ctdcast.reports._manifest.PanelGroup(over: Callable[[Any], Sequence[Any]], panel: Callable[[Any], Panel])[source]
A
Section.panelsentry that expands to N panels at resolution time.Places a data-driven run of panels — one per item — under a single heading without inflating the section count. Use this (not
Expand) when the run is over a numeric or unbounded set, e.g. one panel per isopycnal; reserveExpand’s section-level expansion for a closed editorial vocabulary. Numbering reflects editorial structure, not data cardinality.- Parameters:
over (Callable) – Returns the sequence of items to expand over, in display order.
panel (Callable) – Builds one
Panelper item yielded byover.
- over: Callable[[Any], Sequence[Any]]
- ctdcast.reports._manifest.PanelKind
Panel content kinds.
"figure"renders a base64 PNG (escaped as an imagesrc);"html"and"table"render pre-built markup that the template macro emits with|safe. The discriminator keeps theautoescape=Trueboundary to one auditable branch — a figure payload is never|safe-d.alias of
Literal[‘figure’, ‘html’, ‘table’]
- class ctdcast.reports._manifest.Profile(numbering: str = 'flat', entries: tuple[~ctdcast.reports._manifest.Section | ~ctdcast.reports._manifest.Expand, ...]=<factory>)[source]
An ordered page description plus a numbering policy.
- Parameters:
- numbering: str = 'flat'
- ctdcast.reports._manifest.RenderContext
A page’s render context, threaded opaque through the resolver: the concrete keys are decided by each page’s own registry (grid’s
Ctx, etc.), never by this package-neutral module. Named once so the resolver signatures readctx: RenderContextwith no per-signature noqa, and any future narrowing lands in one place.
- class ctdcast.reports._manifest.ResolvedPanel(id: str, kind: Literal['figure', 'html', 'table'], slot: str | None, payload: str | None, caption: str | None, stub_reason: str | None = None)[source]
A panel after rendering: a payload (figure b64 or markup) or a
.warnstub.- caption: str | None
- id: str
- property is_stub: bool
True when the panel applied but produced no output.
- kind: Literal['figure', 'html', 'table']
- payload: str | None
- slot: str | None
- stub_reason: str | None = None
- class ctdcast.reports._manifest.ResolvedReport(sections: tuple[ResolvedSection, ...], not_applicable: tuple[str, ...])[source]
The full resolved page: numbered sections plus the not-applicable list.
- anchor(section_id: str) str | None[source]
Return the display number for section_id, or None if not rendered.
- not_applicable: tuple[str, ...]
- sections: tuple[ResolvedSection, ...]
- class ctdcast.reports._manifest.ResolvedSection(id: str, number: str, title: str, level: int, intro: str | None, panels: tuple[ResolvedPanel, ...], role: str, layout: str | None = None)[source]
A section after resolution: numbered heading plus resolved panels.
- id: str
- intro: str | None
- layout: str | None = None
- level: int
- number: str
- panels: tuple[ResolvedPanel, ...]
- role: str
- title: str
- class ctdcast.reports._manifest.Section(id: str, title: str, panels: tuple[str | PanelGroup, ...], level: int = 2, intro: str | None = None, applies_to: Callable[[Any], bool] | None = None, role: str = 'content', layout: str | None = None)[source]
A numbered heading and the panels beneath it.
- Parameters:
id (str) – Unique section identifier, e.g.
"hydrography". Doubles as the anchor and the cross-reference key.title (str) – Human-readable heading text (without a number — the number is computed).
panels (tuple of (str or PanelGroup)) – Panel ids, in display order. An entry may instead be a
PanelGroup, which expands to a data-driven run of panels under this one heading.level (int, optional) – Heading level (
2for<h2>).intro (str, optional) – Introductory prose rendered under the heading, before the panels.
applies_to (Callable, optional) – Predicate deciding whether the section is included. When
None(default), the section applies if any of its panels apply. A section dropped here is named in the report’s “not applicable” footer line.role (str, optional) –
"content"(numbered1..N) or"appendix"(numberedA..), so appendix material such as a NetCDF-variable table does not pad the science numbering.layout (str, optional) – Within-section panel arrangement.
None(default) stacks panels vertically;"row"lays them out in a wrapping flex row, each panel in a cell sized by itsslot(so twoslot="half"figures sit side-by-side and aslot="full"one wraps to its own line).
- applies_to: Callable[[Any], bool] | None = None
- id: str
- intro: str | None = None
- layout: str | None = None
- level: int = 2
- panels: tuple[str | PanelGroup, ...]
- role: str = 'content'
- title: str
- ctdcast.reports._manifest.resolve(profile: Profile, ctx: Any, panels: dict[str, Panel], *, drop_stub: bool = False) ResolvedReport[source]
Resolve profile against ctx into a numbered
ResolvedReport.One pass: expand
Expandentries, drop sections whoseapplies_tois false (collecting them for the not-applicable footer), resolve each kept section’s panels (Nonerender →.warnstub), then number the survivors —contentsections1..NandappendixsectionsA...- Parameters:
profile (Profile) – The page’s ordered section description and numbering policy.
ctx (Any) – The render context passed to every predicate and render callable (dataset, config, paths — whatever the page’s panels need).
panels (dict of str to Panel) – The panel registry; section
panelsentries are ids into this map.drop_stub (bool, optional) – When True, a section that applies but whose panels are all stubs (applicable-but-entirely-unavailable) is dropped and the survivors renumber over it, so the page closes cleanly. Default False keeps the heading with its stub, which surfaces the failure rather than hiding it. A section dropped this way is not added to
not_applicable— that list stays reserved for genuinely-not-applicable-to-this-deployment sections, so a plot failure never reads as “not applicable”. A section with any non-stub panel is never dropped.
- Returns:
Numbered, rendered sections plus the titles of any omitted sections.
- Return type:
- Raises:
NotImplementedError – If
profile.numbering == "grouped"(reserved for a later branch).KeyError – If a section references a panel id absent from panels.
Legacy anchor aliases for report pages — a transition shim for the D3 fix.
Section ids became page anchors, so the old hand-authored #s-* anchors change
(#s-profiles/#s-physics/#s-hydro all collapse onto #hydrography).
The committed demo pages under docs/source/_static/demo/ deep-link the old
ids, and external links (papers, issues, cruise reports) may too. For one
release each page emits an empty <span id="old"> for every old anchor whose
new section is rendered on it, so those links keep resolving.
Remove this module and the one template call that uses it once the old links are gone. The set was audited complete against the templates and the demo pages.
- ctdcast.reports._anchors.LEGACY_ANCHORS: dict[str, str] = {'s-aux': 'biogeochemistry', 's-biogeo': 'biogeochemistry', 's-diagnostics': 'diagnostics', 's-hydro': 'hydrography', 's-ladcp': 'velocity', 's-map': 'map', 's-overview': 'overview', 's-physics': 'hydrography', 's-profiles': 'hydrography', 's-sensors': 'sensors', 's-stability': 'stability', 's-ts': 'ts_diagram'}
Old
#s-*anchor -> new section id (which is the new anchor). Multiple old anchors mapping to one new id is the D3 defect being retired: Hydrography wass-profiles(cast),s-physics(section/timeseries) ands-hydro(index); Biogeochemistry wass-aux(cast) ands-biogeo(elsewhere).
- ctdcast.reports._anchors.legacy_anchor_spans(rendered_ids: set[str]) str[source]
Return empty
<span id="old">aliases for rendered sections’ old anchors.- Parameters:
rendered_ids – The set of section ids actually rendered on the page. A legacy anchor is aliased only when its new section is present, so no alias dangles at a section the page does not have.
- Returns:
Concatenated empty spans, one per matching legacy anchor (possibly empty).
- Return type:
str
Page generators
Tier-2: generate a per-cast HTML report page.
- ctdcast.reports._cast.CAST_DEFAULT: Profile = Profile(numbering='flat', entries=(Section(id='overview', title='Overview', panels=('ts_density', 'station_map', 'ts_updown'), level=2, intro='CT · SA · σ₀ profiles, station location, and T–S down-vs-up.', applies_to=None, role='content', layout=None), Section(id='hydrography', title='Hydrography', panels=('ct_sa_sigma0',), level=2, intro='CT · SA · σ₀ vs pressure — downcast in colour, upcast in grey.', applies_to=<function _has_ts>, role='content', layout=None), Section(id='biogeochemistry', title='Biogeochemistry', panels=('aux',), level=2, intro='O₂ saturation · fluorescence · turbidity.', applies_to=<function _has_biogeo>, role='content', layout=None), Section(id='ts_diagram', title='T–S diagram', panels=('ts_diagram',), level=2, intro='Coloured by O₂ saturation — downcast only.', applies_to=<function _has_ts>, role='content', layout=None), Section(id='stability', title='Stability', panels=('stability',), level=2, intro='N² and Turner angle — downcast only.', applies_to=<function _has_ts>, role='content', layout=None), Section(id='velocity', title='Velocity (bottom track)', panels=('ladcp_bottomtrack',), level=2, intro=None, applies_to=<function <lambda>>, role='content', layout=None), Section(id='diagnostics', title='Diagnostics', panels=('pressure_time', 'sensor_diff', 'updown_diff'), level=2, intro=None, applies_to=None, role='content', layout=None), Section(id='qc_flags', title='QC flags', panels=('qc_histogram', 'qc_flags'), level=2, intro='QARTOD flags recorded by the pipeline, read back from the cast file: gross-range thresholds applied and the flag-count breakdown per variable.', applies_to=<function _has_qc>, role='content', layout=None), Section(id='sensors', title='Sensors', panels=('sensors_table',), level=2, intro=None, applies_to=<function <lambda>>, role='appendix', layout=None), Section(id='file_provenance', title='File provenance', panels=('file_provenance',), level=2, intro='File identity read back from the cast file: cast, OceanSITES data mode, processing stage, and the tracking_id lineage to the source CNV.', applies_to=<function _has_file_provenance>, role='appendix', layout=None), Section(id='data_ranges', title='netCDF data ranges', panels=('data_ranges',), level=2, intro='Min · max · valid count for every variable in the cast file on disk.', applies_to=None, role='appendix', layout=None), Section(id='provenance', title='Processing provenance', panels=('provenance',), level=2, intro='What the SBE deck unit and Sea-Bird Data Processing did to this cast before ctdcast read it, recovered from the raw header.', applies_to=<function _has_provenance>, role='appendix', layout=None)))
same figures, same grouping — the manifest only renumbers (closing the D1 gaps), generates the jump-nav, and turns section ids into anchors (the D3 fix).
- Type:
The cast page profile. Conservative port of the current page
- ctdcast.reports._cast.CAST_PANELS: dict[str, Panel] = {'aux': Panel(id='aux', render=<function <lambda>>, kind='figure', slot='full', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'ct_sa_sigma0': Panel(id='ct_sa_sigma0', render=<function <lambda>>, kind='figure', slot='full', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'data_ranges': Panel(id='data_ranges', render=<function <lambda>>, kind='table', slot='full', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'file_provenance': Panel(id='file_provenance', render=<function _render_file_provenance>, kind='table', slot='full', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'ladcp_bottomtrack': Panel(id='ladcp_bottomtrack', render=<function <lambda>>, kind='figure', slot='third', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'pressure_time': Panel(id='pressure_time', render=<function <lambda>>, kind='figure', slot='third', caption='Cast trajectory: pressure vs elapsed time', applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'provenance': Panel(id='provenance', render=<function _render_provenance_panel>, kind='table', slot='full', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'qc_flags': Panel(id='qc_flags', render=<function <lambda>>, kind='table', slot='full', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'qc_histogram': Panel(id='qc_histogram', render=<function <lambda>>, kind='figure', slot='full', caption='Data-value distributions: grey = all data, colour = kept (soak/deck and missing excluded); orange dashed = gross-range suspect threshold. A threshold outside the data flags a mis-set bound.', applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'sensor_diff': Panel(id='sensor_diff', render=<function <lambda>>, kind='figure', slot='twothirds', caption='T₁−T₂, S₁−S₂: primary minus secondary sensor. Ideal: scatter around zero with ±0.01 spread.', applies_to=<function _has_dual_sensors>, unavailable_if=<function _unavailable_never>), 'sensors_table': Panel(id='sensors_table', render=<function <lambda>>, kind='table', slot='full', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'stability': Panel(id='stability', render=<function <lambda>>, kind='figure', slot='twothirds', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'station_map': Panel(id='station_map', render=<function <lambda>>, kind='figure', slot='two-fifths', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'ts_density': Panel(id='ts_density', render=<function <lambda>>, kind='figure', slot='three-fifths', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'ts_diagram': Panel(id='ts_diagram', render=<function <lambda>>, kind='figure', slot='third', caption='Contours: σ₀ (kg m⁻³) — potential density referenced to surface', applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'ts_updown': Panel(id='ts_updown', render=<function <lambda>>, kind='figure', slot='two-fifths', caption=None, applies_to=<function _always>, unavailable_if=<function _unavailable_never>), 'updown_diff': Panel(id='updown_diff', render=<function <lambda>>, kind='figure', slot='full', caption='ΔCT, ΔSA, Δσ₀ downcast minus upcast on 1-dbar grid — measures hysteresis from pump lag or sensor response time', applies_to=<function _always>, unavailable_if=<function _unavailable_never>)}
Cast panel registry — each wraps an existing
_make_*_b64adapter unchanged, reading only fromPageCtx. Slots mirror the current cast.html layout.
- ctdcast.reports._cast.CAST_SECTION_COLUMNS: dict[str, int] = {'overview': 1}
Presentational intra-section layout, kept out of the layout-neutral manifest model. Maps a section id to the panel index from which the trailing panels stack in a right-hand
fig-col(rather than wrapping onto their own row).overview: 1puts the CT·SA·σ₀ profiles on the left and stacks the station map above the T–S down-vs-up plot in the right column.
- class ctdcast.reports._cast.PageCtx(ds: Any, cfg: ReportConfig, lat: float, lon: float, all_meta: list[dict[str, Any]], ladcp_path: Path | None, ladcp_configured: bool, ladcp_exists: bool, sensor_info: list[dict[str, Any]], nc_path: Path)[source]
Per-cast render context: the frozen inputs every cast panel/predicate reads.
Wraps the frozen
ReportConfig(cfg) with the values derived once per cast, so a panel’srender/applies_todepends only on this object. Keeping it frozen and section-independent is what makes “derived context must not depend on section inclusion” enforceable rather than aspirational.- all_meta: list[dict[str, Any]]
- cfg: ReportConfig
- ds: Any
- ladcp_configured: bool
- ladcp_exists: bool
- ladcp_path: Path | None
- lat: float
- lon: float
- nc_path: Path
- sensor_info: list[dict[str, Any]]
- ctdcast.reports._cast.generate_station_page(nc_path: Path, out_dir: Path, all_meta: list[dict[str, Any]], prev_cast_str: str | None = None, next_cast_str: str | None = None, force: bool = False, ladcp_dir: Path | None = None, ladcp_pattern: str | None = None, cast_num_str: str | None = None, sal_range: tuple[float, float] | None = None, trim_soak: bool = False, cast_notes: list[str] | None = None, cruise_info: dict | None = None, drop_stub: bool = False, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Path | None[source]
Generate a per-cast HTML report page and write it to out_dir/casts/.
- Parameters:
nc_path – Path to a single cast
.ncfile.out_dir – Root output directory.
all_meta – List of dicts with keys
lat,lonfor all casts (used for map).prev_cast_str – Full cast identifier string of the previous cast for nav links, e.g.
"010"or"004b"(or None for no previous link).next_cast_str – Full cast identifier string of the next cast for nav links (or None).
force – Overwrite existing file if True.
ladcp_dir – Directory containing processed LADCP
.matfiles namedNNN.matorNNNb.mat. If None or no matching file exists, LADCP panels are omitted.ladcp_pattern – Optional filename glob for non-standard LADCP naming conventions, e.g.
"msm_142_1_*.mat". The*is replaced with the zero-padded cast number. Falls back to glob-based discovery when omitted.cast_num_str – Full cast identifier string, e.g.
"011"or"004b". Derived from nc_path if not provided.sal_range –
(sal_min, sal_max)— records withctd_salinity_1outside this range are excluded from all plots (but the NC file is not modified). The count of excluded records is shown in the page header.trim_soak – If True, apply pre-soak detection via
find_soak_end(). Finds the last record within 10 dbar of the surface before the cast maximum depth, crawls back up to 20 seconds to the shallowest point preceding the real descent, and trims everything up to that point. Applied before sal_range trimming. NC files are not modified.cast_notes – Optional list of free-text notes for this cast (e.g. “SBE43 malfunction”). Rendered as warning banners near the top of the page.
cruise_info (dict | None) – Optional cruise-level metadata used in the page header.
drop_stub (bool) – If True, omit sections whose figure returned None instead of rendering a stub.
cfg (ReportConfig) – Report configuration (styling, paths, display options) threaded to the plotters.
- Return type:
Path to the written HTML file, or None on failure.
- ctdcast.reports._cast.resolve_cast(ctx: PageCtx, *, drop_stub: bool = False) ResolvedReport[source]
Resolve the cast profile against ctx into numbered, rendered sections.
drop_stub (from the
--drop-stubCLI flag) drops an applicable section whose panels all failed to render, instead of keeping its heading with a stub.
Tier-2: generate a per-section HTML report page.
- ctdcast.reports._section.SECTION_DEFAULT: Profile = Profile(numbering='flat', entries=(Section(id='map', title='Map', panels=('section_map',), level=2, intro='The section track and the CTD stations it comprises.', applies_to=None, role='content', layout=None), Section(id='hydrography', title='Hydrography', panels=(PanelGroup(over=<function <lambda>>, panel=<function _field_panel>),), level=2, intro='Sections of conservative temperature (CT), absolute salinity (SA) and potential density (σ₀) against distance along the section. The σ₀ panel carries the 27.7 and 27.8 kg m⁻³ isopycnals (black, labelled). Open triangles along the top of each panel mark the profiles; station numbers are labelled at intervals.', applies_to=None, role='content', layout=None), Section(id='biogeochemistry', title='Biogeochemistry', panels=(PanelGroup(over=<function _biogeo_present>, panel=<function <lambda>>),), level=2, intro='Sections of the biogeochemical sensors present on this section — oxygen, fluorescence and turbidity where available — against distance along the section.', applies_to=None, role='content', layout=None), Section(id='velocity', title='Velocity (U east, V north)', panels=(PanelGroup(over=<function <lambda>>, panel=<function _ladcp_panel>),), level=2, intro='Eastward (U) and northward (V) velocity from the LADCP, against distance along the section.', applies_to=None, role='content', layout=None), Section(id='ts_diagram', title='T–S diagrams', panels=(PanelGroup(over=<function <lambda>>, panel=<function _ts_panel>),), level=2, intro='Water-mass structure of the section in temperature–salinity space.', applies_to=None, role='content', layout=None)))
The section page profile. Same order as the previous hand-authored page — Map, Hydrography, Biogeochemistry, then the former “extra cards” Velocity and T–S — now numbered and anchored by the resolver. Physics/biogeo are PanelGroups over their variables; Velocity is a PanelGroup over the pre-rendered LADCP panels.
- ctdcast.reports._section.SECTION_PANELS: dict[str, Panel] = {'section_map': Panel(id='section_map', render=<function <lambda>>, kind='figure', slot='half', caption=None, applies_to=<function <lambda>>, unavailable_if=<function _unavailable_never>)}
String-addressable section panels. Only the Map is fixed; field panels (physics/biogeo), LADCP and T–S panels are all data-driven PanelGroups.
- class ctdcast.reports._section.SectionPageCtx(ds_sec: Any, x_vals: Any, x_label: str, section_style: str, bathy_depths: Any, bathy_x: Any, cast_labels: list[int], vmin: dict[str, float], vmax: dict[str, float], section_figsize: tuple[float, float], section_slot_key: str, section_name: str, lats: list[float], lons: list[float], map_b64: str | None, ts_panels: tuple[RenderedPanel, ...], ladcp_panels: tuple[RenderedPanel, ...], cfg: ReportConfig)[source]
Per-section render context: the frozen inputs every section panel reads.
Holds the values derived once per section (selected profiles, x-axis, bathy, colour limits, the computed figure geometry) so a panel’s
renderdepends only on this object.- bathy_depths: Any
- bathy_x: Any
- cast_labels: list[int]
- cfg: ReportConfig
- ds_sec: Any
- ladcp_panels: tuple[RenderedPanel, ...]
- lats: list[float]
- lons: list[float]
- map_b64: str | None
- section_figsize: tuple[float, float]
- section_name: str
- section_slot_key: str
- section_style: str
- ts_panels: tuple[RenderedPanel, ...]
- vmax: dict[str, float]
- vmin: dict[str, float]
- x_label: str
- x_vals: Any
- ctdcast.reports._section.generate_section_page(section_name: str, section_cfg: dict[str, Any], profiles_path: Path, out_dir: Path, force: bool = False, section_style: str = 'pcolormesh', vmin_override: dict[str, float] | None = None, vmax_override: dict[str, float] | None = None, ladcp_dir: Path | None = None, ladcp_pattern: str | None = None, dbar_step: int = 1, prev_name: str | None = None, next_name: str | None = None, cruise_info: dict[str, Any] | None = None, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Path | None[source]
Generate a section HTML report page.
- Parameters:
section_name – Key from
ctd_sections.yaml, e.g."KTout".section_cfg – Dict with keys
description,cast_numbers,color.profiles_path – Path to
profiles.nc(built bycnv_build_profiles.py).out_dir – Root output directory.
force – Overwrite existing file if True.
section_style –
"pcolormesh"or"contourf"— passed through to each section figure.vmin_override – Per-variable colormap limit overrides (e.g.
{"SA": 34.5}).vmax_override – Per-variable colormap limit overrides (e.g.
{"SA": 34.5}).ladcp_dir – Directory containing processed LADCP
.matfiles. If None, the LADCP velocity section panel is omitted.ladcp_pattern – Filename pattern for LADCP files, e.g.
"msm_142_1_*.mat". The*is replaced with the zero-padded cast number. Falls back toNNN.matif not given.dbar_step – Subsample the pressure axis by this step before plotting (default 1, no subsampling).
build_profiles()always stores 1-dbar data; this controls plot-time resolution only.prev_name – Name of the preceding section (for the ← nav button). None omits the button.
next_name – Name of the following section (for the → nav button). None omits the button.
cruise_info – Cruise-level metadata mapping used in the page header;
Noneomits it.cfg – Report configuration (styling, paths, display) threaded to the plotters.
- Return type:
Path to the written HTML file, or None on failure.
- ctdcast.reports._section.resolve_section(ctx: SectionPageCtx) ResolvedReport[source]
Resolve the section profile against ctx into numbered, rendered sections.
Tier-2: per-group timeseries HTML report pages.
A timeseries in this codebase is a named group of repeat casts at one location (yoyo CTDs, e.g. 24 h of repeated profiles). Config is analogous to sections: a name, a description, and a list of cast numbers.
The cruise-wide stacked overview plots (all casts, station-number x-axis) live on index.html, generated by _index.py. They are not timeseries pages.
- ctdcast.reports._timeseries.TIMESERIES_DEFAULT: Profile = Profile(numbering='flat', entries=(Section(id='map', title='Map', panels=('ts_location_map',), level=2, intro='Location of the repeat CTD casts in this group.', applies_to=None, role='content', layout=None), Section(id='hydrography', title='Hydrography', panels=(PanelGroup(over=<function <lambda>>, panel=<function _ts_field_panel>),), level=2, intro='Time series of conservative temperature (CT), absolute salinity (SA) and potential density (σ₀) against cast time. The σ₀ panel carries the 27.7 and 27.8 kg m⁻³ isopycnals (black, labelled). Triangles along the top mark each profile — downward for downcasts, upward for upcasts — with station numbers labelled at intervals above the downcasts.', applies_to=None, role='content', layout=None), Section(id='biogeochemistry', title='Biogeochemistry', panels=(PanelGroup(over=<function _ts_biogeo_present>, panel=<function <lambda>>),), level=2, intro='Time series of the biogeochemical sensors present in this group — oxygen, fluorescence and turbidity where available — against cast time.', applies_to=None, role='content', layout=None), Section(id='velocity', title='Velocity (U east, V north)', panels=(PanelGroup(over=<function <lambda>>, panel=<function _ts_ladcp_panel>),), level=2, intro='Eastward (U) and northward (V) velocity from the LADCP, against hours since the first cast.', applies_to=None, role='content', layout=None), Section(id='ts_diagram', title='T–S diagram', panels=(PanelGroup(over=<function <lambda>>, panel=<function _ts_diagram_panel>),), level=2, intro='Water-mass structure over the occupation in temperature–salinity space, profiles coloured by time.', applies_to=None, role='content', layout=None)))
The timeseries page profile — the same shape and section ids as the section page (so the shared anchors and the legacy #s-* aliases resolve identically), but with time-axis panels and cast-count-driven widths.
- ctdcast.reports._timeseries.TIMESERIES_PANELS: dict[str, Panel] = {'ts_location_map': Panel(id='ts_location_map', render=<function <lambda>>, kind='figure', slot='third', caption=None, applies_to=<function <lambda>>, unavailable_if=<function _unavailable_never>)}
Only the location map is fixed; fields, LADCP and T–S are data-driven groups.
- class ctdcast.reports._timeseries.TimeseriesPageCtx(ds_ts: Any, section_style: str, vmin: dict[str, float], vmax: dict[str, float], ts_figw: float, ts_slot_key: str, fig_location_b64: str | None, ts_panels: tuple[RenderedPanel, ...], ladcp_panels: tuple[RenderedPanel, ...], cfg: ReportConfig)[source]
Per-timeseries render context: the frozen inputs every panel reads.
Holds the values derived once per group (the selected profiles, colour limits, the cast-count-driven figure width and slot, and the pre-rendered location map, LADCP and T–S panels) so a panel’s
renderdepends only on this object.- cfg: ReportConfig
- ds_ts: Any
- fig_location_b64: str | None
- ladcp_panels: tuple[RenderedPanel, ...]
- section_style: str
- ts_figw: float
- ts_panels: tuple[RenderedPanel, ...]
- ts_slot_key: str
- vmax: dict[str, float]
- vmin: dict[str, float]
- ctdcast.reports._timeseries.generate_timeseries_page(ts_name: str, ts_cfg: dict, profiles_path: Path, out_dir: Path, force: bool = False, section_style: str = 'contourf', vmin_override: dict[str, float] | None = None, vmax_override: dict[str, float] | None = None, all_meta: list[dict] | None = None, ladcp_dir: Path | None = None, ladcp_pattern: str | None = None, dbar_step: int = 1, prev_name: str | None = None, next_name: str | None = None, cruise_info: dict[str, Any] | None = None, cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Path | None[source]
Generate a per-timeseries HTML page for a named group of repeat casts.
Plots both downcast and upcast profiles sorted by time_start on a shared time x-axis. Output:
<out_dir>/timeseries/timeseries_<ts_name>.html.- Parameters:
ts_name – Group name (used in filename and page title).
ts_cfg – Dict with
cast_numbers,description, andcolorkeys.profiles_path – Path to
profiles.nc.out_dir – Root output directory.
force – Overwrite existing file if True.
section_style –
"pcolormesh"or"contourf".vmin_override – Per-variable colormap limit overrides.
vmax_override – Per-variable colormap limit overrides.
all_meta – List of per-cast metadata dicts (keys
lat,lon,cast_num); used to render a cruise-context location map. If None, no map is shown.ladcp_dir – Directory containing processed LADCP
.matfiles. If None, the LADCP velocity panel is omitted.ladcp_pattern – Filename pattern for LADCP files, e.g.
"msm_142_1_*.mat". The*is replaced with the zero-padded cast number. Falls back toNNN.matif not given.dbar_step – Subsample the pressure axis by this step before plotting (default 1, no subsampling).
build_profiles()always stores 1-dbar data; this controls plot-time resolution only.prev_name – Name of the preceding timeseries group (for the ← nav button). None omits the button.
next_name – Name of the following timeseries group (for the → nav button). None omits the button.
cruise_info – Cruise-level metadata mapping used in the page header;
Noneomits it.cfg – Report configuration (styling, paths, display) threaded to the plotters.
- Return type:
Path to the written HTML file, or None if skipped or failed.
- ctdcast.reports._timeseries.resolve_timeseries(ctx: TimeseriesPageCtx) ResolvedReport[source]
Resolve the timeseries profile against ctx into numbered, rendered sections.
Interactive map
Tier-2: self-contained interactive cruise map using Leaflet.js.
Leaflet JS/CSS (~160 KB) is bundled in ctdcast/reports/leaflet/ as package
data so the generated leaflet.html requires no internet access at either
generation or view time.
If a GEBCO path is configured, GEBCO bathymetry for the cruise region is rendered as an embedded PNG image layer using discrete depth bands (standard oceanographic levels: 0, 100, 200, 500, 1000, 2000, 3000, 4000, 6000 m).
- Interaction:
Hover cast dot or section line → info panel updates (bottom-left).
Click cast dot → navigate directly to station page.
Click section line → navigate directly to section page.
Scroll or +/− buttons to zoom.
Shift+drag to box-zoom (Leaflet built-in).
- ctdcast.reports._leaflet.generate_leaflet_map(all_meta: list[dict[str, Any]], sections_cfg: dict[str, Any], out_dir: Path, force: bool = False, ship_track_nc: Path | None = None, cruise: str = 'UNK', cfg: ReportConfig = ReportConfig(gebco_path=None, clean_spines=True, profile_figsize=(7.0, 10.0), overview_figsize=(9.0, 4.5), map_lat_min=None, map_lat_max=None, map_lon_min=None, map_lon_max=None, var_cmaps=mappingproxy({'ctd_temperature_1': 'RdYlBu_r', 'ctd_temperature_2': 'RdYlBu_r', 'ctd_temperature': 'RdYlBu_r', 'ctd_salinity_1': 'YlGnBu_r', 'ctd_salinity_2': 'YlGnBu_r', 'ctd_salinity': 'YlGnBu_r', 'ctd_oxygen_1': 'RdYlGn', 'ctd_oxygen_2': 'RdYlGn', 'ctd_oxygen': 'RdYlGn', 'oxygen_saturation': 'RdYlGn', 'ctd_fluor': 'YlGn', 'ctd_turbidity': 'YlOrBr', 'conservative_temperature': 'RdYlBu_r', 'absolute_salinity': 'YlGnBu_r', 'sigma0': 'Purples', 'AOU': 'RdBu_r'}))) Path | None[source]
Generate a Leaflet.js interactive cruise map at
<out_dir>/leaflet.html.Always regenerates (force is accepted for API symmetry but ignored). If
ship_track_ncis provided and the file exists, the ship track is loaded, subsampled, and rendered as a grey polyline behind cast markers.cruiseshould be the resolved cruise identifier (fromcruise_infoor the NC file attribute); it is displayed in the map page title. Returns the output path, or None if all_meta is empty.
Analysis helpers
Derived physical quantities computed from raw CTD measurements.
All functions use GSW — the same library oceanographers use directly.
Per-cast (1-D, dim=time) functions
derive_salinity SP from conductivity/temperature/pressure derive_SA Absolute Salinity from SP/pressure/lat/lon derive_CT Conservative Temperature from SA/in-situ-T/pressure derive_sigma0 Potential density anomaly from SA/CT derive_AOU Apparent Oxygen Utilization from oxygen_saturation (% sat) derive_teos10 Convenience: SA + CT + sigma0 + optional O2 unit conversion
Profiles (2-D, dims N_PROF × pressure) functions
derive_teos10_profiles SA + CT + sigma0 for compiled profiles datasets
Output variable names match the VARIABLES registry in
ctdcast.config.parameters: absolute_salinity,
conservative_temperature, sigma0.
Variable resolution
Functions accept both the canonical CCHDO names (ctd_temperature,
ctd_salinity, ctd_oxygen) and the suffixed dual-sensor names
(ctd_temperature_1, ctd_salinity_1, ctd_oxygen_1). Old
pre-rename names (temperature_1, salinity_1, oxygen_1)
are accepted for backward compatibility with NC files written before
the stage1-normalise rename.
- ctdcast.analysis.derive.derive_AOU(ds: Dataset) Dataset[source]
Return ds with AOU added as 100 - oxygen_saturation (O₂ saturation deficit, % sat).
Note: this is a saturation-deficit proxy, not the traditional AOU in µmol/kg, because it uses
oxygen_saturation(% saturation) rather than dissolved O₂ in µmol/kg.Returns ds unchanged if no oxygen saturation variable is present or
AOUalready exists. Acceptsoxygen_saturation(canonical) oroxsat_1(pre-rename name).- Parameters:
ds – Dataset (any dimensionality) with
oxygen_saturationoroxsat_1in % saturation.- Returns:
New Dataset with
AOUadded; input is not mutated.- Return type:
xr.Dataset
- ctdcast.analysis.derive.derive_CT(ds: Dataset) Dataset[source]
Return ds with Conservative Temperature (CT) added.
Requires
ds["absolute_salinity"]to already be present (callderive_SA()first). Usesgsw.CT_from_twith the first available temperature variable (ctd_temperature,ctd_temperature_1, ortemperature_1) andpressure.- Parameters:
ds – Per-cast Dataset (dim=time) with
absolute_salinity, a temperature variable, andpressure.- Returns:
New Dataset with
ds["conservative_temperature"]added; input is not mutated.- Return type:
xr.Dataset
- ctdcast.analysis.derive.derive_SA(ds: Dataset) Dataset[source]
Return ds with Absolute Salinity (SA) added.
Uses
gsw.SA_from_SPwith the first available salinity variable (ctd_salinity,ctd_salinity_1, orsalinity_1),pressure(dbar), and the cast’s median latitude/longitude.- Parameters:
ds – Per-cast Dataset (dim=time) with a salinity variable,
pressure,latitude,longitude.- Returns:
New Dataset with
ds["absolute_salinity"]added; input is not mutated.- Return type:
xr.Dataset
- ctdcast.analysis.derive.derive_salinity(ds: Dataset) Dataset[source]
Re-compute practical salinity from conductivity, temperature, pressure.
Uses
gsw.SP_from_C, which expects conductivity in mS/cm. Stored conductivity is mS/cm since stage1; a file still carrying S/m (unitsnot mS/cm) is converted by ×10. Call this after any conductivity calibration so that salinity reflects the calibrated conductivity.Does nothing if
conductivity_1or any temperature variable is absent.Writes output to
ctd_salinity_1/ctd_salinity_2(CCHDO canonical names). Records the conversion method in the variable’s attrs.- Parameters:
ds – Per-cast Dataset (dim=time) containing at minimum
conductivity_1, a temperature variable, andpressurein their expected units (conductivity in mS/cm, temperature in °C ITS-90, pressure in dbar).- Returns:
New Dataset with updated
ctd_salinity_1(andctd_salinity_2whenconductivity_2is present); input is not mutated.- Return type:
xr.Dataset
- ctdcast.analysis.derive.derive_sigma0(ds: Dataset) Dataset[source]
Return ds with potential density anomaly (sigma0) added.
Requires
ds["absolute_salinity"]andds["conservative_temperature"]to already be present. Usesgsw.sigma0.- Parameters:
ds – Per-cast Dataset (dim=time) with
absolute_salinityandconservative_temperature.- Returns:
New Dataset with
ds["sigma0"]added; input is not mutated.- Return type:
xr.Dataset
- ctdcast.analysis.derive.derive_teos10(ds: Dataset) Dataset[source]
Return ds with SA, CT, sigma0 added (1-D per-cast Dataset, dim=time).
Convenience function that calls
derive_SA()→derive_CT()→derive_sigma0()in order. Also derivesoxygen_saturation(% sat) from the first available oxygen variable (ctd_oxygen,ctd_oxygen_1, oroxygen_1) when that variable carries molar units.- Parameters:
ds – Per-cast Dataset (dim=time) with a salinity variable, a temperature variable,
pressure,latitude,longitude.- Returns:
New Dataset with SA, CT, sigma0 added; input is not mutated.
- Return type:
xr.Dataset
- ctdcast.analysis.derive.derive_teos10_profiles(ds: Dataset) Dataset[source]
Return ds with SA, CT, sigma0 added (2-D profiles Dataset).
Expects
pressureas a 1-D coordinate and a temperature variable, a salinity variable,latitude,longitudewith dims(N_PROF,)or(N_PROF, pressure). Returns ds unchanged if SA, CT, and sigma0 are already present.- Parameters:
ds – Profiles Dataset (dims N_PROF × pressure).
- Returns:
New Dataset with SA, CT, sigma0 added; input is not mutated.
- Return type:
xr.Dataset
Cast geometry: along-track distance and section orientation.
Pure computation — no matplotlib, no HTML.
- ctdcast.analysis.geometry.along_track_km(lats: list[float], lons: list[float]) tuple[ndarray, str][source]
Return (cumulative_distance_km, x_axis_label) for a list of positions.
- ctdcast.analysis.geometry.distance_from_km(key_lat: float, key_lon: float, lats: list[float], lons: list[float]) ndarray[source]
Return great-circle distance in km from a key position to each position.
Used for
key_castsection ordering: each cast’s x-coordinate is its distance from the chosen key cast. Usesgsw.distanceper pair so the convention matchesalong_track_km(). A position identical to the key yields 0. Agsw.distancefailure is allowed to propagate rather than being silently substituted with a fabricated distance.
- ctdcast.analysis.geometry.section_orientation(lats: list[float], lons: list[float]) bool[source]
Return True if the section x-axis should be flipped for geographic convention.
Convention: west on the left for E–W-dominant sections; north on the left for N–S-dominant sections. Dominance is determined by comparing the end-to-end longitude span against the latitude span.
- Parameters:
lats – Latitude of each cast in the section, in cast order.
lons – Longitude of each cast in the section, in cast order.
- Returns:
True if
x_vals(cumulative along-track distance from first cast) should be replaced byx_total - x_valsbefore plotting.- Return type:
bool
GEBCO bathymetry loading and interpolation.
Pure computation — no matplotlib, no HTML. GEBCO stores elevation as negative
below sea level; this module returns depth as positive below sea level
(depth = -elevation).
- ctdcast.analysis.bathymetry.dense_bathy_along_track(lats: list[float], lons: list[float], x_vals: ndarray, path: Path | None = None, n_per_segment: int = 20) tuple[ndarray | None, ndarray | None][source]
Return
(dense_x, dense_depths)interpolated between cast positions.Generates n_per_segment equally-spaced points along each segment between consecutive casts, giving a smooth GEBCO bathymetry fill rather than the stepped appearance produced by one sample per cast. Returns
(None, None)when GEBCO is unavailable or fewer than two cast positions are supplied.
- ctdcast.analysis.bathymetry.interpolate_bathy_at_casts(lats: list[float], lons: list[float], path: Path | None = None) ndarray | None[source]
Return GEBCO water depth (m, positive below sea level) at each cast position.
Uses bilinear interpolation via xarray. Returns
Noneif GEBCO is not available or on any error. Land points (elevation > 0) are clamped to 0.
- ctdcast.analysis.bathymetry.load_gebco(lat_lo: float, lat_hi: float, lon_lo: float, lon_hi: float, margin: float = 0.05, path: Path | None = None, use_cache: bool = True) tuple[ndarray, ndarray, ndarray] | None[source]
Return a GEBCO subset as (lons, lats, depth_m) or None if unavailable.
If
preload_gebcohas been called for path, subsets from the in-memory numpy cache (fast). Otherwise opens the file from disk (slow).- Parameters:
lat_lo (float) – Southern latitude bound of the region (degrees North).
lat_hi (float) – Northern latitude bound of the region (degrees North).
lon_lo (float) – Western longitude bound of the region (degrees East).
lon_hi (float) – Eastern longitude bound of the region (degrees East).
margin (float, default 0.05) – Extra degrees added on every side, so the fill runs to the panel edge.
path – Path to GEBCO_2025.nc. Pass
cfg.gebco_pathfrom the caller. Returns None if not provided or file not found.use_cache – When
False, read the requested region from the file directly instead of subsetting the preloaded cache. Use this for a region wider than the preloaded cruise extent (e.g. the interactive map, which frames far more longitude than the per-cast maps) — the cache would otherwise clip it.
- ctdcast.analysis.bathymetry.preload_gebco(path: Path, lat_lo: float, lat_hi: float, lon_lo: float, lon_hi: float, margin: float = 1.0) bool[source]
Load a GEBCO region into memory once for the cruise area.
Call this at report-generation start with the full lat/lon extent of all casts. Subsequent
load_gebcocalls then subset from numpy arrays (no disk I/O) instead of reopening the file for every map figure.- Parameters:
path – Path to GEBCO netCDF file.
lat_lo – Bounding box of the cruise area.
lat_hi – Bounding box of the cruise area.
lon_lo – Bounding box of the cruise area.
lon_hi – Bounding box of the cruise area.
margin – Extra degrees around the bounding box. Default 1.0 deg.
- Returns:
True if the file was found and cached successfully.
- Return type:
bool
Cast processing
Stage 1 — CNV-to-netCDF conversion.
Defines the CtdBackend Protocol, concrete backend implementations, and
stage1(), the public function that converts a directory of CNV files
to per-cast netCDF. To add a new backend implement CtdBackend and add a
branch in get_ctd_backend() — nothing else changes.
The converters module re-exports these names for backward compatibility.
- class ctdcast.processors.stage1.CtdBackend(*args, **kwargs)[source]
Protocol for per-cast CNV-to-netCDF converters.
- convert_cast(cnv_path: Path, nc_path: Path, *, force: bool = False, cruise_info: dict | None = None, sensor_overrides: SensorOverrides | None = None) bool[source]
Convert one CNV file to netCDF.
- Parameters:
cnv_path – Path to the raw SBE CNV input file.
nc_path – Desired output netCDF path.
force – If True, overwrite an existing nc_path.
cruise_info – The config
cruise_info:mapping, stamped as cruise identity on the output;Nonewrites none.sensor_overrides – The config
sensors:overrides, used to resolve the per-cast sensor catalog;Noneuses an emptySensorOverrides.
- Returns:
True if the file was written; False if skipped.
- Return type:
bool
- ctdcast.processors.stage1.get_ctd_backend(name: str) CtdBackend[source]
Return a CtdBackend instance for the given backend name.
- Parameters:
name – Currently only
"seasenselib".- Raises:
ValueError – If name is not a recognised backend.
ImportError – If the requested backend’s package is not installed.
- ctdcast.processors.stage1.run(cnv_dir: Path, nc_dir: Path, *, force: bool = False, dry_run: bool = False, cast_tags: set[str] | None = None, **kw: object) int[source]
Run stage1 (CNV → netCDF) for explicit input and output directories.
Called by
ctdcast.processors.process()withstage=1orstage="stage1".- Parameters:
cnv_dir – Directory containing raw SBE CNV files.
nc_dir – The CTD stage root; stage 1 writes
stage1/<stem>_stage1.ncunder it (the stage1 directory is created if absent).force – Overwrite existing NC files.
dry_run – Print what would be converted without writing any files.
cast_tags – If given, process only files whose stem contains one of the zero-padded 3-digit cast numbers (e.g.
{"042", "043"}).**kw – Passed to
stage1()(e.g.backend,pattern).
- Returns:
Number of files written (0 for dry_run).
- Return type:
int
- ctdcast.processors.stage1.stage1(cnv_dir: Path, nc_dir: Path, *, backend: str = 'seasenselib', force: bool = False, cast_filter: int | list[int] | None = None, pattern: str = '*.cnv', cruise_info: dict | None = None, sensor_overrides: SensorOverrides | None = None) int[source]
Convert per-cast CNV files to netCDF using the specified backend.
- Parameters:
cnv_dir – Directory containing raw SBE CNV files.
nc_dir – The CTD stage root; stage 1 writes
stage1/<stem>_stage1.ncunder it (the stage1 directory is created if absent).backend – Backend name (currently only
"seasenselib").force – Overwrite existing netCDF files.
cast_filter – If given, convert only files whose stem contains the zero-padded cast number. Accepts a single int or a list of ints for multi-cast filtering.
pattern – Filename glob pattern applied within
cnv_dir(default:"*.cnv").cruise_info – The config
cruise_info:mapping, stamped as cruise identity on each per-cast file;Nonewrites none (the default, so bare conversions are unchanged).sensor_overrides – The config
sensors:overrides, used to resolve the per-cast sensor catalog;Noneuses an emptySensorOverrides.
- Returns:
Number of files written (skipped files not counted).
- Return type:
int
- Raises:
ImportError – If the chosen backend’s package is not installed.
Stage 2 — trim.
Downcast/upcast splitting and soak / back-on-deck detection. This is processing, not analysis: it decides which scans belong to the real cast.
apply_stage2() is the pipeline entry point: it sets QARTOD flag 4 on soak
and post-recovery deck records and records the parameters used in
ds.attrs["history"]. find_soak_end and find_cast_end are kept public
because reports._cast calls them directly for report-time plot trimming (that
coupling will be removed in a later phase when reports read flags from the files).
The turnaround convention (last pressure within 2 dbar of the maximum) and the soak/deck algorithms are deliberate — see the individual docstrings.
- ctdcast.processors.stage2.apply_clock_offset(ds: Dataset, offset_seconds: float, *, n_casts: int | None = None, segment_sd: float | None = None) Dataset[source]
Shift a cast’s
timecoordinate by a clock offset, preserving the uncorrected time.A clock correction is a uniform translation of the time axis — every channel shifted identically, no measured value recomputed — which is why it belongs at stage 2 and is safe on already-binned data at the current sampling.
offset_secondsisNMEA - System(the seconds added to acquisition time to obtain corrected time). The uncorrected time is kept as thetime_origauxiliary coordinate, whose presence also marks the file as corrected; andtime_coverage_*are recomputed so the file does not contradict its own axis.n_castsandsegment_sdare the offset’s evidence — the count and sd of the population that produced it — supplied from the config that recorded the offset. Whichever is given is stamped ontoclock_offset_secondsso the uncertainty travels with the number; when both are absent (a hand-written config), no statistics are recorded rather than a fabricatedsd 0.00.Refuses (raises
ValueError) when the file already carriestime_orig(already corrected — re-run from stage 1), when itstime_coordinate_sourceclassifies as GPS/NMEA (the coordinate is already on GPS, so a correction would introduce an error), or when that source cannot be classified as the System clock (unknown or absent — refuse rather than assume). The clock is classified withctdcast.config.cnv_header.start_time_clock(), not a substring.
- ctdcast.processors.stage2.apply_curated_drop(ds: Dataset, drop_names: list[str] | None = None) Dataset[source]
Drop the curated set of SBE-derived channels stage 1 kept, recording what went.
Stage 1 is a faithful translation and keeps every CNV column (SeaBird-computed quantities under an
sbe_prefix); stage 2 is curated and removes them by default so the working product is not cluttered with duplicates of quantities ctdcast recomputes. The stage-1 file remains the record of what existed, so nothing is lost.- Parameters:
ds – The stage-2 Dataset (post soak/deck flags and clock correction).
drop_names – Explicit names to drop (the config
trim.drop_sbe:list).Noneuses the default: everysbe_*channel present (the known likely-dead-weight set). An unrecognised channel is kept — name it intrim.drop_sbe:to drop it. A name not present is ignored.
- Returns:
A new Dataset without the dropped channels; unchanged when none are present. The dropped names are recorded in
historyand in adropped_channelsattribute, so the file states what it no longer contains.- Return type:
xr.Dataset
- ctdcast.processors.stage2.apply_stage2(ds: Dataset, *, near_surface_dbar: float = 10.0, search_seconds: float = 20.0, deck_window_seconds: float = 20.0, margin_dbar: float = 0.5, max_deck_dbar: float = 20.0) Dataset[source]
Apply QARTOD flag 4 to soak and post-recovery deck records.
Creates
{var}_qcarrays (int8, flag 1=pass) for each physical data variable, then sets flag 4 (fail) on the pre-descent soak window and the post-recovery on-deck window. Records the parameters used inds.attrs["history"]so the treatment is reproducible from the output file alone.- Parameters:
ds – Per-cast Dataset (dim=time).
near_surface_dbar – Passed to
find_soak_end— pressure threshold for last near-surface crossing before the real descent.search_seconds – Passed to
find_soak_end— backward-crawl window width.deck_window_seconds – Passed to
find_cast_end— tail window for on-deck reference pressure.margin_dbar – Passed to
find_cast_end— added to on-deck median to form the cut.max_deck_dbar – Passed to
find_cast_end— if on-deck median exceeds this, no trim.
- Returns:
New Dataset; input is not mutated.
- Return type:
xr.Dataset
- ctdcast.processors.stage2.find_cast_end(pressure: ndarray, times: ndarray, deck_window_seconds: float = 20.0, margin_dbar: float = 0.5, max_deck_dbar: float = 20.0) int[source]
Return the exclusive end index, trimming post-recovery deck records.
Algorithm:
Take the median pressure of the last deck_window_seconds seconds as the on-deck reference pressure. Using a median handles sensor offset (the pressure sensor may not read exactly 0 dbar when the CTD is in the air) and is robust to brief oscillations on deck.
Find the first index after the pressure maximum where pressure falls at or below
p_deck_median + margin_dbar. Trim from that index onward.
Returns
len(pressure)(no trim) if:the record is empty or the CTD never returned near the surface (
p_deck_median > max_deck_dbar), orpressure never drops to the threshold on the upcast.
- Parameters:
pressure – Pressure array in dbar.
times – Time coordinate array (
numpy.datetime64or numeric seconds).deck_window_seconds – Duration of the tail window used to estimate on-deck pressure.
margin_dbar – Added to the on-deck median to form the cut threshold.
max_deck_dbar – If the on-deck median exceeds this value the CTD is considered not to have returned to the surface and no trim is applied.
- Returns:
Exclusive end index; slice with
ds.isel(time=slice(None, idx)).- Return type:
int
- ctdcast.processors.stage2.find_soak_end(pressure: ndarray, times: ndarray, near_surface_dbar: float = 10.0, search_seconds: float = 20.0) int[source]
Return the index at which the real downcast begins (exclusive end of soak).
Algorithm (three steps):
Find
i_max, the index of the global pressure maximum (deepest point of the cast). Searching only inpressure[0:i_max+1]keeps the upcast recovery — when the CTD returns to the surface at the end of the cast — from being confused with the pre-soak position.Within
pressure[0:i_max+1], find the last index wherepressure < near_surface_dbar. For a typical MSM-style cast this falls on the early real descent, just as the CTD passesnear_surface_dbargoing downward.Crawl backward from that index within
search_secondsto find the minimum pressure — the shallowest point (closest to the surface) just before the real descent began. Return the index immediately after that minimum as the start of the real downcast.
In bad-weather conditions where the CTD soaks at depth and is never raised back to the surface, step 3 finds the minimum within the soak window and removes only the first
search_secondsof the soak. The operator- visible effect is a truncation of the pre-soak data, not a clean removal.- Parameters:
pressure – Pressure array in dbar (1-D, same length as times).
times – Time coordinate array. May be
numpy.datetime64or numeric seconds; elapsed time is computed relative totimes[0].near_surface_dbar – Pressure threshold used to find the last near-surface crossing before the main descent. Default is 10 dbar (≈10 m), safely below the typical soak depth of 8–10 m.
search_seconds – Width of the backward-crawl window (seconds) used to find the pre-descent surface minimum. Default is 20 s.
- Returns:
Index of the first record to keep; slice with
ds.isel(time=slice(idx, None)). Returns 0 if the cast never reaches belownear_surface_dbar(no trim applied).- Return type:
int
- ctdcast.processors.stage2.run(root: Path, *, force: bool = False, dry_run: bool = False, cast_tags: set[str] | None = None, cruise_cfg: dict | None = None, **kw: object) int[source]
Apply stage 2 (soak/deck flagging) across the casts under root.
Reads each cast’s stage-1 file, applies
apply_stage2(), and writes a newstage2/<stem>_stage2.nc— never in place, so stage 1 stays frozen and the run is re-runnable. Reads are strict: a cast with no stage-1 file is skipped with a warning, not silently promoted from a lower stage. Called byctdcast.processors.process()withstage=2.- Parameters:
root – The CTD stage root (
ctd_root). Stage files live underroot/stage1/…root/stage2/; an old flatnc_diris read as stage 1 via the compatibility shim inctdcast.processors.stage_layout.force – Rewrite the stage-2 file even when it already exists. Without
force, a cast whosestage2/<stem>_stage2.ncexists is skipped.dry_run – Print what would be processed without writing any output.
cast_tags – If given, process only casts selected by these zero-padded tags (e.g.
{"042"}), matched on the parsed cast identity.cruise_cfg – The
processing:config block, read for the stage-2 curated drop (trim.drop_sbe) and the clock correction;Noneuses defaults.**kw – Tuning forwarded to
apply_stage2()(e.g.near_surface_dbar); keys it does not accept are ignored.
- Returns:
Number of stage-2 files written (0 for dry_run).
- Return type:
int
- Raises:
FileNotFoundError – If root does not exist or is not a directory.
- ctdcast.processors.stage2.split_cast(ds: Dataset) tuple[Dataset, Dataset][source]
Split ds (individual cast file, dim=time) into (downcast, upcast).
Uses the turnaround convention: last index where pressure is within 2 dbar of its maximum.
Stage 3 — QC, calibration, and derived-variable orchestrator.
Stage3 is iterative: re-run it as calibration improves. Each run applies gross-range QC, then any conductivity calibration present in the cruise config, then re-derives salinity from the calibrated conductivity.
Sea-Bird processing-chain calibration (hex-level, frequency coefficients) is Phase 5 scope; this module covers only the post-conversion treatment.
- ctdcast.processors.stage3.run(root: Path, *, force: bool = False, dry_run: bool = False, cast_tags: set[str] | None = None, cruise_info: dict | None = None, **kw: object) int[source]
Apply stage 3 (QC + calibration) across the casts under root.
Reads each cast’s stage-2 file, applies
stage3(), and writes a newstage3/<stem>_stage3.nc— never in place. Reads are strict: a cast with no stage-2 file is skipped with a warning (its soak/deck flags were never applied, so QC’ing it would manufacture a product whose filename lies). Called byctdcast.processors.process()withstage=3.- Parameters:
root – The CTD stage root (
ctd_root). Stage files live underroot/stage2/…root/stage3/.force – Rewrite the stage-3 file even when it already exists. Without
force, a cast whosestage3/<stem>_stage3.ncexists is skipped.dry_run – Print what would be processed without writing any output.
cast_tags – If given, process only casts selected by these zero-padded tags (e.g.
{"042"}), matched on the parsed cast identity.cruise_info – The config
cruise_info:block. When it declaresdata_mode: D, each stage-3 file is stampeddata_mode = "D"; if no calibration was actually applied the delayed-mode claim is unsupported, so the cast is named in a warning and the gap recorded inhistory.**kw – Passed to
stage3()(e.g.cruise_cfg).
- Returns:
Number of stage-3 files written (0 for dry_run).
- Return type:
int
- Raises:
FileNotFoundError – If root does not exist or is not a directory.
- ctdcast.processors.stage3.stage3(ds: Dataset, cruise_cfg: dict | None = None) Dataset[source]
Apply QC and calibration to a per-cast Dataset.
Applies the following in order:
Two-tier gross-range QC (
qc.apply_gross_range) and the QARTOD spike test (qc.apply_spike_test), using the package defaults merged with any overrides fromcruise_cfg["qc"]["gross_range"]andcruise_cfg["qc"]["spike"](each withsuspect/failsub-dicts).Conductivity calibration slope, if
cruise_cfg["calibration"]["conductivity_slope"]is present. Multipliesconductivity_1(andconductivity_2if present) by the slope and records it in the variable’s attributes.Re-derives
salinity_1(andsalinity_2) from the calibrated conductivity usingderive_salinity()— only when a conductivity calibration was applied.
- Parameters:
ds – Per-cast Dataset (dim=time), already through stage1 and stage2.
cruise_cfg – Optional dict with sub-keys
qcandcalibration. Passcfg.get("processing")or the full cruise config dict.
- Returns:
New Dataset; input is not mutated.
- Return type:
xr.Dataset
Stage QC — two-tier gross-range and spike flagging.
Each test has a suspect tier (QARTOD flag 3) and a fail tier (flag 4); the wider
fail bound wins, and a more-severe flag is never downgraded (worst-flag-wins).
Non-finite samples are marked missing (flag 9), not pass. The thresholds applied
are recorded as attributes on each {var}_qc companion so the treatment
reconstructs from the file alone. Operates on per-cast Datasets (dim=time); call
after apply_stage2 so the soak/deck flags are already present.
Note: config overrides are trusted, not validated — a suspect range set wider than its fail range is accepted as given (ioos_qc would reject it).
- ctdcast.processors.qc.GROSS_RANGE_FAIL: dict[str, tuple[float, float]] = {'conductivity_1': (0.0, 75.0), 'conductivity_2': (0.0, 75.0), 'ctd_fluor': (0.0, 50.0), 'ctd_oxygen': (0.0, 450.0), 'ctd_oxygen_1': (0.0, 450.0), 'ctd_oxygen_2': (0.0, 450.0), 'ctd_salinity': (0.0, 40.0), 'ctd_salinity_1': (0.0, 40.0), 'ctd_salinity_2': (0.0, 40.0), 'ctd_temperature': (-2.5, 40.0), 'ctd_temperature_1': (-2.5, 40.0), 'ctd_temperature_2': (-2.5, 40.0), 'ctd_turbidity': (0.0, 50.0), 'fluorescence': (0.0, 50.0), 'oxsat_1': (0.0, 200.0), 'oxygen_1': (0.0, 450.0), 'oxygen_saturation': (0.0, 200.0), 'pressure': (-5.0, 7000.0), 'salinity_1': (0.0, 40.0), 'salinity_2': (0.0, 40.0), 'temperature_1': (-2.5, 40.0), 'temperature_2': (-2.5, 40.0), 'turbidity': (0.0, 50.0)}
outside is instrument malfunction.
- Type:
Gross-range FAIL bounds (flag 4)
- ctdcast.processors.qc.GROSS_RANGE_SUSPECT: dict[str, tuple[float, float]] = {'conductivity_1': (0.0, 65.0), 'conductivity_2': (0.0, 65.0), 'ctd_salinity': (2.0, 40.0), 'ctd_salinity_1': (2.0, 40.0), 'ctd_salinity_2': (2.0, 40.0), 'ctd_temperature': (-2.0, 35.0), 'ctd_temperature_1': (-2.0, 35.0), 'ctd_temperature_2': (-2.0, 35.0), 'pressure': (-0.5, 7000.0), 'salinity_1': (2.0, 40.0), 'salinity_2': (2.0, 40.0), 'temperature_1': (-2.0, 35.0), 'temperature_2': (-2.0, 35.0)}
outside is oceanographically implausible.
- Type:
Gross-range SUSPECT bounds (flag 3)
- ctdcast.processors.qc.QARTOD_SUSPECT = np.int8(3)
QARTOD primary flag values (IOOS QARTOD). The complete vocabulary — every value and its meaning — is encoded in
_qc_attrs(); these name the two flags the ctdcast pipeline sets, so the code that writes a flag, the code that masks on it, and the vocabulary that gives it meaning share one definition.
- ctdcast.processors.qc.SPIKE_FAIL: dict[str, float] = {'conductivity_1': 5.0, 'conductivity_2': 5.0, 'ctd_salinity': 2.0, 'ctd_salinity_1': 2.0, 'ctd_salinity_2': 2.0, 'ctd_temperature': 6.0, 'ctd_temperature_1': 6.0, 'ctd_temperature_2': 6.0, 'pressure': 50.0, 'salinity_1': 2.0, 'salinity_2': 2.0, 'temperature_1': 6.0, 'temperature_2': 6.0}
Spike FAIL thresholds (flag 4). Oxygen has no fail default yet (suspect only).
- ctdcast.processors.qc.SPIKE_SUSPECT: dict[str, float] = {'conductivity_1': 2.0, 'conductivity_2': 2.0, 'ctd_oxygen': 20.0, 'ctd_oxygen_1': 20.0, 'ctd_oxygen_2': 20.0, 'ctd_salinity': 1.0, 'ctd_salinity_1': 1.0, 'ctd_salinity_2': 1.0, 'ctd_temperature': 2.0, 'ctd_temperature_1': 2.0, 'ctd_temperature_2': 2.0, 'oxygen_1': 20.0, 'pressure': 10.0, 'salinity_1': 1.0, 'salinity_2': 1.0, 'temperature_1': 2.0, 'temperature_2': 2.0}
Spike SUSPECT thresholds (flag 3) on
|v[i] - (v[i-1]+v[i+1])/2|, in the variable’s units. Fluorescence and turbidity are excluded — natural fine-scale variability, not instrument spikes.
- ctdcast.processors.qc.apply_gross_range(ds: Dataset, thresholds: dict | None = None) Dataset[source]
Two-tier gross-range QC: flag 3 outside suspect bounds, flag 4 outside fail.
Creates
{var}_qccompanions (int8, 1=pass) as needed, sets suspect (3) then fail (4) so the wider fail bound wins, and records the applied bounds (qc_gross_range_{suspect,fail}_{min,max}) on each companion so the report reads them back. A more-severe flag is never downgraded.- Parameters:
ds – Per-cast Dataset (dim=time); input is not mutated.
thresholds – Config overrides with
suspectand/orfailsub-dicts of{var: (min, max)}, merged overGROSS_RANGE_SUSPECT/GROSS_RANGE_FAIL.
- Returns:
New Dataset with the flags and threshold attributes.
- Return type:
xr.Dataset
- ctdcast.processors.qc.apply_spike_test(ds: Dataset, thresholds: dict | None = None) Dataset[source]
Two-tier QARTOD spike test: flag 3 above the suspect threshold, flag 4 above fail.
The spike metric for an interior sample is
|v[i] - (v[i-1] + v[i+1]) / 2|; endpoints are not evaluated. Records the thresholds (qc_spike_{suspect,fail}_threshold) on each{var}_qc. Fluorescence and turbidity have no spike test (natural fine-scale variability). A more-severe flag is never downgraded.- Parameters:
ds – Per-cast Dataset (dim=time); input is not mutated.
thresholds – Config overrides with
suspectand/orfailsub-dicts of{var: threshold}, merged overSPIKE_SUSPECT/SPIKE_FAIL.
- Returns:
New Dataset with the flags and threshold attributes.
- Return type:
xr.Dataset
Cruise-level profile compiler: per-cast netCDF → profiles.nc.
Reads all per-cast netCDF files in a directory, splits each into downcast and
upcast halves, bins to a common 1-dbar grid, and writes a single
(N_PROF × pressure) netCDF. The converters module re-exports
build_profiles for backward compatibility.
- ctdcast.processors.profiles.build_profiles(nc_dir: Path, profiles_path: Path, *, force: bool = False, gebco_path: Path | None = None, dbar: int = 1, cruise_info: dict | None = None, refuse_catalog_less: bool = False) bool[source]
Compile per-cast netCDF files into a single profiles.nc on a dbar-spaced grid.
Reads all
*.ncfiles in nc_dir, splits each cast into downcast and upcast halves, bins to a common dbar-dbar pressure grid (default 1 dbar), and writes a single (N_PROF × pressure) netCDF. N_PROF is a plain integer index (0, 1, 2, …); cast identity is carried bycast_number,cast_suffix, andcast_directionvariables. The bin spacing is recorded in thepressure_spacing_dbarglobal attribute so the gridding can be reconstructed from the output file alone.Per-cast scalar variables added to the output:
max_pressure_dbar— maximum pressure recorded over the full cast.gebco_depth_m— GEBCO bathymetry depth (m, positive down) at the max-pressure lat/lon position; NaN when gebco_path is None or the file is unavailable.
The
altimeterchannel (when present in the input files) is binned onto the 1-dbar grid as a standard 2-D variable.Samples carrying a QARTOD suspect (3) or fail (4) flag on their
{var}_qccompanion are NaN-masked before binning, so flagged data does not enter the bin means. Each science variable records how many finite input samples it carried (qc_input_samples) and how many were excluded (qc_excluded_samples).- Parameters:
nc_dir – Directory containing per-cast netCDF files.
profiles_path – Output path for the compiled profiles netCDF.
force – Overwrite an existing profiles_path.
gebco_path – Path to a GEBCO_2025.nc file. Used to look up water depth at each cast’s max-pressure position. Pass
cfg.gebco_pathwhen calling from report generation code. Silently omitted when None.dbar – Vertical bin spacing (dbar) of the output grid. Default 1. Use 2 (or more) to average adjacent pressure levels together, reducing per-level noise when the raw scan resolution does not justify a 1-dbar grid.
cruise_info – The
cruise_info:mapping from the cruise config. Supplies the cruise id (so the compiled file is not labelledUNK), the authored discovery fields, people, embargo, and theship/start_datefrom which the EXPOCODE coordinate is derived. Coverage bounds and creation time are computed from the data, not taken from here.refuse_catalog_less – Policy for a CTD cast whose per-cast file carries no
SENSOR_*catalog (it predates sensor provenance). The defaultFalseis the library behaviour: warn once, name the files, and stamp the admission into the output so the gap is visible in the file itself. The pipeline driver (ctdcast process --stage profilesandctdcast run) passesTrueso the compile refuses instead — at the point a product is shipped, re-running stage 1 to build the catalog is the actionable fix.
- Returns:
True if profiles.nc was written; False if skipped (existed, force=False).
- Return type:
bool
- Raises:
ValueError – If no recognised cast files are found in nc_dir, or if refuse_catalog_less is True and any cast carries no sensor catalog.
- ctdcast.processors.profiles.run(nc_dir: Path, profiles_path: Path, *, force: bool = False, dry_run: bool = False, **kw: object) bool[source]
Build
profiles.ncfrom NC files in nc_dir.Called by
ctdcast.processors.process()withstage="profiles".- Parameters:
nc_dir – Directory of per-cast netCDF files.
profiles_path – Output path for the compiled profiles netCDF.
force – Overwrite an existing profiles.nc.
dry_run – Print what would be built without writing any output.
**kw – Passed to
build_profiles()(e.g.gebco_path).
- Returns:
True if profiles.nc was written; False if skipped (or dry_run).
- Return type:
bool
Stamped history notes for the processing stages.
Every stage records what it did — and the parameters it used — as a line in the
CF history global attribute, so the treatment is reconstructable from the
output file alone. Because each stage reads its predecessor and copies its
attributes, history accumulates: stage 3 inherits the stage 1 and stage 2
lines and appends its own.
The format, from one helper so it stays consistent, is:
<timestamp> <producer> <version> <stage>: <note>
joined with newlines. ctdcast’s own stages use the defaults — an ISO-8601 UTC
timestamp, producer ctdcast — giving <ISO-8601 UTC> ctdcast <version> <stage>.
An upstream Sea-Bird step overrides producer ("SBE Data Processing"), timestamp
(the tool’s own verbatim stamp, no Z) and the stage field (the module name, e.g.
celltm), so a stage-1 file’s history shows Sea-Bird’s own lines alongside ctdcast’s.
- ctdcast.processors.history.PL_CONVERTED = 'Instrument data that has been converted to geophysical values'
OceanSITES reference-table-3
processing_levelvalues, verbatim. A variable can carry several at once (converted, then calibrated, then flagged): they are joined with"; "in the order applied byadd_processing_level(), so the attribute doubles as the per-variable procedure sequence. Comma is unusable as a separator —RANGES_FLAGGEDcontains one.
- ctdcast.processors.history.add_processing_level(attrs: MutableMapping[str, object], value: str) None[source]
Append a table-3
processing_levelvalue to attrs, idempotently.The attribute is an ordered set joined with
"; ": value is appended only if it is not already present, so re-running a stage — or two stages that apply the same procedure (stage-2 soak/deck and stage-3 gross-range both flag) — never accumulates a duplicate. Order of first application is preserved, so the attribute reads as the procedure sequence.- Parameters:
attrs – A variable’s attribute mapping (
ds[var].attrs), mutated in place.value – One of the
PL_*table-3 strings.
- ctdcast.processors.history.append_history(attrs: MutableMapping[str, object], note: str, *, stage: str, version: str = '0.2.1', producer: str = 'ctdcast', timestamp: str | None = None, prepend: bool = False) None[source]
Append a stamped
historyline to attrs in place.- Parameters:
attrs – The attribute mapping to extend — an
xarrayds.attrsor the plaindicta builder assembles before constructing its Dataset. Mutated in place: the existinghistory(if any) is preserved and the new line is appended after a newline.note – The human-readable record of the operation, including the parameters it used (e.g.
"gross_range: ctd_salinity_1:[2.0,42.0]").stage – The stage token that did the work —
"stage1","stage2","stage3","profiles","ladcp_profiles"for ctdcast’s own stages, or a Sea-Bird module name ("celltm","binavg", …) for an upstream step.version – The version to stamp; defaults to the installed ctdcast package version.
producer – Who did the work —
"ctdcast"by default. An upstream Sea-Bird step passes"SBE Data Processing"so the line reads as Sea-Bird’s, not ctdcast’s.timestamp – The stamp to use.
None(the default) uses the current ISO-8601 UTC time. A Sea-Bird step passes its own verbatim timestamp, kept as-is — noZis appended, because it comes from a processing workstation whose offset is unknown and a different clock again from the deck unit.prepend – Insert the line at the front of
historyinstead of the end. Used for an upstream step (a Sea-Bird module) that predates the reader’s own line, so the record stays oldest-first; the default appends.
Readers
Reader for LDEO IXv14 LADCP .mat files.
Locates the .mat file for a cast (find_ladcp_file()), loads it with a
single set of scipy.io.loadmat options (read_ladcp()), and maps the
result struct to a single-cast xarray.Dataset on the native 10 m depth
grid (read_ladcp_cast()). The .mat is the LDEO IX velocity solution;
its ~50 fields are mapped to the compiled-dataset schema.
- ctdcast.readers.ladcp.MatField
A field read from a
scipy.io.loadmatmat-struct — an ndarray, amat_struct, or a scalar, depending on how MATLAB stored it. Named once here so the struct accessors below readfield: MatFieldwith no per-signature noqa, the explanation travels into the docs, and a future narrowing happens in one place.
- ctdcast.readers.ladcp.find_ladcp_file(ladcp_dir: Path, cast_num: int, cast_suffix: str = '', ladcp_pattern: str | None = None) Path | None[source]
Return the .mat file for cast_num in ladcp_dir, or
Noneif absent.If ladcp_pattern is given (e.g.
"msm_142_1_*.mat"), the*wildcard is replaced with the zero-padded cast number (and optional suffix) and that name is tried first. Falls back to standard names (NNN.mat,NNNb.mat) then a*_NNN.matglob for cruise-prefixed filenames. The first glob match (lexicographic) is returned when multiple files match.
- ctdcast.readers.ladcp.read_ladcp(path: Path | str) dict[str, Any][source]
Load an LDEO IXv14 LADCP
.matfile.Uses
squeeze_me=Trueandstruct_as_record=Falseso the LADCP result struct is reachable asread_ladcp(path)["dr"]with attribute access.
- ctdcast.readers.ladcp.read_ladcp_cast(path: Path | str, *, cast_num: int, cast_suffix: str = '') Dataset[source]
Map an LDEO LADCP
.matto a single-cast Dataset on the native 10 m grid.Returns a Dataset with dimension
depth(uniform 10 m, positive down), an auxiliarypressure(depth)coordinate (dbar, the bridge toprofiles.nc), the inverse (u/v) and shear (u_shear/v_shear/w_shear) velocity solutions, per-instrument down/up-looker profiles, per-cast scalars (barotropic, bottom-track, position,instrument_config), and the LADCP processing provenance as global attributes. Velocity carries the CFeastward/northward/error_sea_water_velocitystandard names.
Reader for per-cast sensor metadata written by seasenselib.
Parses the raw_metadata global attribute (a JSON blob) into a list of sensor
descriptors for the cast page.
- ctdcast.readers.metadata.parse_sensor_channels(ds: Dataset) list[dict][source]
Return one full descriptor per sensor channel in ds, from the raw header.
A thin reader over
ctdcast.config.cnv_header.parse_sensor_block()(the single parser of the<Sensors>block, authoritative because seasenselib’s per-channelcnv_sensor_Ndicts are lossy for some channels): it resolves the block comment to a canonicalroleand normalises thecalibration_date.Free(unused) channels getrole = Noneand an empty serial.Returns
[]ifraw_metadataor the header sensor block is absent. Each dict haschannel,element,sensor_id,serial,calibration_date(normalised),slope,offsetandrole(canonical role orNone).
- ctdcast.readers.metadata.parse_sensor_info(ds: Dataset) list[dict[str, str]][source]
Extract sensor serial numbers and calibration dates from ds.
Parses the
raw_metadataglobal attribute (a JSON string written by seasenselib) and returns one entry per sensor channel that has both asensor_typeand aserial_number.- Parameters:
ds – Per-cast Dataset as opened from a netCDF file.
- Returns:
Each dict has keys
sensor_type(human-readable label),serial_number, andcalibration_date. Returns[]ifraw_metadatais absent, unparseable, or contains no usable sensors.- Return type:
list[dict[str, str]]
- ctdcast.readers.metadata.source_to_canonical(ds: Dataset, *, lower_keys: bool = False) dict[str, str][source]
Reconstruct a reader’s
{source_name: canonical_name}rename table from ds.Reads each coordinate and data variable’s recorded source name (
_SOURCE_NAME_ATTRS) and maps it to the variable’s canonical name, keeping only variables whose source name actually differs. This is the authoritative, per-cast rename the reader applied — more complete than the staticCNV_ALIASES, which misses unit spellings such asc0mS/cm. lower_keys lower-cases the keys, for case-insensitive lookup against header channel names.