OceanArray Project Structure
This document provides an overview of the oceanarray codebase structure and organisation.
Project Structure Overview
oceanarray/
├── oceanarray/ # Main Python package
│ ├── __init__.py # Public API: process(), STAGES, parameters
│ ├── _version.py # Package version (auto-generated by setuptools-scm)
│ ├── cli.py # [core] Command-line interface (oceanarray init/process/run/report/validate/…)
│ ├── utilities.py # [core] Shared utility helpers used across the pipeline
│ ├── inspect.py # [core] Quick dataset inspection helpers for interactive use
│ ├── logger.py # [core] Logging configuration and helpers
│ ├── paths.py # [core] Filesystem path and filename conventions
│ │
│ ├── processors/ # [core] Processing stages + pipeline registry
│ │ ├── __init__.py # Pipeline stage registry and public process() entry point
│ │ ├── stage1.py # Stage 1: raw data → CF-NetCDF (*_stage1.nc)
│ │ ├── stage2.py # Stage 2: clock correction + deployment trim (*_stage2.nc)
│ │ ├── stage3.py # Stage 3: pressure interpolation + QARTOD QC (*_stage3.nc)
│ │ ├── pressure.py # Pressure interpolation helpers for stage 3
│ │ ├── qc.py # QARTOD QC tests and CTD derivations for stage 3
│ │ ├── coordinate.py # BEAM→ENU transforms and ADCP QC for stage 3
│ │ ├── caldip.py # CalDip calibration-dip corrections (not yet implemented)
│ │ ├── stack.py # MooringStacker: instruments → common time grid (*_stack.nc)
│ │ ├── grid.py # MooringGridder: stack → regular pressure grid (*_grid.nc)
│ │ └── helpers.py # Internal helpers for stack and grid operations
│ │
│ ├── analysis/ # [analysis] Science utilities
│ │ ├── clock.py # Clock-offset analysis functions
│ │ ├── hydrographic.py # Hydrographic analysis utilities
│ │ ├── science.py # QC and dataset processing functions
│ │ ├── spectral.py # Spectral analysis of time series
│ │ ├── temporal.py # Time-series analysis utilities
│ │ └── vector.py # Vector rotation and progressive-vector utilities
│ │
│ ├── plotters/ # [viz] Three-tier plotting package
│ │ ├── __init__.py # Public API
│ │ ├── primitives.py # Tier 1: data-agnostic plotting primitives
│ │ ├── current.py # Tier 2: current/velocity domain plots
│ │ ├── diagnostic.py # Tier 2: T-S, histogram, spectrum, QC diagnostic plots
│ │ ├── hydrography.py # Tier 2: hydrographic section and isopycnal plots
│ │ ├── timeseries.py # Tier 2: time-series and gridded-section plots
│ │ ├── spectrum.py # Tier 2: spectral diagnostics
│ │ ├── ts.py # T-S diagram and thermohaline structure plots
│ │ ├── animation.py # Animated plot functions
│ │ ├── helpers.py # Shared colormap/style helpers
│ │ └── _cli_legacy.py # Legacy CLI plots (pending `oceanarray plot <file>` redesign)
│ │
│ ├── reports/ # [report] HTML report generation
│ │ ├── __init__.py # Public entry-points (MooringReport, etc.)
│ │ ├── _mooring.py # Mooring summary template + MooringReport orchestrator
│ │ ├── _instrument.py # Per-instrument report ({mooring}_{serial}_report.html)
│ │ ├── _stack.py # Stack report (multi-instrument timeseries, tilt panels)
│ │ ├── _grid.py # Grid report (vertical section, spectra, N²)
│ │ ├── _array.py # Array-level multi-mooring summary report
│ │ ├── _plots.py # Tier 3: report-level figure wrappers (base64 PNGs)
│ │ ├── _html_helpers.py # HTML/QC constants, base64 helpers, NC metadata readers
│ │ ├── _recovery_table.py # Per-mooring cruise-report recovery table
│ │ └── _pdf.py # combine_mooring_pdf: HTML pages → single A4 PDF (WeasyPrint)
│ │
│ ├── tools/ # [core] Shared I/O infrastructure
│ │ ├── readers.py # NetCDF and legacy-format instrument readers
│ │ ├── writers.py # NetCDF write helpers (OceanSITES-compliant)
│ │ └── rapid_interp.py # Physics-informed vertical interpolation for sparse profiles
│ │
│ ├── config/ # Configuration
│ │ ├── parameters.py # Package-level defaults (VARIABLES, QC ranges, style)
│ │ ├── validation.py # Mooring YAML validation
│ │ ├── OS1_var_names.yaml # OceanSITES variable name mappings
│ │ ├── OS1_sensor_attrs.yaml # OceanSITES sensor attributes
│ │ ├── logging.yaml # Logging configuration
│ │ └── legacy/ # Legacy RAPID/RODB configuration files
│ │
│ └── legacy/ # Legacy RODB/RAPID processing (deprecated; excluded from lint)
│
├── tests/ # pytest test suite
│ ├── unit/ # seasenselib-free unit tests
│ ├── integration/ # Golden + pipeline tests (needs_seasenselib)
│ ├── fixtures/ # Committed raw + processed dune2_1_2026 fixtures
│ └── legacy/ # Tests for legacy RODB/RAPID processing
│
├── notebooks/ # Demo notebooks
├── docs/ # Sphinx documentation
│ └── source/
│ ├── methods/ # Method documentation (one page per processing step)
│ └── _static/ # Static files, code examples, CSS
│
├── CLAUDE.md # Claude Code guidance (local-only)
├── CITATION.cff # Citation metadata
├── CHANGELOG.md
├── pyproject.toml # Build system, deps, and project metadata
└── README.md
Processing Stages
The pipeline processes raw instrument data through sequential stages, each
producing a CF-NetCDF output file. oceanarray.STAGES is the registry of stages,
and oceanarray.process(mooring, stage=…, proc_dir=…, raw_dir=…) runs any subset.
Stage 1 — Standardisation (processors/stage1.py)
Input: raw instrument files (
.cnv,.rsk,.dat,.hex, RDI raw)Output:
{proc_dir}/{mooring}/{mooring}_{serial}_stage1.ncFaithful to the raw data — no QC, no trimming. Stores the transformation matrix and coordinate system so later stages can rotate correctly.
Stage 2 — Clock correction + trimming (processors/stage2.py)
Input:
*_stage1.nc+ mooring YAML (clock offsets, deployment window)Output:
*_stage2.ncApplies linear clock-offset correction; trims to the deployment window.
Stage 3 — QC, rotation, derived variables (processors/stage3.py)
Input:
*_stage2.ncOutput:
*_stage3.ncGross-range and tilt QC flags, BEAM→ENU rotation (Aquadopp), magnetic declination correction, salinity, density.
Stack — multi-instrument coordination (processors/stack.py)
Input:
*_stage3.ncfiles for all instruments on a mooringOutput:
{mooring}_stack.nc—(N_LEVELS, time)DatasetAligns instruments onto a common time axis; HAB-ordered deepest-first (index 0).
Grid — vertical gridding (processors/grid.py)
Input:
*_stack.ncOutput:
{mooring}_grid.nc—(N_LEVELS, time)on uniform pressure levelsSimple 1-D linear interpolation at each time step (preliminary; no objective mapping).
Data Flow
Raw files
│
▼ oceanarray process --stage 1
*_stage1.nc (faithful copy, CF-NetCDF)
│
▼ oceanarray process --stage 2
*_stage2.nc (clock-corrected, trimmed)
│
▼ oceanarray process --stage 3
*_stage3.nc (QC flagged, ENU velocities, salinity/density)
│
▼ oceanarray process --stage stack
{mooring}_stack.nc (N_LEVELS × time)
│
├──▶ oceanarray report → HTML reports
│
▼ oceanarray process --stage grid
{mooring}_grid.nc (uniform pressure × time)
│
└──▶ oceanarray report → grid HTML report
The full pipeline runs in one command: oceanarray process MOORING --stage 1 2 3 stack grid
(or oceanarray run MOORING).
Plotters Package Architecture
Three-tier architecture (see .claude/plotters_update-20260718.md for the rules):
Tier 1 (
plotters/primitives.py): low-level axes primitives, no domain knowledgeTier 2 (
plotters/current.py,timeseries.py,hydrography.py, etc.): domain functions that know about oceanographic variablesTier 3 (
reports/_plots.py): report wrappers that call Tier-2 functions and return base64 PNG strings for embedding in HTML
Report Package
Report types, each in its own module:
Report |
Module |
Output file |
|---|---|---|
Mooring summary |
|
|
Per-instrument |
|
|
Stack |
|
|
Grid |
|
|
Array |
|
|
All figures are generated by reports/_plots.py (Tier 3) and embedded as base64 PNGs.
Key Design Principles
Data provenance: never silently substitute defaults; store all processing parameters in NetCDF global attributes so treatment can be reconstructed from the file.
CF-compliant: CF conventions for metadata and variable naming throughout.
xarray-based:
xr.Datasetis the primary data structure in all stages.Discrete colorbars: all figures use
_nice_colorbar_bounds+BoundaryNorm; continuous colorbars are not used.Configurable: YAML-driven configuration for QC ranges, clock offsets, deployment windows, and instrument metadata.
Legacy Modules
legacy/ contains the RODB/RAPID-format processing path, kept for backward
compatibility with older datasets. New projects use the stage 1–3 pipeline above.