Skip to content
Mitchel Carson

Research

HYDRA: watershed forecasting research

Motivated by experiencing Hurricane Helene in Boone, HYDRA studies watershed dynamics and forecast reliability: NextGen reforecast generation software and Google Cloud data workflows, plus a deep-learning post-processor for 1–18 hour streamflow forecasts, with LSTM, Transformer, and Mamba-style models compared under leakage-aware evaluation. A results manuscript for Water Resources Research is in preparation, and a software paper for Environmental Modelling & Software is planned.

Ongoing research · In progress

Scope

What is being built and tested

1–18 h

forecast lead times targeted by the post-processing model

In progress

No performance number is reported on this site until the analysis is complete; results will be published with the manuscripts.

Reforecast generation software drives NOAA’s NextGen framework to produce the retrospective forecasts the model learns from, with initialization, lead-time, and version metadata preserved so every training example is traceable.

LSTM, vanilla Transformer, and Mamba-style state-space models are trained as post-processors on identical inputs and splits and compared against the raw reforecasts at every lead time from 1 to 18 hours.

Figure 1 · interactive

Inside the HYDRA pipeline

Reforecast generation

The forecasts to learn from are generated, not scraped.

Reforecasts: NextGen reforecast generation software produces retrospective forecasts with consistent initialization, lead-time, and version metadata. Inputs: weather, streamflow, and forecast data, acquired and validated with Google Cloud workflows and aligned to forecast issue time with traceable provenance.

Consistent initialization, lead-time, and version metadata make every training example traceable to the forecast that produced it.

MarkersNextGen · Reforecasts · USGS observations · Issue time

Figure 1. Interactive trace of the HYDRA pipeline: reforecast generation, the three model families compared, leakage-aware evaluation by lead time, and the outputs.

Method

Architecture, evaluation, constraints, reproducibility

Architecture

  1. Reforecasts: NextGen reforecast generation software produces retrospective forecasts with consistent initialization, lead-time, and version metadata.
  2. Inputs: weather, streamflow, and forecast data, acquired and validated with Google Cloud workflows and aligned to forecast issue time with traceable provenance.
  3. Models: LSTM, vanilla Transformer, and Mamba-style state-space models (with attention) trained as post-processors on identical inputs.
  4. Outputs: Improved forecasts at 1–18 hour lead times, diagnostics, and research artifacts.

Evaluation

  1. Metrics: Hydrologic error and skill metrics (RMSE, NSE, KGE) reported by site and lead time.
  2. Validation: Leakage-aware temporal splits designed around forecast availability.
  3. Reference: Raw NextGen reforecasts at every lead time.

Constraints

  1. Research conclusions remain provisional until the analysis is complete; no performance numbers are reported yet.
  2. Operational claims must respect the timing and availability of every model input.

Reproducibility

  1. Configuration-driven experiments with tracked parameters and artifacts.
  2. Versioned data lineage, manifests, and evaluation outputs.
  3. Environment and run records designed to support scientific review.

Communication

Scientific communication, labeled by status.

Outputs are listed with their current status and updated as milestones are confirmed.

  • December 2026

    Accepted

    Workshop

    Best Practices for AI and Agentic Workflows in Earth Science Research

    AGU26 Annual Meeting · San Francisco · December 7–11, 2026

    Role: Scientific workshop facilitator

    Teaching earth and environmental scientists practical AI methods for their research workflows.

  • Decision pending

    Under review

    Abstract

    HYDRA streamflow-forecasting abstract

    AGU26 · Hydrology session H100 (machine learning in hydrology)

    Abstract on the HYDRA streamflow-forecasting work. Acceptance and scheduling will be posted when confirmed.

    Research details

  • In preparation

    In progress

    Manuscript

    HYDRA results manuscript

    Water Resources Research (target journal)

    Role: Author

    Post-processing results for NextGen streamflow forecasts at 1–18 hour lead times. Analysis and writing are in progress; a link will be added when one exists.

    Research details

  • Planned

    Planned

    Manuscript

    NextGen reforecast generation software

    Environmental Modelling & Software (target journal)

    Role: Author

    Software paper describing the reforecast generation and data tooling, planned alongside the code release.

    Code on GitHub

  • December 2025

    Completed

    Thesis

    Senior Honors Thesis on runoff forecasting with deep learning

    Appalachian State University

    Role: Author

    The deep-learning runoff-forecasting work that became HYDRA.

    Thesis-era code on GitHub

Discussions · office hours

Topics I am glad to talk through.

These are invitations, not past talks. Book 30 minutes or email me if one of these is your problem too.

  • Leakage-safe temporal evaluation

    What “available at forecast time” really means for train, validation, and test splits, and how easy it is to cheat by accident.

  • Post-processing operational forecasts by lead time

    Why learn a correction on top of NextGen reforecasts instead of replacing the model, and what changes between 1 and 18 hours ahead.

  • Sequence models for hydrology: LSTM, Transformer, Mamba

    What a fair comparison between recurrent, attention, and state-space models needs before anyone declares a winner.

  • AI and agentic workflows in earth-science research

    What I am putting in front of scientists at AGU26: what is worth adopting, and what to be cautious about.

  • Reproducible research pipelines

    Reforecast generation you can rerun, versioned data lineage, and metrics by site and lead time instead of one aggregate score.

  • From production software to research code

    What transfers from GraphQL services at USAA to research code, and what had to be unlearned.

  • Bounding LLM autonomy in decision systems

    Why LLMs should advise rather than act, and how Harmony confines them to nodes that can flag or veto while every consequential action waits for a human.

  • Reliability lessons from executive-missions operations

    What 50+ executive airlift missions, including Air Force Two, with zero safety-related incidents taught me about preparation and reliability, and how that shows up in software.

  • From military service into AI

    Why I chose to pursue AI while serving, and what the path from an Air Force flying job to computer science and graduate AI research looked like, for veterans weighing the same move.

Austin, Texas · Central Time