Cloud Run Guide#

Running on your own AWS account

How CryoStack launches, monitors, and retrieves the results of a run on AWS Batch — using an AWS account you connect yourself (bring-your-own-AWS), and the same application configuration you already use locally or on HPC.

Where to go next.

Cards navigate to the sections below. This page describes the platform-wide cloud execution path shared by CryoLauncher and ICESEE — application-specific fields and screens stay in each application's own User Manual.

JM

The cloud run journey

The real sequence: application, configuration, backend, AWS connection, launch, monitor, results.

BK

Local, HPC, or Cloud

What is identical across backends, and what genuinely changes.

AC

Connecting your AWS account

CloudFormation onboarding, ExternalId, and temporary credentials.

PC

Preparing cloud infrastructure

Prepare cloud: storage, container registry, and AWS Batch.

CM

Compute mode

Fargate (default) or EC2 (Advanced): capacity, accelerator, network, and execution options.

RL

Review & Launch

The run estimate, the review card, and what actually launches.

MR

Monitoring & results

Run states, logs, and how outputs come back to your Workspace.

WE

Worked examples verified on AWS

What has actually run end-to-end, and where to find the exact step-by-step walkthrough for each application.

TS

Troubleshooting

Common failure points and how to verify your deployment.

What this page documents.

This describes software that exists today — every button, label, and step name below matches the running application. Nothing here is a roadmap. "Cloud parity" is not one fact — it is five separate ones, and they are not all at the same maturity:

SA

Shared architecture Working

CryoLauncher and ICESEE both provision cloud infrastructure through the same AWSDriver/CloudFormation-onboarding architecture — one implementation, not two.

CP

Configuration portability Working

The same scientific/example configuration (params.yaml, model settings) carries over across Local, Remote/HPC, and Cloud — there is no separate cloud-only configuration to maintain.

CE

Cloud exposure Partial

CryoLauncher and ICESEE currently expose a Cloud execution mode in the UI. This is not true of every CryoStack application — do not assume a third application has a cloud path just because these two do.

VW

Validated workflows Demonstrated

Three configurations have been run and confirmed end-to-end on AWS Batch's Fargate compute mode: ICESEE Lorenz-96 at NP = 1, CryoLauncher/Icepack 04-synthetic-ice-stream-xy, and CryoLauncher/ISSM (MPI-parallel solver execution, postprocessing into CryoStack's shared result package, retrieval, and visualization). CryoLauncher/ISSM has also completed the same end-to-end path on EC2 On-Demand, in both cases reaching the configured Georgia Tech institutional MATLAB license through the same Connector/Relay infrastructure used for Remote access. This validates one institutional Cloud/Connector configuration; it does not establish compatibility with arbitrary institutional license-server arrangements. See MATLAB licensing and Connector setup.

CM

Compute mode Fargate and EC2 On-Demand validated

Fargate is the default AWS Batch compute mode. EC2 (Advanced) is an opt-in alternative; its On-Demand, single-node, CPU configuration has now also run end to end on live AWS — CryoLauncher/Icepack's 00-meshes-functions tutorial and CryoLauncher/ISSM, both submitted to a CryoStack-provisioned managed EC2 compute environment. EC2 Spot, GPU, and multi-node, and EC2 for ICESEE, remain implemented/provisioned but not yet run against live AWS in this repository's evidence. See Compute mode below.

In short: CryoStack has backend/configuration parity at the architecture level today. Application-level operational validation — an actual completed run, end to end, on live AWS — is still expanding, one configuration at a time, and this page names exactly which one. Not yet available in-app teardown of AWS-side infrastructure (stack/role deletion) — see Cleanup below for exactly what that does and does not cover.

The cloud run journey#

CryoStack’s cloud path follows the same shape as Local and Remote/HPC execution — the application and its configuration do not change; only the execution backend changes:

 Application  →  Configuration  →  Compute backend  →  AWS connection
                                                              │
                                                              ▼
                    Results  ←  Monitor  ←  Launch  ←  Prepare cloud
                                                              ▲
                                                              │
                                              AWS Batch  ──┬── Fargate (default)
                                                            └── EC2 (Advanced)
                                                                ├─ On-Demand / Spot
                                                                ├─ Default / Custom network
                                                                └─ Single node / Multi-node*

Both branches converge on the same Prepare cloud → Launch → Monitor → Results path below — EC2 changes what Batch schedules the job onto, not the application, the staging, or the result pipeline. (* guarded/experimental — see Compute mode.)

This is the shared architecture, not a claim that Monitor/Results behave identically in every application today: CryoLauncher uses an active cloud-run controller that polls the job and synchronizes results automatically after a successful completion, while ICESEE currently uses explicit, manual status and result actions instead (details in Monitoring a run and retrieving results).

For CryoLauncher, Basic/Advanced and Remote/Cloud are separate choices — Basic/Advanced is how you configure a run; Remote/Cloud is where it executes. Selecting Cloud does not by itself change which Cloud controls you see:

CryoLauncher Basic  + Cloud  ->  simplified Cloud surface  ->  Fargate-only

CryoLauncher Advanced + Cloud  ->  advanced Cloud controls  ->  supported
    Fargate/EC2 configuration and other exposed controls, subject to the
    documented capability limitations below

Basic mode always submits to Fargate — the entire Advanced Cloud panel (compute-mode choice, EC2 capacity/accelerator/network/execution options) is hidden. Switching to Advanced exposes those controls, but exposing a control is not the same as that capability being AWS-validated — see Compute mode below for exactly which Advanced options have and have not been run against live AWS.

In the application’s own terms, this is:

  1. Open the application (CryoLauncher or ICESEE) and select an example.
  2. Configure it exactly as you would for a local or Remote/HPC run.
  3. In Run settings, set Execution mode to Cloud.
  4. Under Cloud Environment → AWS ACCOUNT, connect your AWS account (once) — see Connecting your AWS account.
  5. Click Prepare cloud to provision what your account needs — see Preparing cloud infrastructure.
  6. Click Review & Launch, then Launch cloud run — see Reviewing and launching a run.
  7. Watch the CLOUD RUN status card, then open Results — see Monitoring a run and retrieving results.

What stays the same, what changes#

Configuration Identical

The example, model parameters, filter/solver settings, ensemble size, and dataset references are the same object regardless of backend. There is no separate "cloud version" of a scientific configuration to create or maintain.

Local

Runs inside the CryoStack server process itself. No connection, no queue — the fastest path for small runs and iteration.

Remote / HPC

Runs on a Linux server or Slurm-managed cluster you already have access to, through the CryoStack Connector or direct SSH. Adds: your HPC identity, a remote working directory, and (for ICESEE) a first-time Spack environment Check/Prepare step.

Cloud

Runs on AWS Batch, in your own AWS account. Adds: connecting the account once, letting CryoStack prepare the AWS infrastructure, and reviewing an estimated cost before launch. It does not add a different scientific workflow.

Compute mode: Fargate (default) or EC2 (Advanced)#

Every run above submits to AWS Batch. Batch itself needs a compute environment to schedule jobs onto, and CryoStack lets you choose which kind under Advanced cloud settings in Cloud Environment:

  • Fargate — the default. No infrastructure to choose or manage; ISSM, Icepack, and ICESEE’s Lorenz-96 have each completed validated runs on it (see Worked examples verified on AWS).

  • EC2 (Advanced) — CryoStack provisions its own EC2-backed Batch compute environment instead. It scales to zero when idle; CryoStack manages the AMI, the ECS instance role, and scaling. Two sizing fields apply to every EC2 environment:

    • Max vCPUs — the ceiling for the compute environment (default 16).

    • Instance types — optimal (default; AWS chooses) or a list of instance families such as c5,m5,r5.

    Selecting EC2 also reveals four further choices, each purpose-built rather than exposing raw AWS knobs:

    • Capacity — On-Demand (default; predictable EC2 capacity) or Spot (lower-cost, interruptible capacity that AWS can reclaim).

    • Accelerator — None (default) or GPU — stages GPU-capable EC2 infrastructure, but CryoStack’s qualified container image has no CUDA runtime, so a GPU job is refused at submission until a GPU-qualified image exists. This is infrastructure preparation, not a working scientific execution mode.

    • Network — Default or Custom / Private — places the EC2 compute environment into a VPC, subnets, and security groups you already control (a vpc-…, subnet-…, sg-… you supply), for example one already routed to an institutional network. CryoStack does not create a VPN, Direct Connect connection, Transit Gateway, or firewall rule itself — the VPC you point it at must already have whatever route it needs. See MATLAB licensing for ISSM cloud runs for the supported Connector alternative; custom networking is not required by the validated Fargate/EC2 path.

    • Execution — Single node (default) or Multi-node — registers an AWS Batch multi-node parallel job definition (EC2 only), but CryoStack’s scientific runners do not yet coordinate distributed MPI across Batch nodes, so this stages the infrastructure ahead of a scientific run rather than enabling one today.

On-Demand, single-node capacity has now run end to end against live AWS for CryoLauncher/Icepack’s 00-meshes-functions tutorial (Advanced Cloud, EC2, On-Demand, single node, CPU, 2 vCPU / 8 GiB) — CryoStack submitted to its own managed EC2 compute environment (cryostack-ec2), the cryostack-ec2-queue job queue, and the cryostack-icepack-ec2 job definition, and the job completed successfully. This does not extend to: Spot capacity, GPU, multi-node execution, custom/private networking, ISSM on EC2, ICESEE on EC2, or every Icepack example — each of those is implemented and exercised locally but not yet run against live AWS in this repository’s evidence, so treat them as advanced, not-yet-AWS-validated configurations rather than a second fully validated backend. GPU and multi-node are additionally guarded at submission specifically so that selecting them stages infrastructure without ever silently attempting a scientific run neither the image nor the runners can actually perform.

Connecting your AWS account (BYO-AWS)#

CryoStack never asks for an AWS access key, secret, or CLI profile. Instead, you create a cross-account IAM role yourself, in your own AWS console, and CryoStack assumes it for short-lived (temporary) credentials only.

  1. In Cloud Environment → AWS ACCOUNT, click Connect AWS Account.

  2. Click Open AWS Setup. This opens a CloudFormation Quick Create page in your own AWS console, pre-filled with a unique ExternalId (a confused-deputy defence — the role can only be assumed with this exact value) and the CryoStack principal it should trust. Review the stack and click Create stack.

  3. The stack creates one IAM role scoped to cryostack-* resources only — no AdministratorAccess, no wildcard actions. Copy the Role ARN from the stack’s Outputs tab.

  4. Paste the Role ARN back into CryoStack and click Verify connection. CryoStack assumes the role with your ExternalId, confirms your account, and shows ● Connected.

Each CryoStack connection gets its own CloudFormation stack and its own IAM role — connecting a second AWS identity, or reconnecting after a previous attempt didn’t finish, never requires deleting an existing, working role first.

Note

Retry connection (after a failed verification) and Change AWS account (to connect a different account without losing the current one) are separate, explicit actions in the AWS ACCOUNT panel — see the CryoLauncher User Manual’s Connect AWS Account walkthrough for every button and state in detail.

Disconnect removes only the connection metadata CryoStack stored locally. There are no long-lived credentials to revoke — STS sessions are short-lived and are never stored. It does not delete anything in your AWS account.

Preparing cloud infrastructure#

Once connected, click Prepare cloud. Using your temporary role session, CryoStack creates whatever is missing in your account — nothing is created in a CryoStack-owned account:

  • an S3 bucket for run input/output (cryostack-runs-<your-account-id>);

  • an ECR repository holding the exact, digest-pinned tested container image;

  • an AWS Batch compute environment, job queue, and job definition.

The Account / Storage / Containers / Compute rows move to Ready as each step completes. Prepare cloud is safe to run again — existing resources are reused, never recreated or duplicated.

Reviewing and launching a run#

Once infrastructure is Ready, a RUN ESTIMATE appears (expected runtime, requested resources, and an estimated AWS cost when pricing is available — on Fargate; EC2 has no cost model implemented yet, so its review honestly shows Estimated cost: unavailable rather than a guessed figure). Click Review & Launch to open the full review card, which shows the experiment, the AWS account and region, the resources, and an infrastructure-readiness checklist. Launch cloud run is only enabled once every check passes — CryoStack never launches a configuration it cannot honestly claim is ready, and it asks you to review again if you change anything after opening the review.

AWS charges apply to your own AWS account for whatever the run actually uses; cost figures shown here are estimates only, not a bill.

Monitoring a run and retrieving results#

The underlying job passes through the same states either way:

Staging → Submitting → Queued → Running → Completed
                                        ↘ Failed
                                        ↘ Cancelled

How you observe those states, and how results come back, currently differs by application — this is a real implementation difference, not an unvalidated version of the same behavior:

  • CryoLauncher shows a CLOUD RUN card that tracks the job in the background and automatically syncs results on completion — no click required before Results is populated. See Connect AWS Account (step 7) for the full View log / View results / Terminate walkthrough.

  • ICESEE has no background poller — you click Check status yourself, and opening a completed run triggers a best-effort result sync before showing Results. See ICESEE Cloud Mode for the exact steps.

View log reads CloudWatch Logs from whichever log group the job’s own Batch job definition actually configured — never a single fixed group — so it works the same way whether the job used the CryoStack-managed log group or AWS Batch’s own default. If the run’s log stream is not available yet (too early after submission) or genuinely absent, CryoStack says so directly in the Run Log rather than presenting it as a run failure — your job and its results are unaffected either way.

Worked examples verified on AWS#

See Validated workflows above for exactly which configurations have been confirmed end-to-end, and Compute mode for the EC2-specific evidence (job queue/definition names). Every validated run follows the identical Connect → Prepare cloud → Review & Launch → Monitor → Results sequence described above; only the application/example you open and, for an EC2 run, selecting Advanced → EC2 → On-Demand under Compute mode before Prepare cloud, differ.

For the exact, button-by-button steps:

  • CryoLauncher — see Connect AWS Account in the CryoLauncher User Manual.

  • ICESEE — see Cloud Mode in the ICESEE User Manual for the full Lorenz-96 walkthrough, including the Processes = 1 restriction (see Verified runtime contracts below) and its own monitoring/retrieval steps.

See Cleanup below before you consider any run finished — nothing about the AWS infrastructure a run used is removed automatically.

Cleanup#

Cleanup is not one action — CryoStack handles some of it, and some of it is entirely manual today:

  • Individual Batch jobs — CryoLauncher’s Terminate button, and ICESEE’s own cloud termination action, stop a running or queued job directly from CryoStack. This works today.

  • AWS job history — nothing to do; AWS retains completed/terminated job records at no ongoing cost. CryoStack does not need to (and does not) clean these up.

  • S3 input/output objects — not deleted automatically. Every run’s staged inputs and outputs persist in your S3 bucket until you remove them yourself.

  • AWS Batch resources (compute environment, job queue, job definitions) — created once by Prepare cloud and reused afterward. There is currently no in-app “unprepare” or teardown action.

  • Disconnecting an AWS account (the Disconnect button) — removes only CryoStack’s local connection record (ExternalId, role ARN, metadata). It does not delete or modify anything in AWS.

  • The CloudFormation stack and its IAM role (from Connecting your AWS account) — not deleted through CryoStack. If you want to tear down a connection’s AWS infrastructure entirely, delete the stack yourself from the AWS Console or CLI; deleting it removes the IAM role with it.

Verified runtime contracts#

ICESEE’s cloud container has been verified, locally and against the exact published image, for exactly:

  • Example: lorenz96

  • Processes (ICESEE_NP): 1

Multi-process execution was tested and found unsafe on this image (the default execution path has no coordination between MPI ranks and races on shared output files), and no ICESEE example other than Lorenz-96 has been run against the cloud container. CryoStack’s Review card enforces this directly for ICESEE — a different example or a higher process count is refused with an explicit reason, never silently changed to a value that would pass. This is a deliberate design choice: the goal is an honest preflight, not a best-effort launch.

An ISSM-coupled ICESEE forecast model is a separate, stronger limitation than the process-count restriction above. ICESEE’s Cloud Environment shows the same MATLAB license field CryoLauncher’s does, for interface consistency, but that field is not currently connected to ICESEE’s own Cloud submission path — no institutional license or Connector tunnel is actually configured for an ICESEE Cloud job. An ISSM-based ICESEE workflow that needs MATLAB is therefore not operational on Cloud today, independent of and in addition to the example/process restriction above. This does not affect ICESEE’s Lorenz-96 or Icepack forecast models, which need no license, or CryoLauncher’s own ISSM Cloud path, which does have this connectivity (see Validated workflows above).

Icepack’s cloud path has confirmed one example end-to-end, 04-synthetic-ice-stream-xy, sharing the same Firedrake export and figure-capture code the Remote/Slurm path uses. Unlike ICESEE, CryoLauncher does not currently refuse a different Icepack example at Review — the code path is model-neutral, so other examples are expected to run, but only this one has actually been confirmed against the cloud container. Treat other Icepack examples on Cloud the same way you would treat them on Remote: architecturally supported, not yet individually verified.

Downloads and reference material#

  • CloudFormation template source — deployment/cloudformation/cryostack-execution-role.json in the CryoStack repository is the exact source of the template your administrator publishes at the URL CryoStack opens during Open AWS Setup. See Preparing cloud infrastructure above for what it creates.

  • CryoStack on GitHub — browse the full source, including the CloudFormation template and the cloud execution code referenced throughout this page.

Troubleshooting and verification#

“Resource of type ‘AWS::IAM::Role’ … already exists” during stack creation. This was a real defect in an earlier template revision (a fixed role name collided across connections) and has been fixed — CloudFormation now generates a unique role name per stack. If you still see this error, your deployment’s published template may be stale; ask your administrator to confirm it matches the checked-in deployment/cloudformation/cryostack-execution-role.json (sha256sum the two files and compare).

A CloudFormation stack is stuck in ROLLBACK_COMPLETE. Retrying with the exact same stack name is refused by AWS once a stack has failed. Use Retry connection, not a manual re-submission of the same Quick Create URL — CryoStack mints a fresh, non-colliding stack name for the same connection without changing your ExternalId or an already-working role.

Launch is disabled and the review lists a reason. This is intentional — read the listed reason (infrastructure not yet Ready, account connection stale, or an unverified example/process count for ICESEE) rather than retrying blindly; each reason names exactly what to fix.