Skip to content

Spheron overview

Spheron is an aggregated GPU cloud that pools capacity from multiple providers and exposes it through a single API and dashboard, at 60-80% lower cost than traditional cloud providers.

What is Spheron?

Spheron is not a blockchain network. It is a GPU cloud platform that aggregates capacity from multiple providers across North America, Europe, and Asia Pacific and exposes it through a unified API and dashboard. You get a single interface across every provider without managing separate accounts, contracts, or billing relationships. Alongside GPUs, Spheron offers CPU-only nodes for workloads that need cores rather than accelerators.

Key features

VM access

Get full root access to your instances from the moment they are deployed. You can install custom drivers, configure the operating system, and set up your software stack exactly the way you need it, with no container restrictions or sandboxed environments.

Bare metal performance

Bare metal instances give your workloads direct access to the physical hardware, with no hypervisor or virtualization layer in between. This means consistent, predictable performance and full utilization of GPU memory and compute resources for your training and inference jobs.

Multi-GPU hosts with high-speed interconnects

Deploy multi-GPU bare-metal instances with NVLink or NVSwitch between the GPUs on a host. SXM offers are purpose-built for large-scale distributed training workloads that require fast GPU-to-GPU communication and low-latency gradient synchronization.

Aggregated provider network

Access GPUs and CPU nodes from multiple providers including Verda, Sesterce, Spheron AI, Spheron ES, Spheron MS, and Massed Compute through a single dashboard and API. Switching between providers does not require separate accounts, contracts, or billing relationships.

Hardware variety

Choose from a wide range of GPU hardware to match your workload:

  • High-end: B300 SXM6, B200 SXM6, and H100 SXM5 machines with NVLink and InfiniBand for large-scale training
  • Largest GPU memory: AMD Instinct MI300X with 192 GB of HBM3e per GPU, running ROCm
  • Mid-tier: A100 GPUs for production workloads
  • Cost-effective: RTX 4090 and other PCIe GPUs for development and testing

Offers and instance cards carry an AMD, NVIDIA, or Intel mark, so the vendor is visible before you deploy.

Shared volumes and persistent storage

Create persistent storage volumes that exist independently from your instances. Attach a volume to a running instance, detach it without losing data, and reattach it to a different instance later. Volumes support multi-instance attachment for shared datasets and model checkpoints across your team.

CPU nodes for work that needs no GPU

Not every job needs an accelerator. CPU Node is a CPU-only instance for build steps, data preparation, schedulers, API workers, and control planes, priced from $0.09 per hour. It has its own page in the dashboard sidebar and its own wizard, with sizes from 4 vCPU with 4 GB of memory across Verda, Sesterce, Spheron AI, and Massed Compute. Spot pricing is available on Verda.

Reserved GPU nodes

Reserve dedicated GPU nodes for long-term commitments to get better rates and guaranteed availability. Reserved instances are suited for teams with predictable, sustained compute needs who want to lock in access and reduce per-hour costs compared to on-demand pricing.

Flexible billing and cost savings

Pay only for what you use with per-second billing and no minimum commitments. Spot instances offer the same GPU hardware at lower prices when you can tolerate occasional interruptions. Combined with 60-80% savings over traditional cloud providers, Spheron significantly reduces your total infrastructure spend.

Team coordination

Manage GPU access for your entire team from a shared account. Teams share credits, SSH keys, and API keys across members. Role-based access controls let you assign owner, admin, or member permissions so each person has the right level of access for their responsibilities.

Cost savings

Spheron reduces GPU costs by 60-80% compared to traditional cloud providers:

  • RTX 4090: ~$0.52/hr (check current prices in the dashboard)
  • Traditional clouds: Typically charge 3-4x more for equivalent GPU resources
  • No hidden fees: Zero ingress/egress charges, transparent billing

Performance notes

  • Cluster (bare metal): No hypervisor layer means direct NVLink access and no virtualization overhead for multi-GPU training
  • Dedicated/Spot (VM): High-performance VMs with guaranteed GPU access; suitable for single-node training and inference

Platform advantages

Reliability

Five providers across dozens of regions mean you can redeploy to a different provider if one has availability issues. No single datacenter dependency.

Scalability

Deploy a single GPU instance or a multi-node H100 cluster. Scale up or down between deployments; no reserved capacity required.

Security

Choose providers with specific compliance certifications for your workload: Verda (ISO 27001, GDPR), Sesterce (SOC 2 Type II, ISO 27001), and Massed Compute (HIPAA, SOC 2 Type II).

Stop instead of destroy

Park an instance you will come back to. Stopping releases the GPU and drops billing to the disk rate while the disk and everything on it stays in place; starting returns the same environment on the same public IP. Available on Spheron AI and Spheron ES. See Instance lifecycle.

Deployment

  • Dashboard and REST API for deployment
  • Real-time metrics and monitoring
  • Pay-per-second billing with no hidden fees

How Spheron compares

FeatureSpheronTraditional CloudsOther GPU Clouds
Root Access✅ Full by default⚠️ Limited⚠️ Container-only (some)
Architecture✅ Bare metal (Cluster) / VM (Dedicated, Spot)❌ Virtualized⚠️ Mixed
Provider Model✅ Aggregated❌ Single vendor❌ Single vendor
High-end GPUs✅ SXM + NVLink⚠️ Limited⚠️ Limited
Pricing✅ 60-80% cheaper❌ Premium⚠️ Moderate

Use cases

  • LLM training and fine-tuning: Single-GPU to 8x H100 NVLink runs with PyTorch DDP or DeepSpeed
  • Production inference: Dedicated instances that cannot be interrupted mid-request
  • Distributed training: Multi-GPU bare-metal instances with NVLink or NVSwitch interconnects
  • Development and testing: RTX 4090 Spot instances at ~$0.52/hr for prototyping and iteration
  • Research: EU data-resident instances (Verda, Sesterce) for GDPR-compliant workloads
  • Support workloads: CPU nodes from $0.09/hr for data prep, build steps, and schedulers that never touch a GPU

Platform primitives

Understand how the platform works before deploying at larger scales:

What's next