# Enterprise GPU Rental | H100, A100, B200, B300 On-Demand | Spheron Docs > Rent NVIDIA H100, A100, B200 GPUs from $0.72/hr. No contracts. Instant deployment from Tier 3/4 data centers. Perfect for AI training, LLM inference & ML workloads. 99% uptime. ## Docs - [API reference](/api-reference): This page covers the Spheron AI REST API for programmatic access to GPU instances, SSH keys, volumes, and account balance. - [API skill for AI agents](/api-skill): This page gives you a ready-made skill file that teaches an AI agent how to use the [Spheron GPU API](/api-reference). Hand it to Claude or ChatGPT, and the agent understands the full deployment flow: which endpoint to call, in what order, how each parameter works, and how to handle errors, provider rules, and instance lifecycles. - [Billing](/billing): Manage credits, monitor usage, and track spending for GPU instances. - [Changelog](/changelog): All notable changes to **Spheron AI** will be documented in this file. - [Cost optimization](/cost-optimization): This page covers strategies for reducing GPU infrastructure costs on Spheron, including tier selection, instance type trade-offs, reserved GPU savings, and spend monitoring. - [General information](/general-info): This page lists official Spheron channels, support options, and answers to common questions. - [Getting started](/getting-started): This guide takes you from account creation to a deployed and verified GPU instance in about 10 minutes. - [Spheron overview](/overview): Spheron is an aggregated GPU cloud that pools capacity from multiple providers and exposes it through a single API and dashboard, at 60-80% lower cost than traditional cloud providers. - [Provider integration guide](/provider-integration): This guide describes the API contract a compute provider's orchestrator must expose for the Spheron AI Marketplace to list, provision, and manage GPU instances on the provider's infrastructure. It is written for the engineers who build and operate that orchestrator. - [Quick Start](/quick-start): Fast GPU deployment for users with an account already configured. Deploy in under 3 minutes. - [Reserved GPUs](/reserved-gpus): Request bulk GPU allocations, specific locations, or preferential pricing for long-term commitments. - [Security best practices](/security): This page covers essential security guidelines for protecting your Spheron account, API credentials, and GPU instances. Apply these practices before deploying in any production environment. - [Templates and images](/templates): Copy-ready cloud-init startup scripts organized by use case. Paste the script into the **Startup Script** field when deploying an instance. - [User settings](/user-settings): Manage your account profile, SSH keys, and API access credentials. - [Quick Guides](/quick-guides): **Training a model?** → Start with [Distributed Training](/quick-guides/training/distributed-training) for multi-GPU, or pick any RTX 4090 Spot instance for single-GPU fine-tuning - [Distributed Training (PyTorch DDP)](/quick-guides/training/distributed-training): Run large-scale distributed training with PyTorch DDP or DeepSpeed on a Voltage Park bare-metal H100 NVLink cluster. - [Training Guides](/quick-guides/training): Guides for running model training workloads on Spheron GPU instances, from single-GPU fine-tuning to large-scale distributed training on bare-metal H100 clusters. - [Gonka AI Node](/quick-guides/nodes/gonka-ai): Deploy a Gonka AI node on a Spheron GPU instance. Gonka is a decentralized AI compute network that uses Proof of Work 2.0, directing GPU compute toward real AI training and inference workloads. Operators earn rewards for providing verifiable compute. - [AI Node Guides](/quick-guides/nodes): Guides for deploying and running AI network nodes on Spheron GPU instances. Participate in decentralized AI compute networks and distributed model training protocols. - [Pluralis Node0-7.5B](/quick-guides/nodes/pluralis-node-0): Deploy a Pluralis Node0-7.5B on a Spheron GPU instance. Pluralis Protocol Learning allows multiple participants to collaboratively train large-scale foundation models without central ownership. Node0-7.5B enables permissionless participation in distributed AI model pretraining with 16GB+ VRAM. - [Chandra OCR](/quick-guides/llms/chandra-ocr): Deploy [Chandra OCR](https://huggingface.co/datalab-to/chandra) on a Spheron GPU instance. Chandra OCR converts images and PDFs into structured Markdown, HTML, or JSON while preserving document layout, hierarchy, and visual elements. It achieves 83.1% accuracy on the olmOCR benchmark, outperforming GPT-4o, Mistral OCR, and DeepSeek OCR. - [DeepSeek R1 & V3](/quick-guides/llms/deepseek-r1): Deploy [DeepSeek R1](https://huggingface.co/deepseek-ai/DeepSeek-R1) and [DeepSeek V3](https://huggingface.co/deepseek-ai/DeepSeek-V3) reasoning models on Spheron GPU instances using vLLM. DeepSeek R1 features chain-of-thought reasoning exposed via `` blocks; distilled variants (7B–32B) run on single GPUs. - [Gemma 3](/quick-guides/llms/gemma-3): Deploy [Gemma 3](https://huggingface.co/google/gemma-3-27b-it) from Google DeepMind on Spheron GPU instances using vLLM. Gemma 3 is released under the [Gemma Terms of Use](https://ai.google.dev/gemma/terms), which permits commercial use and modification after accepting Google's license agreement. - [LLM & AI Guides](/quick-guides/llms): Guides for running language model and AI inference workloads on Spheron GPU instances, from interactive chat interfaces to high-throughput OpenAI-compatible API servers. - [Janus CoderV-8B](/quick-guides/llms/janus-coderv-8b): Deploy [JanusCoderV-8B](https://huggingface.co/internlm/JanusCoderV-8B) on a Spheron GPU instance. JanusCoderV-8B is an 8B multimodal model that generates code from visual inputs including charts, screenshots, and UI mockups. It converts images into HTML, CSS, React components, and data visualization code. - [Llama 3.1 / 3.2 / 3.3](/quick-guides/llms/llama-3): Deploy Meta's Llama 3 family on Spheron GPU instances using vLLM. The Llama 3 series covers 8B through 405B parameters with strong instruction following and function calling capabilities. - [Llama 4 Scout & Maverick](/quick-guides/llms/llama-4-scout): Deploy [Meta Llama 4](https://huggingface.co/meta-llama) Scout and Maverick on Spheron GPU instances using vLLM. Llama 4 introduces a Mixture-of-Experts (MoE) architecture with native multimodal support for text and images. - [Mistral & Mixtral](/quick-guides/llms/mistral-mixtral): Deploy [Mistral AI](https://huggingface.co/mistralai) models on Spheron GPU instances using vLLM. Includes Mistral 7B for single-GPU deployment, Mixtral 8x7B MoE (requires ~90 GB VRAM in bfloat16, needs 2× A100 80GB), and Mistral Small 3.1 24B. - [Phi-4 & Phi-4 Multimodal](/quick-guides/llms/phi-4): Deploy [Microsoft Phi-4](https://huggingface.co/microsoft/phi-4) and [Phi-4-multimodal](https://huggingface.co/microsoft/Phi-4-multimodal-instruct) on Spheron GPU instances using vLLM. Phi-4 is a 14B parameter small language model (SLM) with strong reasoning capabilities released under the MIT license. - [Qwen3 Dense & MoE](/quick-guides/llms/qwen3): Deploy [Qwen3](https://huggingface.co/Qwen) dense and Mixture-of-Experts (MoE) models on Spheron GPU instances using vLLM. Qwen3 introduces a thinking mode that can be toggled at inference time using system prompt tokens. - [SoulX Podcast-1.7B](/quick-guides/llms/soulx-podcast-1-7b): Deploy [SoulX Podcast-1.7B](https://huggingface.co/Soul-AILab/SoulX-Podcast-1.7B) on a Spheron GPU instance. SoulX Podcast-1.7B is a 1.7B parameter speech generation model that produces multi-speaker podcast dialogues with speaker switching, zero-shot voice cloning, and paralinguistic elements such as laughter and sighs. It supports English, Mandarin, and several Chinese dialects. - [Specialized Models](/quick-guides/llms/specialized-models): Guides for deploying task-specific AI models on Spheron GPU instances. These models are purpose-built for document processing, audio generation, and visual code intelligence rather than general-purpose chat. - [Text Models](/quick-guides/llms/text-models): Guides for deploying large language models (LLMs) on Spheron GPU instances. All models are served via vLLM's OpenAI-compatible API unless otherwise noted. - [Baidu ERNIE-4.5-VL-28B-A3B-Thinking](/quick-guides/llms/multimodal/baidu-ernie-4-5-vl-28b-a3b): Deploy [Baidu ERNIE-4.5-VL-28B-A3B-Thinking](https://huggingface.co/baidu/ERNIE-4.5-VL-28B-A3B-Thinking) on a Spheron GPU instance. This multimodal reasoning model uses a Mixture-of-Experts architecture with 28B total parameters and 3B active per token. It supports visual reasoning, STEM problem solving, chart analysis, and video understanding under the Apache 2.0 license. - [Multimodal Models](/quick-guides/llms/multimodal): Guides for deploying vision-language models (VLMs) on Spheron GPU instances. All models accept both text and image inputs and are served via vLLM's OpenAI-compatible multimodal API. - [InternVL3](/quick-guides/llms/multimodal/internvl3): Deploy [InternVL3](https://huggingface.co/OpenGVLab/InternVL3-8B) on Spheron GPU instances using vLLM. InternVL3 is a vision-language model series from 1B to 78B parameters with strong performance on visual question answering and multimodal reasoning benchmarks. - [LLaVA-Next](/quick-guides/llms/multimodal/llava-next): Deploy [LLaVA-NeXT](https://huggingface.co/llava-hf/llava-v1.6-mistral-7b-hf) on Spheron GPU instances using vLLM. LLaVA-NeXT (Large Language and Vision Assistant Next) improves on LLaVA with better visual reasoning, higher image resolution support, and improved OCR capabilities. - [Pixtral-12B](/quick-guides/llms/multimodal/pixtral-12b): Deploy [Pixtral-12B](https://huggingface.co/mistralai/Pixtral-12B-2409) on a Spheron RTX 4090 (24GB) instance using vLLM. Pixtral-12B is Mistral AI's multimodal model built on Mistral-NeMo 12B, with a dedicated 400M visual encoder supporting variable-resolution image inputs. - [Qwen3-Omni-30B-A3B](/quick-guides/llms/multimodal/qwen3-omni-30b-a3b): Deploy [Qwen3-Omni-30B-A3B](https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct) on a Spheron A100 or H100 instance. This multimodal language model processes text, audio, images, and video with a 32K context window (single GPU). It differs from Qwen3-VL, which handles vision and language only. - [Qwen3-VL 4B & 8B](/quick-guides/llms/multimodal/qwen3-vl-4b-8b): Deploy [Qwen3-VL](https://huggingface.co/Qwen/Qwen3-VL-4B-Thinking) on a Spheron GPU instance. These vision-language models process text, images, and video with a 256K native context window (scalable to 1M tokens). Two size variants are available: 4B and 8B, plus an 8B-Thinking variant with enhanced reasoning. - [Inference Frameworks](/quick-guides/llms/frameworks): Choose the right LLM serving stack for your workload. All frameworks listed here expose an OpenAI-compatible `/v1` API unless otherwise noted. - [llama.cpp Server](/quick-guides/llms/frameworks/llama-cpp): Deploy [llama.cpp](https://github.com/ggerganov/llama.cpp) as an OpenAI-compatible HTTP server on Spheron GPU instances. llama.cpp supports GGUF-quantized models and can offload layers between CPU and GPU, making it ideal for consumer-grade GPUs and quantized inference. - [LMDeploy](/quick-guides/llms/frameworks/lmdeploy): Deploy [LMDeploy](https://github.com/InternLM/lmdeploy) with the TurboMind inference engine on Spheron A100 or H100 instances. LMDeploy supports AWQ quantization for memory-efficient inference and exposes an OpenAI-compatible API. - [LocalAI](/quick-guides/llms/frameworks/localai): Deploy [LocalAI](https://github.com/mudler/LocalAI) on a Spheron GPU instance. LocalAI is an OpenAI-compatible drop-in replacement that supports LLMs, Whisper speech-to-text, and Stable Diffusion image generation through a single Docker container with NVIDIA GPU passthrough. - [Ollama + Open WebUI](/quick-guides/llms/frameworks/ollama): Run [Ollama](https://ollama.com) with [Open WebUI](https://github.com/open-webui/open-webui) on an RTX 4090 Spheron instance. Open WebUI provides a browser-based chat interface backed by any model Ollama can load into VRAM. - [SGLang Inference Server](/quick-guides/llms/frameworks/sglang): Deploy an [SGLang](https://github.com/sgl-project/sglang) OpenAI-compatible inference server on Spheron GPU instances. SGLang features RadixAttention for KV cache reuse across requests and native support for constrained decoding and structured output. - [TensorRT-LLM + Triton Inference Server](/quick-guides/llms/frameworks/tensorrt-llm): Deploy [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) with [Triton Inference Server](https://github.com/triton-inference-server/server) on Spheron H100 instances for maximum NVIDIA GPU throughput. TensorRT-LLM compiles model weights into an optimized engine before inference, yielding best-in-class token generation rates. - [vLLM Inference Server](/quick-guides/llms/frameworks/vllm-server): Deploy an OpenAI-compatible inference server using [vLLM](https://github.com/vllm-project/vllm) on Spheron H100 or A100 instances. - [ComfyUI](/quick-guides/image-generation/comfyui): Deploy [ComfyUI](https://github.com/Comfy-Org/ComfyUI) on a Spheron GPU instance using Docker. ComfyUI provides a browser-based node editor for building image generation workflows, plus a JSON API for programmatic access. - [FLUX.1 & FLUX.2](/quick-guides/image-generation/flux-1): Deploy [FLUX.1](https://huggingface.co/black-forest-labs/FLUX.1-dev) and FLUX.2 from [Black Forest Labs](https://blackforestlabs.ai) on Spheron GPU instances. FLUX models deliver state-of-the-art photorealistic text-to-image generation. - [Image Generation Guides](/quick-guides/image-generation): Deploy GPU-accelerated text-to-image generation models on Spheron GPU instances. - [Stable Diffusion 3.5 & SDXL](/quick-guides/image-generation/stable-diffusion-35): Deploy [Stable Diffusion 3.5](https://huggingface.co/stabilityai/stable-diffusion-3.5-large) and [SDXL](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0) from Stability AI on Spheron GPU instances. A FastAPI wrapper exposes a `/generate` endpoint for programmatic image generation. - [CUDA and NVIDIA drivers](/connecting/cuda-nvidia-drivers): This page explains NVIDIA drivers and CUDA on Spheron GPU instances, how they interact with AI/ML frameworks, and how to choose the right version for your workload. - [Startup scripts](/connecting): Run scripts automatically when your instance first boots to automate environment setup and ensure consistent configuration across deployments. - [Jupyter Notebook](/connecting/jupyter): Deploy GPU instances with Jupyter Notebook pre-configured for interactive AI/ML development. - [Kubernetes addon](/connecting/kubernetes): Deploy a managed Kubernetes cluster on a Voltage Park Cluster instance using the Spheron K8s addon. - [PyTorch environment](/connecting/pytorch): Deploy GPU instances with PyTorch pre-configured for AI/ML development. - [SSH connection setup](/connecting/ssh-connection): Generate and configure SSH keys for secure GPU instance access. - [TensorFlow environment](/connecting/tensorflow): Deploy GPU instances with TensorFlow pre-configured for immediate AI/ML development. - [Ubuntu environments](/connecting/ubuntu): Available Ubuntu configurations for Spheron GPU instances. - [VS Code Remote SSH](/connecting/vscode-remote): Connect VS Code directly to a Spheron GPU instance. IntelliSense, debugging, extensions, and the Ports panel run against the remote environment. - [Verda: mounting shared storage](/connecting/volume-mounting/data-crunch): Mount persistent shared filesystems on Verda instances using Network File System (NFS) protocol. - [Mounting shared storage](/connecting/volume-mounting): Mount persistent storage volumes to your GPU instances across different providers. - [Sesterce: mounting shared storage](/connecting/volume-mounting/sesterce): Mount persistent block storage volumes on Sesterce instances using ext4-formatted block devices. - [Spheron AI: mounting shared storage](/connecting/volume-mounting/spheron-ai): Mount persistent block storage volumes on Spheron AI instances using raw block devices and your choice of filesystem. - [Spheron ES: mounting shared storage](/connecting/volume-mounting/spheron-es): Mount persistent shared filesystems on Spheron ES (Extra Supply) instances using the virtiofs protocol. - [Voltage Park: mounting shared storage](/connecting/volume-mounting/voltage-park): Mount persistent storage volumes on Voltage Park instances using Network File System (NFS) protocol. - [Concepts](/concepts): Spheron GPU offerings differ on two dimensions: **interruptibility** and **hardware isolation**. - [Instance Types](/concepts/instance-types): Spheron GPU offerings are organized along two axes: **interruptibility** and **hardware isolation**. - [Networking](/concepts/networking): Every Spheron GPU instance receives a **dedicated public IP address** when deployed. There is no shared IP, NAT, or port forwarding. Each instance has its own routable IP for the lifetime of the deployment. - [Regions & Providers](/concepts/regions-providers): Spheron sources GPU capacity from six providers across North America and Europe. - [Teams](/concepts/teams): Teams let multiple users share a GPU credit pool, SSH keys, volumes, and deployments under a single account.