GPU rental · dedicated hardware
Your own 128 GB GB10 Grace-Blackwell machine in Stockholm, Sweden — a named, single-tenant physical system on monthly rental from hardware AxForge owns and operates. Your model, your traffic, our machine.
Specifications
| System | NVIDIA DGX Spark — dedicated, single-tenant |
|---|---|
| Superchip | NVIDIA GB10 Grace-Blackwell — see the GB10 page |
| Memory | 128 GB unified, shared CPU/GPU |
| CPU architecture | ARM64 |
| Tenancy | Dedicated — a named physical machine assigned to you, not shared capacity |
| Rental term | Monthly — cancel before the next renewal |
| Pricing | €495 / month (launch pricing, excl. VAT) |
| Region | eu-se-1 · Stockholm, Sweden |
Performance
| Example workload | Model | Community benchmarks* |
|---|---|---|
| Fast agents / reasoning (MoE) | Qwen3.6 35B-A3B · NVFP4 + MTP | up to 86.3 t/s |
| Coding (MoE) | Qwen3-Coder 30B-A3B | ~42–61 t/s |
| Dense reasoning | Qwen3.8 27B · NVFP4 + MTP | 25.1 t/s |
| Smaller dense LLM | Qwen3 8B · Q4 | ~42 t/s |
| Large MoE | DeepSeek V4 Flash · 180B, 13B active | ~23 t/s |
*Community results (SparkBench PBM @ 4k, llama.cpp) — labelled, not ours; they vary by runtime, quantization, context length and serving configuration. Full tables, sources and the MoE-vs-dense story on the GB10 page. The same hardware also serves image generation and editing — see the model catalogue.
Fit
| Use case | Why it fits |
|---|---|
| Dedicated inference, ~7B–35B models | 128 GB unified memory holds model weights and KV cache in one pool — the measured numbers above are exactly this workload. |
| Private inference | A single-tenant machine in an EU region. Your model, your traffic, our hardware — prompts never persisted. |
| Dev / staging nodes | A named machine you keep for the month — a stable target for integration, load testing and pre-production serving. |
| Model evaluation | Run bake-offs on the exact hardware class you would serve from, with results that transfer. |
| AI agents | Host agent runtimes next to their model — long-running processes with no per-token surprise from a third party. |
| Fine-tuning experiments | The unified memory pool fits adapters and smaller fine-tunes that would need careful sharding elsewhere. |
| Self-managed serving & containers | Bring your own serving stack — vLLM, llama.cpp, custom containers — and run it your way on your machine. |
Searching for DGX Spark cloud or DGX Spark hosting? Same machine: you rent a named physical system, we host and operate it, you reach it over the network.
ARM64
The GB10 platform is ARM64, not x86. Most modern AI frameworks and container images ship ARM64 builds — our own production serving stack runs on it. x86-only binaries need ARM64 builds or rebuilds. Not sure about your stack? An engineer can review it with you before you commit.
How it works
| 1 | Create an account — or talk to an engineer first if you want the stack reviewed. |
|---|---|
| 2 | Request a DGX Spark from the console's GPU area. |
| 3 | An engineer confirms availability and provisions your machine — provisioning is engineer-led, not automated. |
| 4 | Connect and deploy your models, containers and tools over remote access. |
| 5 | Keep it monthly at €495/month — cancel before the next renewal. |
Data & privacy
Prompts never persisted. Requests to your dedicated DGX Spark are processed in memory in Sweden — not written to disk, not logged, not retained, never used to train anything. We keep only request metadata (token counts, timestamps, status) for billing and operations. Full policy at axforge.ai/privacy.
FAQ
Yes — machines are rented monthly from Sweden (eu-se-1), subject to current capacity. Talk to an engineer to claim one; they confirm availability with you.
Provisioning is engineer-led: you request a machine (or talk to us first), an engineer confirms availability and compatibility, then sets up your dedicated DGX Spark and hands you remote access. From there you deploy your own models, containers and tools.
Anything that fits a 128 GB ARM64 machine you fully control: dedicated inference of ~7B–35B models, agents, fine-tuning experiments, evaluation, dev/staging, and your own serving stack in containers. If you only need per-token inference, the serverless Qwen API may be enough — no machine required.
€495 per month for a dedicated machine (launch pricing). We watch the market and price under it: that is 90% of the lowest listed dedicated DGX Spark rental we found on 2026-08-26. An engineer scopes the configuration with you.
Dense models of roughly 7B–35B — and MoE models far larger: community benchmarks run 80B–180B MoE at usable speeds, because only the active experts are read per token. Full tables and sources on the GB10 page.
On AxForge you rent a named physical machine, hosted and operated by us in an EU region, reachable over the network like any cloud endpoint — but it is your dedicated system, not a shared cloud instance.
The GB10 platform is ARM64. Our own serving stack runs on it in production — the published numbers were measured there. x86-only binaries need ARM64 builds; an engineer can review your stack before you commit.
No. Your model, your traffic, our hardware — prompts never persisted. Only request metadata (token counts, timestamps, status) is kept for billing and operations — see the privacy policy.