Hardware

One supercomputer on solar at Erkowit — an unmanned compute node, dispatched remotely, running the models locally.

Elevation
~1,100–1,350 m
Ambient
~22 °C — free-air cooling
To Port Sudan
~90 km · ~4 h
Grid
12–18 h daily blackouts
Worst-month sun
4.5 PSH (fog)
Staffing
Unmanned — dispatched remotely

The live sizing model

Pick the box, switch loads off, and watch the plant resize. Choosing the $85K machine over the $4K one roughly triples the solar plant — that delta is the whole argument, and here it is live.

The one box

2.8 kWh

The node

4.5 kWh

Site

2.0 kWh

What to run on it

Frontier weights are closed — no hardware buys them. What runs is the open-weight tier, topping out at 71–72% SWE-bench against 80–95% closed. The decisive rule: MoE with low active parameters, never dense. Decode speed is set by memory bandwidth × active params, not by model size.

Llama 3.1 8B FP48B dense~924 tok/s @ batch 128Trivial for the box. Good for classification and enrichment
Qwen3-Coder-30B-A3B FP83B active (MoE)~483 tok/s @ batch 64The target. Low active params is exactly what this hardware wants
80B-class MoE coder~45 GB weightsfits 128 GB with KV headroomBest quality that still fits and still prefills fast
Dense Llama 3.1 70B70B dense~2.7 tok/s decodeNever. Bandwidth-bound — the box looks broken running this

The OS is not a separate decision

The Spark ships DGX OS, built on Ubuntu, with CUDA preloaded — choosing the box chooses the OS. It is right anyway: vLLM's PagedAttention is Linux + CUDA only, while mlx_lm.server keeps a per-request KV cache — exactly the wrong shape for concurrent dispatched jobs.

The roads not taken

Three reasonable alternatives, each rejected for a specific reason rather than a preference. Figures checked August 2026.

Apple — Mac Studio

rejected
96 GB
max unified memory, Aug 2026
M3 Ultra · 819 GB/s · ~200 W
128 / 256 / 512 GB options withdrawn
M5 Ultra ~768 GB reported for late 2026

Less memory than a $4K Spark, and no CUDA — so no media lane and no vLLM. Even an M5 Ultra would not re-qualify. Stays Abdout's dev machine.

Starlink Mini

watch
~720 Wh/day
against 4,320 for licensed VSAT
~17 W steady · 60 W startup peak
~40 ms vs ~600 ms GEO
Licensed in 27–28 African countries

Best technical fit by far — licensing it would delete the VSAT line item entirely. Not licensed in Sudan, no committed date. Revisit when that changes.

Tesla Powerwall 3

rejected
One box
battery and inverter, inseparable
13.5 kWh · NMC · ~5,000 cycles
~11.5 kW inverter — 4× our peak
~$998/kWh US installed vs ~$170–280 landed

Integrating the inverter destroys the 2 × 3 kW N+1 that exists because help is 4 h away. No Tesla service in Sudan, and a 13.5 kWh monolith cannot ride to Port Sudan in a car — a 5 kWh module can.

Buy the class, never the SKU

NVIDIA's roadmap fixes the current box's one weakness on a known schedule: Rubin Spark with LPDDR6 lands 2027–28, and LPDDR5X bandwidth is precisely what caps decode today. Our Tier 3 trigger plausibly fires in 2027 — so the spec reads “Spark-class, current generation at trigger time,” never a part number. Written that way, buying the box last is not only cash discipline; the plan upgrades itself while it waits.

Nobody is there

The nearest hands are four hours away, and Spark-class hardware has no BMC or IPMI to call. So recovery is built out of dumb, reliable parts.

Switched PDU

The remote power button — the last resort, pressable from anywhere

Power-on after loss

Set in BIOS, so every outage self-heals without a human

Separate management path

A second LTE modem on a different carrier reaching only the PDU. You cannot fix the link through the link

Idempotent queue

A hard power-cycle mid-job is normal here, not an incident. Every job must be safe to re-run

Tiers and triggers

Note the shape: Tier 2 is the expensive part, and none of it has an NVIDIA logo on it. The plant, the link and the ability to reach the site remotely cost four times what the computer costs.

Tier 0

Today

$0

Laptops the team already owns, plus cloud. No site.

Trigger: Current state

Tier 1

Blackout resilience

$400–900 / person

Per person in Sudan: 1 kWh power station, 200–400 W folding panel, LTE router, surge strip. Buys 8–10 working hours through an outage.

Trigger: None — affordable now, and it protects delivery today

← This is the one that waits on nothing.

Tier 2

The node, without the supercomputer

$15,000–22,000

6 kWp, 15 kWh, 2 × 3 kW hybrid, mast + LTE, service plane, switched PDU + management modem, security. CRM, Hermes, git, cache and staging move local.

Trigger: First paying school, or non-dilutive funding

Tier 3

The supercomputer

$4,000–5,000

Spark-class box on the plant Tier 2 already built. Local inference and media generation go live.

Trigger: 3 paying schools / ~$3K MRR sustained 3 months

Tier 4

Scale the box

$5,000 or ~$105,000

Second Spark linked over ConnectX-7 for 256 GB, or a DGX Station GB300 with the plant tripled to 16 kWp / 40 kWh.

Trigger: Measured saturation — the queue is waiting on the machine, not on the link

Tier 5

Customer plane migration

Incremental

Selected tenants on owned iron.

Trigger: 12 months of measured node uptime, and a customer who wants on-prem and pays for it

The full reasoning — choosing the box, protection, cooling, connectivity, bill of materials, risks and the site survey — is in the hardware doc.