Hardware

One supercomputer on solar at Erkowit — an unmanned compute node, dispatched remotely, running the models locally.

Elevation
~1,100–1,350 m
Ambient
~22 °C — free-air cooling
To Port Sudan
~90 km · ~4 h
Grid
12–18 h daily blackouts
Worst-month sun
4.5 PSH (fog)
Staffing
Unmanned — dispatched remotely

The live sizing model

Pick the box, switch loads off, and watch the plant resize. Choosing the $85K machine over the $4K one roughly triples the solar plant — that delta is the whole argument, and here it is live.

The one box

2.8 kWh

The node

4.5 kWh

Site

2.0 kWh

What to run on it

Frontier weights are closed — no hardware buys them. What runs is the open-weight tier, topping out at 71–72% SWE-bench against 80–95% closed. The decisive rule: MoE with low active parameters, never dense. Decode speed is set by memory bandwidth × active params, not by model size.

Llama 3.1 8B FP48B dense~924 tok/s @ batch 128Trivial for the box. Good for classification and enrichment
Qwen3-Coder-30B-A3B FP83B active (MoE)~483 tok/s @ batch 64The target. Low active params is exactly what this hardware wants
80B-class MoE coder~45 GB weightsfits 128 GB with KV headroomBest quality that still fits and still prefills fast
Dense Llama 3.1 70B70B dense~2.7 tok/s decodeNever. Bandwidth-bound — the box looks broken running this

Nobody is there

The nearest hands are four hours away, and Spark-class hardware has no BMC or IPMI to call. So recovery is built out of dumb, reliable parts.

Switched PDU

The remote power button — the last resort, pressable from anywhere

Power-on after loss

Set in BIOS, so every outage self-heals without a human

Separate management path

A second LTE modem on a different carrier reaching only the PDU. You cannot fix the link through the link

Idempotent queue

A hard power-cycle mid-job is normal here, not an incident. Every job must be safe to re-run

Tiers and triggers

Note the shape: Tier 2 is the expensive part, and none of it has an NVIDIA logo on it. The plant, the link and the ability to reach the site remotely cost four times what the computer costs.

Tier 0

Today

$0

Laptops the team already owns, plus cloud. No site.

Trigger: Current state

Tier 1

Blackout resilience

$400–900 / person

Per person in Sudan: 1 kWh power station, 200–400 W folding panel, LTE router, surge strip. Buys 8–10 working hours through an outage.

Trigger: None — affordable now, and it protects delivery today

← This is the one that waits on nothing.

Tier 2

The node, without the supercomputer

$15,000–22,000

6 kWp, 15 kWh, 2 × 3 kW hybrid, mast + LTE, service plane, switched PDU + management modem, security. CRM, Hermes, git, cache and staging move local.

Trigger: First paying school, or non-dilutive funding

Tier 3

The supercomputer

$4,000–5,000

Spark-class box on the plant Tier 2 already built. Local inference and media generation go live.

Trigger: 3 paying schools / ~$3K MRR sustained 3 months

Tier 4

Scale the box

$5,000 or ~$105,000

Second Spark linked over ConnectX-7 for 256 GB, or a DGX Station GB300 with the plant tripled to 16 kWp / 40 kWh.

Trigger: Measured saturation — the queue is waiting on the machine, not on the link

Tier 5

Customer plane migration

Incremental

Selected tenants on owned iron.

Trigger: 12 months of measured node uptime, and a customer who wants on-prem and pays for it

The full reasoning — choosing the box, protection, cooling, connectivity, bill of materials, risks and the site survey — is in the hardware doc.