Skip to content
DigitalNeuron

Chips & infrastructure

Accelerators, data centres, memory supply and the power bill behind AI.

6 articles

Analysis: self-hosting an open-weight model — the arithmetic that decides it, and the costs nobody budgets for

Only at high, steady utilisation. Self-hosting converts a variable per-token cost into a fixed hourly cost, so it wins when accelerators stay busy and loses badly when they idle. The honest comparison prices the full stack — accelerator hours, redundancy, engineering time, and the evaluation work needed to confirm the smaller model is good enough — against the API bill for the same traffic. Sovereignty, data residency and latency floors are separate reasons that can justify self-hosting regardless of the arithmetic.

6 min read

Why does AI use so much electricity?

AI accelerators draw far more power per rack than traditional servers, and that power has to be delivered, cooled and paid for continuously. Training a large model is a one-off spike; serving it to millions of users is a permanent load, and inference is what dominates energy use over a deployed model's life.

4 min read