All notes

AI

Jul 24, 2026

Hetzner Is Building LLM Inference Infrastructure

Hetzner, the European budget cloud provider, is developing LLM inference capacity. For cost-sensitive builders currently locked into US-based inference APIs, this matters.

Hetzner is working on LLM inference. The announcement, covered by the Sliplane team, signals that one of Europe's most price-competitive hosting providers is moving into the managed inference space.

Hetzner's existing footprint is significant: dedicated servers, VPS, and object storage priced well below AWS or GCP equivalents. Their infrastructure is popular with European startups and self-hosters who need predictable costs and GDPR-friendly data residency. Adding inference to that stack is a logical extension.

The implications are straightforward. Right now, running inference at scale on European infrastructure means either self-hosting open-weight models on rented GPUs or routing traffic to US-based API providers. Neither is ideal. Self-hosting requires operational overhead. US-based APIs introduce latency, currency exposure, and data-sovereignty concerns for EU-regulated workloads.

If Hetzner ships managed inference with the same pricing discipline they apply to compute, it compresses the cost gap between running your own stack and using a hosted API. That changes the build-vs-buy calculation for solo founders and small engineering teams who have been absorbing OpenAI or Anthropic API costs as a fixed overhead.

The open questions are model selection, throughput guarantees, and whether the offering runs open-weight models, proprietary ones via partnership, or both. Hetzner's history suggests a lean toward commodity infrastructure rather than curated model marketplaces, which would likely mean open-weight support first.

For teams already on Hetzner for compute, consolidating inference onto the same provider simplifies billing and keeps traffic within a single network boundary. For teams not on Hetzner, the pricing alone may warrant a migration conversation once the service reaches general availability.

Watch for capacity details and supported model families before committing architecture decisions to it.