Active

Homelab with segmented networks and high availability

A personal cloud built to be rebuilt: bare-metal provisioning over PXE, segmented networking, and storage kept outside the cluster, so the lab can be wiped and redeployed without losing data.

Stack
  • MAAS
  • Ubuntu Core
  • MicroCloud (LXD)
  • Landscape
  • Juju
  • Canonical OpenStack
  • Ceph
  • VLANs
  • PXE
Server rack with three Minisforum MS-01 systems and an ASUS Ascent.
Problem
I wanted a homelab where I could experiment and break things without worrying about how long it would take to rebuild. Reinstalling an OS across three MS-01s every time got old, and MAAS and Landscape, which manage the machines, need a place to run too.
Approach
A three-tier bootstrap chain. Three Raspberry Pi 5s, the only machines installed by hand, run MAAS and Landscape on Ubuntu Core with MicroCloud. MAAS provisions a Dell Inspiron, which hosts the Juju controller, and three Minisforum MS-01s over PXE.
Outcome
The lab runs Canonical OpenStack, hyperconverged across the three MS-01s. The x86 tier can be rebuilt over PXE without a USB stick, and the NAS stays outside Ceph, so its data survives a full redeployment.

Lessons

  • Most of the failures in this lab have started at Layer 2.
  • The L2/L3 boundary is where I have burned the most hours.
  • Most of the time, I can learn more by improving how the existing systems fit together than by buying another box.
  • The interesting work is not making the cluster bigger; it is making the existing pieces fit together better and contributing what I learn back to the projects that power it.

Next steps

  • Enough hardware and redundancy to keep the services I need available, before depending on the lab for day-to-day workloads.
  • A broader compute layer that includes Arm and AMD systems.
  • The Jetson Nano: a future project involving a small ground robot and, eventually, a drone.
  • Tier 0: the Raspberry Pis

    Three Raspberry Pi 5s run Ubuntu Core with MicroCloud. On that cluster: MAAS, Landscape Server, and Docker inside a system container.

  • Networking

    OAM and management traffic, including PXE, uses the 2.5 GbE links; Ceph replication has its own VLAN over the 10 GbE SFP+ uplinks.

  • Tier 2: the MS-01s

    Three Minisforum MS-01s run Canonical OpenStack in a hyperconverged topology, with control, compute and storage across all three nodes.

  • The DGX Spark

    The ASUS Ascent in the rack is the NVIDIA DGX Spark. It runs MicroK8s with the GPU Operator and time-slicing, and mounts the NAS over NFS.

Today the lab runs a Canonical OpenStack cluster. Three Raspberry Pi 5s, the only machines installed by hand, run MAAS, which provisions a Dell Inspiron and three Minisforum MS-01s over PXE, so the x86 tier can be rebuilt without plugging in a USB stick.

The NAS stays outside the cluster and out of Ceph, so its data survives a full redeployment. The AI/ML tier is an NVIDIA Jetson Nano and an NVIDIA DGX Spark: the ASUS Ascent in the rack, described in Sharing the DGX Spark GPU with MicroK8s.

The full write-up, including where the network has failed and what that taught me, is A Homelab Built to Be Rebuilt.