Building the Ultimate Home Lab: The Staged Implementation Roadmap

Building the Ultimate Home Lab: The Staged Implementation Roadmap

Building a fully containerized, secure, and monitored home lab can feel overwhelming if approached all at once. To prevent cognitive overload, this master roadmap breaks the entire journey into 5 manageable, sequential stages. Each stage builds directly upon the last, taking you from bare hardware to an enterprise-grade self-hosted ecosystem.

Interactive Navigation Map

Stage 1: Foundation & Access

Hardware prep, Ubuntu Server setup, lid-sleep fixes, and Tailscale mesh VPN.

→ Jump to Stage 1 Details

Stage 2: Containers & Core Stacks

Docker Engine, Portainer CE, Pi-hole, Plex, and Home Assistant deployments.

→ Jump to Stage 2 Details

Stage 3: Headless Desktop & UI

TigerVNC systemd automation, XFCE4 desktop ricing, and Nginx reverse proxies.

→ Jump to Stage 3 Details

Stage 4: Telemetry & Alerting

Prometheus, Grafana dashboards, cAdvisor, Scaphandre power tracking, and Telegram alerts.

→ Jump to Stage 4 Details

Stage 5: Security, SIEM & Automation

Wazuh SIEM deployment, AI-powered SOC triage agents, and Ansible Infrastructure as Code (IaC).

→ Jump to Stage 5 Details

Stage 1: Hardware Prep, OS Foundation & Secure Access

Back to: ↑ Return to Navigation Map

Goal: Get your hardware online, prevent laptops from sleeping when closed, and establish secure encrypted remote access.

  • BIOS Toggles: Enable VT-x/AMD-V virtualization, disable Secure Boot, set storage to AHCI.
  • Lid-Sleep Bypass: Edit /etc/systemd/logind.conf → set HandleLidSwitch=ignore, then sudo systemctl restart systemd-logind.
  • Tailscale Mesh VPN: Run curl -fsSL https://tailscale.com/install.sh | sh && sudo tailscale up --ssh.

Stage 2: Containerization & Core Infrastructure Stacks

Back to: ↑ Return to Navigation Map

Goal: Install Docker, Portainer, and core containerized services like Pi-hole, Plex, and Home Assistant.

  • Docker Engine: Add official Docker GPG keys, configure docker.sources, and install via apt.
  • Portainer CE: Deploy graphical container management via Docker (Port 9443).
  • Pi-hole Optimization: Cap database history to 7 days in pihole-FTL.conf (MAXDBDAYS=7) to prevent SD card I/O starvation.
  • Plex & Home Assistant: Deploy Plex with SMB/CIFS mounts and Home Assistant using network_mode: host for local device discovery.

Stage 3: Headless Desktop Environment & Web Security

Back to: ↑ Return to Navigation Map

Goal: Provide a lightweight VNC workspace and secure internal web apps behind authenticated Nginx reverse proxies.

  • TigerVNC + XFCE4: Configure non-hanging systemd service files (omitting legacy PIDFile) running at 16-bit color depth.
  • Desktop Ricing: Install Arc-Dark, Papirus icons, and Whisker Menu; disable the window compositor for maximum network speed.
  • Nginx Proxy & Basic Auth: Bind custom web apps to localhost (127.0.0.1) and wrap them in Nginx protected by .htpasswd.

Stage 4: Telemetry, Hardware Observability & Alerting

Back to: ↑ Return to Navigation Map

Goal: Collect performance metrics, track container behavior, measure power draw, and push notifications to your phone.

  • Prometheus & cAdvisor: Scrape system health metrics and monitor individual container resource consumption.
  • Scaphandre Power Profiling: Read Intel/AMD RAPL hardware registers to track exact milliwatt power draw.
  • Alertmanager & Telegram: Configure Prometheus rules and Alertmanager to push HTML-formatted push notifications to Telegram.

Stage 5: Enterprise Security, SIEM & Infrastructure as Code

Back to: ↑ Return to Navigation Map

Goal: Implement centralized security logging, AI-assisted log analysis, and automated management playbooks.

  • Wazuh SIEM Deployment: Install Wazuh Manager, Indexer, and Dashboard, and deploy endpoint security agents across your fleet.
  • AI SOC Triage: Integrate Wazuh with a LiteLLM gateway for automated threat triage and daily IT hygiene digests.
  • Ansible IaC: Transition from manual shell execution to centralized YAML playbooks to push configuration updates idempotently.

Building the Ultimate Home Lab: From Old Hardware to an Enterprise-Grade Ecosystem

What starts as a simple desire to run a network-wide ad blocker often snowballs into a full-scale IT infrastructure project. Over the course of building this home lab, the environment has evolved from a repurposed headless machine into a highly optimized, secure, and fully monitored Docker ecosystem.

This journey involved stripping away unnecessary overhead from the base OS, securing access via mesh VPNs, deploying an ultra-lightweight remote graphical interface (XFCE via TigerVNC), and containerizing an entire suite of services. To ensure everything runs smoothly, a sophisticated telemetry stack was layered on top, tracking everything from DNS I/O bottlenecks to the exact milliwatt power draw of individual background processes.

The Master Configuration Topics

If you are looking to replicate this environment, these are the core architectural pillars covered in the deployment logs:

  • Networking & Secure Access: Netplan configuration, NetworkManager routing, and Tailscale mesh VPN integration.
  • Remote Administration: TigerVNC systemd automation, D-Bus session wrappers, and XFCE4 desktop optimization for low-bandwidth links.
  • Virtualization & Containerization: Bare-metal Docker Engine deployment, Portainer CE orchestration, and PUID/PGID permission mapping for network shares.
  • Core Services: Pi-hole (DNS optimization and SQLite database capping), Plex Media Server (CIFS/USB mounts), Home Assistant, and Nginx reverse proxies.
  • Security Information and Event Management (SIEM): Wazuh server deployment, automated agent provisioning, and endpoint telemetry.
  • Observability & Power Tuning: Prometheus scraping, cAdvisor container metrics, Scaphandre CPU RAPL power profiling, and unified Grafana dashboards.

Phase 1: Transforming an Old Laptop into a Server

An old laptop is arguably the best entry-level home lab hardware. It has a built-in UPS (the battery), an integrated keyboard/monitor for troubleshooting, and is highly energy-efficient.

1. Install the Base OS

Flash a USB drive with Ubuntu Server LTS and install it on the laptop. Choose the minimal installation option to keep background processes to an absolute minimum.

2. Prevent Sleep on Lid Close

By default, laptops go to sleep when you close the lid. To keep your server running when tucked away on a shelf, modify the systemd logind configuration:

sudo nano /etc/systemd/logind.conf

Find the line #HandleLidSwitch=suspend, uncomment it, and change it to ignore:

HandleLidSwitch=ignore

Apply the changes by restarting the service:

sudo systemctl restart systemd-logind

3. Secure Networking via Tailscale

Instead of opening ports on your home router, install Tailscale to create a secure, encrypted mesh network. This allows you to SSH into your server from anywhere in the world.

curl -fsSL https://tailscale.com/install.sh | sh
sudo tailscale up --ssh

Phase 2: Deploying the Docker Engine

Containerization is the lifeblood of a modern home lab. It isolates your applications, prevents dependency conflicts, and makes backups trivial.

1. Install Official Docker

Install the official Docker repository to ensure you receive the latest engine updates:

sudo install -m 0755 -d /etc/apt/keyrings
sudo curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.asc
sudo chmod a+r /etc/apt/keyrings/docker.asc

sudo tee /etc/apt/sources.list.d/docker.sources <<EOF
Types: deb
URIs: https://download.docker.com/linux/ubuntu
Suites: $(. /etc/os-release && echo "$VERSION_CODENAME")
Components: stable
Architectures: $(dpkg --print-architecture)
Signed-By: /etc/apt/keyrings/docker.asc
EOF

sudo apt update
sudo apt install docker-ce docker-ce-cli containerd.io docker-compose-plugin -y

2. Deploy Portainer CE

Manage your containers visually by deploying Portainer.

docker volume create portainer_data
docker run -d -p 9443:9443 --name portainer --restart=always \
  -v /var/run/docker.sock:/var/run/docker.sock \
  -v portainer_data:/data \
  portainer/portainer-ce:latest

Navigate to https://<YOUR_SERVER_IP>:9443 to configure your admin profile.


Phase 3: Core Services & Automation

With the infrastructure established, you can begin deploying production stacks. Group related services together using Docker Compose.

  • Home Assistant: Centralize your home automation. Deploy this container using network_mode: host to ensure it can seamlessly discover IoT devices, smart plugs, and sensors across your local subnet.
  • Nginx Reverse Proxy: Route traffic to your various internal services. Secure internal toolkits and dashboards by configuring Nginx proxy blocks with HTTP Basic Authentication (htpasswd) to prevent unauthorized local access.
  • Immich & PostgreSQL: Build a self-hosted photo and video backup solution. Deploy the Immich machine learning and web containers backed by a dedicated PostgreSQL database container for metadata storage.
  • Wazuh SIEM: Protect your fleet. Deploy the Wazuh manager to ingest security logs, monitor endpoints over your Tailscale VPN network, and route high-severity security events directly to a Telegram bot.

Phase 4: Global Observability

A home lab is flying blind without proper telemetry. Deploying a Prometheus and Grafana stack allows you to catch hardware bottlenecks and software memory leaks before they crash your server.

1. Scrape Host & Container Metrics

Run Node Exporter on the host OS to track disk I/O, CPU wait times, and RAM saturation. Deploy cAdvisor as a container to monitor the exact CPU and memory footprint of every other Docker container running on the system. If your hardware supports it, deploy Scaphandre to pull direct RAPL wattage metrics from the CPU package.

2. Centralize with Prometheus & Grafana

Configure your prometheus.yml file to scrape these targets every 15 seconds. Connect Grafana to the Prometheus data source to build unified IT hygiene dashboards.

3. Automated Alerting

Write PromQL rules in Prometheus to detect anomalous behavior (e.g., CPU load > 90% for 5 minutes, or a container entering a crash loop). Route these rules through Alertmanager, configuring a telegram_configs receiver to instantly push HTML-formatted incident alerts to your mobile device.

Phase 5: Configuring the Headless VNC Environment

Traditional VNC setups struggle with modern hardware-accelerated desktop environments like GNOME. By pairing TigerVNC with the ultra-lightweight XFCE4 desktop, you eliminate crashes, bypass systemd cgroup conflicts, and drastically reduce idle RAM usage.

1. The Startup Script

Create the TigerVNC configuration directory and the initialization script to launch the desktop inside a dedicated D-Bus session.

mkdir -p ~/.config/tigervnc
nano ~/.config/tigervnc/xstartup

Paste this optimized execution block. The xhost +local: line is critical to prevent sandbox permission errors when launching Snap applications (like Chromium or CCTV viewers) inside the VNC window:

#!/bin/sh
# 1. Clear out host environment session leaks
unset SESSION_MANAGER
unset DBUS_SESSION_BUS_ADDRESS

# 2. Load standard desktop resources
[ -r $HOME/.Xresources ] && xrdb $HOME/.Xresources
xsetroot -solid grey
vncconfig -iconic &

# 3. Grant local sandbox permissions
export XDG_RUNTIME_DIR=/run/user/$(id -u)
xhost +local:

# 4. Launch XFCE inside a private D-Bus container loop
exec dbus-launch --exit-with-session startxfce4

Make the script executable: chmod +x ~/.config/tigervnc/xstartup

2. The Systemd Automation File

To ensure the remote desktop survives power outages, bind it to a system-level service. We use 16-bit color depth (-depth 16) to slash network transmission overhead by 33% without sacrificing readability.

sudo nano /etc/systemd/system/vncserver@.service
[Unit]
Description=Start TigerVNC server at startup
After=syslog.target network.target

[Service]
Type=forking
User=<YOUR_USERNAME>
WorkingDirectory=/home/<YOUR_USERNAME>
Environment=HOME=/home/<YOUR_USERNAME>

ExecStartPre=-/usr/bin/vncserver -kill :%i > /dev/null 2>&1
ExecStart=/usr/bin/vncserver -depth 16 -geometry 1920x1080 -localhost no :%i
ExecStop=/usr/bin/vncserver -kill :%i

[Install]
WantedBy=multi-user.target

Start and enable the service on port 5901 (Display :1):

sudo systemctl daemon-reload
sudo systemctl enable --now vncserver@1.service

Phase 6: Pi-hole Deployment & I/O Optimization

Deploying Pi-hole natively in Docker is simple, but long-term database bloat can cause severe I/O wait bottlenecks, driving the CPU Load Average up and hanging DNS resolutions.

1. The Compose File

Create the Pi-hole stack. We map the host port to 8443 to avoid conflicts with existing web servers (like Nginx) holding port 443.

services:
  pihole:
    container_name: pihole
    image: pihole/pihole:latest
    ports:
      - "53:53/tcp"
      - "53:53/udp"
      - "80:80/tcp"
      - "8443:443/tcp"
    environment:
      TZ: 'Europe/London'
      WEBPASSWORD: '<YOUR_SECURE_PASSWORD>'
      FTLCONF_dns_listeningMode: 'ALL'
    volumes:
      - './etc-pihole:/etc/pihole'
      - './etc-dnsmasq.d:/etc/dnsmasq.d'
    cap_add:
      - SYS_NICE
    restart: unless-stopped

2. Fixing the "High Load Average" Database Freeze

If your Grafana dashboard shows CPU Load Averages spiking past 4.0 while actual CPU utilization is under 50%, your storage disk is choking on the massive pihole-FTL.db history file (which defaults to storing 365 days of queries).

Restrict the database lifespan to 7 days to maintain lightning-fast I/O performance:

sudo systemctl stop pihole-FTL
sudo rm /etc/pihole/pihole-FTL.db
sudo nano /etc/pihole/pihole-FTL.conf

Add this parameter to the config file:

MAXDBDAYS=7

Restart the service to generate a fresh, optimized database: sudo systemctl start pihole-FTL

Phase 7: Plex Media Server (NAS vs USB Storage)

Plex requires host networking to broadcast discovery protocols successfully. This configuration supports securely mounting network shares directly into Docker without editing the host OS /etc/fstab.

Option A: Mounting a CIFS Network Share

Ensure your host has cifs-utils installed. Create a .env file in the same directory as your compose file containing NAS_USER=<USER> and NAS_PASS=<PASS>.

services:
  plex:
    image: lscr.io/linuxserver/plex:latest
    container_name: plex
    network_mode: host
    environment:
      - PUID=1000
      - PGID=1000
      - TZ=Europe/London
      - VERSION=docker
    volumes:
      - ./config:/config
      - nas_media:/data/nas_shared
    restart: unless-stopped

volumes:
  nas_media:
    driver: local
    driver_opts:
      type: cifs
      o: "username=${NAS_USER},password=${NAS_PASS},iocharset=utf8,vers=3.0"
      device: "//<NAS_IP_ADDRESS>/<SHARE_NAME>"

Option B: Direct USB Bind Mount

If utilizing a locally attached USB hard drive, remove the global volumes block entirely and pass the absolute hardware path.

    volumes:
      - ./config:/config
      - /mnt/usb_drive:/data/usb_media

Crucial Fix: If Plex sees the folders but no media populates, ensure the host directory has explicit read/execute permissions: sudo chmod -R 755 /mnt/usb_drive.

Phase 8: Hardware Power Tuning & Scaphandre Telemetry

Mini PCs and laptops often waste electricity polling idle hardware. By forcing PCIe deeper sleep states and mapping physical CPU sensors via Scaphandre, you can slash idle draw by up to 50%.

1. Force PCIe ASPM Sleep States

If the BIOS misreports ASPM availability, manually force the L0s and L1 deep sleep states using the AutoASPM utility.

git clone https://github.com/K0mh0li0/autoaspm.git
sudo nano /etc/systemd/system/autoaspm.service
[Unit]
Description=AutoASPM PCIe Power Saving Script
After=network.target

[Service]
Type=oneshot
ExecStart=/usr/bin/python3 /path/to/autoaspm/pkgs/autoaspm.py
RemainAfterExit=yes

[Install]
WantedBy=multi-user.target

Enable it on boot: sudo systemctl enable --now autoaspm.service

2. Deploy Scaphandre for Power Metrics

Scaphandre reads the Intel/AMD RAPL hardware registers to output exact milliwatt usage into Prometheus.

wget https://github.com/hubblo-org/scaphandre/releases/download/v1.0.2/scaphandre_v1.0.2-deb12_amd64.deb
sudo dpkg -i scaphandre_v1.0.2-deb12_amd64.deb

sudo nano /etc/systemd/system/scaphandre.service
[Unit]
Description=Scaphandre Prometheus Exporter
After=network.target

[Service]
Type=simple
ExecStart=/usr/bin/scaphandre prometheus -p 8081
Restart=always

[Install]
WantedBy=multi-user.target

Start the exporter: sudo systemctl enable --now scaphandre.service. Add <SERVER_IP>:8081 to your prometheus.yml scrape configs.

Phase 9: Real-Time Alerts via Telegram

When an internal service fails or your CPU maxes out, Alertmanager receives the alert from Prometheus and pushes a formatted HTML payload directly to a Telegram bot.

1. Configure Alertmanager

Create the Telegram receiver configuration, passing the full bot API token and the negative integer Chat ID (for group chats).

sudo nano /etc/alertmanager/alertmanager.yml
route:
  group_by: ['alertname', 'instance']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 3h
  receiver: 'telegram-bot'

receivers:
- name: 'telegram-bot'
  telegram_configs:
    - send_resolved: true
      bot_token: '<YOUR_BOT_TOKEN_ID:SECRET_STRING>'
      chat_id: -<YOUR_GROUP_CHAT_ID>
      parse_mode: 'HTML'
      message: |-
        {{ if eq .Status "firing" }} 🔥 FIRING {{ else }} ✅ RESOLVED {{ end }}: {{ .CommonLabels.alertname }}
        Host: {{ .CommonLabels.instance }}
        Details: {{ .CommonAnnotations.description }}

2. Link Prometheus to Alertmanager

Update prometheus.yml to forward alerts to the Alertmanager port (9093).

alerting:
  alertmanagers:
    - static_configs:
        - targets: ['localhost:9093']

Restart the stack: sudo systemctl restart prometheus alertmanager. You now possess an enterprise-grade monitoring, orchestration, and remote-management fleet configured entirely from scratch.

Phase 10: Enterprise SIEM Deployment (Wazuh)

Wazuh provides unified XDR and SIEM protection, aggregating logs and endpoint telemetry. Deploying the manager and agents properly requires strict repository configurations.

1. Install the Wazuh Manager

Use the official automated script to deploy the Wazuh Indexer, Server, and Dashboard on a single node.

curl -sO https://packages.wazuh.com/4.x/wazuh-install.sh
sudo bash wazuh-install.sh -a

Once finished, securely store the generated admin password and access the web interface at https://<YOUR_SERVER_IP>.

2. Deploy the Wazuh Agent (Fixing Repo Errors)

If you encounter 404 Not Found errors when updating apt, it indicates a corrupted repository file pointing to the main marketing site instead of the package mirror. Clean it up and install the agent:

# Remove corrupted lists
sudo rm -f /etc/apt/sources.list.d/wazuh.list
sudo rm -rf /var/lib/apt/lists/*wazuh*

# Import the official cryptographic trust key
curl -s https://packages.wazuh.com/key/GPG-KEY-WAZUH | gpg --no-default-keyring --keyring gnupg-ring:/usr/share/keyrings/wazuh.gpg --import
sudo chmod 644 /usr/share/keyrings/wazuh.gpg

# Add the verified repository URL
echo "deb [signed-by=/usr/share/keyrings/wazuh.gpg] https://packages.wazuh.com/4.x/apt/ stable main" | sudo tee /etc/apt/sources.list.d/wazuh.list

# Update and install the agent, passing the manager IP
sudo apt-get update
export WAZUH_MANAGER="<YOUR_WAZUH_MANAGER_IP>"
sudo -E apt-get install wazuh-agent -y

Enable and start the agent: sudo systemctl enable --now wazuh-agent

3. Configure Grafana OpenSearch Data Source

Because the Wazuh Indexer is built on OpenSearch architecture, you can visualize security telemetry directly in Grafana. Install the OpenSearch plugin via the Grafana CLI:

grafana-cli plugins install grafana-opensearch-datasource
sudo systemctl restart grafana-server

In Grafana, add an OpenSearch data source pointing to https://<YOUR_INDEXER_IP>:9200. Enable Basic Auth with your cluster credentials and toggle Skip TLS Verify if using default self-signed certificates.

Phase 11: Network Diagnostics & Hardware Firmware

A home lab is only as stable as its underlying network and hardware firmware. When DNS lags or hardware exhibits vulnerabilities, use these terminal-native tools to diagnose and patch the system.

1. Automated DNS Performance Benchmarking

Slow DNS lookups introduce silent latencies that mirror hard network drops. Create dns_audit.sh to automatically benchmark your system's default resolver against public Anycast infrastructure (like Cloudflare or Google).

#!/bin/bash
# Ensure environment has diagnostic utilities
if ! command -v dig &> /dev/null; then
    sudo apt update -qq && sudo apt install dnsutils -y -qq
fi

clear
echo "==================================="
echo " PHASE 1: ACTIVE SYSTEM RESOLVER"
echo "==================================="
resolvectl status | grep "DNS Servers" -A 2

echo -e "\n==================================="
echo " PHASE 2: JITTER ASSESSMENT"
echo "==================================="
for i in {1..5}; do
    TIME_MS=$(dig google.com | grep "Query time" | awk '{print $4}')
    echo "-> Query Iteration $i: ${TIME_MS} ms"
    sleep 0.5
done

echo -e "\n==================================="
echo " PHASE 3: PUBLIC RESOLVER BENCHMARK"
echo "==================================="
echo "System Default Link: $(dig example.com | grep "Query time" | awk '{print $4, $5}')"
echo "Cloudflare (1.1.1.1): $(dig example.com @1.1.1.1 | grep "Query time" | awk '{print $4, $5}')"
echo "Google (8.8.8.8): $(dig example.com @8.8.8.8 | grep "Query time" | awk '{print $4, $5}')"

Run the script: chmod +x dns_audit.sh && ./dns_audit.sh

2. Terminal Wi-Fi Management (nmcli)

Manage headless Wi-Fi connections via the NetworkManager Command Line Interface:

  • Scan for Networks: nmcli device wifi list
  • Connect to a Network: nmcli device wifi connect "<SSID>" password "<PASSWORD>"
  • Visual Interface Alternative: Run sudo nmtui for a curses-based terminal GUI. (Note: If you get a Polkit "not authorized" error over SSH, you must create a custom Polkit rule in /etc/polkit-1/rules.d/).

3. Securing System Firmware (fwupdmgr)

Linux manages motherboard BIOS and UEFI dbx updates securely using the Firmware Update Daemon (fwupd). Apply critical security revocation lists to prevent Secure Boot bypass exploits.

# See available updates
fwupdmgr get-upgrades

# Fetch latest signatures from Linux Vendor Firmware Service
fwupdmgr refresh

# Apply the updates (requires confirmation and reboot)
fwupdmgr update

Phase 12: Hardware Auditing & Custom Power Automation

When repurposing consumer hardware like an All-In-One PC or old laptop, physical power management is paramount. Integrated screens that cannot be physically detached will waste massive amounts of electricity if left illuminated 24/7.

1. Hardware Profiling over SSH

Identify your bare-metal architecture without needing physical access to the chassis:

  • Motherboard & BIOS Details: sudo dmidecode -t baseboard
  • CPU Topology & Capabilities: lscpu
  • Comprehensive Hardware Tree: sudo lshw -short (or use inxi -Fz for a colorized summary that anonymizes MAC addresses).

2. Automated Display Power Script

For All-in-One machines or laptops, the physical screen backlight must be overridden at the kernel level via sysfs. Save the following script as setup_screen.sh to automatically detect the backlight interface, inject manual bash aliases, and deploy root-level cron jobs to shut the screen off overnight.

#!/bin/bash
# ALL-IN-ONE UBUNTU SERVER DISPLAY CONFIGURATOR

if [ "$EUID" -ne 0 ]; then
    echo "[-] Please run this script with sudo."
    exit 1
fi

echo "[+] Starting automated display configuration..."

# 1. Detect Backlight Interface
BACKLIGHT_DIR="/sys/class/backlight"
INTERFACE=$(ls "$BACKLIGHT_DIR" | head -n 1)

if [ -z "$INTERFACE" ]; then
    echo "[-] Error: No backlight controller found."
    exit 1
fi

BACKLIGHT_PATH="$BACKLIGHT_DIR/$INTERFACE"
MAX_BRIGHTNESS=$(cat "$BACKLIGHT_PATH/max_brightness")
echo "[+] Detected interface: $INTERFACE (Max: $MAX_BRIGHTNESS)"

# 2. Inject Bash Aliases for Manual Control
REAL_USER=${SUDO_USER:-$USER}
BASHRC_PATH="$(eval echo ~$REAL_USER)/.bashrc"

sed -i '/alias screenoff=/d' "$BASHRC_PATH"
sed -i '/alias screenon=/d' "$BASHRC_PATH"

echo "alias screenoff='echo 0 | sudo tee $BACKLIGHT_PATH/brightness'" >> "$BASHRC_PATH"
echo "alias screenon='echo $MAX_BRIGHTNESS | sudo tee $BACKLIGHT_PATH/brightness'" >> "$BASHRC_PATH"
chown "$REAL_USER":"$REAL_USER" "$BASHRC_PATH"

# 3. Configure Console Idle Blanking
setterm --blank 1 --term linux < /dev/tty1 2>/dev/null || true

# 4. Install Root Cron Jobs for Automation (Off at 22:00, On at 07:00)
TMP_CRON=$(mktemp)
crontab -l > "$TMP_CRON" 2>/dev/null || true
sed -i '\#/sys/class/backlight/#d' "$TMP_CRON"

echo "0 22 * * * echo 0 > $BACKLIGHT_PATH/brightness" >> "$TMP_CRON"
echo "0 7 * * * cat $BACKLIGHT_PATH/max_brightness > $BACKLIGHT_PATH/brightness" >> "$TMP_CRON"

crontab "$TMP_CRON"
rm "$TMP_CRON"

echo "[+] SETUP COMPLETE! Run 'source ~/.bashrc' to activate aliases."

Make it executable and deploy it:

chmod +x setup_screen.sh
sudo ./setup_screen.sh
source ~/.bashrc

Phase 13: Identity Management & Homepage Dashboard

When containers write data to your host machine's file system, they often do so as the root user by default. This causes severe "Permission Denied" (EACCES) errors when your normal user account tries to edit configuration files. To fix this, you must map your Process User ID (PUID) and Process Group ID (PGID).

1. Find Your PUID and PGID

Run the id command in your host terminal to output your user metrics:

id -u
id -g

Note these numbers (typically 1000). You will inject them into your Docker Compose files as environment variables.

2. Deploying Homepage (Resolving EACCES Path Errors)

Homepage is a highly customizable starting screen for your lab. A common deployment mistake is mapping the local volume to /config internally, which causes permission failures. The correct internal path for the Homepage container is /app/config.

services:
  homepage:
    image: ghcr.io/gethomepage/homepage:latest
    container_name: homepage
    ports:
      - "3001:3000"
    volumes:
      # THE FIX: Target /app/config internally
      - ./homepage/config:/app/config
    restart: always
    environment:
      - PUID=1000
      - PGID=1000
      - HOMEPAGE_ALLOWED_HOSTS=localhost:3001,127.0.0.1:3001,<YOUR_SERVER_IP>:3001

If you already encountered an error, ensure your local user reclaims ownership of the created folder before spinning the container back up:

sudo chown -R 1000:1000 ./homepage
chmod -R 755 ./homepage
docker compose up -d

Phase 14: Deploying Home Assistant

Home Assistant is unique because it relies heavily on local network broadcasts (mDNS, UPnP) to automatically discover smart bulbs, TVs, and IoT bridges. Standard Docker bridge networks block these broadcasts. The solution is network_mode: host.

The Optimized Compose Configuration

Create your docker-compose.yml for Home Assistant. Note the absence of a ports: block, as host mode binds directly to your server's NIC on port 8123.

services:
  homeassistant:
    container_name: homeassistant
    image: ghcr.io/home-assistant/home-assistant:stable
    volumes:
      - ./config:/config
      # Mount hardware timezone file for accurate sun/schedule automations
      - /etc/localtime:/etc/localtime:ro
    environment:
      - TZ=Asia/Kolkata
    network_mode: host
    restart: unless-stopped

Start the container and navigate to http://<YOUR_SERVER_IP>:8123 to begin the onboarding wizard.

Phase 15: Advanced Server Triage & Automated Power Recovery

When your server slows down but standard htop shows low CPU usage, you are likely experiencing I/O Wait or Network exhaustion. Equip your terminal with modern diagnostic alternatives.

1. The Modern Terminal Toolkit

  • btop: The most visually stunning and responsive TUI for macro-level system metrics (CPU, RAM, Disks, Network). sudo apt install btop
  • iotop & iostat: Crucial for identifying disk saturation. If iostat -xz 1 shows high w_await times and 100% %util, your mechanical drive is drowning in random writes. Run sudo iotop -o to expose the exact process hogging the disk.
  • iftop & nload: Use sudo nload for a macro-level speedometer of your inbound/outbound traffic, and sudo iftop to identify the specific IP addresses consuming your bandwidth.

2. Automating AC Power Recovery (BIOS Level)

If your home experiences a power outage, your headless server will remain off when power returns unless instructed otherwise by the motherboard. You cannot configure this from Linux; it must be done in the BIOS/UEFI.

  1. Reboot the server and tap Del or F2 to enter the BIOS (e.g., MSI Click BIOS 5).
  2. Switch to Advanced Mode (F7).
  3. Navigate to Settings > Advanced > Power Management Setup.
  4. Locate Restore after AC Power Loss and change it from Power Off to Power On.

Now, the moment your smart plug or UPS restores power, the server will instantly pull up the Ubuntu bootloader without physical intervention.

Phase 16: Tuning Grafana Variables for cgroups v2

When utilizing modern Linux environments running cgroups v2, cAdvisor exports metrics slightly differently than legacy systems. If you import popular Grafana dashboards (like ID 14282 or 193) and encounter "No Data", it is because the dashboard's internal PromQL queries are looking for a name= label, while modern cAdvisor maps container endpoints to the id= label (e.g., /system.slice/docker.service).

1. Fixing Dropdown Variable Queries

Navigate to Dashboard Settings > Variables and update the Classic Query expressions to pull from the correct system slices.

  • Host Variable: label_values(container_cpu_usage_seconds_total, instance)
  • Container Variable: Filter out raw background processes by enforcing the id tag mapping:
    label_values(container_cpu_usage_seconds_total{instance=~"$host", id=~"/system.slice/.*"}, id)

2. Updating Panel Metric Queries

Edit your individual visualization panels to utilize the id grouping so that each line correlates cleanly to a running service.

# CPU Usage Query Fix
sum(rate(container_cpu_usage_seconds_total{instance=~"$host", id=~"$container"}[5m])) by (id) * 100

# Memory Working Set Query Fix
container_memory_working_set_bytes{instance=~"$host", id=~"$container"}

3. Recommended Fleet Dashboards

To avoid spending hours building dashboards from scratch, import these industry-standard IDs via Dashboards > New > Import:

  • ID 1860 (Node Exporter Full): The ultimate single-node deep dive for CPU, RAM, Disk, and Network hardware metrics.
  • ID 15172 (Fleet Overview): A consolidated scoreboard design allowing you to monitor multiple nodes (Proxmox, Raspberry Pi, Ubuntu servers) side-by-side.
  • ID 14282 (cAdvisor Exporter): Cleanly visualizes isolated Docker container performance footprint.

You have successfully traversed the journey from flashing a base OS to commanding a multi-node, monitored, and automated home lab ecosystem. Happy self-hosting!

Phase 17: Fixing Linux LAN Drivers & APT Repositories

When working with older consumer hardware or adding third-party applications, you will inevitably run into driver conflicts and GPG key expirations. Here is how to navigate the most common roadblocks.

1. The Realtek LAN link=no Bug

If your Ethernet cable is plugged in, but running sudo lshw -C network or ethtool shows link=no, you have likely hit a known kernel conflict with Realtek RTL8111/8168 chips. The default open-source r8169 driver frequently fails to negotiate a link. The fix is to compile the official proprietary driver:

sudo apt update
sudo apt install r8168-dkms -y
sudo reboot

2. Fixing GPG NO_PUBKEY and Duplicate Sources

If running apt update throws errors like "configured multiple times" or "NO_PUBKEY" (common with Brave Browser, Docker, or Wazuh), your sources list is corrupted. Clean the slate and re-import the secure token.

# 1. Delete the corrupted files
sudo rm -f /etc/apt/sources.list.d/brave-browser-release.list
sudo rm -f /etc/apt/sources.list.d/brave-browser-release.sources
sudo rm -f /usr/share/keyrings/brave-browser-archive-keyring.gpg

# 2. Re-download the official key
sudo curl -fsSLo /usr/share/keyrings/brave-browser-archive-keyring.gpg https://brave-browser-apt-release.s3.brave.com/brave-browser-archive-keyring.gpg

# 3. Add the clean source
echo "deb [signed-by=/usr/share/keyrings/brave-browser-archive-keyring.gpg] https://brave-browser-apt-release.s3.brave.com/ stable main" | sudo tee /etc/apt/sources.list.d/brave-browser-release.list

sudo apt update

3. Snap Packages Crashing in VNC (Signal 6)

Ubuntu installs Chromium via Snap. Snaps enforce a strict sandbox that requires applications to run inside a user cgroup slice. Because our VNC server runs as a system service, launching a Snap app results in a sudden crash. Bypass this by exporting the runtime directory before execution:

export XDG_RUNTIME_DIR=/run/user/$(id -u)
chromium-browser &

Alternative: Ditch the Snap entirely and download the raw .deb package (e.g., Google Chrome) which ignores cgroup restrictions.

Phase 18: Taming VM Insomnia on Proxmox

If you transition your physical setup into a virtualized Proxmox environment, you might notice VMs pulling 10W-15W at "idle." This is known as VM Insomnia—background polling forces the physical CPU to wake from deep C-states thousands of times per second. Apply these 4 tweaks to slash idle draw.

1. Strip Unnecessary Virtual Hardware

When you create a VM, Proxmox attaches virtual audio controllers and USB hubs. The guest kernel continuously polls these empty devices, burning CPU cycles.

  • In the Proxmox UI, go to the VM's Hardware tab.
  • Remove the Audio Device and USB Controllers (unless explicitly needed).
  • Change the Display to Serial terminal 0 if managing entirely via SSH/Tailscale.

2. Tame Docker Container Polling

Tools like Wazuh and NetAlertX are aggressive network scanners. If they ping the subnet every 5 seconds, the CPU never sleeps.

  • NetAlertX: Increase the subnet scan interval in the Web UI from 2 minutes to 15+ minutes.
  • Wazuh: Open /var/ossec/etc/ossec.conf and change the <syscheck> file integrity monitoring frequency to 12 or 24 hours (43200 or 86400 seconds).

3. Apply PowerTOP Auto-Tune

Because you passed the physical CPU through to the VM (qm set <VM_ID> --cpu host), the guest OS can manage its virtual power states. Install and run PowerTOP inside the guest:

sudo apt update && sudo apt install powertop -y
sudo powertop --auto-tune

(Add @reboot /usr/sbin/powertop --auto-tune to the root crontab to make this persistent).

4. Enable QEMU Guest Agent

The Guest Agent allows the Proxmox host and Ubuntu guest to negotiate power states efficiently. Enable it in the Proxmox VM Options tab, then install it inside the guest:

sudo apt install qemu-guest-agent -y
sudo systemctl enable --now qemu-guest-agent

Conclusion: Your Home Lab is Now Production-Ready

By following this comprehensive series, you have successfully transformed aging hardware into a secure, power-efficient, and highly observable enterprise-grade environment. From fighting APT repository errors and resolving VNC permission sandboxes, to tracking millimeter-level power consumption with Scaphandre and Grafana, your infrastructure is now optimized from the bare metal up to the application layer.

Happy Self-Hosting!

Phase 19: XFCE "Ricing" (Aesthetics & Speed Optimization)

Out of the box, XFCE can look like it belongs in 2004, and its default animations can cripple a VNC network stream[cite: 156, 160]. This blueprint transforms the interface into a sleek, flat, modern workspace while slashing network latency.

1. The Speed Layer (Optimize the Network Pipe)

To make VNC snappy, minimize the number of pixels changing during animations[cite: 156].

  • Kill the Compositor: Open Settings > Window Manager Tweaks > Compositor and uncheck Enable display compositing[cite: 156, 160]. This stops VNC from wasting bandwidth rendering drop shadows and window fades[cite: 156, 160].
  • Solid Wallpaper: High-resolution wallpapers require heavy compression[cite: 156]. Right-click the desktop, select Desktop Settings, and change the Style to None (Solid color) with a dark grey hex code[cite: 156, 160].

2. The Aesthetics Layer (Modernizing the UI)

Install the modern visual assets directly from the terminal[cite: 156, 160]:

sudo apt update
sudo apt install arc-theme papirus-icon-theme xfce4-whiskermenu-plugin -y

Apply the new skin via the XFCE Settings menu[cite: 156, 160]:

  • Main Theme: Settings > Appearance > Style → Arc-Dark[cite: 156, 160].
  • Icons: Settings > Appearance > Icons → Papirus-Dark[cite: 156, 160].
  • Window Borders: Settings > Window Manager > Style → Arc-Dark[cite: 156, 160].
  • Font Smoothing: Settings > Appearance > Fonts → Check Enable Anti-Aliasing (Hinting: Slight, Sub-pixel: RGB)[cite: 156, 160].

3. Panel & Whisker Menu Upgrade

Replace the basic application menu with the Whisker Menu (a ChromeOS/Windows-style launcher)[cite: 156].

  1. Right-click the old panel menu button and select Remove[cite: 156, 160].
  2. Right-click the panel → Panel > Add New Items → Select Whisker Menu[cite: 156, 160].
  3. Move it to the far left. Open Panel Preferences, uncheck Lock panel, drag the taskbar to the bottom of the screen, and increase the Row Size to 38px[cite: 156, 160].

Phase 20: Handling Video Tearing in VNC

Disabling the compositor creates a blazing-fast remote desktop, but it completely breaks Vertical Synchronization (VSync)[cite: 157]. If you attempt to watch a video over VNC, the player will drop frames asynchronously, causing the video to split or "tear" horizontally across the screen[cite: 157].

If you explicitly require video playback on a headless server, you must trade some raw performance for rendering accuracy[cite: 157].

1. Force the Xpresent Blanking Protocol

Re-enable the compositor, but force XFCE to use a low-overhead sync method[cite: 157]:

# Run this in your VNC terminal
xfconf-query -c xfwm4 -p /general/vblank_mode -s xpresent

Then, go back to Window Manager Tweaks > Compositor and re-check Enable display compositing[cite: 157].

2. Disable Player Hardware Acceleration

Hardware acceleration (OpenGL/VA-API) crashes over virtual framebuffers[cite: 157]. You must force your media player to use raw software rendering[cite: 157].

  • VLC Player: Go to Tools > Preferences > Video. Change the Output dropdown from Automatic to X11 video output (XCB)[cite: 157].
  • MPV Player: Launch videos from the terminal using: mpv --vo=x11 --hwdec=no video.mp4[cite: 157].

Note: If server performance is your absolute priority and you do not intend to watch media over the remote link, skip these steps and leave the compositor disabled entirely[cite: 158].

Phase 21: The Modern Sysadmin Terminal Toolkit

When your server feels sluggish but standard htop shows plenty of free CPU and RAM, you are likely suffering from silent disk I/O saturation or bandwidth exhaustion[cite: 116]. Upgrade your terminal triage workflow with these modern diagnostic tools[cite: 115, 116].

1. Modern Resource Monitors (htop alternatives)

Move beyond standard process lists with tools that provide macro-level server health[cite: 115]:

  • btop: The current community favorite. A gorgeous, mouse-responsive TUI with real-time graphs for cores, disks, and networking[cite: 115]. (sudo apt install btop)[cite: 115].
  • glances: A dense Python-based dashboard that fits disk I/O speeds, hardware temperatures, and system alerts on a single screen[cite: 115].
  • atop: A historical troubleshooting tool. It runs a background daemon logging performance every 10 minutes, allowing you to "rewind time" to see what crashed your server overnight[cite: 115].

2. Isolating Disk I/O Bottlenecks

If processes are stuck in "Uninterruptible Sleep," your storage is saturated[cite: 116].

  • iostat: Run iostat -xz 1 to monitor hardware saturation[cite: 116]. If %util is near 100% and w_await (write latency) is > 50ms, the physical drive is failing to keep up with write requests[cite: 116, 117].
  • iotop: Run sudo iotop -o to see a live list of processes actively reading/writing to the disk right now, allowing you to kill the specific app hammering your storage[cite: 116].

3. Catching Network Leaks

If web applications lag but hardware resources are fine, check your network pipe[cite: 116].

  • nload: Provides two simple, real-time ASCII graphs for total incoming and outgoing traffic to verify if you are hitting your ISP's bandwidth ceiling[cite: 116].
  • iftop: Run sudo iftop on your main interface (or Tailscale interface) to display a live table of IP addresses, instantly revealing which remote host is pulling a massive data stream from your server[cite: 116].

Phase 22: Upgrading to the Prometheus React UI

While Grafana is your primary dashboarding tool, writing complex PromQL queries directly inside Grafana can be tedious. The modern Prometheus React UI includes syntax highlighting, auto-complete, and a native dark mode, making it the perfect scratchpad for testing telemetry math before building your final panels[cite: 198, 199].

1. Install the React UI Assets

If you installed Prometheus via standard package managers (like APT on Debian/Ubuntu), it often defaults to the classic HTML interface. You can bypass this restriction using the built-in helper script[cite: 198]:

# Run the Debian helper script to pull the React assets
sudo /usr/share/prometheus/install-ui.sh

# Restart the service to load the new web application
sudo systemctl restart prometheus

2. Testing PromQL Queries

Navigate to http://<YOUR_PROMETHEUS_IP>:9090/graph. You will be greeted by the new React interface[cite: 199].

  • Use Auto-Complete: Start typing a metric (e.g., scaph_) and the UI will drop down all available targets[cite: 199].
  • Filter by Labels: Narrow down global queries to a specific instance before porting them to Grafana: scaph_host_power_microwatts{instance="localhost:8081"} / 1000000[cite: 199].
  • Table vs. Graph: Use the Table tab to verify current instantaneous values and label formatting, or the Graph tab to visualize time-series integrity over the last hour[cite: 199].

Phase 23: Managing the ARP Cache (ip neighbor)

If you recently changed a device's IP address on your home network but your server stubbornly refuses to connect to it, you likely have an ARP (Address Resolution Protocol) cache conflict[cite: 110]. The Linux kernel maps local IP addresses to physical MAC addresses; if a device switches IPs, the old mapping becomes stale[cite: 110].

1. Reading the Neighbor Table

To view your system's current IP-to-MAC mappings, use the modern ip neighbor tool (which replaces the legacy arp -a command)[cite: 110]:

ip n

You will see an output like:
192.168.1.1 dev enp4s0 lladdr 70:5a:xx:xx:xx REACHABLE[cite: 110].

  • REACHABLE: The mapping is valid and recently verified[cite: 110].
  • STALE: The mapping is cached but hasn't been verified in a few minutes[cite: 110].
  • FAILED: The device did not respond (likely offline or IP changed)[cite: 110].

2. Flushing the ARP Cache

If a mapping is corrupted or outdated, force the Linux kernel to drop the entire cache and re-discover the network hardware dynamically[cite: 110]:

sudo ip neighbor flush all

If you want to manually assign a permanent static hardware route to prevent spoofing, use:
sudo ip neighbor add 192.168.1.50 lladdr 00:11:22:33:44:55 dev eth0[cite: 110].

Phase 24: Resolving VNC Desktop Crashes (Signal 6 & Polkit)

When experimenting with different desktop environments (GNOME, LXDE, XFCE) over VNC, it is easy to accidentally cross-contaminate your display manager configurations, resulting in instant grey screens or Signal 6 crashes[cite: 101, 124].

1. The "No Session for PID" Error

If you see a Polkit/D-Bus permissions error like no session for pid, it means multiple desktop components are fighting for control of the same VNC screen[cite: 101]. This usually happens if you run a global /etc/X11/Xsession script alongside a dedicated startxfce4 call[cite: 101].

The Fix: Purge conflicting desktop packages and isolate XFCE[cite: 102]:

# Remove conflicting LXDE or XRDP packages
sudo apt purge lxde-core lxpanel lxterminal xrdp -y
sudo apt autoremove --purge -y

# Reinstall XFCE securely to repair any clipped dependencies
sudo apt install --reinstall xfce4 xfce4-goodies -y

2. The GNOME "Signal 6" Crash

If you attempt to run the modern Ubuntu GNOME shell over a VNC system service, it will almost always fail with The X session died with signal 6![cite: 124, 130]. Modern GNOME requires two things that headless VNC servers lack:

  • Hardware 3D Acceleration: GNOME panels and animations require physical GPU hardware. Over a virtual network frame, it panics and aborts[cite: 130].
  • Systemd User Slices: GNOME relies on systemd --user sessions to load its daemons, which are inaccessible when VNC is run as a global root/system service[cite: 124].

The Solution: Do not use full GNOME for remote headless servers. Fall back to XFCE, which handles 2D software rendering flawlessly and operates perfectly outside of strict systemd user slices[cite: 131, 132]. Ensure your ~/.config/tigervnc/xstartup file contains exec dbus-launch --exit-with-session startxfce4 to properly map the X11 sockets[cite: 131].

Epilogue: The Resilient Home Lab

Building a home lab is an ongoing process of iteration. What began with simple Wi-Fi commands and an old laptop has scaled into a secure, power-tuned, and fully observable DevOps environment.

By replacing bloated GUIs with lightweight XFCE, securing network paths with Tailscale, containerizing services via Docker, and keeping a watchful eye over hardware with Prometheus and Scaphandre, you have engineered an infrastructure that is both remarkably robust and exceptionally efficient.

Stay Curious. Keep Building.

Appendix: The Master Home Lab Navigation Index

To help you navigate this extensive deployment guide or easily return to a specific troubleshooting section, here is the complete master index of all phases covered in this home lab series.

Core Infrastructure

  • Phase 1: Base OS, NetworkManager, and Secure Access (Tailscale)
  • Phase 2: Deploying the Docker Engine & Portainer CE
  • Phase 3: Core Services Overview & Automation Planning
  • Phase 4: Global Observability (Prometheus, Grafana, cAdvisor)

Headless Environments & Access

  • Phase 5: Configuring the Headless VNC Environment (TigerVNC + systemd)
  • Phase 19: XFCE "Ricing" (Aesthetics & Speed Optimization)
  • Phase 20: Handling Video Tearing in VNC (Compositor & VSync tuning)
  • Phase 24: Resolving VNC Desktop Crashes (Signal 6 & Polkit logic)

Container Deployments

  • Phase 6: Pi-hole Deployment & I/O Optimization (Fixing Database Bloat)
  • Phase 7: Plex Media Server (CIFS NAS vs. USB Storage Mounts)
  • Phase 13: Identity Management (PUID/PGID) & Homepage Dashboard
  • Phase 14: Deploying Home Assistant (Host Networking)

Telemetry, Security & Diagnostics

  • Phase 8: Hardware Power Tuning & Scaphandre Telemetry
  • Phase 9: Real-Time Alerts via Telegram (Alertmanager)
  • Phase 10: Enterprise SIEM Deployment (Wazuh & OpenSearch)
  • Phase 11: Network Diagnostics & Hardware Firmware (DNS, fwupdmgr)
  • Phase 12: Hardware Auditing & Custom Power Automation (setup_screen.sh)
  • Phase 15: Advanced Server Triage & Automated Power Recovery (BIOS, btop)
  • Phase 16: Tuning Grafana Variables for cgroups v2
  • Phase 17: Fixing Linux LAN Drivers & APT Repositories
  • Phase 18: Taming VM Insomnia on Proxmox
  • Phase 21: The Modern Sysadmin Terminal Toolkit
  • Phase 22: Upgrading to the Prometheus React UI
  • Phase 23: Managing the ARP Cache (ip neighbor)

Where to Go Next: Future Home Lab Expansions

Once you have implemented the entire 24-phase blueprint, your infrastructure will be running like a well-oiled, enterprise-grade machine. But the beauty of a home lab is that the learning never stops. Here are the logical next steps to continue expanding your ecosystem.

1. Reverse Proxies & SSL Certificates (Traefik / Nginx Proxy Manager)

If you want to access your services using clean domain names (e.g., plex.yourdomain.com) instead of IP addresses and port numbers, deploying a reverse proxy is the next critical step. Tools like Traefik or Nginx Proxy Manager integrate directly with Docker to route traffic and automatically provision free SSL certificates via Let's Encrypt.

2. Automated Backup Pipelines

With persistent volumes storing your Plex metadata, Pi-hole configurations, and Prometheus telemetry, establishing a 3-2-1 backup strategy is vital. Consider deploying Duplicati or BorgBackup containers to automatically encrypt and push your /docker directory to an offsite cloud bucket (like AWS S3 or Backblaze B2) nightly.

3. Transitioning to Infrastructure as Code (IaC)

You have already taken the first steps into IaC by writing docker-compose.yml files. The next evolution is managing your entire host configuration using tools like Ansible. With Ansible, you can write a playbook that automatically installs NetworkManager, deploys Tailscale, configures TigerVNC, and applies all your power-tuning scripts across multiple nodes with a single command.


This concludes the complete home lab deployment and troubleshooting series. The configurations provided here will serve as a stable foundation for years of self-hosted exploration.

Appendix A: Complete Automation & Diagnostic Scripts

For ease of deployment, here are the full, copy-pasteable Bash scripts referenced throughout the series. All sensitive identifiers have been replaced with generic placeholders.

1. Automated DNS Auditor (dns_audit.sh)

Use this script to benchmark your local system resolver against global Anycast networks to diagnose silent web latency.

#!/bin/bash
# =================================================================================
# Linux DNS Performance & Diagnostic Benchmarking Tool
# =================================================================================

if ! command -v dig &> /dev/null; then
    echo "[!] Missing dependency 'dnsutils'. Attempting install..."
    sudo apt update -qq && sudo apt install dnsutils -y -qq
fi

clear
echo "==================================="
echo " PHASE 1: ACTIVE SYSTEM RESOLVER"
echo "==================================="
if command -v resolvectl &> /dev/null; then
    ACTIVE_DNS=$(resolvectl status | grep "DNS Servers" -A 2)
    if [ -n "$ACTIVE_DNS" ]; then
        echo "$ACTIVE_DNS"
    else
        echo "No upstream records found in resolvectl status."
    fi
else
    echo "Fallback Mode (/etc/resolv.conf):"
    grep nameserver /etc/resolv.conf
fi

echo -e "\n==================================="
echo " PHASE 2: BASELINE RESPONSE MATRIX"
echo "==================================="
dig ubuntu.com | grep -E "Query time | SERVER"

echo -e "\n==================================="
echo " PHASE 3: STABILITY (5-CYCLE TEST)"
echo "==================================="
for i in {1..5}; do
    TIME_MS=$(dig google.com | grep "Query time" | awk '{print $4}')
    echo "-> Query Iteration $i: ${TIME_MS} ms"
    sleep 0.5
done

echo -e "\n==================================="
echo " PHASE 4: ANYCAST PUBLIC BENCHMARK"
echo "==================================="
echo "System Default: $(dig example.com | grep "Query time" | awk '{print $4, $5}')"
echo "Cloudflare (1.1.1.1): $(dig example.com @1.1.1.1 | grep "Query time" | awk '{print $4, $5}')"
echo "Google (8.8.8.8): $(dig example.com @8.8.8.8 | grep "Query time" | awk '{print $4, $5}')"
echo "Quad9 (9.9.9.9): $(dig example.com @9.9.9.9 | grep "Query time" | awk '{print $4, $5}')"
echo "==================================="
echo "Optimization Rule: If public platforms outpace your System Link by >20ms, update your interface permanently."

2. All-In-One Display Controller (setup_screen.sh)

For repurposed laptops or All-In-One PCs, this script detects the system's backlight interface, injects screenon and screenoff terminal shortcuts, and sets up cron jobs to power down the monitor at night.

#!/bin/bash
# AIO SERVER DISPLAY CONFIGURATOR

if [ "$EUID" -ne 0 ]; then
    echo "[-] Please run this script with sudo."
    exit 1
fi

echo "[+] Starting automated display configuration..."

# 1. Detect Backlight Interface
BACKLIGHT_DIR="/sys/class/backlight"
INTERFACE=$(ls "$BACKLIGHT_DIR" | head -n 1)

if [ -z "$INTERFACE" ]; then
    echo "[-] Error: No backlight controller found."
    exit 1
fi

BACKLIGHT_PATH="$BACKLIGHT_DIR/$INTERFACE"
MAX_BRIGHTNESS=$(cat "$BACKLIGHT_PATH/max_brightness")
echo "[+] Detected interface: $INTERFACE (Max: $MAX_BRIGHTNESS)"

# 2. Inject Bash Aliases for the Calling User
REAL_USER=${SUDO_USER:-$USER}
REAL_USER_HOME=$(eval echo ~$REAL_USER)
BASHRC_PATH="$REAL_USER_HOME/.bashrc"

if [ -f "$BASHRC_PATH" ]; then
    sed -i '/alias screenoff=/d' "$BASHRC_PATH"
    sed -i '/alias screenon=/d' "$BASHRC_PATH"
    
    echo "alias screenoff='echo 0 | sudo tee $BACKLIGHT_PATH/brightness'" >> "$BASHRC_PATH"
    echo "alias screenon='echo $MAX_BRIGHTNESS | sudo tee $BACKLIGHT_PATH/brightness'" >> "$BASHRC_PATH"
    chown "$REAL_USER":"$REAL_USER" "$BASHRC_PATH"
    echo "[+] Shortcuts 'screenoff' and 'screenon' configured."
fi

# 3. Configure Console Idle Blanking (1 minute)
setterm --blank 1 --term linux < /dev/tty1 2>/dev/null || true

# 4. Install Root Cron Jobs for Automation (Off at 22:00, On at 07:00)
TMP_CRON=$(mktemp)
crontab -l > "$TMP_CRON" 2>/dev/null || true
sed -i '\#/sys/class/backlight/#d' "$TMP_CRON"

echo "0 22 * * * echo 0 > $BACKLIGHT_PATH/brightness" >> "$TMP_CRON"
echo "0 7 * * * cat $BACKLIGHT_PATH/max_brightness > $BACKLIGHT_PATH/brightness" >> "$TMP_CRON"

crontab "$TMP_CRON"
rm "$TMP_CRON"

echo "[+] Automated cron jobs deployed successfully."
echo "[*] Run 'source ~/.bashrc' to activate your aliases immediately."

Appendix B: Bridging Isolated Subnets via SSH Tunnels

When running a robust home lab on hypervisors like Proxmox, you often run into a scenario where the physical host collects data (like Scaphandre hardware metrics on port 8081), but the Prometheus database lives inside an isolated guest Virtual Machine. If direct port binding is blocked by firewalls or routing tables, an SSH tunnel is the most secure way to bridge the gap.

1. Passwordless Authentication

To automate the tunnel, the monitoring VM must be able to SSH into the bare-metal host without a password prompt. Generate an RSA key inside the VM and copy it to the host:

ssh-keygen -t rsa -b 4096
ssh-copy-id root@<PROXMOX_HOST_IP>

2. The Persistent Systemd Tunnel

You must ensure the tunnel re-establishes itself if the network drops. Create a systemd service file on the monitoring VM:

sudo nano /etc/systemd/system/scaphandre-tunnel.service

Paste the following configuration. The ServerAliveInterval and ExitOnForwardFailure parameters are critical for preventing "zombie" connections that block the port without actually transmitting data.

[Unit]
Description=SSH Tunnel to Proxmox Scaphandre
After=network.target ssh.service

[Service]
User=<YOUR_VM_USER>
ExecStart=/usr/bin/ssh -N -T -o ServerAliveInterval=60 -o ExitOnForwardFailure=yes -L 8081:localhost:8081 root@<PROXMOX_HOST_IP>
Restart=always
RestartSec=10

[Install]
WantedBy=multi-user.target

3. Activation & Prometheus Integration

Enable the tunnel to start on boot:

sudo systemctl daemon-reload
sudo systemctl enable --now scaphandre-tunnel.service

Now, open /etc/prometheus/prometheus.yml and add the local end of the tunnel to your scrape targets. Prometheus will query its own localhost port, which the SSH tunnel silently forwards to the physical hypervisor:

  - job_name: 'scaphandre_proxmox'
    scrape_interval: 15s
    static_configs:
      - targets: ['localhost:8081']

Restart Prometheus (sudo systemctl restart prometheus) and verify the target is UP in your Prometheus Dashboard.

Appendix C: Advanced Networking & VNC Polkit Overrides

Ubuntu Server natively uses systemd-networkd to manage connections via Netplan. However, when building a home lab, using NetworkManager (and its nmcli/nmtui utilities) provides much greater flexibility for managing Wi-Fi adapters and static IPs dynamically.

1. Forcing Netplan to Use NetworkManager

To hand control of your network interfaces over to NetworkManager, you must edit your core Netplan configuration file (usually found at /etc/netplan/00-installer-config.yaml or similar).

sudo nano /etc/netplan/00-installer-config.yaml

Set the renderer explicitly. (Note: YAML is strictly space-indented; do not use tabs).

network:
  version: 2
  renderer: NetworkManager

Apply the routing changes to the kernel: sudo netplan apply

2. The Headless Wi-Fi Cheat Sheet

Once NetworkManager is active, use these commands to control wireless hardware directly from your SSH terminal:

  • Turn Wi-Fi Radio On/Off: nmcli radio wifi on | nmcli radio wifi off
  • Scan for Local Networks: nmcli device wifi list
  • Connect to a Secure Network: nmcli device wifi connect "<SSID_NAME>" password "<WIFI_PASSWORD>"
  • Verify Adapter Status: nmcli device status

3. Fixing the "Not Authorized" Polkit Error

If you attempt to use the graphical sudo nmtui interface or nmcli commands over an SSH or VNC connection, Ubuntu's PolicyKit (Polkit) security framework will often block it with an org.freedesktop.networkmanager.network-control request failed: not authorized error. This occurs because Polkit defaults to restricting network changes strictly to active, physical local sessions (someone physically typing at the server's keyboard).

To grant your remote user permission to modify networks over VNC or SSH, create a custom authorization override rule.

sudo nano /etc/polkit-1/rules.d/99-networkmanager-vnc.rules

Paste the following JavaScript-based security policy, which instructs Polkit to instantly authorize network requests as long as the remote user belongs to the sudo administrative group:

polkit.addRule(function(action, subject) {
    if (action.id.indexOf("org.freedesktop.NetworkManager.") == 0 &&
        subject.isInGroup("sudo")) {
        return polkit.Result.YES;
    }
});

Save the file and force the Polkit daemon to parse the new security rule:

sudo systemctl restart polkit

Your remote VNC and SSH sessions now possess full administrative authority to configure, bounce, and manage network interfaces using standard tools.

Documentation Complete

Appendix D: The Unified Docker Compose Stack

Throughout this guide, we deployed services individually. However, for a streamlined home lab, you can consolidate your core media and automation services into a single, unified docker-compose.yml file. This allows you to spin up your entire infrastructure with a single command.

Create a master directory (e.g., ~/homelab) and place this docker-compose.yml inside it. Make sure to create a matching .env file in the same folder for your NAS credentials.

services:
  # 1. Dashboard
  homepage:
    image: ghcr.io/gethomepage/homepage:latest
    container_name: homepage
    ports:
      - "3001:3000"
    volumes:
      - ./homepage/config:/app/config
    environment:
      - PUID=1000
      - PGID=1000
    restart: unless-stopped

  # 2. DNS & Adblocking
  pihole:
    container_name: pihole
    image: pihole/pihole:latest
    ports:
      - "53:53/tcp"
      - "53:53/udp"
      - "80:80/tcp"
      - "8443:443/tcp"
    environment:
      TZ: 'Europe/London'
      WEBPASSWORD: '<YOUR_SECURE_PASSWORD>'
      FTLCONF_dns_listeningMode: 'ALL'
    volumes:
      - './pihole/etc-pihole:/etc/pihole'
      - './pihole/etc-dnsmasq.d:/etc/dnsmasq.d'
    cap_add:
      - SYS_NICE
    restart: unless-stopped

  # 3. Home Automation (Host Network)
  homeassistant:
    container_name: homeassistant
    image: ghcr.io/home-assistant/home-assistant:stable
    volumes:
      - ./homeassistant/config:/config
      - /etc/localtime:/etc/localtime:ro
    environment:
      - TZ=Europe/London
    network_mode: host
    restart: unless-stopped

  # 4. Media Server
  plex:
    image: lscr.io/linuxserver/plex:latest
    container_name: plex
    network_mode: host
    environment:
      - PUID=1000
      - PGID=1000
      - TZ=Europe/London
      - VERSION=docker
    volumes:
      - ./plex/config:/config
      - nas_media:/data/nas_shared
    restart: unless-stopped

volumes:
  nas_media:
    driver: local
    driver_opts:
      type: cifs
      o: "username=${NAS_USER},password=${NAS_PASS},iocharset=utf8,vers=3.0"
      device: "//<NAS_IP_ADDRESS>/<SHARE_NAME>"

To deploy the entire fleet at once, navigate to your ~/homelab directory and run: docker compose up -d

Appendix E: Server Maintenance Cheat Sheet

Maintaining your home lab is just as important as building it. Bookmark this section for quick reference on keeping your host OS and Docker environment secure and optimized over time.

1. Routine Host OS Updates

Keep your Ubuntu Server patched and secure. Run this command sequence monthly to update packages and remove orphaned dependencies:

sudo apt update && sudo apt upgrade -y
sudo apt autoremove --purge -y

2. Updating Docker Containers

To pull the latest versions of your Docker images and recreate your containers with the new code (without losing your persistent data), navigate to your compose directory and run:

docker compose pull
docker compose up -d

3. Cleaning Docker Bloat

Over time, downloading new container versions leaves old, dangling images on your hard drive, which can silently consume gigabytes of storage. Clean up unused images, networks, and stopped containers safely:

# Safely remove unused data (leaves running containers untouched)
docker system prune -a

# View current Docker disk usage stats
docker system df

4. Restarting the VNC Engine

If your remote desktop ever hangs due to a memory leak in a graphical application, you can seamlessly restart the background daemon without rebooting the physical server:

sudo systemctl restart vncserver@1.service
sudo systemctl status vncserver@1.service --no-pager

End of Supplemental Appendices

You now possess the complete, exhaustive reference guide covering every facet of your home lab deployment—from initial hardware tuning to long-term unified maintenance strategies.

Appendix F: Resolving Grafana JSON Validation Errors

When migrating highly complex dashboards (like the multi-node Scaphandre and cAdvisor layouts) using raw JSON payloads, Grafana occasionally throws schema validation errors if the source code was exported from an older or conflicting API version (e.g., dashboard.grafana.app/v2).

1. The GridLayoutItem Conflict

If you encounter an error stating: DashboardSpec.layout.spec.items.5.kind: Invalid value: conflicting values "GridLayoutItem" and "GridLayoutImageItem", it is because Grafana's modern UI strict-typing rejects legacy image layouts inside standard grid matrices[cite: 30].

The Fix: Before importing the JSON string into your UI, open it in a standard text editor and run a "Find and Replace." Replace all instances of "GridLayoutImageItem" with "GridLayoutItem"[cite: 30].

2. Fixing Variable Schema Errors

Another common import blocker is the conflicting values "AdhocVariable" and "QueryVariable" error[cite: 50]. This happens when variable schemas mismatch the expected array length in Grafana 13+.

The Fix: Strip the legacy type: "dashboard" parameter from the variables block in your JSON file before importing[cite: 50, 51]. If the import continues to fail, bypass the JSON entirely and use the official Grafana.com community IDs (e.g., ID 14282 for cAdvisor, or 1860 for Node Exporter) to pull heavily vetted, continuously updated templates directly into your environment[cite: 205, 222].

Appendix G: Security Sanitization Checklist

Before making any home lab documentation public on a blog, it is critical to ensure no sensitive infrastructure data is accidentally exposed. This guide has been pre-sanitized, but if you upload your own terminal logs or screenshots, verify the following:

  • Tailscale IPs: Ensure no 100.x.x.x addresses are visible in your Grafana queries or SSH command logs.
  • Telegram Tokens: Never publish your bot_token or chat_id. Malicious actors can use these to intercept your server alerts or flood your personal chat.
  • MAC Addresses & Hardware IDs: When sharing lshw or ip neighbor outputs, obscure MAC addresses (e.g., 70:5a:xx:xx:xx). If using inxi, always append the -z flag (e.g., inxi -Fz) to automatically scramble security identifiers.
  • NAS Credentials: Verify that your Docker Compose .env files or inline cifs mount strings do not contain plain-text passwords or internal usernames.

End of Guide. All source materials and diagnostics have been fully documented.

Phase 26: Bonus - Preventing System Bloat & Log Rotations

As your home lab scales, background logs can quietly consume gigabytes of storage, eventually leading to No space left on device errors. Just like we capped the Pi-hole database to prevent I/O bottlenecks, we need to apply similar limits to the host OS and Docker.

1. Truncating Systemd Journal Logs

Ubuntu's journalctl keeps extensive logs of every systemd service (like our VNC and Scaphandre daemons). Limit these logs to a manageable size (e.g., 500MB) to protect your storage drive:

# Clear out old logs immediately, keeping only the last 2 days
sudo journalctl --vacuum-time=2d

# Or limit by absolute size
sudo journalctl --vacuum-size=500M

To make this limit permanent, edit the journal configuration:

sudo nano /etc/systemd/journald.conf

Uncomment and change the SystemMaxUse line: SystemMaxUse=500M. Then restart the service: sudo systemctl restart systemd-journald.

2. Enforcing Docker Log Limits

By default, Docker container logs grow infinitely. A single chatty container (like Wazuh or NetAlertX) can generate massive text files. Force Docker to rotate logs automatically by creating a global daemon configuration.

sudo nano /etc/docker/daemon.json

Paste the following JSON to restrict containers to three 10MB log files max:

{
  "log-driver": "json-file",
  "log-opts": {
    "max-size": "10m",
    "max-file": "3"
  }
}

Restart the Docker engine to apply the rules to all future containers: sudo systemctl restart docker.

3. Automated Security Patching

For a hands-off approach to critical security updates (like the fwupdmgr vulnerabilities we patched earlier), ensure the unattended-upgrades package is running. It will silently install security patches in the background without breaking your core software.

sudo apt install unattended-upgrades -y
sudo dpkg-reconfigure --priority=low unattended-upgrades

Select Yes to enable automatic daily security downloads.

About This Series

This master guide was compiled directly from real-world terminal logs, debugging sessions, and architectural deployments. From navigating obscure Realtek driver bugs and strict Snap permissions to orchestrating a fully integrated Prometheus telemetry stack, every step here has been battle-tested on live hardware.

Was this guide helpful?

Feel free to bookmark this page for your future home lab rebuilds, share it with fellow self-hosters, and drop any questions in the comments below!

Phase 29: Securing Kali Linux & Nmap Discovery

When deploying penetration testing distributions like Kali Linux within your lab, remote access protocols are disabled by default to minimize the attack surface. Enable them securely and utilize Nmap for network reconnaissance.

1. Bootstrap the OpenSSH Server

Install and bind the OpenSSH daemon to the system autostart framework. Ensure you immediately change the default user password using passwd to prevent unauthorized lateral movement.

sudo apt update && sudo apt install openssh-server -y
sudo systemctl enable --now ssh
sudo systemctl status ssh --no-pager

2. Nmap Host Discovery & Optimization

Nmap (Network Mapper) offers profound insights into active assets and firewall configurations. Here are the core flags for optimizing your scans across local subnets.

  • Decoy Scans (Evasion): Mask your scanning origin by injecting dummy IPs into the traffic logs.
    nmap -sS -D 10.1.0.1 <TARGET_IP>
  • Skip Host Discovery (No Ping): If a firewall drops ICMP packets, force a port scan regardless of ping response.
    nmap -Pn <TARGET_IP>
  • Aggressive Assessments: Enable OS detection, service versioning, and default scripts simultaneously.
    sudo nmap -A -v -T4 192.168.1.*

Timing Profiles: Adjust packet transmission rates to prevent network flooding.

  • -T3: Normal (Default).
  • -T4: Aggressive (Fast, requires stable internal links).
  • -T5: Insane (Maximum speed, risks packet loss).

Phase 30: Automated Debian & Kali Daily Diagnostics

Consolidate your daily security and performance checks into a unified maintenance shell utility. This script resynchronizes packages, audits network bindings, samples virtual memory, and maps active containers.

#!/bin/bash
# DEBIAN / KALI LINUX DAILY MAINTENANCE UTILITY
if [ "$EUID" -ne 0 ]; then
 echo "[-] Error: Please run this diagnostic script with sudo privileges."
 exit 1
fi

echo "======================================================================"
echo " 🔄 PHASE 1: PACKAGE MANAGER RESYNCHRONIZATION & REPAIR"
echo "======================================================================"
sudo apt update && sudo apt upgrade -y
sudo apt --fix-broken install -y
sudo apt autoremove -y

echo "======================================================================"
echo " 📡 PHASE 2: NETWORK INTERFACES & SECURE TOPOLOGY"
echo "======================================================================"
echo "[+] VNC Socket Bindings:"
sudo netstat -ptnl | grep vnc
echo "[+] Link Layer IP Routing Tables:"
ip addr

echo "======================================================================"
echo " ⚙️ PHASE 3: OS ARCHITECTURE & ENVIRONMENT"
echo "======================================================================"
sudo hostnamectl
sudo lsb_release -a

echo "======================================================================"
echo " 📊 PHASE 4: DISK MANAGEMENT & CONTAINER METRICS"
echo "======================================================================"
if command -v smbstatus &> /dev/null; then
 sudo smbstatus --shares | column -t
fi
lsblk

echo "======================================================================"
echo " 🛠️ PHASE 5: HARDWARE FOOTPRINT & PERFORMANCE"
echo "======================================================================"
lscpu
vmstat

echo "======================================================================"
echo " ✅ SYSTEM DIAGNOSTICS LOG LOOP COMPLETE!"
echo "======================================================================"

Save as dailyscri.sh, apply permissions (chmod +x dailyscri.sh), and execute via sudo ./dailyscri.sh.

Phase 31: Advanced Wazuh Administration & Fixes

Maintaining a Wazuh SIEM deployment occasionally requires manual intervention to break APT package loops during uninstallation or to facilitate anonymous read-only dashboard access in isolated lab environments.

1. Fixing Broken Wazuh-Manager Uninstallation

If purging the wazuh-manager package throws exit status errors (like prerm or postrm script failures), you must bypass the broken maintainer scripts manually to unlock the package manager.

# 1. Clear the Pre-Removal Script
sudo nano /var/lib/dpkg/info/wazuh-manager.prerm
# Replace contents with:
# #!/bin/sh
# exit 0

# 2. Clear the Post-Removal Script
sudo nano /var/lib/dpkg/info/wazuh-manager.postrm
# Replace contents with:
# #!/bin/sh
# exit 0

# 3. Complete the Removal and Purge Leftovers
sudo apt-get autoremove -y && sudo apt-get clean
sudo rm -rf /var/ossec
sudo dpkg --purge wazuh-manager

2. Enabling Anonymous Login on Wazuh Dashboard

Warning: Only enable anonymous authentication in strictly isolated test environments, as it grants full access without password verification.

  1. Enable Auth in Indexer: Open /etc/wazuh-indexer/opensearch-security/config.yml and set anonymous_auth_enabled: true under the http: block.
  2. Map Roles: Open /etc/wazuh-indexer/opensearch-security/roles_mapping.yml and add "opendistro_security_anonymous" to the all_access users and backend_roles lists.
  3. Sync Security Settings:
    sudo OPENSEARCH_JAVA_HOME=/usr/share/wazuh-indexer/jdk /usr/share/wazuh-indexer/plugins/opensearch-security/tools/securityadmin.sh \
      -cd /etc/wazuh-indexer/opensearch-security \
      -nhnv -cacert /etc/wazuh-indexer/certs/root-ca.pem \
      -cert /etc/wazuh-indexer/certs/admin.pem \
      -key /etc/wazuh-indexer/certs/admin-key.pem -p 9200
  4. Configure Dashboard: Open /etc/wazuh-dashboard/opensearch_dashboards.yml and append:
    opensearch_security.auth.type: "basicauth"
    opensearch_security.auth.anonymous_auth_enabled: true

Restart the dashboard (sudo systemctl restart wazuh-dashboard) to expose the "Log in as anonymous" button.

Phase 32: Securing Docker Apps with Nginx Reverse Proxy

When running custom containerized web applications (like a Python dashboard or a Bookmark-Doc-app), they often expose themselves directly on unencrypted ports. Lock down access by restricting Docker to localhost and wrapping the traffic in an Nginx proxy protected by HTTP Basic Authentication.

1. Automated Deployment & Security Script

This script automates the installation of Nginx, generates an .htpasswd credential file, pulls an application from GitHub, spins up the Docker container on localhost, and generates the protective reverse proxy configuration block.

#!/bin/bash

# 1. Install Nginx & Apache Utilities
if ! command -v nginx &> /dev/null; then
    sudo apt update && sudo apt install nginx apache2-utils -y
fi

# 2. Setup Nginx Password File
if [ ! -f /etc/nginx/.htpasswd ]; then
    sudo htpasswd -c /etc/nginx/.htpasswd admin
fi

# 3. Pull/Update Application & Deploy Container
sudo rm -rf Bookmark-Doc-app/
git clone https://github.com/<GITHUB_USER>/Bookmark-Doc-app.git
cd Bookmark-Doc-app/
# Ensure docker-compose.yml maps ports as: "127.0.0.1:5000:5000"
sudo docker compose up --build -d
cd ..

# 4. Setup Nginx Proxy Configuration
if [ ! -f /etc/nginx/sites-available/python_proxy ]; then
    sudo bash -c 'cat > /etc/nginx/sites-available/python_proxy << "EOF"
server {
    listen 9000;
    server_name _;

    auth_basic "Restricted Access";
    auth_basic_user_file /etc/nginx/.htpasswd;

    location / {
        proxy_pass http://127.0.0.1:5000;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
    }
}
EOF'
fi

# 5. Enable site and restart Nginx
sudo ln -s /etc/nginx/sites-available/python_proxy /etc/nginx/sites-enabled/
sudo ufw allow 9000/tcp
sudo nginx -t
sudo systemctl restart nginx

Verification: Accessing the application on port 9000 will now prompt for login credentials before securely passing the traffic to the backend Docker container hidden on localhost.

Phase 33: Self-Hosting Navidrome with an External NAS

If you have a massive library of local audio files, Navidrome is incredibly lightweight, lightning-fast, and compatible with almost all Subsonic clients[cite: 285]. In this phase, we deploy Navidrome using Docker Compose, mount an external NAS folder for music, and safely store the database locally[cite: 285].

1. Prepare Your Directories

Navidrome requires two distinct storage targets: fast local storage for its internal database, and the external NAS mount for your music files[cite: 285].

# Create the Navidrome data folder
mkdir -p /home/<YOUR_USERNAME>/navidrome_data

# Ensure proper permissions (Navidrome runs as user 1000 by default)
sudo chown -R 1000:1000 /home/<YOUR_USERNAME>/navidrome_data

(Ensure your external NAS is already mounted to a path like /mnt/external/NAS/Music)[cite: 285].

2. Create the Docker Compose File

Deploy the server using the following configuration. The :ro flag on the music volume ensures your raw audio files on the NAS are mounted as "Read-Only," preventing accidental deletion via the Navidrome web UI[cite: 285].

services:
  navidrome:
    image: deluan/navidrome:latest
    user: 1000:1000
    ports:
      - "4533:4533"
    restart: always
    environment:
      ND_LOGLEVEL: info 
      ND_SCANSCHEDULE: "@every 1h"
      ND_ENABLEDOWNLOADS: "true"
    volumes:
      - "/home/<YOUR_USERNAME>/navidrome_data:/data"
      - "/mnt/external/NAS/SharedMusic:/music:ro"

Spin up the server (sudo docker compose up -d) and navigate to http://<YOUR_SERVER_IP>:4533 to create your secure admin account (avoid using generic usernames like "admin" or "root")[cite: 285].

Phase 34: Securing Navidrome with a Tailscale Funnel & HA Kill Switch

To listen to your music on strict corporate networks where VPN clients are blocked, you can use Tailscale Funnel to expose port 4533 to the public internet securely[cite: 285]. Leaving a public tunnel open permanently is risky, so we will build a physical "Kill Switch" inside Home Assistant to toggle the gateway on and off instantly[cite: 285].

1. Allow "Silent Sudo" & Generate SSH Keys

Home Assistant needs to execute Tailscale commands on the host machine without a password prompt[cite: 285].

# Run visudo on the host
sudo visudo
# Add to the bottom of the file:
<YOUR_USERNAME> ALL=(ALL) NOPASSWD: /usr/bin/tailscale

Next, generate an SSH key directly inside Home Assistant's mapped configuration directory so the Docker container can SSH into the host[cite: 285]:

sudo mkdir -p /path/to/hass/config/.ssh
sudo ssh-keygen -t rsa -f /path/to/hass/config/.ssh/id_rsa -q -N ""
sudo ssh-copy-id -i /path/to/hass/config/.ssh/id_rsa <YOUR_USERNAME>@<YOUR_HOST_IP>
sudo chmod 700 /path/to/hass/config/.ssh
sudo chmod 600 /path/to/hass/config/.ssh/id_rsa

2. The Home Assistant YAML Configuration

Open your Home Assistant configuration.yaml file and add the command_line switch[cite: 285]. The -q flag and UserKnownHostsFile=/dev/null are critical to prevent SSH warnings from breaking the switch state[cite: 285].

command_line:
  - switch:
      name: "Navidrome Public Tunnel"
      unique_id: navidrome_public_tunnel
      command_on: "ssh -q -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -i /config/.ssh/id_rsa <YOUR_USERNAME>@<YOUR_HOST_IP> 'sudo tailscale funnel --bg 4533'"
      command_off: "ssh -q -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -i /config/.ssh/id_rsa <YOUR_USERNAME>@<YOUR_HOST_IP> 'sudo tailscale serve reset'"
      command_state: "ssh -q -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -i /config/.ssh/id_rsa <YOUR_USERNAME>@<YOUR_HOST_IP> 'sudo tailscale funnel status | grep 4533'"
      icon: mdi:music-network

Restart the Home Assistant container (sudo docker restart homeassistant) and add a Button Card mapped to switch.navidrome_public_tunnel to your dashboard[cite: 285].

Phase 35: Automating the Funnel via Telegram & Safety Timers

To prevent leaving your music server exposed to the public web indefinitely, we will build a two-way Telegram alert system with an interactive inline button, backed by a 2-hour "dead man's switch"[cite: 285].

1. The Sender Automation (Telegram Alert + Button)

In Home Assistant, create a YAML automation to watch the tunnel switch and send an HTML-formatted Telegram alert containing a clickable inline button ("🔒 Close Tunnel Now:/close_tunnel") when opened[cite: 285]:

alias: "Security: Navidrome Tunnel Status Alert"
mode: single
trigger:
  - platform: state
    entity_id: switch.navidrome_public_tunnel
    not_to: ["unknown", "unavailable"]
    not_from: ["unknown", "unavailable"]
action:
  - choose:
      - conditions:
          - condition: state
            entity_id: switch.navidrome_public_tunnel
            state: "on"
        sequence:
          - action: telegram_bot.send_message
            data:
              message: "🟢 <b>Navidrome Tunnel: OPEN</b>\nYour music server is now exposed to the public web."
              parse_mode: html
              inline_keyboard:
                - "🔒 Close Tunnel Now:/close_tunnel"
    default:
      - action: telegram_bot.send_message
        data:
          message: "🔴 <b>Navidrome Tunnel: CLOSED</b>\nThe Tailscale Funnel has been shut down. Public access is fully blocked."
          parse_mode: html

2. The Listener Automation (Catching the Button Press)

Create a second automation that catches the /close_tunnel callback event, turns off the switch, and edits the original Telegram message to confirm the action[cite: 285]:

alias: "Security: Telegram Close Tunnel Callback"
mode: single
trigger:
  - platform: event
    event_type: telegram_callback
    event_data:
      data: "/close_tunnel"
action:
  - action: telegram_bot.answer_callback_query
    data:
      callback_query_id: "{{ trigger.event.data.id }}"
      message: "Closing tunnel..."
  - action: switch.turn_off
    target:
      entity_id: switch.navidrome_public_tunnel
  - action: telegram_bot.edit_message
    data:
      message_id: "{{ trigger.event.data.message.message_id }}"
      chat_id: "{{ trigger.event.data.user_id }}"
      message: "🔴 <b>Navidrome Tunnel: CLOSED via Telegram</b>\nThe Tailscale Funnel has been shut down remotely."
      parse_mode: html
      inline_keyboard: []

3. The 2-Hour Auto-Shutoff Timer

Use a State Duration Trigger to automatically close the tunnel if left open for 2 hours (this is safer than a delay action, which wipes from memory on reboot)[cite: 285]:

alias: "Security: Auto-Close Navidrome Tunnel (2 Hours)"
mode: single
trigger:
  - platform: state
    entity_id: switch.navidrome_public_tunnel
    to: "on"
    for:
      hours: 2
action:
  - action: switch.turn_off
    target:
      entity_id: switch.navidrome_public_tunnel

Phase 36: Proxmox USB Backups (The "Host Owns It" Architecture)

If you pass a raw USB hard drive directly to a Proxmox VM, Proxmox cannot back up that VM because it creates a dependency loop resulting in a target is busy error[cite: 285]. To resolve this, move the physical drive management to the Proxmox host, use it for backups, and securely share the media files back to the VM via an NFS network share[cite: 285].

1. Detach the Drive from the VM

Log into your VM shell, stop active services (Docker/Samba), and unmount the drive. Then, remove the USB hardware passthrough in the Proxmox Web GUI[cite: 285].

sudo systemctl stop docker.socket docker.service
sudo systemctl stop smbd nmbd
sudo umount /mnt/external

2. Mount on the Proxmox Host

Switch to the Proxmox Host shell, identify your drive's UUID (blkid), and configure /etc/fstab to mount it automatically[cite: 285].

mkdir -p /mnt/pve/usb-storage
nano /etc/fstab
# Add: UUID=<YOUR-UUID-HERE> /mnt/pve/usb-storage ntfs-3g defaults,nofail 0 0
mount -a

3. Share the Drive Back to the VM via NFS

Install the NFS kernel server on the Proxmox host and grant the VM exclusive access. The fsid=1 flag is critical for FUSE/NTFS drives[cite: 285]:

apt install nfs-kernel-server -y
nano /etc/exports
# Add: /mnt/pve/usb-storage <VM_IP>(rw,sync,no_subtree_check,no_root_squash,fsid=1)
exportfs -arv
systemctl restart nfs-kernel-server

4. Reconnect the Drive Inside the VM

Install the NFS client inside your VM (sudo apt install nfs-common -y) and update the VM's /etc/fstab[cite: 285]:

<PROXMOX_IP>:/mnt/pve/usb-storage /mnt/external nfs defaults,_netdev,nofail 0 0

Reload the daemon, mount the drive, and start your Docker services[cite: 285]. Proxmox can now safely target the USB directory for VZDump backups while the VM transparently accesses the media files over the internal virtual bridge.

Phase 37: Troubleshooting Wazuh Startup Timeouts

When deploying Wazuh on resource-constrained VMs, the wazuh-manager may refuse to start due to systemd exceeding its default 45-second timeout limit, leaving orphaned processes behind[cite: 285]. Furthermore, the Wazuh Dashboard may stall indefinitely with a "Server is not ready yet" error[cite: 285].

1. Clean Up and Apply the Permanent Systemd Fix

Clear out "sleeping" background processes and the locked startup directory, then increase the timeout threshold to 5 minutes (300 seconds)[cite: 285].

sudo pkill -f wazuh
sudo rm -rf /var/ossec/var/start-script-lock

sudo mkdir -p /etc/systemd/system/wazuh-manager.service.d
echo -e "[Service]\nTimeoutStartSec=300" | sudo tee /etc/systemd/system/wazuh-manager.service.d/override.conf

sudo systemctl daemon-reload
sudo systemctl restart wazuh-manager

2. Fixing the Dashboard "Not Ready" Error

If the dashboard logs (journalctl -u wazuh-dashboard -f) show connect ECONNREFUSED 127.0.0.1:9200, the web frontend cannot reach the backend database[cite: 285]. Start the wazuh-indexer service and wait 60 to 90 seconds for the JVM heap memory to initialize[cite: 285].

sudo systemctl start wazuh-indexer
sudo systemctl status wazuh-indexer

Test the API socket directly with curl -k https://127.0.0.1:9200. Once it returns "cluster_name": "wazuh-cluster", the dashboard will automatically resolve the error loop[cite: 285].

Phase 38: Integrating Wazuh IT Hygiene Dashboards into Grafana

To extract endpoint telemetry (like packages, ports, and hardware inventory) into Grafana, you must authenticate with the Wazuh REST API using long-lived JWT Bearer Tokens and deploy the Infinity Data Source plugin[cite: 285].

1. Generate and Extend API Tokens

Fetch your API credentials from /usr/share/wazuh-dashboard/data/wazuh/config/wazuh.yml[cite: 285]. By default, Wazuh API tokens expire every 15 minutes, which will break live Grafana boards[cite: 285]. Extend the expiration to 1 year (31,536,000 seconds) via the API[cite: 285]:

# 1. Fetch a temporary token
TOKEN=$(curl -u 'wazuh-wui:<YOUR_PASSWORD>' -k -s -X GET "https://127.0.0.1:55000/security/user/authenticate?raw=true")

# 2. Update the API configuration database
curl -k -X PUT "https://127.0.0.1:55000/security/config" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"auth_token_exp_timeout": 31536000}'

# 3. Generate your permanent token
curl -u 'wazuh-wui:<YOUR_PASSWORD>' -k -X GET "https://127.0.0.1:55000/security/user/authenticate?raw=true"

2. Configuring Grafana and Managing Agent Scope

Inside Grafana, add the Yesoreyeram Infinity Data Source plugin. Set the Base URL to https://<YOUR_IP>:55000 and input the long-lived token string under Auth Type: Bearer Token[cite: 285].

Wazuh syscollector endpoints require explicit agent identifiers in the URI path (e.g., /syscollector/${agent_id}/packages) and will return a 404 Not Found if queried globally[cite: 285]. Configure a dashboard variable (agent_id) to fetch available systems via /agents?select=id,name, enable multi-value, and use Grafana's Row Repeating feature to automatically stack IT hygiene panels for every active agent in your fleet[cite: 285].

Phase 39: AI-Powered SOC Triage (Wazuh & LiteLLM)

Piping every SIEM alert into an AI model quickly results in context window limits or massive API bills. By combining Wazuh's detection engine with a LiteLLM Gateway, we can build a dual-lane triage pipeline: a Real-Time Fast Lane for critical alerts, and a Daily Batch Digest using local models[cite: 285].

1. The Real-Time Fast Lane (Critical Alerts)

Configure Wazuh (/var/ossec/etc/ossec.conf) to trigger a custom python script when Level 12+ alerts occur[cite: 285]:

  <integration>
    <name>custom-llm-telegram</name>
    <level>12</level>
    <alert_format>json</alert_format>
  </integration>

The script at /var/ossec/integrations/custom-llm-telegram captures the JSON payload, pushes it to LiteLLM, and sanitizes the output (triage_report.replace('<', '&lt;').replace('>', '&gt;')) to prevent AI markdown from crashing Telegram's strict HTML parser[cite: 285].

2. The Batch Processing Lane (Local AI Swarm)

Dumping 24 hours of raw JSON logs into a local model (like Qwen 1.5b) will crash it. Instead, query the OpenSearch Indexer directly and minify the data before passing it to the LLM[cite: 285].

The daily digest script requests logs matching "rule.level": {"gte": 5} over "now-24h", runs a Counter() on rule descriptions, and pushes the simplified metrics to the local model to write an IT Hygiene summary[cite: 285].

3. Specialized Threat Reporting (ISO 27001)

Using the same minification architecture, you can build compliance reports. Query the indexer for {"wildcard": {"rule.compliance.iso_27001": "*"}} over a 7-day period[cite: 285]. If the network is clean, the script bypasses the AI completely and sends a green checkmark to Telegram (Positive Confirmation), ensuring you know the script ran successfully[cite: 285]. Add these scripts to the root crontab to execute securely on scheduled intervals[cite: 285].

Phase 40: Android ADB Debloating (Samsung & Vivo)

To safely disable telemetric bloatware, aggressive performance throttlers, and sponsored feeds on Samsung One UI and Vivo Funtouch OS, utilize the Android Debug Bridge (ADB)[cite: 284].

1. Troubleshooting ADB Errors

If your device returns an unauthorized status, toggle Revoke USB debugging authorizations in Developer Options and reconnect[cite: 284]. If using the package manager, do not include the package: prefix output by pm list packages[cite: 284].

Correct: .\adb shell pm uninstall --user 0 com.samsung.android.smartsuggestions[cite: 284]

2. Vivo Z1x (Funtouch OS 11) Specific Fixes

Funtouch OS 11 throws DELETE_FAILED_USER_RESTRICTED errors when attempting to uninstall packages. Bypass this by stripping background permissions, clearing app data, and disabling the packages for the current user instead[cite: 284].

.\adb shell "
cmd appops set com.bbk.cloud RUN_IN_BACKGROUND ignore; pm clear com.bbk.cloud;
cmd package disable-user --user 0 com.bbk.cloud;
# Repeat for com.vivo.pushservice, com.vivo.gamecube, etc.
"

Since Funtouch OS leaves empty app shortcuts in the default launcher, install a custom launcher like Nova or Lawnchair and use the "Hide App" feature to clean up the UI[cite: 284].

3. Samsung One UI (Game Optimizing Service & Bloat)

Packages like com.samsung.discover (sponsored feeds) and com.sec.android.mimage.avatarstickers (AR Emoji) are 100% safe to remove[cite: 284]. The Game Optimizing Service (com.samsung.android.game.gos) requires an AppOps clear alongside disabling to neutralize aggressive thermal throttling[cite: 284].

.\adb shell pm disable-user --user 0 com.samsung.android.game.gos
.\adb shell cmd appops set com.samsung.android.game.gos RUN_IN_BACKGROUND ignore
.\adb shell pm clear com.samsung.android.game.gos

Master Guide Complete

Phase 41: Advanced Media Management (Immich & Jellyfin)

Once your core NAS and Plex services are running, the next evolution of a home media lab is moving away from proprietary cloud services for personal photos and customizing your streaming interfaces.

1. Self-Hosting Google Photos Alternative (Immich)

Immich is a high-performance photo and video backup solution. It relies on a PostgreSQL database and machine learning containers for facial recognition. When deploying Immich on Ubuntu, ensure your database backups are properly formatted to avoid index restoration errors.

# Example database restoration command for Immich PostgreSQL
# If you encounter unsupported vector index statements during a restore, 
# ensure you are using the pgvector extension matching your database version.

docker exec -i immich_postgres psql -U postgres -d immich < immich-db-backup.sql

2. Customizing the Jellyfin UI

Jellyfin is a fantastic open-source media server, but its default web interface can feel a bit standard. You can inject custom CSS directly into the dashboard to personalize the login screens and library hover effects using modern accent colors.

  1. Log into your Jellyfin web dashboard as an administrator.
  2. Navigate to Dashboard > General > Custom CSS.
  3. Paste the following CSS to apply a sleek blue and purple theme:
/* Custom Jellyfin Accent Theme */
:root {
  --accent1-light: #00a4dc; /* Vibrant Blue */
  --accent2-light: #aa5cc3; /* Deep Purple */
}

/* Button & Accent Styling */
.paper-icon-button-light:hover,
.raised,
.button-submit {
  background: linear-gradient(135deg, var(--accent1-light), var(--accent2-light)) !important;
  color: #ffffff !important;
}

/* Card Hover Effects */
.card:hover .cardImageContainer {
  border: 2px solid var(--accent1-light);
  box-shadow: 0 4px 15px rgba(0, 164, 220, 0.4);
}

Phase 42: Lightweight Storage & Custom Web Apps

Expanding the lab doesn't always mean massive x86 servers. Single-board computers and custom-coded applications add highly specialized functionality to the network.

1. OpenMediaVault on Raspberry Pi

A Raspberry Pi running OpenMediaVault (OMV) acts as a highly energy-efficient NAS and Docker volume manager. To maintain storage health on the Pi, schedule routine Docker pruning and volume maintenance directly from the OMV terminal:

# Clear dangling images and unused build caches to preserve SD card space
docker image prune -a -f
docker volume prune -f

2. Hosting Custom Python Web Apps (Streamlit)

If you build bespoke data tools—such as a cricket match score tracker or a Google Fit data minification pipeline—Streamlit is the fastest way to deploy them. To host these securely on your own domain, deploy them in Docker behind the Nginx reverse proxy with Basic Authentication.

# Example Docker Compose for a custom Streamlit App
services:
  streamlit-app:
    build: .
    container_name: custom_score_tracker
    ports:
      - "127.0.0.1:8501:8501" # Bound to localhost for Nginx proxy routing
    restart: unless-stopped

By routing Port 8501 through your previously established Nginx .htpasswd proxy block, your personal data analysis pipelines remain entirely hidden from the public internet.

The Homelab Journey Continues

From managing ISO 27001 compliance matrices to customizing media servers and tracking power down to the hardware level, this infrastructure is now a fully realized enterprise environment operating right inside your home. The foundation is set, the monitoring is active, and the automation is taking care of the heavy lifting.

Phase 43: Centralized Telegram Alerts for Homelab Cronjobs

Instead of setting up a heavy centralized monitoring server to track automated background tasks across multiple VMs and Raspberry Pis, use this lightweight Bash Wrapper Agent. Placed on each node, it wraps your existing crontab jobs and pushes formatted execution metrics, exit codes, and error logs directly to your phone.

sudo nano /usr/local/bin/cron_agent.sh

Paste the following code, replacing the credentials with your Telegram Bot Token and Chat ID:

#!/bin/bash
# ==========================================
# Telegram Cronjob Monitoring Agent
# ==========================================

BOT_TOKEN="YOUR_TELEGRAM_BOT_TOKEN"
CHAT_ID="YOUR_TELEGRAM_CHAT_ID"

# Parse arguments
ALERT_SUCCESS=true
if [ "$1" == "--fail-only" ]; then
    ALERT_SUCCESS=false
    shift
fi

SERVER_NAME="$1"
JOB_NAME="$2"
shift 2
COMMAND="$@"

# Start execution and timer
START_TIME=$(date +%s)
OUTPUT=$($COMMAND 2>&1)
EXIT_CODE=$?
END_TIME=$(date +%s)
DURATION=$((END_TIME - START_TIME))

# Determine Status
if [ $EXIT_CODE -eq 0 ]; then
    # If success and fail-only is true, exit silently
    if [ "$ALERT_SUCCESS" = false ]; then
        exit 0
    fi
    STATUS="✅ SUCCESS"
else
    STATUS="❌ FAILED (Code: $EXIT_CODE)"
fi

# Format the Telegram Message
MESSAGE=$(cat <<EOF
$STATUS
🖥 Host: $SERVER_NAME
🛠 Job: $JOB_NAME
⏱ Duration: ${DURATION}s

📄 Output:
$(echo "$OUTPUT" | tail -n 15)
EOF
)

# Send to Telegram
curl -s -X POST "https://api.telegram.org/bot${BOT_TOKEN}/sendMessage" \
    -F chat_id="${CHAT_ID}" \
    -F text="$MESSAGE" > /dev/null

Secure the file: sudo chmod +x /usr/local/bin/cron_agent.sh. You can now wrap any cron job using this syntax: /usr/local/bin/cron_agent.sh [--fail-only] "Host Name" "Job Name" /path/to/script.sh.

Phase 44: Deploying RustDesk over Tailscale

Self-hosting your remote desktop infrastructure using RustDesk provides complete control over your privacy. Running it strictly over your Tailscale mesh network creates an encrypted, private administration backbone.

1. The Docker Compose Configuration

Create your docker-compose.yml file. The -k _ flag enforces encryption, and network_mode: "host" allows it to bind directly to your Tailscale interface.

services:
  hbbs:
    container_name: hbbs
    image: rustdesk/rustdesk-server:latest
    environment:
      - ALWAYS_USE_RELAY=Y
    command: hbbs -r <YOUR_VPN_IP>:21117 -k _
    volumes:
      - ./data:/root
    network_mode: "host"
    depends_on:
      - hbbr
    restart: unless-stopped

  hbbr:
    container_name: hbbr
    image: rustdesk/rustdesk-server:latest
    command: hbbr -k _
    volumes:
      - ./data:/root
    network_mode: "host"
    restart: unless-stopped

2. Extracting the Encryption Key from Distroless Containers

Because the RustDesk image is "distroless" (it lacks standard tools like cat or bash), you cannot execute commands inside it. You must copy the auto-generated ED25519 key out to the host to read it.

sudo docker compose up -d
sudo docker cp hbbs:/root/id_ed25519.pub /tmp/id_ed25519.pub
cat /tmp/id_ed25519.pub

Copy this key into your RustDesk client application settings under the Key field, setting both the ID Server and Relay Server to your Tailscale IP.

Phase 44: Deploying RustDesk over Tailscale

Self-hosting your remote desktop infrastructure using RustDesk provides complete control over your privacy. Running it strictly over your Tailscale mesh network creates an encrypted, private administration backbone.

1. The Docker Compose Configuration

Create your docker-compose.yml file. The -k _ flag enforces encryption, and network_mode: "host" allows it to bind directly to your Tailscale interface.

services:
  hbbs:
    container_name: hbbs
    image: rustdesk/rustdesk-server:latest
    environment:
      - ALWAYS_USE_RELAY=Y
    command: hbbs -r <YOUR_VPN_IP>:21117 -k _
    volumes:
      - ./data:/root
    network_mode: "host"
    depends_on:
      - hbbr
    restart: unless-stopped

  hbbr:
    container_name: hbbr
    image: rustdesk/rustdesk-server:latest
    command: hbbr -k _
    volumes:
      - ./data:/root
    network_mode: "host"
    restart: unless-stopped

2. Extracting the Encryption Key from Distroless Containers

Because the RustDesk image is "distroless" (it lacks standard tools like cat or bash), you cannot execute commands inside it. You must copy the auto-generated ED25519 key out to the host to read it.

sudo docker compose up -d
sudo docker cp hbbs:/root/id_ed25519.pub /tmp/id_ed25519.pub
cat /tmp/id_ed25519.pub

Copy this key into your RustDesk client application settings under the Key field, setting both the ID Server and Relay Server to your Tailscale IP.

Phase 45: Glances Systemd Service & Hardware Sensors

Glances is a powerful system monitor that reads lm-sensors data to provide live motherboard and CPU thermal telemetry. We can configure it to run permanently as a lightweight background web server.

1. Install Dependencies and Scan Hardware

sudo apt update && sudo apt install lm-sensors glances -y
# Auto-detect motherboard sensors (bypass prompts)
sudo sensors-detect --auto

2. Construct the Systemd Daemon

sudo nano /etc/systemd/system/glances.service

Paste the following service layout to bind Glances to port 61208:

[Unit]
Description=Glances Web Server Daemon
After=network.target

[Service]
Type=simple
ExecStart=/usr/bin/glances -w -B 0.0.0.0 -p 61208
Restart=on-failure
RestartSec=5

[Install]
WantedBy=multi-user.target

Initialize the service: sudo systemctl enable --now glances.service.

Optional Bugfix: If the Web UI displays an empty frame on older Ubuntu iterations due to uncompiled assets, run: sudo mkdir -p /usr/lib/python3/dist-packages/glances/outputs/static/public.

Phase 46: Wazuh AI Triage - The Real-Time Fast Lane

To prevent alert fatigue, configure a dual-lane AI pipeline. The Fast Lane intercepts critical Wazuh alerts (Level 12+) and sends them to a fast cloud model via a LiteLLM gateway for instant triage analysis.

Create the integration script on your Wazuh Manager:

sudo nano /var/ossec/integrations/custom-llm-telegram

Paste the Python processing code. The HTML sanitization block ensures AI-generated markdown doesn't crash Telegram's strict parsers.

#!/usr/bin/env python3
import sys, json, urllib.request

# --- CONFIGURATION ---
LITELLM_URL = "http://<YOUR_LITELLM_IP>:4000/v1/chat/completions"
LITELLM_KEY = "YOUR_LITELLM_API_KEY"
TELEGRAM_BOT_TOKEN = "YOUR_TELEGRAM_BOT_TOKEN"
TELEGRAM_CHAT_ID = "YOUR_TELEGRAM_CHAT_ID"
FAST_MODEL = "vibe-coder"
# ---------------------

def send_telegram(text):
    url = f"https://api.telegram.org/bot{TELEGRAM_BOT_TOKEN}/sendMessage"
    payload = {"chat_id": TELEGRAM_CHAT_ID, "text": text, "parse_mode": "HTML"}
    req = urllib.request.Request(url, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json"})
    urllib.request.urlopen(req, timeout=15)

def analyze_with_llm(alert_json):
    system_prompt = (
        "You are an automated SOC triage analyst. Analyze this Level 12+ security alert. "
        "Provide a concise analysis in 3 bullet points: "
        "1. Threat Nature & Impact "
        "2. Likely Attack Vector / Indicator "
        "3. Recommended Containment Step. "
        "Be direct, highly technical, and do not use introductory filler."
    )
    prompt = f"Wazuh Alert Data:\n{json.dumps(alert_json, indent=2)}"
    payload = {
        "model": FAST_MODEL,
        "messages": [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
        "temperature": 0.2
    }
    req = urllib.request.Request(LITELLM_URL, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json", "Authorization": f"Bearer {LITELLM_KEY}"})
    with urllib.request.urlopen(req, timeout=60) as resp:
        result = json.loads(resp.read().decode("utf-8"))
        return result["choices"][0]["message"]["content"]

def main():
    if len(sys.argv) < 2: sys.exit(1)
    with open(sys.argv[1], "r") as f: alert_data = json.load(f)
        
    rule = alert_data.get("rule", {})
    agent = alert_data.get("agent", {})
    level = rule.get("level", 0)
    
    if level < 12: sys.exit(0)

    try:
        triage_report = analyze_with_llm(alert_data)
    except Exception as e:
        triage_report = f"⚠️ LLM Analysis Failed: {str(e)}"

    # CRITICAL: Sanitize output so AI asterisks/brackets don't break Telegram HTML
    safe_report = triage_report.replace('<', '&lt;').replace('>', '&gt;')

    msg = (
        f"🚨 <b>CRITICAL WAZUH ALERT (Level {level})</b>\n\n"
        f"<b>Host:</b> <code>{agent.get('name', 'Unknown')}</code>\n"
        f"<b>Rule:</b> <code>{rule.get('id', 'N/A')}</code> - {rule.get('description', 'N/A')}\n\n"
        f"<b>AI Triage Analysis:</b>\n{safe_report}"
    )
    send_telegram(msg)

if __name__ == "__main__": main()

Assign permissions: sudo chmod 750 /var/ossec/integrations/custom-llm-telegram && sudo chown root:wazuh /var/ossec/integrations/custom-llm-telegram.

Phase 47: Wazuh Batch Processing & ISO 27001 Scripts

To avoid hitting LLM context limits, the "Batch Processing Lane" uses a Python script to query OpenSearch, minify the JSON data down to aggregated counters, and process the results locally using a privacy-first AI swarm.

1. The Daily IT Hygiene Digest

Create /usr/local/bin/wazuh-daily-digest.py:

#!/usr/bin/env python3
import urllib.request, json, ssl, base64, sys
from collections import Counter

# --- CONFIGURATION ---
INDEXER_URL = "https://127.0.0.1:9200/wazuh-alerts-*/_search"
INDEXER_USER = "admin"
INDEXER_PASS = "YOUR_INDEXER_PASSWORD"
LITELLM_URL = "http://<YOUR_LITELLM_IP>:4000/v1/chat/completions"
LITELLM_KEY = "YOUR_LITELLM_API_KEY"
TELEGRAM_BOT_TOKEN = "YOUR_TELEGRAM_BOT_TOKEN"
TELEGRAM_CHAT_ID = "YOUR_TELEGRAM_CHAT_ID"
LOCAL_MODEL = "qwen-cluster"

def fetch_wazuh_alerts():
    query = {
        "query": {"bool": {"must": [{"range": {"timestamp": {"gte": "now-24h", "lte": "now"}}}, {"range": {"rule.level": {"gte": 5}}}]}},
        "size": 1000, "sort": [{"rule.level": "desc"}]
    }
    ctx = ssl.create_default_context()
    ctx.check_hostname = False; ctx.verify_mode = ssl.CERT_NONE
    req = urllib.request.Request(INDEXER_URL, data=json.dumps(query).encode("utf-8"), headers={"Content-Type": "application/json"})
    auth_b64 = base64.b64encode(f"{INDEXER_USER}:{INDEXER_PASS}".encode()).decode()
    req.add_header("Authorization", f"Basic {auth_b64}")
    with urllib.request.urlopen(req, context=ctx) as resp:
        return json.loads(resp.read().decode("utf-8"))

def minify_data(data):
    hits = data.get("hits", {}).get("hits", [])
    if not hits: return None
    rule_counts = Counter()
    for hit in hits:
        source = hit.get("_source", {})
        rule_counts[f"Level {source.get('rule', {}).get('level', 0)}: {source.get('rule', {}).get('description', 'Unknown')}"] += 1
        
    summary = "=== 24H ALERT DATA ===\nTop Alert Types:\n"
    for rule, count in rule_counts.most_common(10): summary += f"- {rule} (Triggered {count} times)\n"
    return summary

def analyze_with_llm(minified_data):
    system_prompt = "You are a SOC manager writing a daily IT Hygiene digest. Review the provided alert counts for the last 24 hours. Provide a summary explaining overall network health and specific remediation actions. Output plain text only."
    payload = {"model": LOCAL_MODEL, "messages": [{"role": "system", "content": system_prompt}, {"role": "user", "content": minified_data}], "temperature": 0.3}
    req = urllib.request.Request(LITELLM_URL, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json", "Authorization": f"Bearer {LITELLM_KEY}"})
    with urllib.request.urlopen(req, timeout=300) as resp:
        return json.loads(resp.read().decode("utf-8"))["choices"][0]["message"]["content"]

def send_telegram(analysis):
    url = f"https://api.telegram.org/bot{TELEGRAM_BOT_TOKEN}/sendMessage"
    clean_analysis = analysis.replace('<', '&lt;').replace('>', '&gt;').replace('*', '').replace('#', '')
    msg = f"📊 <b>Daily IT Hygiene Digest</b>\n\n<b>Swarm Analysis:</b>\n{clean_analysis}"
    payload = {"chat_id": TELEGRAM_CHAT_ID, "text": msg, "parse_mode": "HTML"}
    req = urllib.request.Request(url, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json"})
    urllib.request.urlopen(req, timeout=15)

def main():
    raw_data = fetch_wazuh_alerts()
    minified_data = minify_data(raw_data)
    if minified_data: send_telegram(analyze_with_llm(minified_data))

if __name__ == "__main__": main()

2. ISO 27001 Compliance Query Adjustments

To create a weekly ISO 27001 script, copy the code above and replace the Elasticsearch query and minifier loops:

# Query
{"query": {"bool": {"must": [{"range": {"timestamp": {"gte": "now-7d", "lte": "now"}}}, {"wildcard": {"rule.compliance.iso_27001": "*"}}]}}, "size": 1000}

# Minification
for tag in source.get("rule", {}).get("compliance", {}).get("iso_27001", []):
    cve_list[f"ISO 27001 Control {tag} triggered on host '{agent}'"] += 1

Set them to run via Cron (e.g., 0 8 * * * /usr/local/bin/wazuh-daily-digest.py).

Phase 48: Immich Photo Server & Hardware Acceleration

Deploying Immich natively requires correctly passing your iGPU through to the containers to enable VA-API video transcoding and OpenVINO machine learning tasks. It also requires the pgvectors extension inside PostgreSQL.

The Complete Docker Compose File

Create a docker-compose.yml mapping both your local upload path and your read-only NAS share.

name: immich

services:
  immich-server:
    container_name: immich_server
    image: ghcr.io/immich-app/immich-server:${IMMICH_VERSION:-release}
    devices:
      - /dev/dri:/dev/dri # Pass iGPU for VA-API video transcoding
    volumes:
      - ${UPLOAD_LOCATION}:/data
      - /mnt/nas_photos:/usr/src/app/external/nas_photos:ro
      - /etc/localtime:/etc/localtime:ro
    env_file:
      - .env
    ports:
      - '2283:2283'
    depends_on:
      - redis
      - database
    restart: always

  immich-machine-learning:
    container_name: immich_machine_learning
    image: ghcr.io/immich-app/immich-machine-learning:${IMMICH_VERSION:-release}-openvino
    devices:
      - /dev/dri:/dev/dri # Pass iGPU for OpenVINO acceleration
    volumes:
      - model-cache:/cache
    env_file:
      - .env
    restart: always

  redis:
    container_name: immich_redis
    image: valkey/valkey:8-alpine
    restart: always

  database:
    container_name: immich_postgres
    image: ghcr.io/immich-app/postgres:14-vectorchord0.4.3-pgvectors0.2.0
    environment:
      POSTGRES_PASSWORD: ${DB_PASSWORD}
      POSTGRES_USER: ${DB_USERNAME}
      POSTGRES_DB: ${DB_DATABASE_NAME}
      POSTGRES_INITDB_ARGS: '--data-checksums'
    volumes:
      - ${DB_DATA_LOCATION}:/var/lib/postgresql/data
    shm_size: 128mb
    restart: always

volumes:
  model-cache:

Database Note: Ensure your .env file's DB_PASSWORD contains only alphanumeric characters (A-Za-z0-9) without hyphens to prevent initialization crashes.

Phase 49: Android ADB Debloat Automation Scripts

To safely disable telemetric bloatware, aggressive performance throttlers (like GOS), and sponsored feeds on Samsung One UI and Vivo Funtouch OS, utilize the Android Debug Bridge (ADB) via PowerShell.

1. Vivo Z1x (Funtouch OS 11) Debloat Script

Funtouch OS 11 throws DELETE_FAILED_USER_RESTRICTED errors when attempting to uninstall packages. Bypass this by stripping background permissions and disabling the packages instead.

.\adb shell "
cmd appops set com.bbk.cloud RUN_IN_BACKGROUND ignore; pm clear com.bbk.cloud;
cmd appops set com.vivo.pushservice RUN_IN_BACKGROUND ignore; pm clear com.vivo.pushservice;
cmd appops set com.vivo.gamecube RUN_IN_BACKGROUND ignore; pm clear com.vivo.gamecube;
cmd appops set com.vivo.game RUN_IN_BACKGROUND ignore; pm clear com.vivo.game;

cmd package disable-user --user 0 com.bbk.cloud;
cmd package disable-user --user 0 com.vivo.pushservice;
cmd package disable-user --user 0 com.vivo.gamecube;
cmd package disable-user --user 0 com.vivo.game;
"

2. Samsung One UI (Game Optimizing Service & Smart Suggestions)

The Game Optimizing Service requires an AppOps clear alongside a standard disable command to neutralize aggressive thermal throttling. Use this block for your Galaxy devices:

# Remove standard bloat
.\adb shell "
pm uninstall -k --user 0 com.facebook.system;
pm uninstall -k --user 0 com.facebook.appmanager;
pm uninstall -k --user 0 com.facebook.services;
pm uninstall -k --user 0 com.samsung.android.bixby.wakeup;
pm uninstall --user 0 com.samsung.discover;
pm uninstall --user 0 com.sec.android.mimage.avatarstickers;
"

# Freeze & Neutralize Game Optimizing Service (GOS)
.\adb shell pm disable-user --user 0 com.samsung.android.game.gos
.\adb shell cmd appops set com.samsung.android.game.gos RUN_IN_BACKGROUND ignore
.\adb shell pm clear com.samsung.android.game.gos

# Freeze & Neutralize Smart Suggestions
.\adb shell pm disable-user --user 0 com.samsung.android.smartsuggestions
.\adb shell cmd appops set com.samsung.android.smartsuggestions RUN_IN_BACKGROUND ignore
.\adb shell pm clear com.samsung.android.smartsuggestions

Documentation Framework Finalized

Phase 50: Monitoring Docker Containers with Wazuh

Wazuh uses a built-in docker-listener wodle to monitor container lifecycle events (start, stop, pause, pull)[cite: 284]. This configuration works for both a standalone Wazuh Agent and a local Wazuh Manager host listening to /var/run/docker.sock[cite: 284].

1. Install Python Docker SDK & Grant Permissions

The Docker listener module requires the Python Docker library to interact directly with the local Docker API socket[cite: 284]. It also requires the Wazuh service account to hold sufficient read/write permissions[cite: 284].

# Install the Python Docker SDK
sudo apt update && sudo apt install python3-pip python3-docker -y

# Add the wazuh system user to the docker group
sudo usermod -aG docker wazuh

2. Enable the Docker Listener Module

Open the main configuration file on your target host (/var/ossec/etc/ossec.conf) and add the <wodle name="docker-listener"> block inside the main <ossec_config> tag[cite: 284]:

<wodle name="docker-listener">
  <disabled>no</disabled>
  <interval>10m</interval>
  <attempts>5</attempts>
  <run_on_start>yes</run_on_start>
</wodle>

3. Restart Service & Verify

Restart the active daemon (sudo systemctl restart wazuh-manager or wazuh-agent) to initialize the module[cite: 284]. Stream the log file to confirm success:

sudo tail -n 30 /var/ossec/logs/ossec.log | grep -i docker

To test it, run a temporary container (docker run --rm hello-world). Open your Wazuh Dashboard, navigate to Modules > Docker (or use the Discover tab with the filter rule.groups: docker), and you will see the container creation and termination events logged in real-time[cite: 284].

Phase 51: Essential Sysadmin CLI Extras & MacBook Drivers

Rounding out your terminal toolkit requires a few extra utilities for system fetching, package management, and niche hardware driver fixes[cite: 284].

1. System Fetching & Inline Browsing

  • Visual System Fetch: Use fastfetch --logo none (or neofetch) to print a clean hardware and OS summary to your terminal[cite: 284].
  • Terminal Web Browser: If you need to read documentation or verify web server output entirely from a headless shell, install Lynx[cite: 284]:
    sudo apt install lynx -y

2. Flatpak Integration

If you are running a graphical environment and need to install sandboxed applications outside of the apt or snap ecosystems, initialize Flatpak[cite: 284]:

sudo apt install flatpak gnome-software-plugin-flatpak
flatpak remote-add --if-not-exists flathub https://dl.flathub.org/repo/flathub.flatpakrepo

3. MacBook Pro Linux Fixes (HFS+ & Drivers)

If you installed Linux on an old Intel MacBook, you must manually install specific drivers for thermals, keyboard backlights, and file systems[cite: 284].

# Grant Write Access to macOS HFS+ Drives
sudo apt install hfsprogs
sudo mount -t hfsplus -o rw,force /dev/sdX /mnt/mac_drive/

# Install Broadcom Wi-Fi Drivers
sudo apt install firmware-b43-installer broadcom-sta-dkms -y

# Fix Fan Curves & Power Management
sudo apt install -y tlp tlp-rdw mbpfan
sudo systemctl enable --now tlp
sudo systemctl enable --now mbpfan

# Fix Apple Backlit Keyboard
sudo apt install pommed

Phase 52: Native Remote Desktop Access (Remmina)

If you deployed the full Ubuntu GNOME desktop (rather than the headless XFCE/TigerVNC environment) and want to connect to it locally using the built-in screen sharing protocol, you can use the Remmina client[cite: 284].

Configuring Native Screen Sharing

  1. On the Ubuntu host machine, open Settings > Sharing and enable the global sharing toggle[cite: 284].
  2. Click on Remote Desktop and enable both Remote Desktop and Remote Control.
  3. Under the Authentication section, click the hamburger menu (three dots) or the security dropdown, and explicitly select Require a password (VNC Auth)[cite: 284].
  4. Set a secure password.

On your client machine, open Remmina, select the VNC protocol, enter the host's IP address, and authenticate with the password you just configured. This bypasses the need for manual systemd services by utilizing Ubuntu's built-in vino or gnome-remote-desktop servers[cite: 284].

Archive Complete

Every fragment of documentation, troubleshooting logs, and configuration scripts provided has now been exhaustively processed, sanitized, and published. Your home lab deployment manual is officially complete in its entirety.

Phase 53: Infrastructure as Code - Automating Nodes with Ansible

Once you have mastered manual deployments across your Proxmox VMs, Raspberry Pis, and remote servers, the ultimate evolution of a home lab is Infrastructure as Code (IaC). Instead of running shell scripts or manual configuration commands on each node individually, Ansible allows you to define your desired server state in centralized YAML playbooks and push changes across your entire fleet instantly.

1. Setting Up the Control Node

Install Ansible on your primary management machine (your administrative workstation or management VM):

sudo apt update
sudo apt install ansible -y

2. Defining the Inventory File

Create a working directory for your Ansible project (e.g., ~/homelab-ansible) and create an inventory file named hosts.ini. This file groups your servers so you can target them individually or as a fleet.

[proxmox_hosts]
pve-node ansible_host=<PROXMOX_IP> ansible_user=root

[docker_nodes]
server1454 ansible_host=<SERVER_IP> ansible_user=server1454

[pis]
node_pi ansible_host=<RASPI_IP> ansible_user=pi

[all:vars]
ansible_ssh_private_key_file=~/.ssh/id_ed25519

3. Writing a Master Setup Playbook

Create a playbook named site.yml to automate essential tasks across your Linux nodes—such as ensuring NetworkManager is running, updating packages, and installing diagnostic tools like btop and curl.

---
- name: Configure and Harden Homelab Nodes
  hosts: all
  become: yes
  tasks:
    - name: Update apt repo cache
      apt:
        update_cache: yes
        cache_valid_time: 3600

    - name: Install essential sysadmin utilities
      apt:
        name:
          - btop
          - curl
          - git
          - ufw
          - cifs-utils
        state: present

    - name: Ensure UFW firewall is active
      ufw:
        state: enabled
        policy: allow

4. Executing the Playbook

Test your connectivity and push the configuration across your entire infrastructure using a single command:

# Test SSH connection to all inventory hosts
ansible all -i hosts.ini -m ping

# Run the master configuration playbook
ansible-playbook -i hosts.ini site.yml

Ansible executes tasks idempotently, meaning it only applies changes if the target state differs from your playbook definition, ensuring your servers remain stable and predictably configured across every run.

Comments

Popular Posts

Marriage Registration Online steps [Tamil Nadu]

Chennai :MTC complaint cell Customer Care No.:+91-9445030516 /Toll Free : 18005991500

CBI arrests Indian mastermind behind Hire-a-Hacker service on FBI tip-off