Building the Ultimate Home Lab: The Staged Implementation Roadmap
Building the Ultimate Home Lab: The Staged Implementation Roadmap
Building a fully containerized, secure, and monitored home lab can feel overwhelming if approached all at once. To prevent cognitive overload, this master roadmap breaks the entire journey into 5 manageable, sequential stages. Each stage builds directly upon the last, taking you from bare hardware to an enterprise-grade self-hosted ecosystem.
Interactive Navigation Map
Stage 1: Foundation & Access
Hardware prep, Ubuntu Server setup, lid-sleep fixes, and Tailscale mesh VPN.
→ Jump to Stage 1 DetailsStage 2: Containers & Core Stacks
Docker Engine, Portainer CE, Pi-hole, Plex, and Home Assistant deployments.
→ Jump to Stage 2 DetailsStage 3: Headless Desktop & UI
TigerVNC systemd automation, XFCE4 desktop ricing, and Nginx reverse proxies.
→ Jump to Stage 3 DetailsStage 4: Telemetry & Alerting
Prometheus, Grafana dashboards, cAdvisor, Scaphandre power tracking, and Telegram alerts.
→ Jump to Stage 4 DetailsStage 5: Security, SIEM & Automation
Wazuh SIEM deployment, AI-powered SOC triage agents, and Ansible Infrastructure as Code (IaC).
→ Jump to Stage 5 DetailsStage 1: Hardware Prep, OS Foundation & Secure Access
Back to: ↑ Return to Navigation Map
Goal: Get your hardware online, prevent laptops from sleeping when closed, and establish secure encrypted remote access.
- BIOS Toggles: Enable VT-x/AMD-V virtualization, disable Secure Boot, set storage to AHCI.
- Lid-Sleep Bypass: Edit
/etc/systemd/logind.conf→ setHandleLidSwitch=ignore, thensudo systemctl restart systemd-logind. - Tailscale Mesh VPN: Run
curl -fsSL https://tailscale.com/install.sh | sh && sudo tailscale up --ssh.
Stage 2: Containerization & Core Infrastructure Stacks
Back to: ↑ Return to Navigation Map
Goal: Install Docker, Portainer, and core containerized services like Pi-hole, Plex, and Home Assistant.
- Docker Engine: Add official Docker GPG keys, configure
docker.sources, and install viaapt. - Portainer CE: Deploy graphical container management via Docker (Port 9443).
- Pi-hole Optimization: Cap database history to 7 days in
pihole-FTL.conf(MAXDBDAYS=7) to prevent SD card I/O starvation. - Plex & Home Assistant: Deploy Plex with SMB/CIFS mounts and Home Assistant using
network_mode: hostfor local device discovery.
Stage 3: Headless Desktop Environment & Web Security
Back to: ↑ Return to Navigation Map
Goal: Provide a lightweight VNC workspace and secure internal web apps behind authenticated Nginx reverse proxies.
- TigerVNC + XFCE4: Configure non-hanging systemd service files (omitting legacy
PIDFile) running at 16-bit color depth. - Desktop Ricing: Install Arc-Dark, Papirus icons, and Whisker Menu; disable the window compositor for maximum network speed.
- Nginx Proxy & Basic Auth: Bind custom web apps to localhost (
127.0.0.1) and wrap them in Nginx protected by.htpasswd.
Stage 4: Telemetry, Hardware Observability & Alerting
Back to: ↑ Return to Navigation Map
Goal: Collect performance metrics, track container behavior, measure power draw, and push notifications to your phone.
- Prometheus & cAdvisor: Scrape system health metrics and monitor individual container resource consumption.
- Scaphandre Power Profiling: Read Intel/AMD RAPL hardware registers to track exact milliwatt power draw.
- Alertmanager & Telegram: Configure Prometheus rules and Alertmanager to push HTML-formatted push notifications to Telegram.
Stage 5: Enterprise Security, SIEM & Infrastructure as Code
Back to: ↑ Return to Navigation Map
Goal: Implement centralized security logging, AI-assisted log analysis, and automated management playbooks.
- Wazuh SIEM Deployment: Install Wazuh Manager, Indexer, and Dashboard, and deploy endpoint security agents across your fleet.
- AI SOC Triage: Integrate Wazuh with a LiteLLM gateway for automated threat triage and daily IT hygiene digests.
- Ansible IaC: Transition from manual shell execution to centralized YAML playbooks to push configuration updates idempotently.
Building the Ultimate Home Lab: From Old Hardware to an Enterprise-Grade Ecosystem
What starts as a simple desire to run a network-wide ad blocker often snowballs into a full-scale IT infrastructure project. Over the course of building this home lab, the environment has evolved from a repurposed headless machine into a highly optimized, secure, and fully monitored Docker ecosystem.
This journey involved stripping away unnecessary overhead from the base OS, securing access via mesh VPNs, deploying an ultra-lightweight remote graphical interface (XFCE via TigerVNC), and containerizing an entire suite of services. To ensure everything runs smoothly, a sophisticated telemetry stack was layered on top, tracking everything from DNS I/O bottlenecks to the exact milliwatt power draw of individual background processes.
The Master Configuration Topics
If you are looking to replicate this environment, these are the core architectural pillars covered in the deployment logs:
- Networking & Secure Access: Netplan configuration, NetworkManager routing, and Tailscale mesh VPN integration.
- Remote Administration: TigerVNC systemd automation, D-Bus session wrappers, and XFCE4 desktop optimization for low-bandwidth links.
- Virtualization & Containerization: Bare-metal Docker Engine deployment, Portainer CE orchestration, and PUID/PGID permission mapping for network shares.
- Core Services: Pi-hole (DNS optimization and SQLite database capping), Plex Media Server (CIFS/USB mounts), Home Assistant, and Nginx reverse proxies.
- Security Information and Event Management (SIEM): Wazuh server deployment, automated agent provisioning, and endpoint telemetry.
- Observability & Power Tuning: Prometheus scraping, cAdvisor container metrics, Scaphandre CPU RAPL power profiling, and unified Grafana dashboards.
Phase 1: Transforming an Old Laptop into a Server
An old laptop is arguably the best entry-level home lab hardware. It has a built-in UPS (the battery), an integrated keyboard/monitor for troubleshooting, and is highly energy-efficient.
1. Install the Base OS
Flash a USB drive with Ubuntu Server LTS and install it on the laptop. Choose the minimal installation option to keep background processes to an absolute minimum.
2. Prevent Sleep on Lid Close
By default, laptops go to sleep when you close the lid. To keep your server running when tucked away on a shelf, modify the systemd logind configuration:
sudo nano /etc/systemd/logind.conf
Find the line #HandleLidSwitch=suspend, uncomment it, and change it to ignore:
HandleLidSwitch=ignore
Apply the changes by restarting the service:
sudo systemctl restart systemd-logind
3. Secure Networking via Tailscale
Instead of opening ports on your home router, install Tailscale to create a secure, encrypted mesh network. This allows you to SSH into your server from anywhere in the world.
curl -fsSL https://tailscale.com/install.sh | sh
sudo tailscale up --ssh
Phase 2: Deploying the Docker Engine
Containerization is the lifeblood of a modern home lab. It isolates your applications, prevents dependency conflicts, and makes backups trivial.
1. Install Official Docker
Install the official Docker repository to ensure you receive the latest engine updates:
sudo install -m 0755 -d /etc/apt/keyrings
sudo curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.asc
sudo chmod a+r /etc/apt/keyrings/docker.asc
sudo tee /etc/apt/sources.list.d/docker.sources <<EOF
Types: deb
URIs: https://download.docker.com/linux/ubuntu
Suites: $(. /etc/os-release && echo "$VERSION_CODENAME")
Components: stable
Architectures: $(dpkg --print-architecture)
Signed-By: /etc/apt/keyrings/docker.asc
EOF
sudo apt update
sudo apt install docker-ce docker-ce-cli containerd.io docker-compose-plugin -y
2. Deploy Portainer CE
Manage your containers visually by deploying Portainer.
docker volume create portainer_data
docker run -d -p 9443:9443 --name portainer --restart=always \
-v /var/run/docker.sock:/var/run/docker.sock \
-v portainer_data:/data \
portainer/portainer-ce:latest
Navigate to https://<YOUR_SERVER_IP>:9443 to configure your admin profile.
Phase 3: Core Services & Automation
With the infrastructure established, you can begin deploying production stacks. Group related services together using Docker Compose.
- Home Assistant: Centralize your home automation. Deploy this container using
network_mode: hostto ensure it can seamlessly discover IoT devices, smart plugs, and sensors across your local subnet. - Nginx Reverse Proxy: Route traffic to your various internal services. Secure internal toolkits and dashboards by configuring Nginx proxy blocks with HTTP Basic Authentication (htpasswd) to prevent unauthorized local access.
- Immich & PostgreSQL: Build a self-hosted photo and video backup solution. Deploy the Immich machine learning and web containers backed by a dedicated PostgreSQL database container for metadata storage.
- Wazuh SIEM: Protect your fleet. Deploy the Wazuh manager to ingest security logs, monitor endpoints over your Tailscale VPN network, and route high-severity security events directly to a Telegram bot.
Phase 4: Global Observability
A home lab is flying blind without proper telemetry. Deploying a Prometheus and Grafana stack allows you to catch hardware bottlenecks and software memory leaks before they crash your server.
1. Scrape Host & Container Metrics
Run Node Exporter on the host OS to track disk I/O, CPU wait times, and RAM saturation. Deploy cAdvisor as a container to monitor the exact CPU and memory footprint of every other Docker container running on the system. If your hardware supports it, deploy Scaphandre to pull direct RAPL wattage metrics from the CPU package.
2. Centralize with Prometheus & Grafana
Configure your prometheus.yml file to scrape these targets every 15 seconds. Connect Grafana to the Prometheus data source to build unified IT hygiene dashboards.
3. Automated Alerting
Write PromQL rules in Prometheus to detect anomalous behavior (e.g., CPU load > 90% for 5 minutes, or a container entering a crash loop). Route these rules through Alertmanager, configuring a telegram_configs receiver to instantly push HTML-formatted incident alerts to your mobile device.
Phase 5: Configuring the Headless VNC Environment
Traditional VNC setups struggle with modern hardware-accelerated desktop environments like GNOME. By pairing TigerVNC with the ultra-lightweight XFCE4 desktop, you eliminate crashes, bypass systemd cgroup conflicts, and drastically reduce idle RAM usage.
1. The Startup Script
Create the TigerVNC configuration directory and the initialization script to launch the desktop inside a dedicated D-Bus session.
mkdir -p ~/.config/tigervnc
nano ~/.config/tigervnc/xstartup
Paste this optimized execution block. The xhost +local: line is critical to prevent sandbox permission errors when launching Snap applications (like Chromium or CCTV viewers) inside the VNC window:
#!/bin/sh
# 1. Clear out host environment session leaks
unset SESSION_MANAGER
unset DBUS_SESSION_BUS_ADDRESS
# 2. Load standard desktop resources
[ -r $HOME/.Xresources ] && xrdb $HOME/.Xresources
xsetroot -solid grey
vncconfig -iconic &
# 3. Grant local sandbox permissions
export XDG_RUNTIME_DIR=/run/user/$(id -u)
xhost +local:
# 4. Launch XFCE inside a private D-Bus container loop
exec dbus-launch --exit-with-session startxfce4
Make the script executable: chmod +x ~/.config/tigervnc/xstartup
2. The Systemd Automation File
To ensure the remote desktop survives power outages, bind it to a system-level service. We use 16-bit color depth (-depth 16) to slash network transmission overhead by 33% without sacrificing readability.
sudo nano /etc/systemd/system/vncserver@.service
[Unit]
Description=Start TigerVNC server at startup
After=syslog.target network.target
[Service]
Type=forking
User=<YOUR_USERNAME>
WorkingDirectory=/home/<YOUR_USERNAME>
Environment=HOME=/home/<YOUR_USERNAME>
ExecStartPre=-/usr/bin/vncserver -kill :%i > /dev/null 2>&1
ExecStart=/usr/bin/vncserver -depth 16 -geometry 1920x1080 -localhost no :%i
ExecStop=/usr/bin/vncserver -kill :%i
[Install]
WantedBy=multi-user.target
Start and enable the service on port 5901 (Display :1):
sudo systemctl daemon-reload
sudo systemctl enable --now vncserver@1.service
Phase 6: Pi-hole Deployment & I/O Optimization
Deploying Pi-hole natively in Docker is simple, but long-term database bloat can cause severe I/O wait bottlenecks, driving the CPU Load Average up and hanging DNS resolutions.
1. The Compose File
Create the Pi-hole stack. We map the host port to 8443 to avoid conflicts with existing web servers (like Nginx) holding port 443.
services:
pihole:
container_name: pihole
image: pihole/pihole:latest
ports:
- "53:53/tcp"
- "53:53/udp"
- "80:80/tcp"
- "8443:443/tcp"
environment:
TZ: 'Europe/London'
WEBPASSWORD: '<YOUR_SECURE_PASSWORD>'
FTLCONF_dns_listeningMode: 'ALL'
volumes:
- './etc-pihole:/etc/pihole'
- './etc-dnsmasq.d:/etc/dnsmasq.d'
cap_add:
- SYS_NICE
restart: unless-stopped
2. Fixing the "High Load Average" Database Freeze
If your Grafana dashboard shows CPU Load Averages spiking past 4.0 while actual CPU utilization is under 50%, your storage disk is choking on the massive pihole-FTL.db history file (which defaults to storing 365 days of queries).
Restrict the database lifespan to 7 days to maintain lightning-fast I/O performance:
sudo systemctl stop pihole-FTL
sudo rm /etc/pihole/pihole-FTL.db
sudo nano /etc/pihole/pihole-FTL.conf
Add this parameter to the config file:
MAXDBDAYS=7
Restart the service to generate a fresh, optimized database: sudo systemctl start pihole-FTL
Phase 7: Plex Media Server (NAS vs USB Storage)
Plex requires host networking to broadcast discovery protocols successfully. This configuration supports securely mounting network shares directly into Docker without editing the host OS /etc/fstab.
Option A: Mounting a CIFS Network Share
Ensure your host has cifs-utils installed. Create a .env file in the same directory as your compose file containing NAS_USER=<USER> and NAS_PASS=<PASS>.
services:
plex:
image: lscr.io/linuxserver/plex:latest
container_name: plex
network_mode: host
environment:
- PUID=1000
- PGID=1000
- TZ=Europe/London
- VERSION=docker
volumes:
- ./config:/config
- nas_media:/data/nas_shared
restart: unless-stopped
volumes:
nas_media:
driver: local
driver_opts:
type: cifs
o: "username=${NAS_USER},password=${NAS_PASS},iocharset=utf8,vers=3.0"
device: "//<NAS_IP_ADDRESS>/<SHARE_NAME>"
Option B: Direct USB Bind Mount
If utilizing a locally attached USB hard drive, remove the global volumes block entirely and pass the absolute hardware path.
volumes:
- ./config:/config
- /mnt/usb_drive:/data/usb_media
Crucial Fix: If Plex sees the folders but no media populates, ensure the host directory has explicit read/execute permissions: sudo chmod -R 755 /mnt/usb_drive.
Phase 8: Hardware Power Tuning & Scaphandre Telemetry
Mini PCs and laptops often waste electricity polling idle hardware. By forcing PCIe deeper sleep states and mapping physical CPU sensors via Scaphandre, you can slash idle draw by up to 50%.
1. Force PCIe ASPM Sleep States
If the BIOS misreports ASPM availability, manually force the L0s and L1 deep sleep states using the AutoASPM utility.
git clone https://github.com/K0mh0li0/autoaspm.git
sudo nano /etc/systemd/system/autoaspm.service
[Unit]
Description=AutoASPM PCIe Power Saving Script
After=network.target
[Service]
Type=oneshot
ExecStart=/usr/bin/python3 /path/to/autoaspm/pkgs/autoaspm.py
RemainAfterExit=yes
[Install]
WantedBy=multi-user.target
Enable it on boot: sudo systemctl enable --now autoaspm.service
2. Deploy Scaphandre for Power Metrics
Scaphandre reads the Intel/AMD RAPL hardware registers to output exact milliwatt usage into Prometheus.
wget https://github.com/hubblo-org/scaphandre/releases/download/v1.0.2/scaphandre_v1.0.2-deb12_amd64.deb
sudo dpkg -i scaphandre_v1.0.2-deb12_amd64.deb
sudo nano /etc/systemd/system/scaphandre.service
[Unit]
Description=Scaphandre Prometheus Exporter
After=network.target
[Service]
Type=simple
ExecStart=/usr/bin/scaphandre prometheus -p 8081
Restart=always
[Install]
WantedBy=multi-user.target
Start the exporter: sudo systemctl enable --now scaphandre.service. Add <SERVER_IP>:8081 to your prometheus.yml scrape configs.
Phase 9: Real-Time Alerts via Telegram
When an internal service fails or your CPU maxes out, Alertmanager receives the alert from Prometheus and pushes a formatted HTML payload directly to a Telegram bot.
1. Configure Alertmanager
Create the Telegram receiver configuration, passing the full bot API token and the negative integer Chat ID (for group chats).
sudo nano /etc/alertmanager/alertmanager.yml
route:
group_by: ['alertname', 'instance']
group_wait: 30s
group_interval: 5m
repeat_interval: 3h
receiver: 'telegram-bot'
receivers:
- name: 'telegram-bot'
telegram_configs:
- send_resolved: true
bot_token: '<YOUR_BOT_TOKEN_ID:SECRET_STRING>'
chat_id: -<YOUR_GROUP_CHAT_ID>
parse_mode: 'HTML'
message: |-
{{ if eq .Status "firing" }} 🔥 FIRING {{ else }} ✅ RESOLVED {{ end }}: {{ .CommonLabels.alertname }}
Host: {{ .CommonLabels.instance }}
Details: {{ .CommonAnnotations.description }}
2. Link Prometheus to Alertmanager
Update prometheus.yml to forward alerts to the Alertmanager port (9093).
alerting:
alertmanagers:
- static_configs:
- targets: ['localhost:9093']
Restart the stack: sudo systemctl restart prometheus alertmanager. You now possess an enterprise-grade monitoring, orchestration, and remote-management fleet configured entirely from scratch.
Phase 10: Enterprise SIEM Deployment (Wazuh)
Wazuh provides unified XDR and SIEM protection, aggregating logs and endpoint telemetry. Deploying the manager and agents properly requires strict repository configurations.
1. Install the Wazuh Manager
Use the official automated script to deploy the Wazuh Indexer, Server, and Dashboard on a single node.
curl -sO https://packages.wazuh.com/4.x/wazuh-install.sh
sudo bash wazuh-install.sh -a
Once finished, securely store the generated admin password and access the web interface at https://<YOUR_SERVER_IP>.
2. Deploy the Wazuh Agent (Fixing Repo Errors)
If you encounter 404 Not Found errors when updating apt, it indicates a corrupted repository file pointing to the main marketing site instead of the package mirror. Clean it up and install the agent:
# Remove corrupted lists
sudo rm -f /etc/apt/sources.list.d/wazuh.list
sudo rm -rf /var/lib/apt/lists/*wazuh*
# Import the official cryptographic trust key
curl -s https://packages.wazuh.com/key/GPG-KEY-WAZUH | gpg --no-default-keyring --keyring gnupg-ring:/usr/share/keyrings/wazuh.gpg --import
sudo chmod 644 /usr/share/keyrings/wazuh.gpg
# Add the verified repository URL
echo "deb [signed-by=/usr/share/keyrings/wazuh.gpg] https://packages.wazuh.com/4.x/apt/ stable main" | sudo tee /etc/apt/sources.list.d/wazuh.list
# Update and install the agent, passing the manager IP
sudo apt-get update
export WAZUH_MANAGER="<YOUR_WAZUH_MANAGER_IP>"
sudo -E apt-get install wazuh-agent -y
Enable and start the agent: sudo systemctl enable --now wazuh-agent
3. Configure Grafana OpenSearch Data Source
Because the Wazuh Indexer is built on OpenSearch architecture, you can visualize security telemetry directly in Grafana. Install the OpenSearch plugin via the Grafana CLI:
grafana-cli plugins install grafana-opensearch-datasource
sudo systemctl restart grafana-server
In Grafana, add an OpenSearch data source pointing to https://<YOUR_INDEXER_IP>:9200. Enable Basic Auth with your cluster credentials and toggle Skip TLS Verify if using default self-signed certificates.
Phase 11: Network Diagnostics & Hardware Firmware
A home lab is only as stable as its underlying network and hardware firmware. When DNS lags or hardware exhibits vulnerabilities, use these terminal-native tools to diagnose and patch the system.
1. Automated DNS Performance Benchmarking
Slow DNS lookups introduce silent latencies that mirror hard network drops. Create dns_audit.sh to automatically benchmark your system's default resolver against public Anycast infrastructure (like Cloudflare or Google).
#!/bin/bash
# Ensure environment has diagnostic utilities
if ! command -v dig &> /dev/null; then
sudo apt update -qq && sudo apt install dnsutils -y -qq
fi
clear
echo "==================================="
echo " PHASE 1: ACTIVE SYSTEM RESOLVER"
echo "==================================="
resolvectl status | grep "DNS Servers" -A 2
echo -e "\n==================================="
echo " PHASE 2: JITTER ASSESSMENT"
echo "==================================="
for i in {1..5}; do
TIME_MS=$(dig google.com | grep "Query time" | awk '{print $4}')
echo "-> Query Iteration $i: ${TIME_MS} ms"
sleep 0.5
done
echo -e "\n==================================="
echo " PHASE 3: PUBLIC RESOLVER BENCHMARK"
echo "==================================="
echo "System Default Link: $(dig example.com | grep "Query time" | awk '{print $4, $5}')"
echo "Cloudflare (1.1.1.1): $(dig example.com @1.1.1.1 | grep "Query time" | awk '{print $4, $5}')"
echo "Google (8.8.8.8): $(dig example.com @8.8.8.8 | grep "Query time" | awk '{print $4, $5}')"
Run the script: chmod +x dns_audit.sh && ./dns_audit.sh
2. Terminal Wi-Fi Management (nmcli)
Manage headless Wi-Fi connections via the NetworkManager Command Line Interface:
- Scan for Networks:
nmcli device wifi list - Connect to a Network:
nmcli device wifi connect "<SSID>" password "<PASSWORD>" - Visual Interface Alternative: Run
sudo nmtuifor a curses-based terminal GUI. (Note: If you get a Polkit "not authorized" error over SSH, you must create a custom Polkit rule in/etc/polkit-1/rules.d/).
3. Securing System Firmware (fwupdmgr)
Linux manages motherboard BIOS and UEFI dbx updates securely using the Firmware Update Daemon (fwupd). Apply critical security revocation lists to prevent Secure Boot bypass exploits.
# See available updates
fwupdmgr get-upgrades
# Fetch latest signatures from Linux Vendor Firmware Service
fwupdmgr refresh
# Apply the updates (requires confirmation and reboot)
fwupdmgr update
Phase 12: Hardware Auditing & Custom Power Automation
When repurposing consumer hardware like an All-In-One PC or old laptop, physical power management is paramount. Integrated screens that cannot be physically detached will waste massive amounts of electricity if left illuminated 24/7.
1. Hardware Profiling over SSH
Identify your bare-metal architecture without needing physical access to the chassis:
- Motherboard & BIOS Details:
sudo dmidecode -t baseboard - CPU Topology & Capabilities:
lscpu - Comprehensive Hardware Tree:
sudo lshw -short(or useinxi -Fzfor a colorized summary that anonymizes MAC addresses).
2. Automated Display Power Script
For All-in-One machines or laptops, the physical screen backlight must be overridden at the kernel level via sysfs. Save the following script as setup_screen.sh to automatically detect the backlight interface, inject manual bash aliases, and deploy root-level cron jobs to shut the screen off overnight.
#!/bin/bash
# ALL-IN-ONE UBUNTU SERVER DISPLAY CONFIGURATOR
if [ "$EUID" -ne 0 ]; then
echo "[-] Please run this script with sudo."
exit 1
fi
echo "[+] Starting automated display configuration..."
# 1. Detect Backlight Interface
BACKLIGHT_DIR="/sys/class/backlight"
INTERFACE=$(ls "$BACKLIGHT_DIR" | head -n 1)
if [ -z "$INTERFACE" ]; then
echo "[-] Error: No backlight controller found."
exit 1
fi
BACKLIGHT_PATH="$BACKLIGHT_DIR/$INTERFACE"
MAX_BRIGHTNESS=$(cat "$BACKLIGHT_PATH/max_brightness")
echo "[+] Detected interface: $INTERFACE (Max: $MAX_BRIGHTNESS)"
# 2. Inject Bash Aliases for Manual Control
REAL_USER=${SUDO_USER:-$USER}
BASHRC_PATH="$(eval echo ~$REAL_USER)/.bashrc"
sed -i '/alias screenoff=/d' "$BASHRC_PATH"
sed -i '/alias screenon=/d' "$BASHRC_PATH"
echo "alias screenoff='echo 0 | sudo tee $BACKLIGHT_PATH/brightness'" >> "$BASHRC_PATH"
echo "alias screenon='echo $MAX_BRIGHTNESS | sudo tee $BACKLIGHT_PATH/brightness'" >> "$BASHRC_PATH"
chown "$REAL_USER":"$REAL_USER" "$BASHRC_PATH"
# 3. Configure Console Idle Blanking
setterm --blank 1 --term linux < /dev/tty1 2>/dev/null || true
# 4. Install Root Cron Jobs for Automation (Off at 22:00, On at 07:00)
TMP_CRON=$(mktemp)
crontab -l > "$TMP_CRON" 2>/dev/null || true
sed -i '\#/sys/class/backlight/#d' "$TMP_CRON"
echo "0 22 * * * echo 0 > $BACKLIGHT_PATH/brightness" >> "$TMP_CRON"
echo "0 7 * * * cat $BACKLIGHT_PATH/max_brightness > $BACKLIGHT_PATH/brightness" >> "$TMP_CRON"
crontab "$TMP_CRON"
rm "$TMP_CRON"
echo "[+] SETUP COMPLETE! Run 'source ~/.bashrc' to activate aliases."
Make it executable and deploy it:
chmod +x setup_screen.sh
sudo ./setup_screen.sh
source ~/.bashrc
Phase 13: Identity Management & Homepage Dashboard
When containers write data to your host machine's file system, they often do so as the root user by default. This causes severe "Permission Denied" (EACCES) errors when your normal user account tries to edit configuration files. To fix this, you must map your Process User ID (PUID) and Process Group ID (PGID).
1. Find Your PUID and PGID
Run the id command in your host terminal to output your user metrics:
id -u
id -g
Note these numbers (typically 1000). You will inject them into your Docker Compose files as environment variables.
2. Deploying Homepage (Resolving EACCES Path Errors)
Homepage is a highly customizable starting screen for your lab. A common deployment mistake is mapping the local volume to /config internally, which causes permission failures. The correct internal path for the Homepage container is /app/config.
services:
homepage:
image: ghcr.io/gethomepage/homepage:latest
container_name: homepage
ports:
- "3001:3000"
volumes:
# THE FIX: Target /app/config internally
- ./homepage/config:/app/config
restart: always
environment:
- PUID=1000
- PGID=1000
- HOMEPAGE_ALLOWED_HOSTS=localhost:3001,127.0.0.1:3001,<YOUR_SERVER_IP>:3001
If you already encountered an error, ensure your local user reclaims ownership of the created folder before spinning the container back up:
sudo chown -R 1000:1000 ./homepage
chmod -R 755 ./homepage
docker compose up -d
Phase 14: Deploying Home Assistant
Home Assistant is unique because it relies heavily on local network broadcasts (mDNS, UPnP) to automatically discover smart bulbs, TVs, and IoT bridges. Standard Docker bridge networks block these broadcasts. The solution is network_mode: host.
The Optimized Compose Configuration
Create your docker-compose.yml for Home Assistant. Note the absence of a ports: block, as host mode binds directly to your server's NIC on port 8123.
services:
homeassistant:
container_name: homeassistant
image: ghcr.io/home-assistant/home-assistant:stable
volumes:
- ./config:/config
# Mount hardware timezone file for accurate sun/schedule automations
- /etc/localtime:/etc/localtime:ro
environment:
- TZ=Asia/Kolkata
network_mode: host
restart: unless-stopped
Start the container and navigate to http://<YOUR_SERVER_IP>:8123 to begin the onboarding wizard.
Phase 15: Advanced Server Triage & Automated Power Recovery
When your server slows down but standard htop shows low CPU usage, you are likely experiencing I/O Wait or Network exhaustion. Equip your terminal with modern diagnostic alternatives.
1. The Modern Terminal Toolkit
- btop: The most visually stunning and responsive TUI for macro-level system metrics (CPU, RAM, Disks, Network).
sudo apt install btop - iotop & iostat: Crucial for identifying disk saturation. If
iostat -xz 1shows highw_awaittimes and 100%%util, your mechanical drive is drowning in random writes. Runsudo iotop -oto expose the exact process hogging the disk. - iftop & nload: Use
sudo nloadfor a macro-level speedometer of your inbound/outbound traffic, andsudo iftopto identify the specific IP addresses consuming your bandwidth.
2. Automating AC Power Recovery (BIOS Level)
If your home experiences a power outage, your headless server will remain off when power returns unless instructed otherwise by the motherboard. You cannot configure this from Linux; it must be done in the BIOS/UEFI.
- Reboot the server and tap
DelorF2to enter the BIOS (e.g., MSI Click BIOS 5). - Switch to Advanced Mode (F7).
- Navigate to Settings > Advanced > Power Management Setup.
- Locate Restore after AC Power Loss and change it from Power Off to Power On.
Now, the moment your smart plug or UPS restores power, the server will instantly pull up the Ubuntu bootloader without physical intervention.
Phase 16: Tuning Grafana Variables for cgroups v2
When utilizing modern Linux environments running cgroups v2, cAdvisor exports metrics slightly differently than legacy systems. If you import popular Grafana dashboards (like ID 14282 or 193) and encounter "No Data", it is because the dashboard's internal PromQL queries are looking for a name= label, while modern cAdvisor maps container endpoints to the id= label (e.g., /system.slice/docker.service).
1. Fixing Dropdown Variable Queries
Navigate to Dashboard Settings > Variables and update the Classic Query expressions to pull from the correct system slices.
- Host Variable:
label_values(container_cpu_usage_seconds_total, instance) - Container Variable: Filter out raw background processes by enforcing the
idtag mapping:
label_values(container_cpu_usage_seconds_total{instance=~"$host", id=~"/system.slice/.*"}, id)
2. Updating Panel Metric Queries
Edit your individual visualization panels to utilize the id grouping so that each line correlates cleanly to a running service.
# CPU Usage Query Fix
sum(rate(container_cpu_usage_seconds_total{instance=~"$host", id=~"$container"}[5m])) by (id) * 100
# Memory Working Set Query Fix
container_memory_working_set_bytes{instance=~"$host", id=~"$container"}
3. Recommended Fleet Dashboards
To avoid spending hours building dashboards from scratch, import these industry-standard IDs via Dashboards > New > Import:
- ID 1860 (Node Exporter Full): The ultimate single-node deep dive for CPU, RAM, Disk, and Network hardware metrics.
- ID 15172 (Fleet Overview): A consolidated scoreboard design allowing you to monitor multiple nodes (Proxmox, Raspberry Pi, Ubuntu servers) side-by-side.
- ID 14282 (cAdvisor Exporter): Cleanly visualizes isolated Docker container performance footprint.
You have successfully traversed the journey from flashing a base OS to commanding a multi-node, monitored, and automated home lab ecosystem. Happy self-hosting!
Phase 17: Fixing Linux LAN Drivers & APT Repositories
When working with older consumer hardware or adding third-party applications, you will inevitably run into driver conflicts and GPG key expirations. Here is how to navigate the most common roadblocks.
1. The Realtek LAN link=no Bug
If your Ethernet cable is plugged in, but running sudo lshw -C network or ethtool shows link=no, you have likely hit a known kernel conflict with Realtek RTL8111/8168 chips. The default open-source r8169 driver frequently fails to negotiate a link. The fix is to compile the official proprietary driver:
sudo apt update
sudo apt install r8168-dkms -y
sudo reboot
2. Fixing GPG NO_PUBKEY and Duplicate Sources
If running apt update throws errors like "configured multiple times" or "NO_PUBKEY" (common with Brave Browser, Docker, or Wazuh), your sources list is corrupted. Clean the slate and re-import the secure token.
# 1. Delete the corrupted files
sudo rm -f /etc/apt/sources.list.d/brave-browser-release.list
sudo rm -f /etc/apt/sources.list.d/brave-browser-release.sources
sudo rm -f /usr/share/keyrings/brave-browser-archive-keyring.gpg
# 2. Re-download the official key
sudo curl -fsSLo /usr/share/keyrings/brave-browser-archive-keyring.gpg https://brave-browser-apt-release.s3.brave.com/brave-browser-archive-keyring.gpg
# 3. Add the clean source
echo "deb [signed-by=/usr/share/keyrings/brave-browser-archive-keyring.gpg] https://brave-browser-apt-release.s3.brave.com/ stable main" | sudo tee /etc/apt/sources.list.d/brave-browser-release.list
sudo apt update
3. Snap Packages Crashing in VNC (Signal 6)
Ubuntu installs Chromium via Snap. Snaps enforce a strict sandbox that requires applications to run inside a user cgroup slice. Because our VNC server runs as a system service, launching a Snap app results in a sudden crash. Bypass this by exporting the runtime directory before execution:
export XDG_RUNTIME_DIR=/run/user/$(id -u)
chromium-browser &
Alternative: Ditch the Snap entirely and download the raw .deb package (e.g., Google Chrome) which ignores cgroup restrictions.
Phase 18: Taming VM Insomnia on Proxmox
If you transition your physical setup into a virtualized Proxmox environment, you might notice VMs pulling 10W-15W at "idle." This is known as VM Insomnia—background polling forces the physical CPU to wake from deep C-states thousands of times per second. Apply these 4 tweaks to slash idle draw.
1. Strip Unnecessary Virtual Hardware
When you create a VM, Proxmox attaches virtual audio controllers and USB hubs. The guest kernel continuously polls these empty devices, burning CPU cycles.
- In the Proxmox UI, go to the VM's Hardware tab.
- Remove the Audio Device and USB Controllers (unless explicitly needed).
- Change the Display to Serial terminal 0 if managing entirely via SSH/Tailscale.
2. Tame Docker Container Polling
Tools like Wazuh and NetAlertX are aggressive network scanners. If they ping the subnet every 5 seconds, the CPU never sleeps.
- NetAlertX: Increase the subnet scan interval in the Web UI from 2 minutes to 15+ minutes.
- Wazuh: Open
/var/ossec/etc/ossec.confand change the<syscheck>file integrity monitoring frequency to 12 or 24 hours (43200or86400seconds).
3. Apply PowerTOP Auto-Tune
Because you passed the physical CPU through to the VM (qm set <VM_ID> --cpu host), the guest OS can manage its virtual power states. Install and run PowerTOP inside the guest:
sudo apt update && sudo apt install powertop -y
sudo powertop --auto-tune
(Add @reboot /usr/sbin/powertop --auto-tune to the root crontab to make this persistent).
4. Enable QEMU Guest Agent
The Guest Agent allows the Proxmox host and Ubuntu guest to negotiate power states efficiently. Enable it in the Proxmox VM Options tab, then install it inside the guest:
sudo apt install qemu-guest-agent -y
sudo systemctl enable --now qemu-guest-agent
Conclusion: Your Home Lab is Now Production-Ready
By following this comprehensive series, you have successfully transformed aging hardware into a secure, power-efficient, and highly observable enterprise-grade environment. From fighting APT repository errors and resolving VNC permission sandboxes, to tracking millimeter-level power consumption with Scaphandre and Grafana, your infrastructure is now optimized from the bare metal up to the application layer.
Happy Self-Hosting!
Phase 19: XFCE "Ricing" (Aesthetics & Speed Optimization)
Out of the box, XFCE can look like it belongs in 2004, and its default animations can cripple a VNC network stream[cite: 156, 160]. This blueprint transforms the interface into a sleek, flat, modern workspace while slashing network latency.
1. The Speed Layer (Optimize the Network Pipe)
To make VNC snappy, minimize the number of pixels changing during animations[cite: 156].
- Kill the Compositor: Open Settings > Window Manager Tweaks > Compositor and uncheck Enable display compositing[cite: 156, 160]. This stops VNC from wasting bandwidth rendering drop shadows and window fades[cite: 156, 160].
- Solid Wallpaper: High-resolution wallpapers require heavy compression[cite: 156]. Right-click the desktop, select Desktop Settings, and change the Style to None (Solid color) with a dark grey hex code[cite: 156, 160].
2. The Aesthetics Layer (Modernizing the UI)
Install the modern visual assets directly from the terminal[cite: 156, 160]:
sudo apt update
sudo apt install arc-theme papirus-icon-theme xfce4-whiskermenu-plugin -y
Apply the new skin via the XFCE Settings menu[cite: 156, 160]:
- Main Theme: Settings > Appearance > Style → Arc-Dark[cite: 156, 160].
- Icons: Settings > Appearance > Icons → Papirus-Dark[cite: 156, 160].
- Window Borders: Settings > Window Manager > Style → Arc-Dark[cite: 156, 160].
- Font Smoothing: Settings > Appearance > Fonts → Check Enable Anti-Aliasing (Hinting: Slight, Sub-pixel: RGB)[cite: 156, 160].
3. Panel & Whisker Menu Upgrade
Replace the basic application menu with the Whisker Menu (a ChromeOS/Windows-style launcher)[cite: 156].
- Right-click the old panel menu button and select Remove[cite: 156, 160].
- Right-click the panel → Panel > Add New Items → Select Whisker Menu[cite: 156, 160].
- Move it to the far left. Open Panel Preferences, uncheck Lock panel, drag the taskbar to the bottom of the screen, and increase the Row Size to 38px[cite: 156, 160].
Phase 20: Handling Video Tearing in VNC
Disabling the compositor creates a blazing-fast remote desktop, but it completely breaks Vertical Synchronization (VSync)[cite: 157]. If you attempt to watch a video over VNC, the player will drop frames asynchronously, causing the video to split or "tear" horizontally across the screen[cite: 157].
If you explicitly require video playback on a headless server, you must trade some raw performance for rendering accuracy[cite: 157].
1. Force the Xpresent Blanking Protocol
Re-enable the compositor, but force XFCE to use a low-overhead sync method[cite: 157]:
# Run this in your VNC terminal
xfconf-query -c xfwm4 -p /general/vblank_mode -s xpresent
Then, go back to Window Manager Tweaks > Compositor and re-check Enable display compositing[cite: 157].
2. Disable Player Hardware Acceleration
Hardware acceleration (OpenGL/VA-API) crashes over virtual framebuffers[cite: 157]. You must force your media player to use raw software rendering[cite: 157].
- VLC Player: Go to Tools > Preferences > Video. Change the Output dropdown from Automatic to X11 video output (XCB)[cite: 157].
- MPV Player: Launch videos from the terminal using:
mpv --vo=x11 --hwdec=no video.mp4[cite: 157].
Note: If server performance is your absolute priority and you do not intend to watch media over the remote link, skip these steps and leave the compositor disabled entirely[cite: 158].
Phase 21: The Modern Sysadmin Terminal Toolkit
When your server feels sluggish but standard htop shows plenty of free CPU and RAM, you are likely suffering from silent disk I/O saturation or bandwidth exhaustion[cite: 116]. Upgrade your terminal triage workflow with these modern diagnostic tools[cite: 115, 116].
1. Modern Resource Monitors (htop alternatives)
Move beyond standard process lists with tools that provide macro-level server health[cite: 115]:
- btop: The current community favorite. A gorgeous, mouse-responsive TUI with real-time graphs for cores, disks, and networking[cite: 115]. (
sudo apt install btop)[cite: 115]. - glances: A dense Python-based dashboard that fits disk I/O speeds, hardware temperatures, and system alerts on a single screen[cite: 115].
- atop: A historical troubleshooting tool. It runs a background daemon logging performance every 10 minutes, allowing you to "rewind time" to see what crashed your server overnight[cite: 115].
2. Isolating Disk I/O Bottlenecks
If processes are stuck in "Uninterruptible Sleep," your storage is saturated[cite: 116].
- iostat: Run
iostat -xz 1to monitor hardware saturation[cite: 116]. If%utilis near 100% andw_await(write latency) is > 50ms, the physical drive is failing to keep up with write requests[cite: 116, 117]. - iotop: Run
sudo iotop -oto see a live list of processes actively reading/writing to the disk right now, allowing you to kill the specific app hammering your storage[cite: 116].
3. Catching Network Leaks
If web applications lag but hardware resources are fine, check your network pipe[cite: 116].
- nload: Provides two simple, real-time ASCII graphs for total incoming and outgoing traffic to verify if you are hitting your ISP's bandwidth ceiling[cite: 116].
- iftop: Run
sudo iftopon your main interface (or Tailscale interface) to display a live table of IP addresses, instantly revealing which remote host is pulling a massive data stream from your server[cite: 116].
Phase 22: Upgrading to the Prometheus React UI
While Grafana is your primary dashboarding tool, writing complex PromQL queries directly inside Grafana can be tedious. The modern Prometheus React UI includes syntax highlighting, auto-complete, and a native dark mode, making it the perfect scratchpad for testing telemetry math before building your final panels[cite: 198, 199].
1. Install the React UI Assets
If you installed Prometheus via standard package managers (like APT on Debian/Ubuntu), it often defaults to the classic HTML interface. You can bypass this restriction using the built-in helper script[cite: 198]:
# Run the Debian helper script to pull the React assets
sudo /usr/share/prometheus/install-ui.sh
# Restart the service to load the new web application
sudo systemctl restart prometheus
2. Testing PromQL Queries
Navigate to http://<YOUR_PROMETHEUS_IP>:9090/graph. You will be greeted by the new React interface[cite: 199].
- Use Auto-Complete: Start typing a metric (e.g.,
scaph_) and the UI will drop down all available targets[cite: 199]. - Filter by Labels: Narrow down global queries to a specific instance before porting them to Grafana:
scaph_host_power_microwatts{instance="localhost:8081"} / 1000000[cite: 199]. - Table vs. Graph: Use the Table tab to verify current instantaneous values and label formatting, or the Graph tab to visualize time-series integrity over the last hour[cite: 199].
Phase 23: Managing the ARP Cache (ip neighbor)
If you recently changed a device's IP address on your home network but your server stubbornly refuses to connect to it, you likely have an ARP (Address Resolution Protocol) cache conflict[cite: 110]. The Linux kernel maps local IP addresses to physical MAC addresses; if a device switches IPs, the old mapping becomes stale[cite: 110].
1. Reading the Neighbor Table
To view your system's current IP-to-MAC mappings, use the modern ip neighbor tool (which replaces the legacy arp -a command)[cite: 110]:
ip n
You will see an output like:
192.168.1.1 dev enp4s0 lladdr 70:5a:xx:xx:xx REACHABLE[cite: 110].
- REACHABLE: The mapping is valid and recently verified[cite: 110].
- STALE: The mapping is cached but hasn't been verified in a few minutes[cite: 110].
- FAILED: The device did not respond (likely offline or IP changed)[cite: 110].
2. Flushing the ARP Cache
If a mapping is corrupted or outdated, force the Linux kernel to drop the entire cache and re-discover the network hardware dynamically[cite: 110]:
sudo ip neighbor flush all
If you want to manually assign a permanent static hardware route to prevent spoofing, use:
sudo ip neighbor add 192.168.1.50 lladdr 00:11:22:33:44:55 dev eth0[cite: 110].
Phase 24: Resolving VNC Desktop Crashes (Signal 6 & Polkit)
When experimenting with different desktop environments (GNOME, LXDE, XFCE) over VNC, it is easy to accidentally cross-contaminate your display manager configurations, resulting in instant grey screens or Signal 6 crashes[cite: 101, 124].
1. The "No Session for PID" Error
If you see a Polkit/D-Bus permissions error like no session for pid, it means multiple desktop components are fighting for control of the same VNC screen[cite: 101]. This usually happens if you run a global /etc/X11/Xsession script alongside a dedicated startxfce4 call[cite: 101].
The Fix: Purge conflicting desktop packages and isolate XFCE[cite: 102]:
# Remove conflicting LXDE or XRDP packages
sudo apt purge lxde-core lxpanel lxterminal xrdp -y
sudo apt autoremove --purge -y
# Reinstall XFCE securely to repair any clipped dependencies
sudo apt install --reinstall xfce4 xfce4-goodies -y
2. The GNOME "Signal 6" Crash
If you attempt to run the modern Ubuntu GNOME shell over a VNC system service, it will almost always fail with The X session died with signal 6![cite: 124, 130]. Modern GNOME requires two things that headless VNC servers lack:
- Hardware 3D Acceleration: GNOME panels and animations require physical GPU hardware. Over a virtual network frame, it panics and aborts[cite: 130].
- Systemd User Slices: GNOME relies on
systemd --usersessions to load its daemons, which are inaccessible when VNC is run as a global root/system service[cite: 124].
The Solution: Do not use full GNOME for remote headless servers. Fall back to XFCE, which handles 2D software rendering flawlessly and operates perfectly outside of strict systemd user slices[cite: 131, 132]. Ensure your ~/.config/tigervnc/xstartup file contains exec dbus-launch --exit-with-session startxfce4 to properly map the X11 sockets[cite: 131].
Epilogue: The Resilient Home Lab
Building a home lab is an ongoing process of iteration. What began with simple Wi-Fi commands and an old laptop has scaled into a secure, power-tuned, and fully observable DevOps environment.
By replacing bloated GUIs with lightweight XFCE, securing network paths with Tailscale, containerizing services via Docker, and keeping a watchful eye over hardware with Prometheus and Scaphandre, you have engineered an infrastructure that is both remarkably robust and exceptionally efficient.
Stay Curious. Keep Building.
Appendix: The Master Home Lab Navigation Index
To help you navigate this extensive deployment guide or easily return to a specific troubleshooting section, here is the complete master index of all phases covered in this home lab series.
Core Infrastructure
- Phase 1: Base OS, NetworkManager, and Secure Access (Tailscale)
- Phase 2: Deploying the Docker Engine & Portainer CE
- Phase 3: Core Services Overview & Automation Planning
- Phase 4: Global Observability (Prometheus, Grafana, cAdvisor)
Headless Environments & Access
- Phase 5: Configuring the Headless VNC Environment (TigerVNC + systemd)
- Phase 19: XFCE "Ricing" (Aesthetics & Speed Optimization)
- Phase 20: Handling Video Tearing in VNC (Compositor & VSync tuning)
- Phase 24: Resolving VNC Desktop Crashes (Signal 6 & Polkit logic)
Container Deployments
- Phase 6: Pi-hole Deployment & I/O Optimization (Fixing Database Bloat)
- Phase 7: Plex Media Server (CIFS NAS vs. USB Storage Mounts)
- Phase 13: Identity Management (PUID/PGID) & Homepage Dashboard
- Phase 14: Deploying Home Assistant (Host Networking)
Telemetry, Security & Diagnostics
- Phase 8: Hardware Power Tuning & Scaphandre Telemetry
- Phase 9: Real-Time Alerts via Telegram (Alertmanager)
- Phase 10: Enterprise SIEM Deployment (Wazuh & OpenSearch)
- Phase 11: Network Diagnostics & Hardware Firmware (DNS, fwupdmgr)
- Phase 12: Hardware Auditing & Custom Power Automation (setup_screen.sh)
- Phase 15: Advanced Server Triage & Automated Power Recovery (BIOS, btop)
- Phase 16: Tuning Grafana Variables for cgroups v2
- Phase 17: Fixing Linux LAN Drivers & APT Repositories
- Phase 18: Taming VM Insomnia on Proxmox
- Phase 21: The Modern Sysadmin Terminal Toolkit
- Phase 22: Upgrading to the Prometheus React UI
- Phase 23: Managing the ARP Cache (ip neighbor)
Where to Go Next: Future Home Lab Expansions
Once you have implemented the entire 24-phase blueprint, your infrastructure will be running like a well-oiled, enterprise-grade machine. But the beauty of a home lab is that the learning never stops. Here are the logical next steps to continue expanding your ecosystem.
1. Reverse Proxies & SSL Certificates (Traefik / Nginx Proxy Manager)
If you want to access your services using clean domain names (e.g., plex.yourdomain.com) instead of IP addresses and port numbers, deploying a reverse proxy is the next critical step. Tools like Traefik or Nginx Proxy Manager integrate directly with Docker to route traffic and automatically provision free SSL certificates via Let's Encrypt.
2. Automated Backup Pipelines
With persistent volumes storing your Plex metadata, Pi-hole configurations, and Prometheus telemetry, establishing a 3-2-1 backup strategy is vital. Consider deploying Duplicati or BorgBackup containers to automatically encrypt and push your /docker directory to an offsite cloud bucket (like AWS S3 or Backblaze B2) nightly.
3. Transitioning to Infrastructure as Code (IaC)
You have already taken the first steps into IaC by writing docker-compose.yml files. The next evolution is managing your entire host configuration using tools like Ansible. With Ansible, you can write a playbook that automatically installs NetworkManager, deploys Tailscale, configures TigerVNC, and applies all your power-tuning scripts across multiple nodes with a single command.
This concludes the complete home lab deployment and troubleshooting series. The configurations provided here will serve as a stable foundation for years of self-hosted exploration.
Appendix A: Complete Automation & Diagnostic Scripts
For ease of deployment, here are the full, copy-pasteable Bash scripts referenced throughout the series. All sensitive identifiers have been replaced with generic placeholders.
1. Automated DNS Auditor (dns_audit.sh)
Use this script to benchmark your local system resolver against global Anycast networks to diagnose silent web latency.
#!/bin/bash
# =================================================================================
# Linux DNS Performance & Diagnostic Benchmarking Tool
# =================================================================================
if ! command -v dig &> /dev/null; then
echo "[!] Missing dependency 'dnsutils'. Attempting install..."
sudo apt update -qq && sudo apt install dnsutils -y -qq
fi
clear
echo "==================================="
echo " PHASE 1: ACTIVE SYSTEM RESOLVER"
echo "==================================="
if command -v resolvectl &> /dev/null; then
ACTIVE_DNS=$(resolvectl status | grep "DNS Servers" -A 2)
if [ -n "$ACTIVE_DNS" ]; then
echo "$ACTIVE_DNS"
else
echo "No upstream records found in resolvectl status."
fi
else
echo "Fallback Mode (/etc/resolv.conf):"
grep nameserver /etc/resolv.conf
fi
echo -e "\n==================================="
echo " PHASE 2: BASELINE RESPONSE MATRIX"
echo "==================================="
dig ubuntu.com | grep -E "Query time | SERVER"
echo -e "\n==================================="
echo " PHASE 3: STABILITY (5-CYCLE TEST)"
echo "==================================="
for i in {1..5}; do
TIME_MS=$(dig google.com | grep "Query time" | awk '{print $4}')
echo "-> Query Iteration $i: ${TIME_MS} ms"
sleep 0.5
done
echo -e "\n==================================="
echo " PHASE 4: ANYCAST PUBLIC BENCHMARK"
echo "==================================="
echo "System Default: $(dig example.com | grep "Query time" | awk '{print $4, $5}')"
echo "Cloudflare (1.1.1.1): $(dig example.com @1.1.1.1 | grep "Query time" | awk '{print $4, $5}')"
echo "Google (8.8.8.8): $(dig example.com @8.8.8.8 | grep "Query time" | awk '{print $4, $5}')"
echo "Quad9 (9.9.9.9): $(dig example.com @9.9.9.9 | grep "Query time" | awk '{print $4, $5}')"
echo "==================================="
echo "Optimization Rule: If public platforms outpace your System Link by >20ms, update your interface permanently."
2. All-In-One Display Controller (setup_screen.sh)
For repurposed laptops or All-In-One PCs, this script detects the system's backlight interface, injects screenon and screenoff terminal shortcuts, and sets up cron jobs to power down the monitor at night.
#!/bin/bash
# AIO SERVER DISPLAY CONFIGURATOR
if [ "$EUID" -ne 0 ]; then
echo "[-] Please run this script with sudo."
exit 1
fi
echo "[+] Starting automated display configuration..."
# 1. Detect Backlight Interface
BACKLIGHT_DIR="/sys/class/backlight"
INTERFACE=$(ls "$BACKLIGHT_DIR" | head -n 1)
if [ -z "$INTERFACE" ]; then
echo "[-] Error: No backlight controller found."
exit 1
fi
BACKLIGHT_PATH="$BACKLIGHT_DIR/$INTERFACE"
MAX_BRIGHTNESS=$(cat "$BACKLIGHT_PATH/max_brightness")
echo "[+] Detected interface: $INTERFACE (Max: $MAX_BRIGHTNESS)"
# 2. Inject Bash Aliases for the Calling User
REAL_USER=${SUDO_USER:-$USER}
REAL_USER_HOME=$(eval echo ~$REAL_USER)
BASHRC_PATH="$REAL_USER_HOME/.bashrc"
if [ -f "$BASHRC_PATH" ]; then
sed -i '/alias screenoff=/d' "$BASHRC_PATH"
sed -i '/alias screenon=/d' "$BASHRC_PATH"
echo "alias screenoff='echo 0 | sudo tee $BACKLIGHT_PATH/brightness'" >> "$BASHRC_PATH"
echo "alias screenon='echo $MAX_BRIGHTNESS | sudo tee $BACKLIGHT_PATH/brightness'" >> "$BASHRC_PATH"
chown "$REAL_USER":"$REAL_USER" "$BASHRC_PATH"
echo "[+] Shortcuts 'screenoff' and 'screenon' configured."
fi
# 3. Configure Console Idle Blanking (1 minute)
setterm --blank 1 --term linux < /dev/tty1 2>/dev/null || true
# 4. Install Root Cron Jobs for Automation (Off at 22:00, On at 07:00)
TMP_CRON=$(mktemp)
crontab -l > "$TMP_CRON" 2>/dev/null || true
sed -i '\#/sys/class/backlight/#d' "$TMP_CRON"
echo "0 22 * * * echo 0 > $BACKLIGHT_PATH/brightness" >> "$TMP_CRON"
echo "0 7 * * * cat $BACKLIGHT_PATH/max_brightness > $BACKLIGHT_PATH/brightness" >> "$TMP_CRON"
crontab "$TMP_CRON"
rm "$TMP_CRON"
echo "[+] Automated cron jobs deployed successfully."
echo "[*] Run 'source ~/.bashrc' to activate your aliases immediately."
Appendix B: Bridging Isolated Subnets via SSH Tunnels
When running a robust home lab on hypervisors like Proxmox, you often run into a scenario where the physical host collects data (like Scaphandre hardware metrics on port 8081), but the Prometheus database lives inside an isolated guest Virtual Machine. If direct port binding is blocked by firewalls or routing tables, an SSH tunnel is the most secure way to bridge the gap.
1. Passwordless Authentication
To automate the tunnel, the monitoring VM must be able to SSH into the bare-metal host without a password prompt. Generate an RSA key inside the VM and copy it to the host:
ssh-keygen -t rsa -b 4096
ssh-copy-id root@<PROXMOX_HOST_IP>
2. The Persistent Systemd Tunnel
You must ensure the tunnel re-establishes itself if the network drops. Create a systemd service file on the monitoring VM:
sudo nano /etc/systemd/system/scaphandre-tunnel.service
Paste the following configuration. The ServerAliveInterval and ExitOnForwardFailure parameters are critical for preventing "zombie" connections that block the port without actually transmitting data.
[Unit]
Description=SSH Tunnel to Proxmox Scaphandre
After=network.target ssh.service
[Service]
User=<YOUR_VM_USER>
ExecStart=/usr/bin/ssh -N -T -o ServerAliveInterval=60 -o ExitOnForwardFailure=yes -L 8081:localhost:8081 root@<PROXMOX_HOST_IP>
Restart=always
RestartSec=10
[Install]
WantedBy=multi-user.target
3. Activation & Prometheus Integration
Enable the tunnel to start on boot:
sudo systemctl daemon-reload
sudo systemctl enable --now scaphandre-tunnel.service
Now, open /etc/prometheus/prometheus.yml and add the local end of the tunnel to your scrape targets. Prometheus will query its own localhost port, which the SSH tunnel silently forwards to the physical hypervisor:
- job_name: 'scaphandre_proxmox'
scrape_interval: 15s
static_configs:
- targets: ['localhost:8081']
Restart Prometheus (sudo systemctl restart prometheus) and verify the target is UP in your Prometheus Dashboard.
Appendix C: Advanced Networking & VNC Polkit Overrides
Ubuntu Server natively uses systemd-networkd to manage connections via Netplan. However, when building a home lab, using NetworkManager (and its nmcli/nmtui utilities) provides much greater flexibility for managing Wi-Fi adapters and static IPs dynamically.
1. Forcing Netplan to Use NetworkManager
To hand control of your network interfaces over to NetworkManager, you must edit your core Netplan configuration file (usually found at /etc/netplan/00-installer-config.yaml or similar).
sudo nano /etc/netplan/00-installer-config.yaml
Set the renderer explicitly. (Note: YAML is strictly space-indented; do not use tabs).
network:
version: 2
renderer: NetworkManager
Apply the routing changes to the kernel: sudo netplan apply
2. The Headless Wi-Fi Cheat Sheet
Once NetworkManager is active, use these commands to control wireless hardware directly from your SSH terminal:
- Turn Wi-Fi Radio On/Off:
nmcli radio wifi on|nmcli radio wifi off - Scan for Local Networks:
nmcli device wifi list - Connect to a Secure Network:
nmcli device wifi connect "<SSID_NAME>" password "<WIFI_PASSWORD>" - Verify Adapter Status:
nmcli device status
3. Fixing the "Not Authorized" Polkit Error
If you attempt to use the graphical sudo nmtui interface or nmcli commands over an SSH or VNC connection, Ubuntu's PolicyKit (Polkit) security framework will often block it with an org.freedesktop.networkmanager.network-control request failed: not authorized error. This occurs because Polkit defaults to restricting network changes strictly to active, physical local sessions (someone physically typing at the server's keyboard).
To grant your remote user permission to modify networks over VNC or SSH, create a custom authorization override rule.
sudo nano /etc/polkit-1/rules.d/99-networkmanager-vnc.rules
Paste the following JavaScript-based security policy, which instructs Polkit to instantly authorize network requests as long as the remote user belongs to the sudo administrative group:
polkit.addRule(function(action, subject) {
if (action.id.indexOf("org.freedesktop.NetworkManager.") == 0 &&
subject.isInGroup("sudo")) {
return polkit.Result.YES;
}
});
Save the file and force the Polkit daemon to parse the new security rule:
sudo systemctl restart polkit
Your remote VNC and SSH sessions now possess full administrative authority to configure, bounce, and manage network interfaces using standard tools.
Documentation Complete
Appendix D: The Unified Docker Compose Stack
Throughout this guide, we deployed services individually. However, for a streamlined home lab, you can consolidate your core media and automation services into a single, unified docker-compose.yml file. This allows you to spin up your entire infrastructure with a single command.
Create a master directory (e.g., ~/homelab) and place this docker-compose.yml inside it. Make sure to create a matching .env file in the same folder for your NAS credentials.
services:
# 1. Dashboard
homepage:
image: ghcr.io/gethomepage/homepage:latest
container_name: homepage
ports:
- "3001:3000"
volumes:
- ./homepage/config:/app/config
environment:
- PUID=1000
- PGID=1000
restart: unless-stopped
# 2. DNS & Adblocking
pihole:
container_name: pihole
image: pihole/pihole:latest
ports:
- "53:53/tcp"
- "53:53/udp"
- "80:80/tcp"
- "8443:443/tcp"
environment:
TZ: 'Europe/London'
WEBPASSWORD: '<YOUR_SECURE_PASSWORD>'
FTLCONF_dns_listeningMode: 'ALL'
volumes:
- './pihole/etc-pihole:/etc/pihole'
- './pihole/etc-dnsmasq.d:/etc/dnsmasq.d'
cap_add:
- SYS_NICE
restart: unless-stopped
# 3. Home Automation (Host Network)
homeassistant:
container_name: homeassistant
image: ghcr.io/home-assistant/home-assistant:stable
volumes:
- ./homeassistant/config:/config
- /etc/localtime:/etc/localtime:ro
environment:
- TZ=Europe/London
network_mode: host
restart: unless-stopped
# 4. Media Server
plex:
image: lscr.io/linuxserver/plex:latest
container_name: plex
network_mode: host
environment:
- PUID=1000
- PGID=1000
- TZ=Europe/London
- VERSION=docker
volumes:
- ./plex/config:/config
- nas_media:/data/nas_shared
restart: unless-stopped
volumes:
nas_media:
driver: local
driver_opts:
type: cifs
o: "username=${NAS_USER},password=${NAS_PASS},iocharset=utf8,vers=3.0"
device: "//<NAS_IP_ADDRESS>/<SHARE_NAME>"
To deploy the entire fleet at once, navigate to your ~/homelab directory and run: docker compose up -d
Appendix E: Server Maintenance Cheat Sheet
Maintaining your home lab is just as important as building it. Bookmark this section for quick reference on keeping your host OS and Docker environment secure and optimized over time.
1. Routine Host OS Updates
Keep your Ubuntu Server patched and secure. Run this command sequence monthly to update packages and remove orphaned dependencies:
sudo apt update && sudo apt upgrade -y
sudo apt autoremove --purge -y
2. Updating Docker Containers
To pull the latest versions of your Docker images and recreate your containers with the new code (without losing your persistent data), navigate to your compose directory and run:
docker compose pull
docker compose up -d
3. Cleaning Docker Bloat
Over time, downloading new container versions leaves old, dangling images on your hard drive, which can silently consume gigabytes of storage. Clean up unused images, networks, and stopped containers safely:
# Safely remove unused data (leaves running containers untouched)
docker system prune -a
# View current Docker disk usage stats
docker system df
4. Restarting the VNC Engine
If your remote desktop ever hangs due to a memory leak in a graphical application, you can seamlessly restart the background daemon without rebooting the physical server:
sudo systemctl restart vncserver@1.service
sudo systemctl status vncserver@1.service --no-pager
End of Supplemental Appendices
You now possess the complete, exhaustive reference guide covering every facet of your home lab deployment—from initial hardware tuning to long-term unified maintenance strategies.
Appendix F: Resolving Grafana JSON Validation Errors
When migrating highly complex dashboards (like the multi-node Scaphandre and cAdvisor layouts) using raw JSON payloads, Grafana occasionally throws schema validation errors if the source code was exported from an older or conflicting API version (e.g., dashboard.grafana.app/v2).
1. The GridLayoutItem Conflict
If you encounter an error stating: DashboardSpec.layout.spec.items.5.kind: Invalid value: conflicting values "GridLayoutItem" and "GridLayoutImageItem", it is because Grafana's modern UI strict-typing rejects legacy image layouts inside standard grid matrices[cite: 30].
The Fix: Before importing the JSON string into your UI, open it in a standard text editor and run a "Find and Replace." Replace all instances of "GridLayoutImageItem" with "GridLayoutItem"[cite: 30].
2. Fixing Variable Schema Errors
Another common import blocker is the conflicting values "AdhocVariable" and "QueryVariable" error[cite: 50]. This happens when variable schemas mismatch the expected array length in Grafana 13+.
The Fix: Strip the legacy type: "dashboard" parameter from the variables block in your JSON file before importing[cite: 50, 51]. If the import continues to fail, bypass the JSON entirely and use the official Grafana.com community IDs (e.g., ID 14282 for cAdvisor, or 1860 for Node Exporter) to pull heavily vetted, continuously updated templates directly into your environment[cite: 205, 222].
Appendix G: Security Sanitization Checklist
Before making any home lab documentation public on a blog, it is critical to ensure no sensitive infrastructure data is accidentally exposed. This guide has been pre-sanitized, but if you upload your own terminal logs or screenshots, verify the following:
- Tailscale IPs: Ensure no
100.x.x.xaddresses are visible in your Grafana queries or SSH command logs. - Telegram Tokens: Never publish your
bot_tokenorchat_id. Malicious actors can use these to intercept your server alerts or flood your personal chat. - MAC Addresses & Hardware IDs: When sharing
lshworip neighboroutputs, obscure MAC addresses (e.g.,70:5a:xx:xx:xx). If usinginxi, always append the-zflag (e.g.,inxi -Fz) to automatically scramble security identifiers. - NAS Credentials: Verify that your Docker Compose
.envfiles or inlinecifsmount strings do not contain plain-text passwords or internal usernames.
End of Guide. All source materials and diagnostics have been fully documented.
Phase 26: Bonus - Preventing System Bloat & Log Rotations
As your home lab scales, background logs can quietly consume gigabytes of storage, eventually leading to No space left on device errors. Just like we capped the Pi-hole database to prevent I/O bottlenecks, we need to apply similar limits to the host OS and Docker.
1. Truncating Systemd Journal Logs
Ubuntu's journalctl keeps extensive logs of every systemd service (like our VNC and Scaphandre daemons). Limit these logs to a manageable size (e.g., 500MB) to protect your storage drive:
# Clear out old logs immediately, keeping only the last 2 days
sudo journalctl --vacuum-time=2d
# Or limit by absolute size
sudo journalctl --vacuum-size=500M
To make this limit permanent, edit the journal configuration:
sudo nano /etc/systemd/journald.conf
Uncomment and change the SystemMaxUse line: SystemMaxUse=500M. Then restart the service: sudo systemctl restart systemd-journald.
2. Enforcing Docker Log Limits
By default, Docker container logs grow infinitely. A single chatty container (like Wazuh or NetAlertX) can generate massive text files. Force Docker to rotate logs automatically by creating a global daemon configuration.
sudo nano /etc/docker/daemon.json
Paste the following JSON to restrict containers to three 10MB log files max:
{
"log-driver": "json-file",
"log-opts": {
"max-size": "10m",
"max-file": "3"
}
}
Restart the Docker engine to apply the rules to all future containers: sudo systemctl restart docker.
3. Automated Security Patching
For a hands-off approach to critical security updates (like the fwupdmgr vulnerabilities we patched earlier), ensure the unattended-upgrades package is running. It will silently install security patches in the background without breaking your core software.
sudo apt install unattended-upgrades -y
sudo dpkg-reconfigure --priority=low unattended-upgrades
Select Yes to enable automatic daily security downloads.
About This Series
This master guide was compiled directly from real-world terminal logs, debugging sessions, and architectural deployments. From navigating obscure Realtek driver bugs and strict Snap permissions to orchestrating a fully integrated Prometheus telemetry stack, every step here has been battle-tested on live hardware.
Was this guide helpful?
Feel free to bookmark this page for your future home lab rebuilds, share it with fellow self-hosters, and drop any questions in the comments below!
Phase 29: Securing Kali Linux & Nmap Discovery
When deploying penetration testing distributions like Kali Linux within your lab, remote access protocols are disabled by default to minimize the attack surface. Enable them securely and utilize Nmap for network reconnaissance.
1. Bootstrap the OpenSSH Server
Install and bind the OpenSSH daemon to the system autostart framework. Ensure you immediately change the default user password using passwd to prevent unauthorized lateral movement.
sudo apt update && sudo apt install openssh-server -y
sudo systemctl enable --now ssh
sudo systemctl status ssh --no-pager
2. Nmap Host Discovery & Optimization
Nmap (Network Mapper) offers profound insights into active assets and firewall configurations. Here are the core flags for optimizing your scans across local subnets.
- Decoy Scans (Evasion): Mask your scanning origin by injecting dummy IPs into the traffic logs.
nmap -sS -D 10.1.0.1 <TARGET_IP> - Skip Host Discovery (No Ping): If a firewall drops ICMP packets, force a port scan regardless of ping response.
nmap -Pn <TARGET_IP> - Aggressive Assessments: Enable OS detection, service versioning, and default scripts simultaneously.
sudo nmap -A -v -T4 192.168.1.*
Timing Profiles: Adjust packet transmission rates to prevent network flooding.
-T3: Normal (Default).-T4: Aggressive (Fast, requires stable internal links).-T5: Insane (Maximum speed, risks packet loss).
Phase 30: Automated Debian & Kali Daily Diagnostics
Consolidate your daily security and performance checks into a unified maintenance shell utility. This script resynchronizes packages, audits network bindings, samples virtual memory, and maps active containers.
#!/bin/bash
# DEBIAN / KALI LINUX DAILY MAINTENANCE UTILITY
if [ "$EUID" -ne 0 ]; then
echo "[-] Error: Please run this diagnostic script with sudo privileges."
exit 1
fi
echo "======================================================================"
echo " 🔄 PHASE 1: PACKAGE MANAGER RESYNCHRONIZATION & REPAIR"
echo "======================================================================"
sudo apt update && sudo apt upgrade -y
sudo apt --fix-broken install -y
sudo apt autoremove -y
echo "======================================================================"
echo " 📡 PHASE 2: NETWORK INTERFACES & SECURE TOPOLOGY"
echo "======================================================================"
echo "[+] VNC Socket Bindings:"
sudo netstat -ptnl | grep vnc
echo "[+] Link Layer IP Routing Tables:"
ip addr
echo "======================================================================"
echo " ⚙️ PHASE 3: OS ARCHITECTURE & ENVIRONMENT"
echo "======================================================================"
sudo hostnamectl
sudo lsb_release -a
echo "======================================================================"
echo " 📊 PHASE 4: DISK MANAGEMENT & CONTAINER METRICS"
echo "======================================================================"
if command -v smbstatus &> /dev/null; then
sudo smbstatus --shares | column -t
fi
lsblk
echo "======================================================================"
echo " 🛠️ PHASE 5: HARDWARE FOOTPRINT & PERFORMANCE"
echo "======================================================================"
lscpu
vmstat
echo "======================================================================"
echo " ✅ SYSTEM DIAGNOSTICS LOG LOOP COMPLETE!"
echo "======================================================================"
Save as dailyscri.sh, apply permissions (chmod +x dailyscri.sh), and execute via sudo ./dailyscri.sh.
Phase 31: Advanced Wazuh Administration & Fixes
Maintaining a Wazuh SIEM deployment occasionally requires manual intervention to break APT package loops during uninstallation or to facilitate anonymous read-only dashboard access in isolated lab environments.
1. Fixing Broken Wazuh-Manager Uninstallation
If purging the wazuh-manager package throws exit status errors (like prerm or postrm script failures), you must bypass the broken maintainer scripts manually to unlock the package manager.
# 1. Clear the Pre-Removal Script
sudo nano /var/lib/dpkg/info/wazuh-manager.prerm
# Replace contents with:
# #!/bin/sh
# exit 0
# 2. Clear the Post-Removal Script
sudo nano /var/lib/dpkg/info/wazuh-manager.postrm
# Replace contents with:
# #!/bin/sh
# exit 0
# 3. Complete the Removal and Purge Leftovers
sudo apt-get autoremove -y && sudo apt-get clean
sudo rm -rf /var/ossec
sudo dpkg --purge wazuh-manager
2. Enabling Anonymous Login on Wazuh Dashboard
Warning: Only enable anonymous authentication in strictly isolated test environments, as it grants full access without password verification.
- Enable Auth in Indexer: Open
/etc/wazuh-indexer/opensearch-security/config.ymland setanonymous_auth_enabled: trueunder thehttp:block. - Map Roles: Open
/etc/wazuh-indexer/opensearch-security/roles_mapping.ymland add"opendistro_security_anonymous"to theall_accessusers and backend_roles lists. - Sync Security Settings:
sudo OPENSEARCH_JAVA_HOME=/usr/share/wazuh-indexer/jdk /usr/share/wazuh-indexer/plugins/opensearch-security/tools/securityadmin.sh \ -cd /etc/wazuh-indexer/opensearch-security \ -nhnv -cacert /etc/wazuh-indexer/certs/root-ca.pem \ -cert /etc/wazuh-indexer/certs/admin.pem \ -key /etc/wazuh-indexer/certs/admin-key.pem -p 9200 - Configure Dashboard: Open
/etc/wazuh-dashboard/opensearch_dashboards.ymland append:opensearch_security.auth.type: "basicauth" opensearch_security.auth.anonymous_auth_enabled: true
Restart the dashboard (sudo systemctl restart wazuh-dashboard) to expose the "Log in as anonymous" button.
Phase 32: Securing Docker Apps with Nginx Reverse Proxy
When running custom containerized web applications (like a Python dashboard or a Bookmark-Doc-app), they often expose themselves directly on unencrypted ports. Lock down access by restricting Docker to localhost and wrapping the traffic in an Nginx proxy protected by HTTP Basic Authentication.
1. Automated Deployment & Security Script
This script automates the installation of Nginx, generates an .htpasswd credential file, pulls an application from GitHub, spins up the Docker container on localhost, and generates the protective reverse proxy configuration block.
#!/bin/bash
# 1. Install Nginx & Apache Utilities
if ! command -v nginx &> /dev/null; then
sudo apt update && sudo apt install nginx apache2-utils -y
fi
# 2. Setup Nginx Password File
if [ ! -f /etc/nginx/.htpasswd ]; then
sudo htpasswd -c /etc/nginx/.htpasswd admin
fi
# 3. Pull/Update Application & Deploy Container
sudo rm -rf Bookmark-Doc-app/
git clone https://github.com/<GITHUB_USER>/Bookmark-Doc-app.git
cd Bookmark-Doc-app/
# Ensure docker-compose.yml maps ports as: "127.0.0.1:5000:5000"
sudo docker compose up --build -d
cd ..
# 4. Setup Nginx Proxy Configuration
if [ ! -f /etc/nginx/sites-available/python_proxy ]; then
sudo bash -c 'cat > /etc/nginx/sites-available/python_proxy << "EOF"
server {
listen 9000;
server_name _;
auth_basic "Restricted Access";
auth_basic_user_file /etc/nginx/.htpasswd;
location / {
proxy_pass http://127.0.0.1:5000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
}
EOF'
fi
# 5. Enable site and restart Nginx
sudo ln -s /etc/nginx/sites-available/python_proxy /etc/nginx/sites-enabled/
sudo ufw allow 9000/tcp
sudo nginx -t
sudo systemctl restart nginx
Verification: Accessing the application on port 9000 will now prompt for login credentials before securely passing the traffic to the backend Docker container hidden on localhost.
Phase 33: Self-Hosting Navidrome with an External NAS
If you have a massive library of local audio files, Navidrome is incredibly lightweight, lightning-fast, and compatible with almost all Subsonic clients[cite: 285]. In this phase, we deploy Navidrome using Docker Compose, mount an external NAS folder for music, and safely store the database locally[cite: 285].
1. Prepare Your Directories
Navidrome requires two distinct storage targets: fast local storage for its internal database, and the external NAS mount for your music files[cite: 285].
# Create the Navidrome data folder
mkdir -p /home/<YOUR_USERNAME>/navidrome_data
# Ensure proper permissions (Navidrome runs as user 1000 by default)
sudo chown -R 1000:1000 /home/<YOUR_USERNAME>/navidrome_data
(Ensure your external NAS is already mounted to a path like /mnt/external/NAS/Music)[cite: 285].
2. Create the Docker Compose File
Deploy the server using the following configuration. The :ro flag on the music volume ensures your raw audio files on the NAS are mounted as "Read-Only," preventing accidental deletion via the Navidrome web UI[cite: 285].
services:
navidrome:
image: deluan/navidrome:latest
user: 1000:1000
ports:
- "4533:4533"
restart: always
environment:
ND_LOGLEVEL: info
ND_SCANSCHEDULE: "@every 1h"
ND_ENABLEDOWNLOADS: "true"
volumes:
- "/home/<YOUR_USERNAME>/navidrome_data:/data"
- "/mnt/external/NAS/SharedMusic:/music:ro"
Spin up the server (sudo docker compose up -d) and navigate to http://<YOUR_SERVER_IP>:4533 to create your secure admin account (avoid using generic usernames like "admin" or "root")[cite: 285].
Phase 34: Securing Navidrome with a Tailscale Funnel & HA Kill Switch
To listen to your music on strict corporate networks where VPN clients are blocked, you can use Tailscale Funnel to expose port 4533 to the public internet securely[cite: 285]. Leaving a public tunnel open permanently is risky, so we will build a physical "Kill Switch" inside Home Assistant to toggle the gateway on and off instantly[cite: 285].
1. Allow "Silent Sudo" & Generate SSH Keys
Home Assistant needs to execute Tailscale commands on the host machine without a password prompt[cite: 285].
# Run visudo on the host
sudo visudo
# Add to the bottom of the file:
<YOUR_USERNAME> ALL=(ALL) NOPASSWD: /usr/bin/tailscale
Next, generate an SSH key directly inside Home Assistant's mapped configuration directory so the Docker container can SSH into the host[cite: 285]:
sudo mkdir -p /path/to/hass/config/.ssh
sudo ssh-keygen -t rsa -f /path/to/hass/config/.ssh/id_rsa -q -N ""
sudo ssh-copy-id -i /path/to/hass/config/.ssh/id_rsa <YOUR_USERNAME>@<YOUR_HOST_IP>
sudo chmod 700 /path/to/hass/config/.ssh
sudo chmod 600 /path/to/hass/config/.ssh/id_rsa
2. The Home Assistant YAML Configuration
Open your Home Assistant configuration.yaml file and add the command_line switch[cite: 285]. The -q flag and UserKnownHostsFile=/dev/null are critical to prevent SSH warnings from breaking the switch state[cite: 285].
command_line:
- switch:
name: "Navidrome Public Tunnel"
unique_id: navidrome_public_tunnel
command_on: "ssh -q -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -i /config/.ssh/id_rsa <YOUR_USERNAME>@<YOUR_HOST_IP> 'sudo tailscale funnel --bg 4533'"
command_off: "ssh -q -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -i /config/.ssh/id_rsa <YOUR_USERNAME>@<YOUR_HOST_IP> 'sudo tailscale serve reset'"
command_state: "ssh -q -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -i /config/.ssh/id_rsa <YOUR_USERNAME>@<YOUR_HOST_IP> 'sudo tailscale funnel status | grep 4533'"
icon: mdi:music-network
Restart the Home Assistant container (sudo docker restart homeassistant) and add a Button Card mapped to switch.navidrome_public_tunnel to your dashboard[cite: 285].
Phase 35: Automating the Funnel via Telegram & Safety Timers
To prevent leaving your music server exposed to the public web indefinitely, we will build a two-way Telegram alert system with an interactive inline button, backed by a 2-hour "dead man's switch"[cite: 285].
1. The Sender Automation (Telegram Alert + Button)
In Home Assistant, create a YAML automation to watch the tunnel switch and send an HTML-formatted Telegram alert containing a clickable inline button ("🔒 Close Tunnel Now:/close_tunnel") when opened[cite: 285]:
alias: "Security: Navidrome Tunnel Status Alert"
mode: single
trigger:
- platform: state
entity_id: switch.navidrome_public_tunnel
not_to: ["unknown", "unavailable"]
not_from: ["unknown", "unavailable"]
action:
- choose:
- conditions:
- condition: state
entity_id: switch.navidrome_public_tunnel
state: "on"
sequence:
- action: telegram_bot.send_message
data:
message: "🟢 <b>Navidrome Tunnel: OPEN</b>\nYour music server is now exposed to the public web."
parse_mode: html
inline_keyboard:
- "🔒 Close Tunnel Now:/close_tunnel"
default:
- action: telegram_bot.send_message
data:
message: "🔴 <b>Navidrome Tunnel: CLOSED</b>\nThe Tailscale Funnel has been shut down. Public access is fully blocked."
parse_mode: html
2. The Listener Automation (Catching the Button Press)
Create a second automation that catches the /close_tunnel callback event, turns off the switch, and edits the original Telegram message to confirm the action[cite: 285]:
alias: "Security: Telegram Close Tunnel Callback"
mode: single
trigger:
- platform: event
event_type: telegram_callback
event_data:
data: "/close_tunnel"
action:
- action: telegram_bot.answer_callback_query
data:
callback_query_id: "{{ trigger.event.data.id }}"
message: "Closing tunnel..."
- action: switch.turn_off
target:
entity_id: switch.navidrome_public_tunnel
- action: telegram_bot.edit_message
data:
message_id: "{{ trigger.event.data.message.message_id }}"
chat_id: "{{ trigger.event.data.user_id }}"
message: "🔴 <b>Navidrome Tunnel: CLOSED via Telegram</b>\nThe Tailscale Funnel has been shut down remotely."
parse_mode: html
inline_keyboard: []
3. The 2-Hour Auto-Shutoff Timer
Use a State Duration Trigger to automatically close the tunnel if left open for 2 hours (this is safer than a delay action, which wipes from memory on reboot)[cite: 285]:
alias: "Security: Auto-Close Navidrome Tunnel (2 Hours)"
mode: single
trigger:
- platform: state
entity_id: switch.navidrome_public_tunnel
to: "on"
for:
hours: 2
action:
- action: switch.turn_off
target:
entity_id: switch.navidrome_public_tunnel
Phase 36: Proxmox USB Backups (The "Host Owns It" Architecture)
If you pass a raw USB hard drive directly to a Proxmox VM, Proxmox cannot back up that VM because it creates a dependency loop resulting in a target is busy error[cite: 285]. To resolve this, move the physical drive management to the Proxmox host, use it for backups, and securely share the media files back to the VM via an NFS network share[cite: 285].
1. Detach the Drive from the VM
Log into your VM shell, stop active services (Docker/Samba), and unmount the drive. Then, remove the USB hardware passthrough in the Proxmox Web GUI[cite: 285].
sudo systemctl stop docker.socket docker.service
sudo systemctl stop smbd nmbd
sudo umount /mnt/external
2. Mount on the Proxmox Host
Switch to the Proxmox Host shell, identify your drive's UUID (blkid), and configure /etc/fstab to mount it automatically[cite: 285].
mkdir -p /mnt/pve/usb-storage
nano /etc/fstab
# Add: UUID=<YOUR-UUID-HERE> /mnt/pve/usb-storage ntfs-3g defaults,nofail 0 0
mount -a
3. Share the Drive Back to the VM via NFS
Install the NFS kernel server on the Proxmox host and grant the VM exclusive access. The fsid=1 flag is critical for FUSE/NTFS drives[cite: 285]:
apt install nfs-kernel-server -y
nano /etc/exports
# Add: /mnt/pve/usb-storage <VM_IP>(rw,sync,no_subtree_check,no_root_squash,fsid=1)
exportfs -arv
systemctl restart nfs-kernel-server
4. Reconnect the Drive Inside the VM
Install the NFS client inside your VM (sudo apt install nfs-common -y) and update the VM's /etc/fstab[cite: 285]:
<PROXMOX_IP>:/mnt/pve/usb-storage /mnt/external nfs defaults,_netdev,nofail 0 0
Reload the daemon, mount the drive, and start your Docker services[cite: 285]. Proxmox can now safely target the USB directory for VZDump backups while the VM transparently accesses the media files over the internal virtual bridge.
Phase 37: Troubleshooting Wazuh Startup Timeouts
When deploying Wazuh on resource-constrained VMs, the wazuh-manager may refuse to start due to systemd exceeding its default 45-second timeout limit, leaving orphaned processes behind[cite: 285]. Furthermore, the Wazuh Dashboard may stall indefinitely with a "Server is not ready yet" error[cite: 285].
1. Clean Up and Apply the Permanent Systemd Fix
Clear out "sleeping" background processes and the locked startup directory, then increase the timeout threshold to 5 minutes (300 seconds)[cite: 285].
sudo pkill -f wazuh
sudo rm -rf /var/ossec/var/start-script-lock
sudo mkdir -p /etc/systemd/system/wazuh-manager.service.d
echo -e "[Service]\nTimeoutStartSec=300" | sudo tee /etc/systemd/system/wazuh-manager.service.d/override.conf
sudo systemctl daemon-reload
sudo systemctl restart wazuh-manager
2. Fixing the Dashboard "Not Ready" Error
If the dashboard logs (journalctl -u wazuh-dashboard -f) show connect ECONNREFUSED 127.0.0.1:9200, the web frontend cannot reach the backend database[cite: 285]. Start the wazuh-indexer service and wait 60 to 90 seconds for the JVM heap memory to initialize[cite: 285].
sudo systemctl start wazuh-indexer
sudo systemctl status wazuh-indexer
Test the API socket directly with curl -k https://127.0.0.1:9200. Once it returns "cluster_name": "wazuh-cluster", the dashboard will automatically resolve the error loop[cite: 285].
Phase 38: Integrating Wazuh IT Hygiene Dashboards into Grafana
To extract endpoint telemetry (like packages, ports, and hardware inventory) into Grafana, you must authenticate with the Wazuh REST API using long-lived JWT Bearer Tokens and deploy the Infinity Data Source plugin[cite: 285].
1. Generate and Extend API Tokens
Fetch your API credentials from /usr/share/wazuh-dashboard/data/wazuh/config/wazuh.yml[cite: 285]. By default, Wazuh API tokens expire every 15 minutes, which will break live Grafana boards[cite: 285]. Extend the expiration to 1 year (31,536,000 seconds) via the API[cite: 285]:
# 1. Fetch a temporary token
TOKEN=$(curl -u 'wazuh-wui:<YOUR_PASSWORD>' -k -s -X GET "https://127.0.0.1:55000/security/user/authenticate?raw=true")
# 2. Update the API configuration database
curl -k -X PUT "https://127.0.0.1:55000/security/config" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"auth_token_exp_timeout": 31536000}'
# 3. Generate your permanent token
curl -u 'wazuh-wui:<YOUR_PASSWORD>' -k -X GET "https://127.0.0.1:55000/security/user/authenticate?raw=true"
2. Configuring Grafana and Managing Agent Scope
Inside Grafana, add the Yesoreyeram Infinity Data Source plugin. Set the Base URL to https://<YOUR_IP>:55000 and input the long-lived token string under Auth Type: Bearer Token[cite: 285].
Wazuh syscollector endpoints require explicit agent identifiers in the URI path (e.g., /syscollector/${agent_id}/packages) and will return a 404 Not Found if queried globally[cite: 285]. Configure a dashboard variable (agent_id) to fetch available systems via /agents?select=id,name, enable multi-value, and use Grafana's Row Repeating feature to automatically stack IT hygiene panels for every active agent in your fleet[cite: 285].
Phase 39: AI-Powered SOC Triage (Wazuh & LiteLLM)
Piping every SIEM alert into an AI model quickly results in context window limits or massive API bills. By combining Wazuh's detection engine with a LiteLLM Gateway, we can build a dual-lane triage pipeline: a Real-Time Fast Lane for critical alerts, and a Daily Batch Digest using local models[cite: 285].
1. The Real-Time Fast Lane (Critical Alerts)
Configure Wazuh (/var/ossec/etc/ossec.conf) to trigger a custom python script when Level 12+ alerts occur[cite: 285]:
<integration>
<name>custom-llm-telegram</name>
<level>12</level>
<alert_format>json</alert_format>
</integration>
The script at /var/ossec/integrations/custom-llm-telegram captures the JSON payload, pushes it to LiteLLM, and sanitizes the output (triage_report.replace('<', '<').replace('>', '>')) to prevent AI markdown from crashing Telegram's strict HTML parser[cite: 285].
2. The Batch Processing Lane (Local AI Swarm)
Dumping 24 hours of raw JSON logs into a local model (like Qwen 1.5b) will crash it. Instead, query the OpenSearch Indexer directly and minify the data before passing it to the LLM[cite: 285].
The daily digest script requests logs matching "rule.level": {"gte": 5} over "now-24h", runs a Counter() on rule descriptions, and pushes the simplified metrics to the local model to write an IT Hygiene summary[cite: 285].
3. Specialized Threat Reporting (ISO 27001)
Using the same minification architecture, you can build compliance reports. Query the indexer for {"wildcard": {"rule.compliance.iso_27001": "*"}} over a 7-day period[cite: 285]. If the network is clean, the script bypasses the AI completely and sends a green checkmark to Telegram (Positive Confirmation), ensuring you know the script ran successfully[cite: 285]. Add these scripts to the root crontab to execute securely on scheduled intervals[cite: 285].
Phase 40: Android ADB Debloating (Samsung & Vivo)
To safely disable telemetric bloatware, aggressive performance throttlers, and sponsored feeds on Samsung One UI and Vivo Funtouch OS, utilize the Android Debug Bridge (ADB)[cite: 284].
1. Troubleshooting ADB Errors
If your device returns an unauthorized status, toggle Revoke USB debugging authorizations in Developer Options and reconnect[cite: 284]. If using the package manager, do not include the package: prefix output by pm list packages[cite: 284].
Correct: .\adb shell pm uninstall --user 0 com.samsung.android.smartsuggestions[cite: 284]
2. Vivo Z1x (Funtouch OS 11) Specific Fixes
Funtouch OS 11 throws DELETE_FAILED_USER_RESTRICTED errors when attempting to uninstall packages. Bypass this by stripping background permissions, clearing app data, and disabling the packages for the current user instead[cite: 284].
.\adb shell "
cmd appops set com.bbk.cloud RUN_IN_BACKGROUND ignore; pm clear com.bbk.cloud;
cmd package disable-user --user 0 com.bbk.cloud;
# Repeat for com.vivo.pushservice, com.vivo.gamecube, etc.
"
Since Funtouch OS leaves empty app shortcuts in the default launcher, install a custom launcher like Nova or Lawnchair and use the "Hide App" feature to clean up the UI[cite: 284].
3. Samsung One UI (Game Optimizing Service & Bloat)
Packages like com.samsung.discover (sponsored feeds) and com.sec.android.mimage.avatarstickers (AR Emoji) are 100% safe to remove[cite: 284]. The Game Optimizing Service (com.samsung.android.game.gos) requires an AppOps clear alongside disabling to neutralize aggressive thermal throttling[cite: 284].
.\adb shell pm disable-user --user 0 com.samsung.android.game.gos
.\adb shell cmd appops set com.samsung.android.game.gos RUN_IN_BACKGROUND ignore
.\adb shell pm clear com.samsung.android.game.gos
Master Guide Complete
Phase 41: Advanced Media Management (Immich & Jellyfin)
Once your core NAS and Plex services are running, the next evolution of a home media lab is moving away from proprietary cloud services for personal photos and customizing your streaming interfaces.
1. Self-Hosting Google Photos Alternative (Immich)
Immich is a high-performance photo and video backup solution. It relies on a PostgreSQL database and machine learning containers for facial recognition. When deploying Immich on Ubuntu, ensure your database backups are properly formatted to avoid index restoration errors.
# Example database restoration command for Immich PostgreSQL
# If you encounter unsupported vector index statements during a restore,
# ensure you are using the pgvector extension matching your database version.
docker exec -i immich_postgres psql -U postgres -d immich < immich-db-backup.sql
2. Customizing the Jellyfin UI
Jellyfin is a fantastic open-source media server, but its default web interface can feel a bit standard. You can inject custom CSS directly into the dashboard to personalize the login screens and library hover effects using modern accent colors.
- Log into your Jellyfin web dashboard as an administrator.
- Navigate to Dashboard > General > Custom CSS.
- Paste the following CSS to apply a sleek blue and purple theme:
/* Custom Jellyfin Accent Theme */
:root {
--accent1-light: #00a4dc; /* Vibrant Blue */
--accent2-light: #aa5cc3; /* Deep Purple */
}
/* Button & Accent Styling */
.paper-icon-button-light:hover,
.raised,
.button-submit {
background: linear-gradient(135deg, var(--accent1-light), var(--accent2-light)) !important;
color: #ffffff !important;
}
/* Card Hover Effects */
.card:hover .cardImageContainer {
border: 2px solid var(--accent1-light);
box-shadow: 0 4px 15px rgba(0, 164, 220, 0.4);
}
Phase 42: Lightweight Storage & Custom Web Apps
Expanding the lab doesn't always mean massive x86 servers. Single-board computers and custom-coded applications add highly specialized functionality to the network.
1. OpenMediaVault on Raspberry Pi
A Raspberry Pi running OpenMediaVault (OMV) acts as a highly energy-efficient NAS and Docker volume manager. To maintain storage health on the Pi, schedule routine Docker pruning and volume maintenance directly from the OMV terminal:
# Clear dangling images and unused build caches to preserve SD card space
docker image prune -a -f
docker volume prune -f
2. Hosting Custom Python Web Apps (Streamlit)
If you build bespoke data tools—such as a cricket match score tracker or a Google Fit data minification pipeline—Streamlit is the fastest way to deploy them. To host these securely on your own domain, deploy them in Docker behind the Nginx reverse proxy with Basic Authentication.
# Example Docker Compose for a custom Streamlit App
services:
streamlit-app:
build: .
container_name: custom_score_tracker
ports:
- "127.0.0.1:8501:8501" # Bound to localhost for Nginx proxy routing
restart: unless-stopped
By routing Port 8501 through your previously established Nginx .htpasswd proxy block, your personal data analysis pipelines remain entirely hidden from the public internet.
The Homelab Journey Continues
From managing ISO 27001 compliance matrices to customizing media servers and tracking power down to the hardware level, this infrastructure is now a fully realized enterprise environment operating right inside your home. The foundation is set, the monitoring is active, and the automation is taking care of the heavy lifting.
Phase 43: Centralized Telegram Alerts for Homelab Cronjobs
Instead of setting up a heavy centralized monitoring server to track automated background tasks across multiple VMs and Raspberry Pis, use this lightweight Bash Wrapper Agent. Placed on each node, it wraps your existing crontab jobs and pushes formatted execution metrics, exit codes, and error logs directly to your phone.
sudo nano /usr/local/bin/cron_agent.sh
Paste the following code, replacing the credentials with your Telegram Bot Token and Chat ID:
#!/bin/bash
# ==========================================
# Telegram Cronjob Monitoring Agent
# ==========================================
BOT_TOKEN="YOUR_TELEGRAM_BOT_TOKEN"
CHAT_ID="YOUR_TELEGRAM_CHAT_ID"
# Parse arguments
ALERT_SUCCESS=true
if [ "$1" == "--fail-only" ]; then
ALERT_SUCCESS=false
shift
fi
SERVER_NAME="$1"
JOB_NAME="$2"
shift 2
COMMAND="$@"
# Start execution and timer
START_TIME=$(date +%s)
OUTPUT=$($COMMAND 2>&1)
EXIT_CODE=$?
END_TIME=$(date +%s)
DURATION=$((END_TIME - START_TIME))
# Determine Status
if [ $EXIT_CODE -eq 0 ]; then
# If success and fail-only is true, exit silently
if [ "$ALERT_SUCCESS" = false ]; then
exit 0
fi
STATUS="✅ SUCCESS"
else
STATUS="❌ FAILED (Code: $EXIT_CODE)"
fi
# Format the Telegram Message
MESSAGE=$(cat <<EOF
$STATUS
🖥 Host: $SERVER_NAME
🛠Job: $JOB_NAME
⏱ Duration: ${DURATION}s
📄 Output:
$(echo "$OUTPUT" | tail -n 15)
EOF
)
# Send to Telegram
curl -s -X POST "https://api.telegram.org/bot${BOT_TOKEN}/sendMessage" \
-F chat_id="${CHAT_ID}" \
-F text="$MESSAGE" > /dev/null
Secure the file: sudo chmod +x /usr/local/bin/cron_agent.sh. You can now wrap any cron job using this syntax: /usr/local/bin/cron_agent.sh [--fail-only] "Host Name" "Job Name" /path/to/script.sh.
Phase 44: Deploying RustDesk over Tailscale
Self-hosting your remote desktop infrastructure using RustDesk provides complete control over your privacy. Running it strictly over your Tailscale mesh network creates an encrypted, private administration backbone.
1. The Docker Compose Configuration
Create your docker-compose.yml file. The -k _ flag enforces encryption, and network_mode: "host" allows it to bind directly to your Tailscale interface.
services:
hbbs:
container_name: hbbs
image: rustdesk/rustdesk-server:latest
environment:
- ALWAYS_USE_RELAY=Y
command: hbbs -r <YOUR_VPN_IP>:21117 -k _
volumes:
- ./data:/root
network_mode: "host"
depends_on:
- hbbr
restart: unless-stopped
hbbr:
container_name: hbbr
image: rustdesk/rustdesk-server:latest
command: hbbr -k _
volumes:
- ./data:/root
network_mode: "host"
restart: unless-stopped
2. Extracting the Encryption Key from Distroless Containers
Because the RustDesk image is "distroless" (it lacks standard tools like cat or bash), you cannot execute commands inside it. You must copy the auto-generated ED25519 key out to the host to read it.
sudo docker compose up -d
sudo docker cp hbbs:/root/id_ed25519.pub /tmp/id_ed25519.pub
cat /tmp/id_ed25519.pub
Copy this key into your RustDesk client application settings under the Key field, setting both the ID Server and Relay Server to your Tailscale IP.
Phase 44: Deploying RustDesk over Tailscale
Self-hosting your remote desktop infrastructure using RustDesk provides complete control over your privacy. Running it strictly over your Tailscale mesh network creates an encrypted, private administration backbone.
1. The Docker Compose Configuration
Create your docker-compose.yml file. The -k _ flag enforces encryption, and network_mode: "host" allows it to bind directly to your Tailscale interface.
services:
hbbs:
container_name: hbbs
image: rustdesk/rustdesk-server:latest
environment:
- ALWAYS_USE_RELAY=Y
command: hbbs -r <YOUR_VPN_IP>:21117 -k _
volumes:
- ./data:/root
network_mode: "host"
depends_on:
- hbbr
restart: unless-stopped
hbbr:
container_name: hbbr
image: rustdesk/rustdesk-server:latest
command: hbbr -k _
volumes:
- ./data:/root
network_mode: "host"
restart: unless-stopped
2. Extracting the Encryption Key from Distroless Containers
Because the RustDesk image is "distroless" (it lacks standard tools like cat or bash), you cannot execute commands inside it. You must copy the auto-generated ED25519 key out to the host to read it.
sudo docker compose up -d
sudo docker cp hbbs:/root/id_ed25519.pub /tmp/id_ed25519.pub
cat /tmp/id_ed25519.pub
Copy this key into your RustDesk client application settings under the Key field, setting both the ID Server and Relay Server to your Tailscale IP.
Phase 45: Glances Systemd Service & Hardware Sensors
Glances is a powerful system monitor that reads lm-sensors data to provide live motherboard and CPU thermal telemetry. We can configure it to run permanently as a lightweight background web server.
1. Install Dependencies and Scan Hardware
sudo apt update && sudo apt install lm-sensors glances -y
# Auto-detect motherboard sensors (bypass prompts)
sudo sensors-detect --auto
2. Construct the Systemd Daemon
sudo nano /etc/systemd/system/glances.service
Paste the following service layout to bind Glances to port 61208:
[Unit]
Description=Glances Web Server Daemon
After=network.target
[Service]
Type=simple
ExecStart=/usr/bin/glances -w -B 0.0.0.0 -p 61208
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
Initialize the service: sudo systemctl enable --now glances.service.
Optional Bugfix: If the Web UI displays an empty frame on older Ubuntu iterations due to uncompiled assets, run: sudo mkdir -p /usr/lib/python3/dist-packages/glances/outputs/static/public.
Phase 46: Wazuh AI Triage - The Real-Time Fast Lane
To prevent alert fatigue, configure a dual-lane AI pipeline. The Fast Lane intercepts critical Wazuh alerts (Level 12+) and sends them to a fast cloud model via a LiteLLM gateway for instant triage analysis.
Create the integration script on your Wazuh Manager:
sudo nano /var/ossec/integrations/custom-llm-telegram
Paste the Python processing code. The HTML sanitization block ensures AI-generated markdown doesn't crash Telegram's strict parsers.
#!/usr/bin/env python3
import sys, json, urllib.request
# --- CONFIGURATION ---
LITELLM_URL = "http://<YOUR_LITELLM_IP>:4000/v1/chat/completions"
LITELLM_KEY = "YOUR_LITELLM_API_KEY"
TELEGRAM_BOT_TOKEN = "YOUR_TELEGRAM_BOT_TOKEN"
TELEGRAM_CHAT_ID = "YOUR_TELEGRAM_CHAT_ID"
FAST_MODEL = "vibe-coder"
# ---------------------
def send_telegram(text):
url = f"https://api.telegram.org/bot{TELEGRAM_BOT_TOKEN}/sendMessage"
payload = {"chat_id": TELEGRAM_CHAT_ID, "text": text, "parse_mode": "HTML"}
req = urllib.request.Request(url, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json"})
urllib.request.urlopen(req, timeout=15)
def analyze_with_llm(alert_json):
system_prompt = (
"You are an automated SOC triage analyst. Analyze this Level 12+ security alert. "
"Provide a concise analysis in 3 bullet points: "
"1. Threat Nature & Impact "
"2. Likely Attack Vector / Indicator "
"3. Recommended Containment Step. "
"Be direct, highly technical, and do not use introductory filler."
)
prompt = f"Wazuh Alert Data:\n{json.dumps(alert_json, indent=2)}"
payload = {
"model": FAST_MODEL,
"messages": [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
"temperature": 0.2
}
req = urllib.request.Request(LITELLM_URL, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json", "Authorization": f"Bearer {LITELLM_KEY}"})
with urllib.request.urlopen(req, timeout=60) as resp:
result = json.loads(resp.read().decode("utf-8"))
return result["choices"][0]["message"]["content"]
def main():
if len(sys.argv) < 2: sys.exit(1)
with open(sys.argv[1], "r") as f: alert_data = json.load(f)
rule = alert_data.get("rule", {})
agent = alert_data.get("agent", {})
level = rule.get("level", 0)
if level < 12: sys.exit(0)
try:
triage_report = analyze_with_llm(alert_data)
except Exception as e:
triage_report = f"⚠️ LLM Analysis Failed: {str(e)}"
# CRITICAL: Sanitize output so AI asterisks/brackets don't break Telegram HTML
safe_report = triage_report.replace('<', '<').replace('>', '>')
msg = (
f"🚨 <b>CRITICAL WAZUH ALERT (Level {level})</b>\n\n"
f"<b>Host:</b> <code>{agent.get('name', 'Unknown')}</code>\n"
f"<b>Rule:</b> <code>{rule.get('id', 'N/A')}</code> - {rule.get('description', 'N/A')}\n\n"
f"<b>AI Triage Analysis:</b>\n{safe_report}"
)
send_telegram(msg)
if __name__ == "__main__": main()
Assign permissions: sudo chmod 750 /var/ossec/integrations/custom-llm-telegram && sudo chown root:wazuh /var/ossec/integrations/custom-llm-telegram.
Phase 47: Wazuh Batch Processing & ISO 27001 Scripts
To avoid hitting LLM context limits, the "Batch Processing Lane" uses a Python script to query OpenSearch, minify the JSON data down to aggregated counters, and process the results locally using a privacy-first AI swarm.
1. The Daily IT Hygiene Digest
Create /usr/local/bin/wazuh-daily-digest.py:
#!/usr/bin/env python3
import urllib.request, json, ssl, base64, sys
from collections import Counter
# --- CONFIGURATION ---
INDEXER_URL = "https://127.0.0.1:9200/wazuh-alerts-*/_search"
INDEXER_USER = "admin"
INDEXER_PASS = "YOUR_INDEXER_PASSWORD"
LITELLM_URL = "http://<YOUR_LITELLM_IP>:4000/v1/chat/completions"
LITELLM_KEY = "YOUR_LITELLM_API_KEY"
TELEGRAM_BOT_TOKEN = "YOUR_TELEGRAM_BOT_TOKEN"
TELEGRAM_CHAT_ID = "YOUR_TELEGRAM_CHAT_ID"
LOCAL_MODEL = "qwen-cluster"
def fetch_wazuh_alerts():
query = {
"query": {"bool": {"must": [{"range": {"timestamp": {"gte": "now-24h", "lte": "now"}}}, {"range": {"rule.level": {"gte": 5}}}]}},
"size": 1000, "sort": [{"rule.level": "desc"}]
}
ctx = ssl.create_default_context()
ctx.check_hostname = False; ctx.verify_mode = ssl.CERT_NONE
req = urllib.request.Request(INDEXER_URL, data=json.dumps(query).encode("utf-8"), headers={"Content-Type": "application/json"})
auth_b64 = base64.b64encode(f"{INDEXER_USER}:{INDEXER_PASS}".encode()).decode()
req.add_header("Authorization", f"Basic {auth_b64}")
with urllib.request.urlopen(req, context=ctx) as resp:
return json.loads(resp.read().decode("utf-8"))
def minify_data(data):
hits = data.get("hits", {}).get("hits", [])
if not hits: return None
rule_counts = Counter()
for hit in hits:
source = hit.get("_source", {})
rule_counts[f"Level {source.get('rule', {}).get('level', 0)}: {source.get('rule', {}).get('description', 'Unknown')}"] += 1
summary = "=== 24H ALERT DATA ===\nTop Alert Types:\n"
for rule, count in rule_counts.most_common(10): summary += f"- {rule} (Triggered {count} times)\n"
return summary
def analyze_with_llm(minified_data):
system_prompt = "You are a SOC manager writing a daily IT Hygiene digest. Review the provided alert counts for the last 24 hours. Provide a summary explaining overall network health and specific remediation actions. Output plain text only."
payload = {"model": LOCAL_MODEL, "messages": [{"role": "system", "content": system_prompt}, {"role": "user", "content": minified_data}], "temperature": 0.3}
req = urllib.request.Request(LITELLM_URL, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json", "Authorization": f"Bearer {LITELLM_KEY}"})
with urllib.request.urlopen(req, timeout=300) as resp:
return json.loads(resp.read().decode("utf-8"))["choices"][0]["message"]["content"]
def send_telegram(analysis):
url = f"https://api.telegram.org/bot{TELEGRAM_BOT_TOKEN}/sendMessage"
clean_analysis = analysis.replace('<', '<').replace('>', '>').replace('*', '').replace('#', '')
msg = f"📊 <b>Daily IT Hygiene Digest</b>\n\n<b>Swarm Analysis:</b>\n{clean_analysis}"
payload = {"chat_id": TELEGRAM_CHAT_ID, "text": msg, "parse_mode": "HTML"}
req = urllib.request.Request(url, data=json.dumps(payload).encode("utf-8"), headers={"Content-Type": "application/json"})
urllib.request.urlopen(req, timeout=15)
def main():
raw_data = fetch_wazuh_alerts()
minified_data = minify_data(raw_data)
if minified_data: send_telegram(analyze_with_llm(minified_data))
if __name__ == "__main__": main()
2. ISO 27001 Compliance Query Adjustments
To create a weekly ISO 27001 script, copy the code above and replace the Elasticsearch query and minifier loops:
# Query
{"query": {"bool": {"must": [{"range": {"timestamp": {"gte": "now-7d", "lte": "now"}}}, {"wildcard": {"rule.compliance.iso_27001": "*"}}]}}, "size": 1000}
# Minification
for tag in source.get("rule", {}).get("compliance", {}).get("iso_27001", []):
cve_list[f"ISO 27001 Control {tag} triggered on host '{agent}'"] += 1
Set them to run via Cron (e.g., 0 8 * * * /usr/local/bin/wazuh-daily-digest.py).
Phase 48: Immich Photo Server & Hardware Acceleration
Deploying Immich natively requires correctly passing your iGPU through to the containers to enable VA-API video transcoding and OpenVINO machine learning tasks. It also requires the pgvectors extension inside PostgreSQL.
The Complete Docker Compose File
Create a docker-compose.yml mapping both your local upload path and your read-only NAS share.
name: immich
services:
immich-server:
container_name: immich_server
image: ghcr.io/immich-app/immich-server:${IMMICH_VERSION:-release}
devices:
- /dev/dri:/dev/dri # Pass iGPU for VA-API video transcoding
volumes:
- ${UPLOAD_LOCATION}:/data
- /mnt/nas_photos:/usr/src/app/external/nas_photos:ro
- /etc/localtime:/etc/localtime:ro
env_file:
- .env
ports:
- '2283:2283'
depends_on:
- redis
- database
restart: always
immich-machine-learning:
container_name: immich_machine_learning
image: ghcr.io/immich-app/immich-machine-learning:${IMMICH_VERSION:-release}-openvino
devices:
- /dev/dri:/dev/dri # Pass iGPU for OpenVINO acceleration
volumes:
- model-cache:/cache
env_file:
- .env
restart: always
redis:
container_name: immich_redis
image: valkey/valkey:8-alpine
restart: always
database:
container_name: immich_postgres
image: ghcr.io/immich-app/postgres:14-vectorchord0.4.3-pgvectors0.2.0
environment:
POSTGRES_PASSWORD: ${DB_PASSWORD}
POSTGRES_USER: ${DB_USERNAME}
POSTGRES_DB: ${DB_DATABASE_NAME}
POSTGRES_INITDB_ARGS: '--data-checksums'
volumes:
- ${DB_DATA_LOCATION}:/var/lib/postgresql/data
shm_size: 128mb
restart: always
volumes:
model-cache:
Database Note: Ensure your .env file's DB_PASSWORD contains only alphanumeric characters (A-Za-z0-9) without hyphens to prevent initialization crashes.
Phase 49: Android ADB Debloat Automation Scripts
To safely disable telemetric bloatware, aggressive performance throttlers (like GOS), and sponsored feeds on Samsung One UI and Vivo Funtouch OS, utilize the Android Debug Bridge (ADB) via PowerShell.
1. Vivo Z1x (Funtouch OS 11) Debloat Script
Funtouch OS 11 throws DELETE_FAILED_USER_RESTRICTED errors when attempting to uninstall packages. Bypass this by stripping background permissions and disabling the packages instead.
.\adb shell "
cmd appops set com.bbk.cloud RUN_IN_BACKGROUND ignore; pm clear com.bbk.cloud;
cmd appops set com.vivo.pushservice RUN_IN_BACKGROUND ignore; pm clear com.vivo.pushservice;
cmd appops set com.vivo.gamecube RUN_IN_BACKGROUND ignore; pm clear com.vivo.gamecube;
cmd appops set com.vivo.game RUN_IN_BACKGROUND ignore; pm clear com.vivo.game;
cmd package disable-user --user 0 com.bbk.cloud;
cmd package disable-user --user 0 com.vivo.pushservice;
cmd package disable-user --user 0 com.vivo.gamecube;
cmd package disable-user --user 0 com.vivo.game;
"
2. Samsung One UI (Game Optimizing Service & Smart Suggestions)
The Game Optimizing Service requires an AppOps clear alongside a standard disable command to neutralize aggressive thermal throttling. Use this block for your Galaxy devices:
# Remove standard bloat
.\adb shell "
pm uninstall -k --user 0 com.facebook.system;
pm uninstall -k --user 0 com.facebook.appmanager;
pm uninstall -k --user 0 com.facebook.services;
pm uninstall -k --user 0 com.samsung.android.bixby.wakeup;
pm uninstall --user 0 com.samsung.discover;
pm uninstall --user 0 com.sec.android.mimage.avatarstickers;
"
# Freeze & Neutralize Game Optimizing Service (GOS)
.\adb shell pm disable-user --user 0 com.samsung.android.game.gos
.\adb shell cmd appops set com.samsung.android.game.gos RUN_IN_BACKGROUND ignore
.\adb shell pm clear com.samsung.android.game.gos
# Freeze & Neutralize Smart Suggestions
.\adb shell pm disable-user --user 0 com.samsung.android.smartsuggestions
.\adb shell cmd appops set com.samsung.android.smartsuggestions RUN_IN_BACKGROUND ignore
.\adb shell pm clear com.samsung.android.smartsuggestions
Documentation Framework Finalized
Phase 50: Monitoring Docker Containers with Wazuh
Wazuh uses a built-in docker-listener wodle to monitor container lifecycle events (start, stop, pause, pull)[cite: 284]. This configuration works for both a standalone Wazuh Agent and a local Wazuh Manager host listening to /var/run/docker.sock[cite: 284].
1. Install Python Docker SDK & Grant Permissions
The Docker listener module requires the Python Docker library to interact directly with the local Docker API socket[cite: 284]. It also requires the Wazuh service account to hold sufficient read/write permissions[cite: 284].
# Install the Python Docker SDK
sudo apt update && sudo apt install python3-pip python3-docker -y
# Add the wazuh system user to the docker group
sudo usermod -aG docker wazuh
2. Enable the Docker Listener Module
Open the main configuration file on your target host (/var/ossec/etc/ossec.conf) and add the <wodle name="docker-listener"> block inside the main <ossec_config> tag[cite: 284]:
<wodle name="docker-listener">
<disabled>no</disabled>
<interval>10m</interval>
<attempts>5</attempts>
<run_on_start>yes</run_on_start>
</wodle>
3. Restart Service & Verify
Restart the active daemon (sudo systemctl restart wazuh-manager or wazuh-agent) to initialize the module[cite: 284]. Stream the log file to confirm success:
sudo tail -n 30 /var/ossec/logs/ossec.log | grep -i docker
To test it, run a temporary container (docker run --rm hello-world). Open your Wazuh Dashboard, navigate to Modules > Docker (or use the Discover tab with the filter rule.groups: docker), and you will see the container creation and termination events logged in real-time[cite: 284].
Phase 51: Essential Sysadmin CLI Extras & MacBook Drivers
Rounding out your terminal toolkit requires a few extra utilities for system fetching, package management, and niche hardware driver fixes[cite: 284].
1. System Fetching & Inline Browsing
- Visual System Fetch: Use
fastfetch --logo none(orneofetch) to print a clean hardware and OS summary to your terminal[cite: 284]. - Terminal Web Browser: If you need to read documentation or verify web server output entirely from a headless shell, install Lynx[cite: 284]:
sudo apt install lynx -y
2. Flatpak Integration
If you are running a graphical environment and need to install sandboxed applications outside of the apt or snap ecosystems, initialize Flatpak[cite: 284]:
sudo apt install flatpak gnome-software-plugin-flatpak
flatpak remote-add --if-not-exists flathub https://dl.flathub.org/repo/flathub.flatpakrepo
3. MacBook Pro Linux Fixes (HFS+ & Drivers)
If you installed Linux on an old Intel MacBook, you must manually install specific drivers for thermals, keyboard backlights, and file systems[cite: 284].
# Grant Write Access to macOS HFS+ Drives
sudo apt install hfsprogs
sudo mount -t hfsplus -o rw,force /dev/sdX /mnt/mac_drive/
# Install Broadcom Wi-Fi Drivers
sudo apt install firmware-b43-installer broadcom-sta-dkms -y
# Fix Fan Curves & Power Management
sudo apt install -y tlp tlp-rdw mbpfan
sudo systemctl enable --now tlp
sudo systemctl enable --now mbpfan
# Fix Apple Backlit Keyboard
sudo apt install pommed
Phase 52: Native Remote Desktop Access (Remmina)
If you deployed the full Ubuntu GNOME desktop (rather than the headless XFCE/TigerVNC environment) and want to connect to it locally using the built-in screen sharing protocol, you can use the Remmina client[cite: 284].
Configuring Native Screen Sharing
- On the Ubuntu host machine, open Settings > Sharing and enable the global sharing toggle[cite: 284].
- Click on Remote Desktop and enable both Remote Desktop and Remote Control.
- Under the Authentication section, click the hamburger menu (three dots) or the security dropdown, and explicitly select Require a password (VNC Auth)[cite: 284].
- Set a secure password.
On your client machine, open Remmina, select the VNC protocol, enter the host's IP address, and authenticate with the password you just configured. This bypasses the need for manual systemd services by utilizing Ubuntu's built-in vino or gnome-remote-desktop servers[cite: 284].
Archive Complete
Every fragment of documentation, troubleshooting logs, and configuration scripts provided has now been exhaustively processed, sanitized, and published. Your home lab deployment manual is officially complete in its entirety.
Phase 53: Infrastructure as Code - Automating Nodes with Ansible
Once you have mastered manual deployments across your Proxmox VMs, Raspberry Pis, and remote servers, the ultimate evolution of a home lab is Infrastructure as Code (IaC). Instead of running shell scripts or manual configuration commands on each node individually, Ansible allows you to define your desired server state in centralized YAML playbooks and push changes across your entire fleet instantly.
1. Setting Up the Control Node
Install Ansible on your primary management machine (your administrative workstation or management VM):
sudo apt update
sudo apt install ansible -y
2. Defining the Inventory File
Create a working directory for your Ansible project (e.g., ~/homelab-ansible) and create an inventory file named hosts.ini. This file groups your servers so you can target them individually or as a fleet.
[proxmox_hosts]
pve-node ansible_host=<PROXMOX_IP> ansible_user=root
[docker_nodes]
server1454 ansible_host=<SERVER_IP> ansible_user=server1454
[pis]
node_pi ansible_host=<RASPI_IP> ansible_user=pi
[all:vars]
ansible_ssh_private_key_file=~/.ssh/id_ed25519
3. Writing a Master Setup Playbook
Create a playbook named site.yml to automate essential tasks across your Linux nodes—such as ensuring NetworkManager is running, updating packages, and installing diagnostic tools like btop and curl.
---
- name: Configure and Harden Homelab Nodes
hosts: all
become: yes
tasks:
- name: Update apt repo cache
apt:
update_cache: yes
cache_valid_time: 3600
- name: Install essential sysadmin utilities
apt:
name:
- btop
- curl
- git
- ufw
- cifs-utils
state: present
- name: Ensure UFW firewall is active
ufw:
state: enabled
policy: allow
4. Executing the Playbook
Test your connectivity and push the configuration across your entire infrastructure using a single command:
# Test SSH connection to all inventory hosts
ansible all -i hosts.ini -m ping
# Run the master configuration playbook
ansible-playbook -i hosts.ini site.yml
Ansible executes tasks idempotently, meaning it only applies changes if the target state differs from your playbook definition, ensuring your servers remain stable and predictably configured across every run.
Comments
Post a Comment