From 64b20a54dc5841f39a109dc66dcf5c8d2e79dc77 Mon Sep 17 00:00:00 2001 From: Lutz Finsterle Date: Sat, 22 Aug 2026 09:50:30 +0200 Subject: [PATCH] Ship an encrypted DR bundle offsite to the NAS daily Local snapshots lived only on the Pi running OpenBAO, so a dead SD card took both the service and its backups. This was the top remaining hardening item. Scope note: fids2 is 192.168.0.234, same LAN as the Pi. This is genuinely OFF-HOST but not off-site -- it covers SD-card death, hardware failure and bad upgrades, not fire, theft, or LAN-wide ransomware. Encrypted with GPG to HOME_SECURE (07E23DC55C0FCF76) before anything touches the share. Two reasons: - The CIFS mount is uid=1000,file_mode=0664, so the root-only 0600 on /var/backups/openbao is LOST on arrival. Ciphertext makes the share's permissions irrelevant, which is what makes it safe to include the unseal material and ship a genuinely restorable DR set. - NOT the transit engine, deliberately: you would need a working OpenBAO to decrypt the backup you are restoring because OpenBAO is broken. Only the public key is on the Pi (committed here; verified no private-key blocks). The script never talks to OpenBAO, so it still runs while OpenBAO is down. Bundle = both raft snapshots + init-output.json + unsealer-init.json + a generated RESTORE.md carrying the restore ORDER (unsealer first, then main) and the no-downgrade warning, so the recovery instructions travel inside the backup rather than living only in a repo the Pi might take with it. Verified by full round-trip from the NAS copy: decrypt, extract, gzip -t and sha256sum -c both snapshots, and confirmed the recovered unseal keys are byte-identical to the live ones. Plaintext was shredded afterwards. Two guards, both tested to actually fire: - Refuses to ship if the newest local snapshot is older than 48h, rather than quietly uploading a stale DR copy. This is the exact failure that went unnoticed for 24 days (tested: 120h-old snapshot -> exit 1). - Verifies the destination is really on a cifs/smb filesystem. Found during testing: as root a bare `mkdir -p` SUCCEEDS when the automount is down, creating a local directory under the mount point, so every "offsite" backup would silently land on the same Pi. Checking that some path is a mountpoint was not enough -- it now stats the filesystem type of the destination's deepest existing ancestor (tested: ext2/ext3 -> exit 1). OnFailure mails an alert like the other units. Retains 30 bundles (~143KB each). Daily at 03:40, a clear gap after the 02:30 local snapshot. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01NFtVLA7VVqXL5G2S18c4Jk --- keys/openbao-backup-recipient.asc | 41 +++++++ scripts/offsite-backup.sh | 170 ++++++++++++++++++++++++++++++ systemd/openbao-offsite.service | 11 ++ systemd/openbao-offsite.timer | 13 +++ 4 files changed, 235 insertions(+) create mode 100644 keys/openbao-backup-recipient.asc create mode 100755 scripts/offsite-backup.sh create mode 100644 systemd/openbao-offsite.service create mode 100644 systemd/openbao-offsite.timer diff --git a/keys/openbao-backup-recipient.asc b/keys/openbao-backup-recipient.asc new file mode 100644 index 0000000..0002baf --- /dev/null +++ b/keys/openbao-backup-recipient.asc @@ -0,0 +1,41 @@ +-----BEGIN PGP PUBLIC KEY BLOCK----- + +mQGNBGcxBS4BDADO4C2YrddiphCk0S7ZevxawkLonz1qFGIN5eCO5BNv5Slv+IqG +7rr0oepQ2ZEOrb6eytV/vD4SBC2IlhpGI2HNkrZfrJgPuR3cjW+KhPUEL9s11jTu +8fPxXO+B6Ps5mdwmR6BcLLKp/mOYUR10+1ExChl5nlYN9ZCIJM/Hn4cXlV2NqInX +v35+qQvfvYkjUO0Y5JV7asDj3dA3YJJVfjQwggR/MEhapIBX0VHF7nEeANnHV0dT +rlgj1krJYGrPuHHXvEK76v2jBsm24070j1jwVzqmTUohYVYaOKLpQA/V4OOTuJDQ ++NN04yrq5R8jWGtvmNcr97UZHVt7K/7XLZkfV4wp2BGfzhdgTqj+bZZJWR5EE1+4 +O4UYmd1Xt29cS2+hy48HHXF0rzBwLd9Enb/USLIjO6L5o7N25umCfc6ck7WRq2Cd +zIM/bdveH5z3eloscxIC/ROLToTpyf+E86iXoN7F2EkAhwZShHUrx8NVHTAGpxWq +EZsZ1R+M+3BHs30AEQEAAbQoSE9NRV9TRUNVUkUgPGx1dHouZmluc3RlcmxlQHQt +b25saW5lLmRlPokBzgQTAQoAOBYhBKWn2wAPyhwkPzWhKAfiPcVcD892BQJnMQUu +AhsDBQsJCAcCBhUKCQgLAgQWAgMBAh4BAheAAAoJEAfiPcVcD892VykL/2tBpAfY +gmsjr+UZPWWnL+S32/5ao4WAkOWJj95eeixDQtuCCCfJxkW4n3O1B97EzHpyTsUa +U/kMujHXznyLLxwJfPiwVv3IQLr4Kd57Rf0pEjJLeDnnc91gCAtfxgu6aEC31fBJ +bD+kNJuJlBEySMHIC/4pMptlx4AtauqwQcdJBmZKTJ6BhLeKXS+ZDsszK/l4jbg8 +h2Xo9/VqjZWLnzhzEn7eoexPtRiILYJJmf+50p5l5pJ7+okmwHHNRb8ASw14XcYW +lnA24oe6EVot/IHLT1nOcIUrE1qJrGaqCopgV9nkF5jFoAz3S06DPTVSWDUgw+lQ +AmyVMXjHys35CQ8gUUg+liS0xYfPupVy1H+Durch12cLEXg3J09blQyRgm4bkgEq +bP9bXpb80SEe/lhNukytSVzc0gdPWA56nZjvT1ClToPEWMe1NCIZxKviaqCJNrOF +R8/bPann5Jq5+Cneq3q+6UZzQ/194Sv3P1UqIVjJwG6sKcr2opOq87JSAbkBjQRn +MQUuAQwAukFh6xHN+6n0KKN+027H6yfcaBGXPY1d3YGwgY0lH5CuxmJ+EvUg1Og2 +U+gYoMVuP+SnC9iwH/6ST8Zbir4D22sC/6yJuRBCOJfQSVGLBUymBHzj6o7Yjy4T +4ugjDH/LkPZbw+/89T29IQiN48NaQGG7z1013LjFLqDg7u65LZwDZ98tGOeXlAyw +vGtjYpc7WeYAzR53tv2XaUelloTPxcx0oGIHUcv8+XwTfV6K8BnNzZKJKg9WgvIC +bcMDaQYbhn8GRNfUDAcIXnwgHwb58Id5zWZbaWycqmgUgZ/qsbVKanIplxtvkv64 +cF+dgqx0nCFldm0/bGDShHh4bCkJMhpgZuw0KYmeI++MD60Hwv1nkLHWJba5BoLi +zgwAxmLq88nT35Qt0HIcHwqxvWPpaPd37GdxKNtohBc5LS1IfyS8iEwFbOzgP5J+ +VLmPHHTsiVI3eNHIXbxV9C84wX0uyzHNYAnbaLybr+OMnZZHRWrjTKm6cNY82/H4 +DlelR6UJABEBAAGJAbYEGAEKACAWIQSlp9sAD8ocJD81oSgH4j3FXA/PdgUCZzEF +LgIbDAAKCRAH4j3FXA/Pdsp7DACHAFABWazH/7pdSmguyJ6I0yXFIXpyXBFNweKL +8eDJpJZ+0ezjZvJ/tWKJ7pSjzTBYtzIfG9h9+jtHi95VOTMc8mTBed2fYzq8OQUG +GpMF2YPe9OL/U21fwKYiDEkcy1us0vtH2ZzH7MZbCvJpUUzk4bert9btqTps0Q7g +TnIOi8QxXTVb0L6rtyrZfkDjQSYBvs2zdMY8QjG1YqbRjAJaJ53sXgPDwHjE2Wvq +ZMZcF1JHyycO/PnMaJYsVZYQ6OZZsfsrnyEYTS8tlqyvUIQf03qWvYcgYnZq8SYx +D3F97/bLC891NAdSS96tT4LNq8xQbhB8oqNzUuy8OMeDsKfY755wqd01RFmp2RR2 +zkMuY0IPHxmTrMCTAQ/NDEJJ5xBeZAuokPFSLyQI9Dum8d+pO56oEPk2/dUamhhs +igt1u1dV/J/4x5WLr1/rRF80e38+GQn67ortvitmUBKnCaxae5SuhuiLJkL3Sgum +MN8dCQmDviD+FWnBQKDpdJ9/vV8= +=Onln +-----END PGP PUBLIC KEY BLOCK----- diff --git a/scripts/offsite-backup.sh b/scripts/offsite-backup.sh new file mode 100755 index 0000000..98069e8 --- /dev/null +++ b/scripts/offsite-backup.sh @@ -0,0 +1,170 @@ +#!/usr/bin/env bash +# Ship an ENCRYPTED OpenBAO disaster-recovery bundle to the NAS (fids2). +# +# Local snapshots live on the same Pi as OpenBAO, so a dead SD card takes both. +# This copies them off the host. Note fids2 is on the same LAN (192.168.0.234), +# so this is genuinely OFF-HOST but not off-site: it does not protect against +# fire, theft, or ransomware that reaches the whole LAN. +# +# WHY ENCRYPTED: the CIFS share is mounted uid=1000,file_mode=0664, so the +# root-only 0600 protection on /var/backups/openbao is LOST the moment a file +# lands there. The bundle is therefore GPG-encrypted to HOME_SECURE before it +# ever touches the share -- the NAS only ever holds ciphertext, which is also +# what makes it safe to include the unseal material. +# +# NOT encrypted with OpenBAO's transit engine, deliberately: you would need a +# working OpenBAO to decrypt the backup you are restoring because OpenBAO is +# broken. GPG keeps the decryption path independent of the thing being backed up. +# +# This script never talks to OpenBAO, so it still works while OpenBAO is down. +# It ships whatever the newest local snapshot is -- and FAILS LOUDLY if that +# snapshot is stale, which is the exact failure that went unnoticed for 24 days. +set -uo pipefail + +SRC_ROOT="${SRC_ROOT:-/var/backups/openbao}" +PROJECT_DIR="/home/lutz/Projects/OpenBAO" +SHARE_ROOT="${SHARE_ROOT:-/home/lutz/nfs_projects}" +DEST_DIR="${DEST_DIR:-${SHARE_ROOT}/backups/openbao}" +RECIPIENT_KEY="/etc/openbao-backup-recipient.asc" +RECIPIENT="07E23DC55C0FCF76" # HOME_SECURE +KEEP="${KEEP:-30}" # encrypted bundles to retain on the NAS +MAX_SNAP_AGE_H="${MAX_SNAP_AGE_H:-48}" # fail if newest local snapshot older than this +STAMP="$(date '+%Y%m%d-%H%M%S')" + +log() { printf '%s [offsite] %s\n' "$(date '+%F %T')" "$*"; } +die() { log "ERROR: $*"; exit 1; } + +WORK="" +cleanup() { + [ -n "$WORK" ] && [ -d "$WORK" ] && rm -rf "$WORK" +} +trap cleanup EXIT INT TERM + +command -v gpg >/dev/null || die "gpg not found" +[ -r "$RECIPIENT_KEY" ] || die "recipient public key $RECIPIENT_KEY not readable" + +# --- 1. locate newest local snapshots, and refuse to ship stale ones -------- +newest() { ls -1t "${SRC_ROOT}/$1"/openbao-"$1"-*.snap 2>/dev/null | head -1; } +MAIN_SNAP="$(newest main)" +UNSEALER_SNAP="$(newest unsealer)" +[ -n "$MAIN_SNAP" ] || die "no main snapshot found under ${SRC_ROOT}/main" +[ -n "$UNSEALER_SNAP" ] || die "no unsealer snapshot found under ${SRC_ROOT}/unsealer" + +for s in "$MAIN_SNAP" "$UNSEALER_SNAP"; do + age_h=$(( ( $(date +%s) - $(stat -c %Y "$s") ) / 3600 )) + [ "$age_h" -le "$MAX_SNAP_AGE_H" ] \ + || die "newest snapshot $(basename "$s") is ${age_h}h old (limit ${MAX_SNAP_AGE_H}h) -- the LOCAL backup is broken; fix that first, shipping a stale DR copy would give false confidence" +done +log "local snapshots fresh: $(basename "$MAIN_SNAP"), $(basename "$UNSEALER_SNAP")" + +# --- 2. assemble the DR set in a root-only workdir (never on the share) ----- +WORK="$(mktemp -d /root/.openbao-offsite.XXXXXX)" || die "cannot create workdir" +chmod 0700 "$WORK" +BUNDLE="openbao-dr-${STAMP}" +STAGE="${WORK}/${BUNDLE}" +mkdir -p "$STAGE" || die "cannot create staging dir" + +cp -p "$MAIN_SNAP" "$STAGE/" || die "cannot stage main snapshot" +cp -p "$UNSEALER_SNAP" "$STAGE/" || die "cannot stage unsealer snapshot" +for f in init-output.json unsealer-init.json; do + [ -r "${PROJECT_DIR}/${f}" ] || die "missing ${PROJECT_DIR}/${f} -- the DR set is incomplete without it" + cp -p "${PROJECT_DIR}/${f}" "$STAGE/" || die "cannot stage $f" +done + +cat > "$STAGE/RESTORE.md" </dev/null \ + || die "could not import recipient key" + +ENC="${WORK}/${BUNDLE}.tar.gz.gpg" +tar -czf - -C "$WORK" "$BUNDLE" \ + | GNUPGHOME="$GNUPGHOME_TMP" gpg --batch --quiet --trust-model always \ + --recipient "$RECIPIENT" --encrypt --output "$ENC" \ + || die "tar/encrypt pipeline failed" +[ -s "$ENC" ] || die "encrypted bundle is empty" + +# Plaintext staging is no longer needed — remove before touching the network. +rm -rf "$STAGE" + +ENC_SIZE=$(stat -c %s "$ENC") +ENC_SHA=$(sha256sum "$ENC" | cut -d' ' -f1) +log "encrypted bundle built (${ENC_SIZE}B, sha256 ${ENC_SHA:0:16}...)" + +# Sanity: it must actually be a GPG message, not a tar that slipped through. +head -c 3 "$ENC" | grep -q $'\x85\|\x84\|\x8c' 2>/dev/null || true +file_type="$(file -b "$ENC" 2>/dev/null || echo unknown)" +case "$file_type" in + *PGP*|*GPG*|*encrypted*) : ;; + *) die "refusing to upload: bundle does not look encrypted ($file_type)" ;; +esac + +# --- 4. ship to the NAS ---------------------------------------------------- +# Touch the automount first so the share is live before we probe it. +ls "$SHARE_ROOT" >/dev/null 2>&1 + +# CRITICAL: prove the destination really is the CIFS share before writing. +# Running as root, a bare `mkdir -p "$DEST_DIR"` SUCCEEDS even when the share +# is not mounted -- it just creates a local directory under the automount +# point, and every "offsite" backup silently lands on the same Pi we are +# trying to survive the loss of. So check the filesystem type of the deepest +# existing ancestor of DEST_DIR, not merely that some path is a mountpoint. +ancestor="$DEST_DIR" +while [ ! -d "$ancestor" ] && [ "$ancestor" != "/" ]; do + ancestor="$(dirname "$ancestor")" +done +fstype="$(stat -f -c %T "$ancestor" 2>/dev/null || echo unknown)" +case "$fstype" in + cifs|smb2|smb3) : ;; + *) die "destination $DEST_DIR resolves to a '$fstype' filesystem, not the CIFS share -- refusing to write a fake 'offsite' copy onto local disk (is the automount for $SHARE_ROOT down?)" ;; +esac + +mkdir -p "$DEST_DIR" || die "cannot create $DEST_DIR on the share" + +DEST="${DEST_DIR}/${BUNDLE}.tar.gz.gpg" +cp "$ENC" "${DEST}.part" || die "copy to NAS failed" +mv -f "${DEST}.part" "$DEST" || die "atomic rename on NAS failed" + +# --- 5. verify the copy that actually landed ------------------------------- +REMOTE_SHA=$(sha256sum "$DEST" | cut -d' ' -f1) +[ "$REMOTE_SHA" = "$ENC_SHA" ] \ + || die "checksum mismatch after upload (local ${ENC_SHA:0:16} vs NAS ${REMOTE_SHA:0:16})" +log "uploaded + verified: $DEST" + +# --- 6. retention on the NAS ---------------------------------------------- +ls -1t "${DEST_DIR}"/openbao-dr-*.tar.gz.gpg 2>/dev/null | tail -n +$((KEEP + 1)) | while read -r old; do + rm -f -- "$old" && log "pruned old offsite bundle $(basename "$old")" +done + +COUNT=$(ls -1 "${DEST_DIR}"/openbao-dr-*.tar.gz.gpg 2>/dev/null | wc -l) +log "done; ${COUNT} encrypted bundle(s) offsite, retaining ${KEEP}" diff --git a/systemd/openbao-offsite.service b/systemd/openbao-offsite.service new file mode 100644 index 0000000..7cb5904 --- /dev/null +++ b/systemd/openbao-offsite.service @@ -0,0 +1,11 @@ +[Unit] +Description=Ship an encrypted OpenBAO DR bundle to the NAS (fids2) +After=network-online.target home-lutz-nfs_projects.automount +Wants=network-online.target +OnFailure=openbao-alert@%n.service + +[Service] +Type=oneshot +# Runs as root: reads /var/backups/openbao (0700 root) and the 0600 unseal +# material. Encrypts to HOME_SECURE before anything touches the share. +ExecStart=/home/lutz/Projects/OpenBAO/scripts/offsite-backup.sh diff --git a/systemd/openbao-offsite.timer b/systemd/openbao-offsite.timer new file mode 100644 index 0000000..dc8a2ee --- /dev/null +++ b/systemd/openbao-offsite.timer @@ -0,0 +1,13 @@ +[Unit] +Description=Daily offsite copy of the OpenBAO DR bundle + +[Timer] +# Local snapshots run at 02:30 (+ up to 15m jitter), so 03:40 leaves a clear +# gap and always ships that morning's snapshot. The script independently +# refuses to run if the newest local snapshot is older than 48h. +OnCalendar=*-*-* 03:40:00 +Persistent=true +RandomizedDelaySec=15m + +[Install] +WantedBy=timers.target