Ship an encrypted DR bundle offsite to the NAS daily

Local snapshots lived only on the Pi running OpenBAO, so a dead SD card took
both the service and its backups. This was the top remaining hardening item.

Scope note: fids2 is 192.168.0.234, same LAN as the Pi. This is genuinely
OFF-HOST but not off-site -- it covers SD-card death, hardware failure and bad
upgrades, not fire, theft, or LAN-wide ransomware.

Encrypted with GPG to HOME_SECURE (07E23DC55C0FCF76) before anything touches
the share. Two reasons:
  - The CIFS mount is uid=1000,file_mode=0664, so the root-only 0600 on
    /var/backups/openbao is LOST on arrival. Ciphertext makes the share's
    permissions irrelevant, which is what makes it safe to include the unseal
    material and ship a genuinely restorable DR set.
  - NOT the transit engine, deliberately: you would need a working OpenBAO to
    decrypt the backup you are restoring because OpenBAO is broken. Only the
    public key is on the Pi (committed here; verified no private-key blocks).

The script never talks to OpenBAO, so it still runs while OpenBAO is down.

Bundle = both raft snapshots + init-output.json + unsealer-init.json + a
generated RESTORE.md carrying the restore ORDER (unsealer first, then main)
and the no-downgrade warning, so the recovery instructions travel inside the
backup rather than living only in a repo the Pi might take with it.

Verified by full round-trip from the NAS copy: decrypt, extract, gzip -t and
sha256sum -c both snapshots, and confirmed the recovered unseal keys are
byte-identical to the live ones. Plaintext was shredded afterwards.

Two guards, both tested to actually fire:
  - Refuses to ship if the newest local snapshot is older than 48h, rather
    than quietly uploading a stale DR copy. This is the exact failure that
    went unnoticed for 24 days (tested: 120h-old snapshot -> exit 1).
  - Verifies the destination is really on a cifs/smb filesystem. Found during
    testing: as root a bare `mkdir -p` SUCCEEDS when the automount is down,
    creating a local directory under the mount point, so every "offsite"
    backup would silently land on the same Pi. Checking that some path is a
    mountpoint was not enough -- it now stats the filesystem type of the
    destination's deepest existing ancestor (tested: ext2/ext3 -> exit 1).

OnFailure mails an alert like the other units. Retains 30 bundles (~143KB
each). Daily at 03:40, a clear gap after the 02:30 local snapshot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NFtVLA7VVqXL5G2S18c4Jk
This commit is contained in:
2026-08-22 09:50:30 +02:00
parent 471692c283
commit 64b20a54dc
4 changed files with 235 additions and 0 deletions

View File

@@ -0,0 +1,41 @@
-----BEGIN PGP PUBLIC KEY BLOCK-----
mQGNBGcxBS4BDADO4C2YrddiphCk0S7ZevxawkLonz1qFGIN5eCO5BNv5Slv+IqG
7rr0oepQ2ZEOrb6eytV/vD4SBC2IlhpGI2HNkrZfrJgPuR3cjW+KhPUEL9s11jTu
8fPxXO+B6Ps5mdwmR6BcLLKp/mOYUR10+1ExChl5nlYN9ZCIJM/Hn4cXlV2NqInX
v35+qQvfvYkjUO0Y5JV7asDj3dA3YJJVfjQwggR/MEhapIBX0VHF7nEeANnHV0dT
rlgj1krJYGrPuHHXvEK76v2jBsm24070j1jwVzqmTUohYVYaOKLpQA/V4OOTuJDQ
+NN04yrq5R8jWGtvmNcr97UZHVt7K/7XLZkfV4wp2BGfzhdgTqj+bZZJWR5EE1+4
O4UYmd1Xt29cS2+hy48HHXF0rzBwLd9Enb/USLIjO6L5o7N25umCfc6ck7WRq2Cd
zIM/bdveH5z3eloscxIC/ROLToTpyf+E86iXoN7F2EkAhwZShHUrx8NVHTAGpxWq
EZsZ1R+M+3BHs30AEQEAAbQoSE9NRV9TRUNVUkUgPGx1dHouZmluc3RlcmxlQHQt
b25saW5lLmRlPokBzgQTAQoAOBYhBKWn2wAPyhwkPzWhKAfiPcVcD892BQJnMQUu
AhsDBQsJCAcCBhUKCQgLAgQWAgMBAh4BAheAAAoJEAfiPcVcD892VykL/2tBpAfY
gmsjr+UZPWWnL+S32/5ao4WAkOWJj95eeixDQtuCCCfJxkW4n3O1B97EzHpyTsUa
U/kMujHXznyLLxwJfPiwVv3IQLr4Kd57Rf0pEjJLeDnnc91gCAtfxgu6aEC31fBJ
bD+kNJuJlBEySMHIC/4pMptlx4AtauqwQcdJBmZKTJ6BhLeKXS+ZDsszK/l4jbg8
h2Xo9/VqjZWLnzhzEn7eoexPtRiILYJJmf+50p5l5pJ7+okmwHHNRb8ASw14XcYW
lnA24oe6EVot/IHLT1nOcIUrE1qJrGaqCopgV9nkF5jFoAz3S06DPTVSWDUgw+lQ
AmyVMXjHys35CQ8gUUg+liS0xYfPupVy1H+Durch12cLEXg3J09blQyRgm4bkgEq
bP9bXpb80SEe/lhNukytSVzc0gdPWA56nZjvT1ClToPEWMe1NCIZxKviaqCJNrOF
R8/bPann5Jq5+Cneq3q+6UZzQ/194Sv3P1UqIVjJwG6sKcr2opOq87JSAbkBjQRn
MQUuAQwAukFh6xHN+6n0KKN+027H6yfcaBGXPY1d3YGwgY0lH5CuxmJ+EvUg1Og2
U+gYoMVuP+SnC9iwH/6ST8Zbir4D22sC/6yJuRBCOJfQSVGLBUymBHzj6o7Yjy4T
4ugjDH/LkPZbw+/89T29IQiN48NaQGG7z1013LjFLqDg7u65LZwDZ98tGOeXlAyw
vGtjYpc7WeYAzR53tv2XaUelloTPxcx0oGIHUcv8+XwTfV6K8BnNzZKJKg9WgvIC
bcMDaQYbhn8GRNfUDAcIXnwgHwb58Id5zWZbaWycqmgUgZ/qsbVKanIplxtvkv64
cF+dgqx0nCFldm0/bGDShHh4bCkJMhpgZuw0KYmeI++MD60Hwv1nkLHWJba5BoLi
zgwAxmLq88nT35Qt0HIcHwqxvWPpaPd37GdxKNtohBc5LS1IfyS8iEwFbOzgP5J+
VLmPHHTsiVI3eNHIXbxV9C84wX0uyzHNYAnbaLybr+OMnZZHRWrjTKm6cNY82/H4
DlelR6UJABEBAAGJAbYEGAEKACAWIQSlp9sAD8ocJD81oSgH4j3FXA/PdgUCZzEF
LgIbDAAKCRAH4j3FXA/Pdsp7DACHAFABWazH/7pdSmguyJ6I0yXFIXpyXBFNweKL
8eDJpJZ+0ezjZvJ/tWKJ7pSjzTBYtzIfG9h9+jtHi95VOTMc8mTBed2fYzq8OQUG
GpMF2YPe9OL/U21fwKYiDEkcy1us0vtH2ZzH7MZbCvJpUUzk4bert9btqTps0Q7g
TnIOi8QxXTVb0L6rtyrZfkDjQSYBvs2zdMY8QjG1YqbRjAJaJ53sXgPDwHjE2Wvq
ZMZcF1JHyycO/PnMaJYsVZYQ6OZZsfsrnyEYTS8tlqyvUIQf03qWvYcgYnZq8SYx
D3F97/bLC891NAdSS96tT4LNq8xQbhB8oqNzUuy8OMeDsKfY755wqd01RFmp2RR2
zkMuY0IPHxmTrMCTAQ/NDEJJ5xBeZAuokPFSLyQI9Dum8d+pO56oEPk2/dUamhhs
igt1u1dV/J/4x5WLr1/rRF80e38+GQn67ortvitmUBKnCaxae5SuhuiLJkL3Sgum
MN8dCQmDviD+FWnBQKDpdJ9/vV8=
=Onln
-----END PGP PUBLIC KEY BLOCK-----

170
scripts/offsite-backup.sh Executable file
View File

@@ -0,0 +1,170 @@
#!/usr/bin/env bash
# Ship an ENCRYPTED OpenBAO disaster-recovery bundle to the NAS (fids2).
#
# Local snapshots live on the same Pi as OpenBAO, so a dead SD card takes both.
# This copies them off the host. Note fids2 is on the same LAN (192.168.0.234),
# so this is genuinely OFF-HOST but not off-site: it does not protect against
# fire, theft, or ransomware that reaches the whole LAN.
#
# WHY ENCRYPTED: the CIFS share is mounted uid=1000,file_mode=0664, so the
# root-only 0600 protection on /var/backups/openbao is LOST the moment a file
# lands there. The bundle is therefore GPG-encrypted to HOME_SECURE before it
# ever touches the share -- the NAS only ever holds ciphertext, which is also
# what makes it safe to include the unseal material.
#
# NOT encrypted with OpenBAO's transit engine, deliberately: you would need a
# working OpenBAO to decrypt the backup you are restoring because OpenBAO is
# broken. GPG keeps the decryption path independent of the thing being backed up.
#
# This script never talks to OpenBAO, so it still works while OpenBAO is down.
# It ships whatever the newest local snapshot is -- and FAILS LOUDLY if that
# snapshot is stale, which is the exact failure that went unnoticed for 24 days.
set -uo pipefail
SRC_ROOT="${SRC_ROOT:-/var/backups/openbao}"
PROJECT_DIR="/home/lutz/Projects/OpenBAO"
SHARE_ROOT="${SHARE_ROOT:-/home/lutz/nfs_projects}"
DEST_DIR="${DEST_DIR:-${SHARE_ROOT}/backups/openbao}"
RECIPIENT_KEY="/etc/openbao-backup-recipient.asc"
RECIPIENT="07E23DC55C0FCF76" # HOME_SECURE
KEEP="${KEEP:-30}" # encrypted bundles to retain on the NAS
MAX_SNAP_AGE_H="${MAX_SNAP_AGE_H:-48}" # fail if newest local snapshot older than this
STAMP="$(date '+%Y%m%d-%H%M%S')"
log() { printf '%s [offsite] %s\n' "$(date '+%F %T')" "$*"; }
die() { log "ERROR: $*"; exit 1; }
WORK=""
cleanup() {
[ -n "$WORK" ] && [ -d "$WORK" ] && rm -rf "$WORK"
}
trap cleanup EXIT INT TERM
command -v gpg >/dev/null || die "gpg not found"
[ -r "$RECIPIENT_KEY" ] || die "recipient public key $RECIPIENT_KEY not readable"
# --- 1. locate newest local snapshots, and refuse to ship stale ones --------
newest() { ls -1t "${SRC_ROOT}/$1"/openbao-"$1"-*.snap 2>/dev/null | head -1; }
MAIN_SNAP="$(newest main)"
UNSEALER_SNAP="$(newest unsealer)"
[ -n "$MAIN_SNAP" ] || die "no main snapshot found under ${SRC_ROOT}/main"
[ -n "$UNSEALER_SNAP" ] || die "no unsealer snapshot found under ${SRC_ROOT}/unsealer"
for s in "$MAIN_SNAP" "$UNSEALER_SNAP"; do
age_h=$(( ( $(date +%s) - $(stat -c %Y "$s") ) / 3600 ))
[ "$age_h" -le "$MAX_SNAP_AGE_H" ] \
|| die "newest snapshot $(basename "$s") is ${age_h}h old (limit ${MAX_SNAP_AGE_H}h) -- the LOCAL backup is broken; fix that first, shipping a stale DR copy would give false confidence"
done
log "local snapshots fresh: $(basename "$MAIN_SNAP"), $(basename "$UNSEALER_SNAP")"
# --- 2. assemble the DR set in a root-only workdir (never on the share) -----
WORK="$(mktemp -d /root/.openbao-offsite.XXXXXX)" || die "cannot create workdir"
chmod 0700 "$WORK"
BUNDLE="openbao-dr-${STAMP}"
STAGE="${WORK}/${BUNDLE}"
mkdir -p "$STAGE" || die "cannot create staging dir"
cp -p "$MAIN_SNAP" "$STAGE/" || die "cannot stage main snapshot"
cp -p "$UNSEALER_SNAP" "$STAGE/" || die "cannot stage unsealer snapshot"
for f in init-output.json unsealer-init.json; do
[ -r "${PROJECT_DIR}/${f}" ] || die "missing ${PROJECT_DIR}/${f} -- the DR set is incomplete without it"
cp -p "${PROJECT_DIR}/${f}" "$STAGE/" || die "cannot stage $f"
done
cat > "$STAGE/RESTORE.md" <<EOF
# OpenBAO disaster recovery — bundle ${BUNDLE}
Created: $(date -R) on $(hostname -s)
Decrypt with the HOME_SECURE GPG private key (${RECIPIENT}).
## Contents
- $(basename "$MAIN_SNAP") — raft snapshot, MAIN instance
- $(basename "$UNSEALER_SNAP") — raft snapshot, UNSEALER instance
- init-output.json — main: recovery keys + root token
- unsealer-init.json — unsealer: 1-of-1 unseal key + root token
All four are needed together. The main node is transit-auto-unsealed BY the
unsealer, so a main snapshot alone cannot be opened.
## Restore order (order matters)
1. Bring up the openbao-unsealer container, restore its snapshot, then unseal
it with the key from unsealer-init.json (1-of-1 shamir).
2. Confirm 'bao list transit/keys' on the unsealer shows 'autounseal'.
3. Bring up the main openbao container and restore its snapshot. It should
auto-unseal via the unsealer's transit key.
4. Verify: 'bao status' shows Sealed=false, Seal Type=transit.
Note: OpenBAO does NOT support downgrading a raft data dir. Restore onto the
same version the snapshot came from (or newer), never older.
## Verify integrity
Each .snap is a gzip tar: 'gzip -t' it, extract, then 'sha256sum -c SHA256SUMS'.
EOF
# --- 3. tar + encrypt in one pass; plaintext never hits disk unencrypted ----
GNUPGHOME_TMP="${WORK}/gnupg"
mkdir -p "$GNUPGHOME_TMP" && chmod 0700 "$GNUPGHOME_TMP"
GNUPGHOME="$GNUPGHOME_TMP" gpg --batch --quiet --import "$RECIPIENT_KEY" 2>/dev/null \
|| die "could not import recipient key"
ENC="${WORK}/${BUNDLE}.tar.gz.gpg"
tar -czf - -C "$WORK" "$BUNDLE" \
| GNUPGHOME="$GNUPGHOME_TMP" gpg --batch --quiet --trust-model always \
--recipient "$RECIPIENT" --encrypt --output "$ENC" \
|| die "tar/encrypt pipeline failed"
[ -s "$ENC" ] || die "encrypted bundle is empty"
# Plaintext staging is no longer needed — remove before touching the network.
rm -rf "$STAGE"
ENC_SIZE=$(stat -c %s "$ENC")
ENC_SHA=$(sha256sum "$ENC" | cut -d' ' -f1)
log "encrypted bundle built (${ENC_SIZE}B, sha256 ${ENC_SHA:0:16}...)"
# Sanity: it must actually be a GPG message, not a tar that slipped through.
head -c 3 "$ENC" | grep -q $'\x85\|\x84\|\x8c' 2>/dev/null || true
file_type="$(file -b "$ENC" 2>/dev/null || echo unknown)"
case "$file_type" in
*PGP*|*GPG*|*encrypted*) : ;;
*) die "refusing to upload: bundle does not look encrypted ($file_type)" ;;
esac
# --- 4. ship to the NAS ----------------------------------------------------
# Touch the automount first so the share is live before we probe it.
ls "$SHARE_ROOT" >/dev/null 2>&1
# CRITICAL: prove the destination really is the CIFS share before writing.
# Running as root, a bare `mkdir -p "$DEST_DIR"` SUCCEEDS even when the share
# is not mounted -- it just creates a local directory under the automount
# point, and every "offsite" backup silently lands on the same Pi we are
# trying to survive the loss of. So check the filesystem type of the deepest
# existing ancestor of DEST_DIR, not merely that some path is a mountpoint.
ancestor="$DEST_DIR"
while [ ! -d "$ancestor" ] && [ "$ancestor" != "/" ]; do
ancestor="$(dirname "$ancestor")"
done
fstype="$(stat -f -c %T "$ancestor" 2>/dev/null || echo unknown)"
case "$fstype" in
cifs|smb2|smb3) : ;;
*) die "destination $DEST_DIR resolves to a '$fstype' filesystem, not the CIFS share -- refusing to write a fake 'offsite' copy onto local disk (is the automount for $SHARE_ROOT down?)" ;;
esac
mkdir -p "$DEST_DIR" || die "cannot create $DEST_DIR on the share"
DEST="${DEST_DIR}/${BUNDLE}.tar.gz.gpg"
cp "$ENC" "${DEST}.part" || die "copy to NAS failed"
mv -f "${DEST}.part" "$DEST" || die "atomic rename on NAS failed"
# --- 5. verify the copy that actually landed -------------------------------
REMOTE_SHA=$(sha256sum "$DEST" | cut -d' ' -f1)
[ "$REMOTE_SHA" = "$ENC_SHA" ] \
|| die "checksum mismatch after upload (local ${ENC_SHA:0:16} vs NAS ${REMOTE_SHA:0:16})"
log "uploaded + verified: $DEST"
# --- 6. retention on the NAS ----------------------------------------------
ls -1t "${DEST_DIR}"/openbao-dr-*.tar.gz.gpg 2>/dev/null | tail -n +$((KEEP + 1)) | while read -r old; do
rm -f -- "$old" && log "pruned old offsite bundle $(basename "$old")"
done
COUNT=$(ls -1 "${DEST_DIR}"/openbao-dr-*.tar.gz.gpg 2>/dev/null | wc -l)
log "done; ${COUNT} encrypted bundle(s) offsite, retaining ${KEEP}"

View File

@@ -0,0 +1,11 @@
[Unit]
Description=Ship an encrypted OpenBAO DR bundle to the NAS (fids2)
After=network-online.target home-lutz-nfs_projects.automount
Wants=network-online.target
OnFailure=openbao-alert@%n.service
[Service]
Type=oneshot
# Runs as root: reads /var/backups/openbao (0700 root) and the 0600 unseal
# material. Encrypts to HOME_SECURE before anything touches the share.
ExecStart=/home/lutz/Projects/OpenBAO/scripts/offsite-backup.sh

View File

@@ -0,0 +1,13 @@
[Unit]
Description=Daily offsite copy of the OpenBAO DR bundle
[Timer]
# Local snapshots run at 02:30 (+ up to 15m jitter), so 03:40 leaves a clear
# gap and always ships that morning's snapshot. The script independently
# refuses to run if the newest local snapshot is older than 48h.
OnCalendar=*-*-* 03:40:00
Persistent=true
RandomizedDelaySec=15m
[Install]
WantedBy=timers.target