Files
Software/artemis-backup-20260614-1920/wolfstack-keepalived/README.md

107 lines
4.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# WolfStack Keepalived HA Setup
## Files
| File | Node | Priority | Intended State |
|------|------|----------|---------------|
| `keepalived-node1-primary.conf` | 192.168.5.214 | 150 | MASTER |
| `keepalived-node2-secondary.conf` | `192.168.15.182` | BACKUP | 130 |
| `keepalived-node3-tertiary.conf` | 192.168.50.41 | 110 | BACKUP |
Shared VIP: `192.168.63.254` on interface `eth0`
## Quick Install (per node)
```bash
sudo apt install keepalived netcat-openbsd -y
sudo cp keepalived-nodeX-*.conf /etc/keepalived/keepalived.conf
sudo systemctl enable keepalived --now
sudo systemctl status keepalived
```
Watch the syslog:
```bash
sudo journalctl -u keepalived -f
```
## Network Confirmation
All three nodes (`.5.x`, `.50.x`) and the VIP (`.63.254`) fall within the same `192.168.0.0/18` network (`192.168.0.0` `192.168.63.255`). This means they share Layer 2/L3 broadcast domain and VRRP unicast failover works cleanly across all nodes — no VLAN bridging or routing tricks required. ✅
### What Happens In Practice
| Scenario | Outcome |
|----------|---------|
| Node 1 dies, Node 2 healthy | VIP moves to Node 2 ✅ |
| Node 1 + Node 2 die, Node 3 alive | VIP moves to Node 3 ✅ (clients reach it directly) |
| Node 3 is primary but 1 & 2 alive | Node 3 will NOT win because priority 110 < 150 / 130 ✅ |
| All 3 nodes lose peer connectivity | **Split-brain risk** — each may claim VIP. Use fencing/STONITH if this is production-critical. |
## Track Script Explained
```bash
script "/usr/bin/nc -z -w 2 127.0.0.1 8553"
```
This checks whether the **local** WolfStack UI (port 8553) is accepting TCP connections.
- **If WolfStack crashes / stops** on the MASTER, keepalived subtracts `weight -40` from that node's priority.
- Priority drops from 150 → 110, which is below Node 2 (130). Node 2 promotes to MASTER and takes the VIP.
- If Node 1 comes back online *and* WolfStack is healthy, its priority returns to 150. The VIP will move back **unless** you keep `nopreempt`.
### Preempt Behaviour
| Setting | Behaviour |
|---------|-----------|
| `nopreempt` present | Once VIP is on a backup node, it stays there even when higher-priority node recovers. Good for stability. Manual failback required. |
| `nopreempt` removed | Highest priority node reclaims VIP automatically. Good for deterministic primary/secondary but causes brief flap. |
## ⚠️ WolfStack-Specific Caveats (from research)
1. **No native UI HA**: WolfStack's web UI is single-node. If you point browsers at the VIP and Node 2 takes over, **users must re-authenticate** (sessions are local to each node).
2. **No reverse-proxy awareness**: The UI doesn't have a "behind proxy" setting. If you terminate SSL on a proxy and forward to `http://backend:8553`, you may hit mixed-content or CSP issues.
3. **WebSockets**: The dashboard uses websockets for the terminal and real-time graphs. Make sure your proxy (or keepalived direct-access path) supports websocket upgrades if not hitting the VIP directly.
4. **WolfProxy**: WolfStack ships a sibling `WolfProxy` (NGINX-compatible), but that is a **separate** reverse proxy for *other services*, not for making the WolfStack UI highly available.
## Suggested Topology
```
Clients / Browsers
|
v
+---------------------+
| VIP: 192.168.63.254 |
| (keepalived floating) |
+----------+----------+
|
+----------------+----------------+
| |
Node 1 (MASTER) Node 2 (BACKUP)
192.168.5.214:8553 192.168.15.182:8553
priority 150 priority 130
| |
WolfStack WolfStack
WolfNet Mesh WolfNet Mesh
+------------+ 192.168.50.41 +---+
Node 3 (BACKUP)
priority 110
(same VLAN?)
```
## Testing Failover
1. From any client, `curl -k https://192.168.63.254:8553` should show the WolfStack UI.
2. On the MASTER node: `sudo systemctl stop wolfstack` (simulating WolfStack crash).
3. Watch `journalctl -u keepalived -f` on Node 2 — it should promote to MASTER within ~6 seconds (3 x 2s intervals).
4. `ip addr show eth0` on Node 2 should now show `192.168.63.254/32`.
5. Client traffic should continue via the VIP.
6. Restart WolfStack on Node 1. If `nopreempt` is enabled, VIP stays on Node 2. If removed, VIP moves back to Node 1.
## Rollback
```bash
sudo systemctl stop keepalived
sudo ip addr del 192.168.63.254 dev eth0 # if needed manually
sudo systemctl disable keepalived
```