MongoDB Replica Set — VM Cluster Setup
Three VMs: one Primary, one Secondary, one Arbiter. This guide walks every step from bare OS to a monitored, production-ready replica set — firewall rules, mongod.conf, user creation, replica set bootstrap, write-concern tuning, and the Prometheus exporter.
Cluster Architecture
graph TD
classDef primary fill:#e74c3c,stroke:#c0392b,color:#fff
classDef secondary fill:#3498db,stroke:#2980b9,color:#fff
classDef arbiter fill:#f39c12,stroke:#d68910,color:#fff
classDef client fill:#2ecc71,stroke:#27ae60,color:#fff
classDef exporter fill:#9b59b6,stroke:#8e44ad,color:#fff
subgraph "VM 1 — Primary (10.0.0.1)"
P["mongod :27017<br/>PRIMARY<br/>reads + writes<br/>holds full oplog"]:::primary
EXP["mongodb_exporter :9216<br/>Prometheus metrics"]:::exporter
end
subgraph "VM 2 — Secondary (10.0.0.2)"
S["mongod :27017<br/>SECONDARY<br/>tails oplog from primary<br/>read-only (optional)"]:::secondary
end
subgraph "VM 3 — Arbiter (10.0.0.3)"
A["mongod :27017<br/>ARBITER<br/>vote-only, no data<br/>breaks election ties"]:::arbiter
end
APP["Application"]:::client
PROM["Prometheus"]:::client
P -->|"oplog stream (async)"| S
P -->|"heartbeat :27017"| A
S -->|"heartbeat :27017"| A
APP -->|"w:majority writes"| P
APP -->|"readPreference: secondary"| S
PROM -->|"scrape :9216"| EXP
EXP -->|"auth :27017"| P
Minimum viable quorum: primary + arbiter = 2 votes. That's enough to elect a new primary if the secondary dies, without paying for a third full data copy. An arbiter holds no data and is cheap to run on a small VM.
The secondary VM is completely down. Can the primary keep accepting writes? Can the arbiter become primary?
1. Firewall Rules
MongoDB members communicate on port 27017. Every node must allow inbound 27017 from the other two nodes. Open this before bootstrapping the replica set — misconfigured firewall rules cause silent heartbeat failures that look like replica set bugs.
What to open (all three nodes)
| Source | Destination | Port | Why |
|---|---|---|---|
| Secondary (10.0.0.2) | Primary (10.0.0.1) | 27017/tcp | oplog pull, heartbeat |
| Arbiter (10.0.0.3) | Primary (10.0.0.1) | 27017/tcp | heartbeat, vote |
| Primary (10.0.0.1) | Secondary (10.0.0.2) | 27017/tcp | heartbeat (bidirectional) |
| Primary (10.0.0.1) | Arbiter (10.0.0.3) | 27017/tcp | heartbeat |
| Secondary (10.0.0.2) | Arbiter (10.0.0.3) | 27017/tcp | heartbeat |
| App servers / Prometheus | Primary (10.0.0.1) | 27017/tcp | client connections |
| Prometheus | Primary (10.0.0.1) | 9216/tcp | exporter scrape |
# Run on PRIMARY — allow secondary and arbiter
PRIMARY_IP="10.0.0.1"
SECONDARY_IP="10.0.0.2"
ARBITER_IP="10.0.0.3"
firewall-cmd --permanent --add-rich-rule="rule family=ipv4 source address=${SECONDARY_IP} port port=27017 protocol=tcp accept"
firewall-cmd --permanent --add-rich-rule="rule family=ipv4 source address=${ARBITER_IP} port port=27017 protocol=tcp accept"
# Allow Prometheus scrape of exporter
firewall-cmd --permanent --add-rich-rule="rule family=ipv4 port port=9216 protocol=tcp accept"
firewall-cmd --reload
# Run on SECONDARY — allow primary + arbiter heartbeats
firewall-cmd --permanent --add-rich-rule="rule family=ipv4 source address=${PRIMARY_IP} port port=27017 protocol=tcp accept"
firewall-cmd --permanent --add-rich-rule="rule family=ipv4 source address=${ARBITER_IP} port port=27017 protocol=tcp accept"
firewall-cmd --reload
# Run on ARBITER — allow primary + secondary heartbeats
firewall-cmd --permanent --add-rich-rule="rule family=ipv4 source address=${PRIMARY_IP} port port=27017 protocol=tcp accept"
firewall-cmd --permanent --add-rich-rule="rule family=ipv4 source address=${SECONDARY_IP} port port=27017 protocol=tcp accept"
firewall-cmd --reload
# Verify
firewall-cmd --list-rich-rules
</div>
<div class="tab-panel" data-tab-panel="ufw">
# PRIMARY
ufw allow from 10.0.0.2 to any port 27017 proto tcp comment "mongo secondary"
ufw allow from 10.0.0.3 to any port 27017 proto tcp comment "mongo arbiter"
ufw allow 9216/tcp comment "mongodb-exporter"
ufw reload
# SECONDARY
ufw allow from 10.0.0.1 to any port 27017 proto tcp comment "mongo primary"
ufw allow from 10.0.0.3 to any port 27017 proto tcp comment "mongo arbiter"
ufw reload
# ARBITER
ufw allow from 10.0.0.1 to any port 27017 proto tcp comment "mongo primary"
ufw allow from 10.0.0.2 to any port 27017 proto tcp comment "mongo secondary"
ufw reload
# Verify
ufw status numbered
</div>
<div class="tab-panel" data-tab-panel="iptables">
# PRIMARY — accept from secondary and arbiter
iptables -A INPUT -s 10.0.0.2 -p tcp --dport 27017 -j ACCEPT
iptables -A INPUT -s 10.0.0.3 -p tcp --dport 27017 -j ACCEPT
iptables -A INPUT -p tcp --dport 9216 -j ACCEPT
# Persist (Debian/Ubuntu)
apt-get install -y iptables-persistent
iptables-save > /etc/iptables/rules.v4
# Persist (RHEL/Rocky)
service iptables save
</div>
Connection check — before you continue
# From SECONDARY VM — must succeed before any mongod config
nc -zv 10.0.0.1 27017
# Expected: Connection to 10.0.0.1 27017 port [tcp] succeeded!
# From ARBITER VM
nc -zv 10.0.0.1 27017
# Expected: Connection to 10.0.0.1 27017 port [tcp] succeeded!
# If nc is not available
curl -s telnet://10.0.0.1:27017 --max-time 2 || echo "FAILED — check firewall"
⚠️ Do not proceed to mongod configuration until
nc -zvsucceeds from both secondary and arbiter. A replica set member that cannot reach its peers will continuously retry and never reach a healthy state. Fix the network first.
You run nc -zv 10.0.0.1 27017 from the secondary and get "Connection refused". The firewall rules look correct. What else could cause this?
2. Keyfile (Shared Secret)
All members of a replica set authenticate to each other using a shared keyfile. Generate it once on the primary and copy it to every node.
# Generate on primary
openssl rand -base64 756 > /etc/mongodb/keyfile
chmod 400 /etc/mongodb/keyfile
chown mongod:mongod /etc/mongodb/keyfile
# Copy to secondary and arbiter
scp /etc/mongodb/keyfile user@10.0.0.2:/etc/mongodb/keyfile
scp /etc/mongodb/keyfile user@10.0.0.3:/etc/mongodb/keyfile
# Set permissions on each remote node
ssh user@10.0.0.2 "chmod 400 /etc/mongodb/keyfile && chown mongod:mongod /etc/mongodb/keyfile"
ssh user@10.0.0.3 "chmod 400 /etc/mongodb/keyfile && chown mongod:mongod /etc/mongodb/keyfile"
The keyfile must be identical on all members, readable only by the
mongoduser, and at least 6 characters long (typically 756 bytes of base64).
3. Primary — mongod.conf
# /etc/mongod.conf (Primary: 10.0.0.1)
systemLog:
destination: file
logAppend: true
path: /var/log/mongodb/mongod.log
storage:
dbPath: /var/lib/mongo
journal:
enabled: true
engine: wiredTiger
wiredTiger:
engineConfig:
# Rule: ~50% of available RAM, leaving room for OS page cache.
# Example: 16 GB RAM → cacheSizeGB: 6 (not 8, so page cache + OS has room)
# Default if unset: 50% of (RAM - 1 GB), min 256 MB
cacheSizeGB: 6
net:
port: 27017
bindIp: 0.0.0.0 # listens on all interfaces; restrict to specific IPs if preferred
maxIncomingConnections: 65536
replication:
replSetName: "rs0"
# oplogSizeMB default: 5% of disk, min 990 MB. Increase for busy primaries.
oplogSizeMB: 5120 # 5 GB — enough for ~24h of lag tolerance on moderate workloads
security:
authorization: enabled
keyFile: /etc/mongodb/keyfile
processManagement:
timeZoneInfo: /usr/share/zoneinfo
cacheSizeGB — how to size it
free -g or cat /proc/meminfo | grep MemTotal. Example: 32 GB total.
(RAM - 1GB) × 0.5. In practice for a dedicated MongoDB VM, set it to roughly 40–50% of RAM — e.g., 32 GB RAM → cacheSizeGB: 12 to 14. Leave the rest for the OS page cache, oplog in memory, and connection overhead.
db.serverStatus().wiredTiger.cacheWatch "pages evicted because they exceeded the in-memory maximum" and "tracked dirty bytes in the cache". If eviction is high, increase
cacheSizeGB.
Your primary has 64 GB of RAM. You set cacheSizeGB: 60 to maximize MongoDB cache. What problem will this cause in production?
4. Bootstrap the Replica Set (Primary only)
Start mongod without security first to create the admin user, then enable auth.
# 1. Start mongod (security disabled for initial user creation)
# Temporarily comment out the security: block in mongod.conf, then:
systemctl start mongod
# 2. Connect locally
mongosh --host 127.0.0.1 --port 27017
// 3. Initiate the replica set — PRIMARY only, with just itself
rs.initiate({
_id: "rs0",
members: [
{ _id: 0, host: "10.0.0.1:27017", priority: 2 } // higher priority = prefers to stay primary
]
})
// Confirm primary is elected (may take a few seconds)
rs.status()
// Look for: "stateStr" : "PRIMARY"
Create Users
// Switch to admin db
use admin
// 1. Admin (superuser — for ops/DBA work only)
db.createUser({
user: "admin",
pwd: "ch@ngeM3!", // use a strong generated password in production
roles: [{ role: "root", db: "admin" }]
})
// 2. Reader (read-only access to application DB)
db.createUser({
user: "reader",
pwd: "R3ad0nly!",
roles: [
{ role: "read", db: "myapp" }
]
})
// 3. Writer (read+write access to application DB)
db.createUser({
user: "writer",
pwd: "Wr1teUser!",
roles: [
{ role: "readWrite", db: "myapp" }
]
})
// Verify all users were created
db.getUsers()
find, listCollections, aggregate (read-only pipelines), and count. No writes, no schema changes. Suitable for analytics queries, reporting tools, and read-only dashboards.
insert, update, delete, and createCollection/Index. Use this for your application's primary connection. It is scoped to one database — the writer cannot touch other databases.
# 4. Enable security — re-enable the security block in mongod.conf, then restart
systemctl restart mongod
# 5. Verify auth works
mongosh --host 10.0.0.1 -u admin -p 'ch@ngeM3!' --authenticationDatabase admin
Your application connects with the writer user scoped to the myapp database. It tries to run a query on the logs database. What happens?
Unauthorized error. The readWrite role was granted on the myapp database only. MongoDB roles are per-database — a role on myapp gives no access to logs. To grant access to multiple databases, either create additional roles scoped to each, or grant readWriteAnyDatabase on the admin database (which is much broader — scope carefully).
5. Add Secondary and Arbiter
// On PRIMARY — connect as admin
mongosh --host 10.0.0.1 -u admin -p 'ch@ngeM3!' --authenticationDatabase admin
// Add secondary (full data-bearing member)
rs.add({ host: "10.0.0.2:27017", priority: 1 })
// Add arbiter (vote-only, no data)
rs.addArb("10.0.0.3:27017")
// ─── IMPORTANT: check status ───────────────────────────────────────────
rs.status()
rs.status() — what to look for
// Healthy output (abbreviated)
{
"set": "rs0",
"myState": 1, // 1 = PRIMARY
"members": [
{
"name": "10.0.0.1:27017",
"stateStr": "PRIMARY",
"health": 1,
"optime": { ... }
},
{
"name": "10.0.0.2:27017",
"stateStr": "SECONDARY",
"health": 1,
"optimeDate": ISODate("..."),
"lastHeartbeatMessage": "", // empty = healthy
"syncSourceHost": "10.0.0.1:27017"
},
{
"name": "10.0.0.3:27017",
"stateStr": "ARBITER",
"health": 1
}
]
}
health: 1. A 0 here means the member is unreachable from the primary — check firewall and whether mongod is running on that VM.
STARTUP2 while it performs initial sync (copying the primary's data). Wait for it to transition to SECONDARY before routing any reads to it.
optimeDate should be close to the primary's. A large gap means the secondary is falling behind — check secondary VM resources (disk I/O, CPU) and oplog size.
10.0.0.1:27017). In a larger cluster, secondaries can sync from other secondaries (chained replication) — fine in most cases, can increase lag.
Right after running rs.addArb("10.0.0.3:27017"), rs.status() shows the arbiter with stateStr: "UNKNOWN" and health: 0. What's the most likely cause?
mongod is not running on the arbiter VM, (2) the firewall blocks the primary from reaching port 27017 on the arbiter, or (3) the keyfile on the arbiter is missing, wrong permissions, or differs from the primary's keyfile. The arbiter cannot join the set if authentication fails or the port is unreachable. Run nc -zv 10.0.0.3 27017 from the primary to isolate networking vs. auth.
6. Secondary — mongod.conf
The secondary's config is nearly identical to the primary. Key difference: no special initialization is needed — it joins via rs.add() from the primary and syncs automatically.
# /etc/mongod.conf (Secondary: 10.0.0.2)
systemLog:
destination: file
logAppend: true
path: /var/log/mongodb/mongod.log
storage:
dbPath: /var/lib/mongo
journal:
enabled: true
engine: wiredTiger
wiredTiger:
engineConfig:
cacheSizeGB: 6 # same sizing logic as primary
net:
port: 27017
bindIp: 0.0.0.0 # primary's IP must be able to reach this
maxIncomingConnections: 65536
replication:
replSetName: "rs0" # MUST match primary exactly (case-sensitive)
security:
authorization: enabled
keyFile: /etc/mongodb/keyfile # same keyfile as primary
processManagement:
timeZoneInfo: /usr/share/zoneinfo
# Start and enable on boot
systemctl enable --now mongod
# Verify it's running and listening on 27017
ss -tlnp | grep 27017
# tcp LISTEN 0.0.0.0:27017 ... users:(("mongod",...))
The secondary's replSetName is set to RS0 (uppercase) while the primary uses rs0. The secondary appears stuck in STARTUP. Why?
replSetName is case-sensitive. A secondary with RS0 cannot join a set named rs0 — it refuses to participate because the names don't match. This is one of the most common config typos. Fix the mongod.conf on the secondary, restart mongod, and verify with rs.status() on the primary.
7. Arbiter — mongod.conf
The arbiter runs a real mongod process but stores no application data. Its config is minimal — it only needs the replica set name and the keyfile to authenticate. Give it the smallest VM you can justify (1–2 vCPU, 1–2 GB RAM is fine).
# /etc/mongod.conf (Arbiter: 10.0.0.3)
systemLog:
destination: file
logAppend: true
path: /var/log/mongodb/mongod.log
storage:
dbPath: /var/lib/mongo/arbiter # keep separate from any data you might have
journal:
enabled: true
# Explicitly set small — arbiter stores almost no data
wiredTiger:
engineConfig:
cacheSizeGB: 0.25
net:
port: 27017
bindIp: 0.0.0.0
replication:
replSetName: "rs0" # must match primary
security:
authorization: enabled
keyFile: /etc/mongodb/keyfile # same keyfile as primary
processManagement:
timeZoneInfo: /usr/share/zoneinfo
# Create the arbiter data directory
mkdir -p /var/lib/mongo/arbiter
chown -R mongod:mongod /var/lib/mongo/arbiter
systemctl enable --now mongod
⚠️ Do not run the arbiter on the same VM as the primary or secondary. An arbiter's entire value is providing an independent vote during an election. Co-located, if that VM fails you lose both the primary and the tiebreaker simultaneously — exactly the failure you were guarding against.
Should you add cacheSizeGB to the arbiter's config? What happens if you leave it at its default?
cacheSizeGB: 0.25 (256 MB) is good practice — it signals intent and prevents the cache from grabbing half the RAM on a small VM where memory is scarce.
Automated Setup Script
The steps above (conf file, systemd unit, data directory, daemon-reload) are mechanical enough to script. Save this as setup-mongo-arbiter.sh on the arbiter VM and run it as root.
#!/usr/bin/env bash
# -----------------------------------------------------------------------------
# setup-mongo-arbiter.sh
#
# Sets up a MongoDB arbiter instance on this VM:
# - Writes /etc/mongod-<db-name>.conf
# - Writes /etc/systemd/system/mongod-<db-name>-arbiter.service
# - Creates /data/mongodb-<db-name> (owned by mongodb:mongodb)
# - Runs systemd daemon-reload, enable, and start
#
# Usage:
# sudo ./setup-mongo-arbiter.sh # interactive, applies changes
# sudo ./setup-mongo-arbiter.sh --dry-run # interactive, only prints what would happen
# -----------------------------------------------------------------------------
set -euo pipefail
die() { echo "ERROR: $*" >&2; exit 1; }
info() { echo "==> $*"; }
# ── dry-run flag ──────────────────────────────────────────────────────────────
DRY_RUN=false
for arg in "$@"; do
case "$arg" in
--dry-run) DRY_RUN=true ;;
*) die "Unknown argument: $arg. Usage: $0 [--dry-run]" ;;
esac
done
# Wrapper: in dry-run mode print the command instead of running it
run() {
if $DRY_RUN; then
echo " [dry-run] $*"
else
"$@"
fi
}
# Write a file: in dry-run mode show the content that would be written
write_file() {
local path="$1"
local content="$2"
local perms="$3"
if $DRY_RUN; then
echo " [dry-run] would write ${path} (chmod ${perms}):"
echo "$content" | sed 's/^/ /'
echo
else
echo "$content" > "$path"
chmod "$perms" "$path"
fi
}
# ── must run as root ──────────────────────────────────────────────────────────
[[ "$EUID" -eq 0 ]] || die "This script must be run as root (use sudo)"
# ── header ────────────────────────────────────────────────────────────────────
echo "============================================"
if $DRY_RUN; then
echo " MongoDB Arbiter Setup [DRY RUN]"
else
echo " MongoDB Arbiter Setup"
fi
echo "============================================"
echo
# ── interactive prompts ───────────────────────────────────────────────────────
# Replicaset name
while true; do
read -rp "Enter replicaset name: " REPLSET
[[ -n "$REPLSET" ]] && break
echo " Replicaset name cannot be empty. Please try again."
done
# Port
while true; do
read -rp "Enter port number for the arbiter: " PORT
if [[ "$PORT" =~ ^[0-9]+$ ]] && (( PORT >= 1024 && PORT <= 65535 )); then
break
fi
echo " Invalid port. Must be a number between 1024 and 65535. Please try again."
done
# ── derive db name ────────────────────────────────────────────────────────────
# Strips trailing -rs/<rs suffix> and leading rs prefix, lowercases,
# and normalises underscores to hyphens.
# e.g. myapp-rs → myapp
# rs0-auth → auth
DERIVED_DB_NAME=$(echo "$REPLSET" \
| sed -E 's/-[Rr][Ss][0-9]*$//;s/^[Rr][Ss][0-9]*-//' \
| sed 's/^[-_]*//;s/[-_]*$//' \
| tr '[:upper:]' '[:lower:]' \
| tr '_' '-')
# Fall back to the full lowercased replicaset name if nothing was left
if [[ -z "$DERIVED_DB_NAME" ]]; then
DERIVED_DB_NAME=$(echo "$REPLSET" | tr '[:upper:]' '[:lower:]' | tr '_' '-')
fi
# Let the user confirm or override
echo
read -rp "DB name derived from replicaset [${DERIVED_DB_NAME}] (press Enter to accept or type to override): " DB_NAME_INPUT
DB_NAME="${DB_NAME_INPUT:-$DERIVED_DB_NAME}"
# ── summary & confirmation ────────────────────────────────────────────────────
CONF_FILE="/etc/mongod-${DB_NAME}.conf"
SERVICE_FILE="/etc/systemd/system/mongod-${DB_NAME}-arbiter.service"
SERVICE_NAME="mongod-${DB_NAME}-arbiter.service"
DATA_DIR="/data/mongodb-${DB_NAME}"
echo
echo "--------------------------------------------"
if $DRY_RUN; then
echo " Mode : DRY RUN — no changes will be made"
fi
echo " Replicaset : $REPLSET"
echo " Port : $PORT"
echo " DB name : $DB_NAME"
echo " Config : $CONF_FILE"
echo " Service : $SERVICE_FILE"
echo " Data dir : $DATA_DIR"
echo "--------------------------------------------"
echo
read -rp "Proceed? [y/N]: " CONFIRM
case "$CONFIRM" in
[yY]|[yY][eE][sS]) ;;
*) echo "Aborted."; exit 0 ;;
esac
echo
# ── guard: abort if targets already exist (skip in dry-run) ──────────────────
if ! $DRY_RUN; then
for target in "$CONF_FILE" "$SERVICE_FILE"; do
if [[ -e "$target" ]]; then
die "$target already exists — remove it manually if you want to recreate it"
fi
done
fi
# ── 1. Write mongod config ────────────────────────────────────────────────────
info "Writing $CONF_FILE"
MONGOD_CONF="storage:
dbPath: ${DATA_DIR}
engine: wiredTiger
systemLog:
destination: file
path: ${DATA_DIR}/mongod.log
logAppend: true
net:
port: ${PORT}
bindIp: 0.0.0.0
replication:
replSetName: ${REPLSET}
security:
authorization: enabled
keyFile: /etc/mongodb-keyfile"
write_file "$CONF_FILE" "$MONGOD_CONF" "640"
$DRY_RUN || info " Done → $CONF_FILE"
# ── 2. Write systemd service ──────────────────────────────────────────────────
info "Writing $SERVICE_FILE"
MONGOD_SERVICE="[Unit]
Description=MongoDB instance ${DB_NAME} arbiter
After=network.target
[Service]
User=mongodb
Group=mongodb
ExecStart=/usr/bin/mongod --config /etc/mongod-${DB_NAME}.conf
Restart=always
RestartSec=5
TimeoutStartSec=180
LimitNOFILE=64000
LimitNPROC=64000
StandardOutput=journal
StandardError=journal
SyslogIdentifier=mongod-${DB_NAME}-arbiter
[Install]
WantedBy=multi-user.target"
write_file "$SERVICE_FILE" "$MONGOD_SERVICE" "644"
$DRY_RUN || info " Done → $SERVICE_FILE"
# ── 3. Create data directory ──────────────────────────────────────────────────
info "Creating data directory $DATA_DIR"
if $DRY_RUN; then
echo " [dry-run] would mkdir -p $DATA_DIR"
echo " [dry-run] would chown -R mongodb:mongodb $DATA_DIR"
echo " [dry-run] would chmod 755 $DATA_DIR"
else
if [[ -d "$DATA_DIR" ]]; then
info " Directory already exists — skipping mkdir"
else
mkdir -p "$DATA_DIR"
info " Created $DATA_DIR"
fi
chown -R mongodb:mongodb "$DATA_DIR"
chmod 755 "$DATA_DIR"
info " Ownership set to mongodb:mongodb on $DATA_DIR"
fi
# ── 4. Reload systemd & start service ────────────────────────────────────────
info "Running systemctl daemon-reload"
run systemctl daemon-reload
info "Enabling $SERVICE_NAME"
run systemctl enable "$SERVICE_NAME"
info "Starting $SERVICE_NAME"
run systemctl start "$SERVICE_NAME"
# ── 5. Done ───────────────────────────────────────────────────────────────────
echo
if $DRY_RUN; then
info "Service status check (skipped in dry-run)"
echo
echo "✓ Dry run complete — no changes were made."
echo " Run without --dry-run to apply."
else
info "Service status:"
systemctl status "$SERVICE_NAME" --no-pager --lines=10 || true
echo
echo "✓ Arbiter setup complete for replicaset '${REPLSET}' (db: ${DB_NAME}) on port ${PORT}"
fi
echo " Config : $CONF_FILE"
echo " Service : $SERVICE_FILE"
echo " Data dir : $DATA_DIR"
What the script does:
- Prompts for replicaset name and port — validates both before proceeding
- Derives a short
DB_NAMEfrom the replicaset name by stripping trailing-rs/-rs0suffixes and leadingrs0-prefixes (e.g.myapp-rs→myapp) - Shows a full summary and asks for confirmation before touching anything
- Guards against overwriting existing conf/service files — fails fast rather than silently clobbering
--dry-runprints exactly what would be written and run, without touching the filesystem or systemd
# Make executable and run
chmod +x setup-mongo-arbiter.sh
# Preview first
sudo ./setup-mongo-arbiter.sh --dry-run
# Apply
sudo ./setup-mongo-arbiter.sh
What the Arbiter Actually Does
People often treat the arbiter as a black box — "it votes". Here's what it's actually doing at every moment.
heartbeatIntervalMillis: 2000). It tracks each member's health, state (PRIMARY/SECONDARY), and optime. Members also heartbeat back to it — the arbiter is a full participant in the mesh, just without data.
local database like any mongod, but it only holds replica set metadata (local.system.replset). It never receives oplog entries, never applies operations, and has no local.oplog.rs with user data in it. Its dbPath stays near-empty.
electionTimeoutMillis (10 seconds default), it concludes the primary is down. It then triggers an election — or supports a candidacy from a secondary that also noticed.
w: "majority" write concern. Write concern counts data-bearing members only. In a PSA set that means primary + secondary — the arbiter is invisible to write acknowledgment. See the Write Concern section for the full implications.
The arbiter VM loses its network connection to both the primary and secondary. From the primary's perspective, is quorum still intact?
8. Async Replication — The Oplog
When you write to the primary with w: 1, MongoDB confirms the write immediately and replicates to secondaries in the background. That background mechanism is the oplog.
sequenceDiagram
participant APP as Application
participant P as Primary (10.0.0.1)
participant OPLOG as local.oplog.rs<br/>(primary)
participant S as Secondary (10.0.0.2)
APP->>P: db.orders.insertOne({...})
P->>P: write to WiredTiger data files
P->>OPLOG: append oplog entry {op:"i", ns:"myapp.orders", o:{...}}
P-->>APP: WriteResult OK ← client unblocked here (w:1)
Note over S: tailing cursor on primary's oplog (long-poll)
OPLOG-->>S: new entry available
S->>S: apply operation to local data files
S->>S: append same entry to own local.oplog.rs
Note over S: secondary's optime advances
w: 1 flow in plain terms: The primary writes to its data files and its own oplog, then tells the client "done" — before the secondary has seen anything. The secondary is always catching up asynchronously. The gap between the primary's latest optime and the secondary's optime is the replication lag.
The Oplog — What's Inside
// Inspect the oplog on the primary
use local
db.oplog.rs.find().sort({ $natural: -1 }).limit(3).pretty()
{
"ts": { "$timestamp": { "t": 1723334400, "i": 1 } },
"t": 1, // election term — increments on each election
"op": "i", // operation: i=insert, u=update, d=delete, c=command, n=noop
"ns": "myapp.orders",
"ui": "<collection UUID>",
"wall": "2026-08-11T10:00:00Z",
"o": { "_id": "...", "item": "widget", "qty": 100 }
}
| Field | Meaning |
|---|---|
ts |
Timestamp + increment counter. Secondaries use this to track where they are in the oplog |
t |
Election term. Lets members detect stale oplog entries from a previous primary |
op |
i insert · u update · d delete · c command (DDL) · n noop (heartbeat tick) |
ns |
Namespace: db.collection |
o / o2 |
The operation: o is the document or update spec; o2 is the query filter for updates |
Oplog entries are idempotent. MongoDB rewrites updates (even $inc) into full replacement forms before writing to the oplog. This means an oplog entry can be applied multiple times and produce the same result — critical for safe replication and crash recovery.
How the Secondary Follows
local.oplog.rs of its sync source (usually the primary). This is a blocking read — the cursor waits at the end of the oplog until a new entry appears, then immediately returns it. No polling, no sleep loops.
local.oplog.rs — this is how the secondary's oplog grows too.
optime — the timestamp of the last entry it applied. This is what rs.status() reports as optimeDate. The primary tracks each member's optime via heartbeat responses.
allowChaining: true (default), MongoDB may automatically route a secondary to sync from another secondary if that path has lower latency. This reduces load on the primary but can add an extra hop of lag.
Oplog Window — the Most Overlooked Tuning Knob
The oplog is a capped collection — fixed size, oldest entries roll off when it fills. The oplog window is how far back in time those entries go. If a secondary falls behind further than the window, it cannot catch up by replaying entries — there are none left. It must do a full initial sync (copy all data from scratch).
// Check the oplog window on the primary
rs.printReplicationInfo()
// configured oplog size: 5120 MB
// log length start to end: 86723 secs (24.09 hrs) ← this is your window
// oplog first event time: 2026-08-10 10:00:00 UTC
// oplog last event time: 2026-08-11 10:05:00 UTC
// Check lag per secondary
rs.printSecondaryReplicationInfo()
// source: 10.0.0.2:27017
// syncedTo: 2026-08-11 10:04:55 UTC
// 0 secs (0 hrs) behind the primary ← healthy
# mongod.conf (primary) — tune oplog size to cover your maintenance window
replication:
replSetName: "rs0"
oplogSizeMB: 10240 # 10 GB — covers ~48h of lag for moderate write workloads
# Rule: oplog window should be >= longest expected secondary downtime
# (planned maintenance, patch reboots, etc.)
⚠️ You cannot shrink the oplog after it's created without a full resync of each secondary. Size it generously from the start. A 5–10 GB oplog is cheap on modern disks; an unplanned initial sync on a 500 GB dataset is hours of downtime.
The secondary was down for planned maintenance for 30 hours. The oplog window on the primary is 24 hours. What happens when the secondary restarts?
stateStr: RECOVERING, log: "too stale to catch up") and forces a full initial sync: it wipes its own data and copies everything fresh from the primary. Depending on dataset size, this can take hours. This is why oplogSizeMB should be sized to cover your longest expected downtime, not just current write rate.
9. Write Concern
Write concern controls how many replica set members must acknowledge a write before MongoDB returns success to the client. Getting this wrong is the most common cause of data loss after a failover.
// Set default write concern on the replica set (MongoDB 5.0+)
// Run on primary as admin
db.adminCommand({
setDefaultRWConcern: 1,
defaultWriteConcern: {
w: "majority", // majority of voting members must acknowledge
j: true, // writes must be written to journal (on-disk) — not just memory
wtimeout: 5000 // fail if not confirmed in 5 seconds (prevents infinite hangs)
}
})
// Per-operation write concern (application level — overrides default)
db.orders.insertOne(
{ item: "widget", qty: 100 },
{ writeConcern: { w: "majority", j: true, wtimeout: 5000 } }
)
How w: "majority" works with an arbiter — the PSA trap:
w: "majority" counts data-bearing members only. The arbiter never counts for write concern, regardless of its voting weight in elections.
PSA set — 3 voting members, but only 2 data-bearing:
For ELECTIONS (voting majority):
primary (1 vote) + secondary (1 vote) + arbiter (1 vote) = 3 votes total
majority = 2 → any 2 members can elect a primary
For w: "majority" (write concern majority):
data-bearing members: primary + secondary = 2 total
majority of data-bearing = 2 → BOTH must acknowledge the write
arbiter: invisible to write concern, never counted
Secondary DOWN, arbiter UP:
Only 1 data-bearing member can acknowledge (just the primary)
1 of 2 is NOT a majority → w:"majority" writes BLOCK until wtimeout
This is the PSA trap — w: "majority" effectively becomes "primary AND secondary must confirm" in a 3-member PSA set. If the secondary goes down (maintenance, crash, lag), your writes will block and timeout until it recovers.
w: "majority" is satisfied (2 of 2 data-bearing). Arbiter is doing heartbeats in the background — invisible to write flow. Replication lag is low; the secondary is tailing the oplog.
w: "majority" requires 2 of 2 data-bearing — it can't be met. Writes block until wtimeout (e.g. 5 seconds), then fail with WriteConcernFailed. The cluster is still up and the primary is still primary (arbiter + primary = election quorum), but writes with w: majority are unavailable. Use w: 1 to allow writes to continue with reduced durability, or restore the secondary.
w: "majority" of 1 data-bearing member = 1 of 1, so the new primary alone can satisfy it. Writes resume after ~10s election window.
PSA recommendation: Use
w: "majority"withwtimeoutset so writes fail fast rather than block forever. Have your application handleWriteConcernFailedby falling back tow: 1during secondary downtime if business requirements allow it, or switch to a PSS topology if you need guaranteedw: "majority"availability even when one node is down.
With w: "majority" in a PSA set, the secondary goes down for a 2-hour OS patch. What happens to writes during those 2 hours?
WriteConcernFailed after wtimeout expires. The arbiter does NOT count for write concern — it's 1 of 2 data-bearing nodes, not 2 of 2. The cluster itself stays up (election quorum = primary + arbiter = 2 votes), but w: "majority" is not satisfiable until the secondary returns. Options during the maintenance window: temporarily lower the app's write concern to w: 1, or use a PSS topology (two full secondaries) so one going down still leaves primary + one secondary = 2 of 3 data-bearing = majority.
10. mongodb-exporter (Prometheus)
The Percona mongodb_exporter exposes MongoDB's internal metrics in Prometheus format. Run it on the primary VM alongside mongod.
Create the monitoring user
// Connect to primary as admin
mongosh --host 10.0.0.1 -u admin -p 'ch@ngeM3!' --authenticationDatabase admin
use admin
db.createUser({
user: "mongodb_exporter",
pwd: "Exporter$ecret!",
roles: [
{ role: "clusterMonitor", db: "admin" }, // rs.status(), serverStatus(), etc.
{ role: "read", db: "local" }, // oplog stats
{ role: "read", db: "admin" }, // profile collection
{ role: "readAnyDatabase", db: "admin" } // collection stats across DBs
]
})
Install the correct binary
# RHEL/Rocky/CentOS — Percona repo
yum install -y https://repo.percona.com/yum/percona-release-latest.noarch.rpm
percona-release enable-only tools
yum install -y mongodb_exporter
# Ubuntu/Debian
wget https://repo.percona.com/apt/percona-release_latest.generic_all.deb
dpkg -i percona-release_latest.generic_all.deb
percona-release enable-only tools
apt-get update && apt-get install -y mongodb-exporter
# Confirm installed binary
which mongodb_exporter
mongodb_exporter --version
</div>
<div class="tab-panel" data-tab-panel="binary">
# Get the latest release from GitHub
# https://github.com/percona/mongodb_exporter/releases
VERSION="0.40.0"
ARCH="linux-amd64"
wget "https://github.com/percona/mongodb_exporter/releases/download/v${VERSION}/mongodb_exporter-${VERSION}.${ARCH}.tar.gz"
tar -xzf "mongodb_exporter-${VERSION}.${ARCH}.tar.gz"
mv mongodb_exporter /usr/local/bin/mongodb_exporter
chmod +x /usr/local/bin/mongodb_exporter
# Verify
mongodb_exporter --version
Use the Percona
mongodb_exporter, not the oldprometheus-community/mongodb_exporter— the Percona fork is actively maintained and supports MongoDB 5.0+. The old one has known metric gaps and is no longer updated.
</div>
Connection string
# Format:
# mongodb://<username>:<password>@<PRIMARY_IP>:<PORT>/admin?authSource=admin
MONGODB_URI="mongodb://mongodb_exporter:Exporter%24ecret%21@10.0.0.1:27017/admin?authSource=admin"
# URL-encode special chars in the password: ! → %21, $ → %24, @ → %40
Run as a systemd service
# /etc/systemd/system/mongodb_exporter.service
[Unit]
Description=Percona MongoDB Exporter
After=network.target mongod.service
[Service]
User=mongod
Group=mongod
ExecStart=/usr/local/bin/mongodb_exporter \
--mongodb.uri="mongodb://mongodb_exporter:Exporter%24ecret%21@10.0.0.1:27017/admin?authSource=admin" \
--web.listen-address=":9216" \
--collect-all \
--compatible-mode
Restart=on-failure
RestartSec=5s
[Install]
WantedBy=multi-user.target
systemctl daemon-reload
systemctl enable --now mongodb_exporter
# Verify metrics are being served
curl -s http://localhost:9216/metrics | grep -E "^mongodb_up|^mongodb_rs_members"
# mongodb_up 1
# mongodb_rs_members_state{member="10.0.0.1:27017",state="PRIMARY"} 1
# mongodb_rs_members_state{member="10.0.0.2:27017",state="SECONDARY"} 2
# mongodb_rs_members_state{member="10.0.0.3:27017",state="ARBITER"} 7
The exporter fails to connect and logs Authentication failed. You've triple-checked the password. What's the most overlooked cause?
Exporter$ecret! contains $ (must become %24) and ! (must become %21). The raw password passed in a URI string is parsed as URL path components — an un-encoded $ or @ will be misinterpreted and the credentials will arrive at mongod garbled. Always URL-encode the username and password segments of a MongoDB connection string.
11. Prometheus Scrape Config
# /etc/prometheus/prometheus.yml (add this job)
scrape_configs:
- job_name: "mongodb"
static_configs:
- targets: ["10.0.0.1:9216"] # primary's exporter
relabel_configs:
- source_labels: [__address__]
target_label: instance
Key metrics to alert on:
| Metric | Alert condition | Meaning |
|---|---|---|
mongodb_up |
== 0 |
Exporter cannot reach mongod |
mongodb_rs_members_state |
any member != 1 (PRIMARY) or != 2 (SECONDARY) |
Unexpected state change |
mongodb_mongod_op_latencies_latency_total |
rate spike | High read/write latency |
mongodb_mongod_wiredtiger_cache_bytes{type="currently in cache"} |
> cacheSizeGB × 0.9 |
Cache near saturation |
mongodb_mongod_replset_member_replication_lag |
> 30s |
Secondary falling behind |
12. Operational Runbook
Health checks at a glance
// ── Connect (always specify replica set name for driver-aware routing)
mongosh "mongodb://admin:ch@ngeM3!@10.0.0.1:27017,10.0.0.2:27017,10.0.0.3:27017/admin?replicaSet=rs0&authSource=admin"
// ── Replica set status
rs.status() // full picture: states, health, lag, oplog
rs.isMaster() // quick: who is primary, what's the set name
rs.printReplicationInfo() // primary: oplog size & coverage window
rs.printSecondaryReplicationInfo() // lag per secondary
// ── Server stats
db.serverStatus().connections // current, available, totalCreated
db.serverStatus().wiredTiger.cache // cache hit ratio, dirty bytes
db.serverStatus().repl // replication info
Common issues and fixes
rs.status() shows all members in SECONDARY or UNKNOWN, no primary.Cause: Fewer than majority of votes are reachable (e.g. primary + secondary both down).
Fix: Restore at least 2 of the 3 VMs. If the arbiter is up and one data-bearing node is up, they form a majority — an election will run. Never force a single node into primary with
rs.reconfigForce() unless you understand rollback risk.
rs.printSecondaryReplicationInfo() shows a large lag; secondary optimeDate is far behind primary.Causes: Disk I/O saturation on secondary, network bandwidth, secondary CPU, or oplog too small (secondary needed data that already rolled off the oplog — forces full resync).
Fix: Check
iostat -x 1 on secondary. Increase oplogSizeMB on primary if the window is too narrow. For severe lag, trigger a manual resync: stop mongod on secondary, wipe dbPath, restart — it will perform initial sync again.
STARTUP2 for an unusually long time after being added.What's happening:
STARTUP2 is the initial sync state — the secondary is copying the primary's data. This is expected and can take a long time for large datasets. Check progress with rs.status() — look at initialSyncStatus.If it's stuck: Check
/var/log/mongodb/mongod.log on the secondary for errors. A keyfile mismatch or network interruption mid-sync can cause it to fail silently and retry indefinitely.
What happened: The old primary had writes that were never replicated before it failed. The new primary's oplog diverged. When the old primary rejoins, those un-replicated writes are rolled back — moved to
/var/lib/mongo/rollback/ as BSON files.Fix: Examine the rollback files with
bsondump. Decide if those writes need to be replayed manually. This is a data recovery operation — prevent it by using w: "majority" write concern.
You decide to step down the primary manually for maintenance. You run rs.stepDown(). What happens to in-flight writes during the ~10-second election window?
readPreference: primaryPreferred will fall back to a secondary.
13. Setup Checklist
nc -zv <peer_ip> 27017 from each node. Primary allows inbound 9216 from Prometheus.
openssl rand -base64 756, mode 400, owner mongod. Identical copy deployed to secondary and arbiter with the same permissions.
replSetName (case-sensitive), correct cacheSizeGB, keyfile path correct, bindIp: 0.0.0.0 (or specific IPs). Mongod started and enabled with systemctl.
rs.initiate() on primary only, with priority: 2 on the primary member. Secondary added with rs.add(). Arbiter added with rs.addArb().
health: 1. Primary is PRIMARY, secondary is SECONDARY (not STARTUP2), arbiter is ARBITER. No lastHeartbeatMessage errors.
db.getUsers(). Security block re-enabled in mongod.conf. Mongod restarted. Auth tested.
w: "majority", j: true, wtimeout: 5000 via setDefaultRWConcern. Application connection string includes replicaSet=rs0.
:9216/metrics, mongodb_up 1 confirmed. Prometheus scrape configured. Key alerts in place.