This is the end-to-end operator runbook: install, run, validate,
reload, observe, and upgrade. It's intentionally smaller than the M-
series interop matrix in INTEROP.md — that document
proves rustbgpd interops cleanly with FRR / BIRD / GoBGP, this one gets
your daemon onto your box.
For deeper references:
CONFIGURATION.md— every config field, with examples.reload-matrix.md— which fields hot-apply vs. need a restart.OPERATIONS.md— debugging, log filtering, failure modes.SECURITY.md— firewalling, gRPC mTLS, authorization tiers.
rustbgpd is public alpha overall. The exception is the narrow,
machine-pinned v1 route-server / route-reflector contract,
which covers only its listed config fields, native gRPC signatures, CLI/JSON
contracts, and control-plane roles. Unlisted TOML and API surfaces can still
change between minor versions. See CHANGELOG.md for
migration notes and run rustbgpd --check <new config> against the new binary
before swapping it in.
It runs on Linux. Other platforms are not tested.
Suitable for: lab pilots, data-center fabric pilots, IX route-server pilots, automation-heavy control-plane work where the API surface evolution is acceptable. Not yet suitable for fully unattended production deployments where the operator can't react to a CHANGELOG note.
Tagged releases publish rustbgpd-linux-amd64.tar.gz and
rustbgpd-linux-arm64.tar.gz under
GitHub Releases. Each
ships rustbgpd (the daemon), rbgp (the CLI), and
rs-config-render (the route-server config renderer), plus man pages
and shell completions under share/:
rustbgpd, rbgp, rs-config-render binaries
LICENSE-MIT, LICENSE-APACHE licenses
rustbgpd.schema.json config JSON Schema
share/man/man1/rbgp.1 CLI man page
share/man/man8/rustbgpd.8 daemon man page
share/completions/rbgp.{bash,zsh,fish} shell completions
The filename is the same on every release; releases/latest/download/
always resolves to the current tag, so this snippet never needs a
version bump.
# Pick the right arch.
SUFFIX=linux-amd64 # or linux-arm64
TARBALL=rustbgpd-${SUFFIX}.tar.gz
curl -fL -o "$TARBALL" \
"https://github.com/lance0/rustbgpd/releases/latest/download/${TARBALL}"
tar -xzf "$TARBALL"
sudo install -m 0755 rustbgpd rbgp rs-config-render /usr/local/bin/
sudo install -m 0644 share/man/man1/rbgp.1 /usr/local/share/man/man1/
sudo install -m 0644 share/man/man8/rustbgpd.8 /usr/local/share/man/man8/
sudo install -m 0644 share/completions/rbgp.bash \
/usr/share/bash-completion/completions/rbgpThe man pages and completions are also generated on demand by the
binaries themselves (rbgp man, rustbgpd --man,
rbgp completions bash|zsh|fish), so an installed binary can always
regenerate them.
To pin to a specific tag for reproducibility, swap latest for the
version, e.g. releases/download/v0.45.0/${TARBALL}. SHA-256
checksums are published alongside each tarball as
checksums-${SUFFIX}.txt.
Verify:
rustbgpd --version
rbgp --version
rs-config-render --versionrs-config-render matters only if you run an IXP route server — it turns
arouteserver template-context output into rustbgpd configuration
(tools/rs-config-render/). It is
in the archive either way, so install it now rather than discovering it is
missing from a cron refresh later.
# Prerequisites: Rust ≥ 1.95, protobuf-compiler
sudo apt-get install -y protobuf-compiler # Debian/Ubuntu
git clone https://github.com/lance0/rustbgpd
cd rustbgpd
# Builds exactly what a deployment ships: the daemon, the rbgp CLI, and
# the route-server config renderer. (`--workspace` would also build
# dev/bench helpers you don't need.)
cargo build --release -p rustbgpd -p rustbgpctl -p rs-config-render
sudo install -m 0755 \
target/release/rustbgpd target/release/rbgp target/release/rs-config-render \
/usr/local/bin/A container image is built on every tagged release and published to
GHCR. Three tag flavors are available per the
docker/metadata-action rules in
.github/workflows/container.yml:
| Tag | Resolves to | Updates on |
|---|---|---|
:X.Y.Z |
exact version | nothing (immutable) |
:X.Y |
latest patch in the X.Y minor | each X.Y.z release |
:latest |
latest non-prerelease release | each minor or patch release |
Major-minor is the usual operator default — auto-receives bug-fix
releases but pins against minor-version churn. The examples in this
document use :latest so they stay copy-pasteable across releases;
substitute the major-minor tag of the series you standardize on:
docker pull ghcr.io/lance0/rustbgpd:latestThe published image is the Dockerfile's default runtime target:
lean, daemon + rbgp only, running as a nonroot rustbgpd user. No
dev/test/bench helpers are included. rs-config-render is deliberately
left out too: the IXP refresh loop is a host-side cron job that runs
arouteserver and the renderer next to each other and hands the daemon a
finished config directory, so the renderer belongs on the host, not in the
daemon's image. Install it from the tarball. If you'd rather build the
image locally:
docker build -t rustbgpd:latest .The development image — adds evpn-tester / evpn-monitor,
iproute2, and the interop start script, and runs as root for lab
plumbing — is the dev target. It's what the M-series interop suite
and the soak harnesses run against:
docker build --target dev -t rustbgpd:dev .A hardened unit lives at
examples/systemd/rustbgpd.service.
It runs the daemon unprivileged by default: a dedicated rustbgpd
system user, the standard sandbox set (NoNewPrivileges,
ProtectSystem=strict, ProtectHome, PrivateTmp, StateDirectory /
RuntimeDirectory), and a capability set of exactly
CAP_NET_BIND_SERVICE (bind port 179 as non-root). That is everything
a route-reflector / control-plane-only deployment needs.
Notes on the sandbox:
CAP_NET_BIND_SERVICElets the daemon bind port 179 without running as root. Nothing else is granted by default — in particular notCAP_NET_ADMIN, so the unprivileged unit cannot program the kernel.ProtectSystem=strict+ explicitReadWritePathsconfines the daemon to/var/lib/rustbgpd(state) and/etc/rustbgpd(config).ExecReload=kill -HUPis the supported reload path. See the reload matrix for which fields hot-apply vs. need a restart.Restart=on-failureis load-bearing. The daemon exits0only on an operator-initiated shutdown (SIGINT/SIGTERM, theShutdownRPC) and1on a component failure it cannot recover from in place — the BGP listener failing to bind, or the gRPC server exiting unexpectedly. Neither listener is rebound without a restart, so the daemon exits rather than run on deaf; the supervisor's retry is the recovery path. WithRestartSec=5a transient bind failure clears on the next attempt, and a permanent one (port held by another speaker, missingCAP_NET_BIND_SERVICE) repeats the bind error in the journal on every attempt.rbgp doctorreports the same cause against the down daemon through itsbgp.listenercheck.
sudo useradd --system --home-dir /var/lib/rustbgpd \
--shell /usr/sbin/nologin rustbgpd
sudo install -m 0644 examples/systemd/rustbgpd.service \
/etc/systemd/system/rustbgpd.service
sudo install -d -m 0755 /etc/rustbgpd
sudo install -m 0640 examples/minimal/config.toml /etc/rustbgpd/config.toml
sudo chgrp rustbgpd /etc/rustbgpd/config.toml # readable by the service user
$EDITOR /etc/rustbgpd/config.toml # set ASN, router_id, neighbors
sudo systemctl daemon-reload
sudo systemctl enable --now rustbgpdIf — and only if — the config programs the kernel ([[fib_tables]]
per ADR-0061, install_blackhole_discard, or the EVPN VTEP/IRB
dataplane), the daemon needs CAP_NET_ADMIN for rtnetlink route /
nexthop programming, VXLAN + bridge FDB writes, and the
RTNLGRP_NEIGH subscription behind local-MAC learning. Grant it via
the shipped drop-in
examples/systemd/rustbgpd-dataplane.conf
instead of editing the base unit — systemd merges capability sets, so
the drop-in adds CAP_NET_ADMIN on top of the unprivileged default:
sudo mkdir -p /etc/systemd/system/rustbgpd.service.d
sudo install -m 0644 examples/systemd/rustbgpd-dataplane.conf \
/etc/systemd/system/rustbgpd.service.d/rustbgpd-dataplane.conf
sudo systemctl daemon-reload
sudo systemctl restart rustbgpdRR-only boxes should not install this drop-in — the unprivileged default is the point.
Dataplane host caveat: if rustbgpd programs the kernel on a host
that also runs systemd-networkd, set ManageForeignRoutes=no,
ManageForeignRoutingPolicyRules=no, and (where supported)
ManageForeignNextHops=no in networkd.conf. The defaults are
yes, and networkd will delete routes and next-hop groups it did
not create — including rustbgpd's, especially across a daemon
restart (ADR-0079). Co-residency with another daemon that claims
proto bgp kernel state (e.g. FRR zebra) is unsupported for the
same reason.
When the EVPN VTEP/IRB dataplane is enabled, rustbgpd exclusively
owns four kernel nexthop-ID ranges, distinguished by the ID's high
nibble (deliberately offset from FRR's 0x1000/0x2000 tags so both
daemons can't collide on NLM_F_REPLACE):
| Range | Use |
|---|---|
0x3000_0000–0x3FFF_FFFF |
L2 per-VTEP FDB nexthop members (ADR-0059 aliasing ECMP) |
0x4000_0000–0x4FFF_FFFF |
L2 FDB nexthop groups (ADR-0059) |
0x5000_0000–0x5FFF_FFFF |
L3VXLAN per-VTEP nexthop members (all-active Type 5) |
0x6000_0000–0x6FFF_FFFF |
L3VXLAN FDB nexthop groups (all-active Type 5) |
This is a deployment contract: no co-resident netlink writer may
create nexthop objects with IDs in these ranges (ip nexthop add id …, another routing daemon, orchestration tooling). rustbgpd
adopts objects it finds there across restarts and garbage-collects
the ones nothing references — an unrelated agent's object parked in
the range would be treated as rustbgpd state.
rustbgpd validates the shape of every in-range object it dumps at
startup adoption and during drift recovery (FDB flag, member vs
group structure matching the tag). An in-range object that does not
look like anything rustbgpd writes fails closed per ID: the object is
left untouched, excluded from adoption and from every reap/delete
path, its ID is quarantined so the allocator can never hand it out
(and thus never NLM_F_REPLACE over it), and the violation is
logged once and counted on the
evpn_foreign_nhid_range_conflicts_total Prometheus counter. A
nonzero counter means a co-resident writer is violating this
contract — find and remove that writer; the quarantined IDs are not
reclaimed until a daemon restart after the foreign objects are gone.
Note this shape check cannot distinguish a foreign object that is
byte-for-byte identical to what rustbgpd writes; stamping
nh_protocol on owned nexthop objects as a provenance marker is a
possible future defense-in-depth, not a substitute for this
contract (any netlink writer can set the same protocol value — it is
provenance, not kernel-enforced exclusivity).
systemctl status rustbgpd # is it running?
journalctl -u rustbgpd -f # follow logs
systemctl reload rustbgpd # SIGHUP → config reload
systemctl restart rustbgpd # full restart (state drained)A worked starting point at
examples/docker-compose/
brings up rustbgpd alongside an FRR peer that advertises sample
prefixes — no real routers required.
cd examples/docker-compose
docker compose up -d --build
docker compose exec rustbgpd \
rbgp -s http://127.0.0.1:50051 neighbor
docker compose downFor your own deployment:
-
State: mount a volume at the daemon's
runtime_state_dirso the GR restart marker and FIB ownership receipt survive container restarts. The daemon runs as uid/gid 999 and rewrites its config in place, so both host directories must be owned by that uid and the config must already exist — the image's default command isrustbgpd /etc/rustbgpd/config.tomland it will not create one:# One-time host prep. sudo install -d -o 999 -g 999 /etc/rustbgpd /var/lib/rustbgpd docker run --rm ghcr.io/lance0/rustbgpd:latest \ rustbgpd --init-config edge --stdout > config.toml # Edit config.toml for your ASN, router ID, peers, and policy. sudo install -o 999 -g 999 -m 0600 config.toml /etc/rustbgpd/config.toml docker run --rm -d \ --name rustbgpd \ -v /etc/rustbgpd:/etc/rustbgpd \ -v /var/lib/rustbgpd:/var/lib/rustbgpd \ -p 179:179 \ -p 9179:9179 \ --ulimit nofile=65536:524288 \ --cap-add=NET_BIND_SERVICE \ --cap-add=NET_ADMIN \ ghcr.io/lance0/rustbgpd:latest
The
edgeprofile is used because itsruntime_state_dir,grpc_udspath, and0.0.0.0metrics bind are already the container forms; thelabprofile's/tmpstate directory and127.0.0.1metrics bind both need editing first. -
File descriptors:
--ulimit nofileis required, not tuning. The Docker default soft limit is 1024, well under the 4096 floorrbgp doctorenforces, and peers exhaust descriptors at scale. -
Writable config directory: mount
/etc/rustbgpdread-write and make it owned by the daemon user (chownthe host directory to the container'srustbgpduid) if you use runtime mutation —rbgp neighbor add, policy edits, gNMISet,rbgp config apply. Config persistence rewrites the file with a temp-file + rename, so the directory, not just the file, has to be writable, and every mutating RPC is rejected without it. Mount it:roonly when the config is managed entirely from outside and reloaded with SIGHUP. -
Logs: structured JSON when
[global.telemetry] log_format = "json"is set; pipe to your log aggregator. -
Networking: Linux FIB integration and BFD require the container to share the host network namespace (or otherwise have access to the routing table you intend to program). The interop suite runs rustbgpd as a containerlab
kind: linuxnode, which is the cleanest reference setup.
Containerlab is the easiest way to get a
working rustbgpd ↔ FRR session on your laptop with no real routers.
The simplest topology in the repo is
tests/interop/m0-frr.clab.yml:
name: m0-frr
topology:
nodes:
rustbgpd:
kind: linux
image: rustbgpd:dev
cmd: sleep infinity
frr:
kind: linux
image: quay.io/frrouting/frr:10.3.1
links:
- endpoints: ["rustbgpd:eth1", "frr:eth1"]Bring it up:
# Build the dev image
docker build --target dev -t rustbgpd:dev .
# Deploy
sudo containerlab deploy -t tests/interop/m0-frr.clab.yml
# Inspect (the test driver scripts under tests/interop/scripts/ show
# the typical incantations for FRR vtysh + rbgp gRPC).
# Tear down
sudo containerlab destroy -t tests/interop/m0-frr.clab.yml --cleanupFor more topologies (route-reflector, EVPN, FlowSpec, BFD, etc.), see
INTEROP.md.
The smallest deployment that exercises the full operator loop — build → validate → reload → observe — without depending on more than one peer. This is your "did I install this correctly" gate before scaling out.
-
One rustbgpd instance + one FRR peer (or any compliant BGP peer). Provision the peer first; record its address + AS.
-
Prometheus scrape on
:9179. Add the scrape target in your Prometheus config:- job_name: rustbgpd static_configs: - targets: ["10.0.0.1:9179"]
-
Config validation. Build the config, then:
rustbgpd --check /etc/rustbgpd/config.toml
Errors are rustc-style with file + line + carets; expect
config OKon success. Run this every time before swapping the live config.A check that finds nothing to flag prints
config OK. A check that flags something still exits 0, but the summary readsconfig VALID, <n> WARNINGS — NOT a clean checkand the warnings are framed on stderr above it. The one warning today is an eBGP neighbor that resolves no explicit policy in a direction: unfiltered when[global] ebgp_requires_policyis off, and carrying no routes in that direction when it is on. Neither is rejected — a permit-all route server is a legitimate configuration — but neither should reach production unnoticed. -
First start.
sudo systemctl start rustbgpd journalctl -u rustbgpd -f
Watch for the session to reach
Established:rbgp neighbor
-
Edit + reload cycle. Edit
/etc/rustbgpd/config.toml, then dry-run the diff:rustbgpd --diff /etc/rustbgpd/config.toml
--diffcalls into the sameConfigDiffmachinery the reload path uses; what it reports is what reload will do. Cross-reference against the reload matrix to confirm which changes hot-apply and which are restart-required. Apply:sudo systemctl reload rustbgpd
-
Restart cycle. Confirm the FRR peer reconverges within the GR timer (GR is on by default):
sudo systemctl restart rustbgpd # FRR should retain the routes during the restart window; rustbgpd # advertises R=1 in the GR capability on the next OPEN.
-
Observability sanity-check. Read at least one Prometheus counter and one neighbor field:
curl -s http://10.0.0.1:9179/metrics \ | grep -E "bgp_session_established_total|bgp_messages_received_total" rbgp neighbor 10.0.0.2
If all six steps work end-to-end, your install is sound. Scale from there: add more peers, add policy chains, wire BMP / gNMI / MRT according to your needs.
These cover the config lifecycle, from bootstrap to reload:
| Command | What it does |
|---|---|
rustbgpd --init-config <lab|edge> --stdout |
Print a curated, commented starter TOML to stdout and exit (file output is not yet supported). lab is a minimal single-box profile; edge is an eBGP edge skeleton with a default-route-dropping import chain and a default-deny export chain to fill in. Both use a mode-0600 local UDS whose filesystem permissions authenticate access and whose stable operator principal is tier-authorized, set [global] ebgp_requires_policy = true, and pass --check --strict as emitted. Cannot be combined with --check / --diff. |
rustbgpd --check <file> |
Parse + validate; print config OK, config VALID, <n> WARNINGS — NOT a clean check (warnings framed on stderr; still exit 0), or a rustc-style diagnostic (exit 1). Does not start the daemon. |
rustbgpd --check --strict <file> |
The same check, but any warning exits 1 instead of 0 — for CI and deployment gates that must not accept a valid-but-risky config. A clean check still exits 0. --strict without --check is an error (exit 2). |
rustbgpd --diff <file> |
Compute the diff against the running daemon's view; print per-section change list with expected reload class. |
systemctl reload rustbgpd (or kill -HUP $(pidof rustbgpd)) |
Apply the diff. Live fields hot-apply; restart-required fields are pinned and logged at ERROR (the live values are kept). |
The validation pipeline is the same in all three places: TOML parse
→ validate() → ConfigDiff. A config that passes --check will
not error on --diff; a clean --diff will not error on reload.
The exporter binds at [global.telemetry] prometheus_addr. The same listener
serves /livez and /readyz for orchestrators. Key counters operators watch:
- Session liveness —
bgp_session_established_total,bgp_session_flaps_total,bgp_messages_received_total,bgp_messages_sent_total(all by peer). - Loop / leak detection —
bgp_as_path_loop_detected_total,bgp_rr_loop_detected_total,bgp_otc_routes_blocked_total{peer, reason}(RFC 9234 / ADR-0071),bgp_role_mismatch_total{peer, local_role, remote_role}. Canonicalreasonvalues are documented indocs/OPERATIONS.md("Ingress rejection / route-leak detection"). - Route processing —
bgp_rib_prefixes{peer, afi_safi},bgp_rib_loc_prefixes{afi_safi},bgp_max_prefix_exceeded_total. - GR / FIB —
bgp_gr_active_peers,bgp_gr_stale_routes,bgp_fib_routes_installed_total,bgp_fib_kernel_failures_total. - Allocator —
jemalloc_allocated_bytes,jemalloc_active_bytes,jemalloc_resident_bytes,jemalloc_mapped_bytes. Present in builds with thejemallocfeature (the published container image and release tarballs); refreshed at scrape time.allocatedis live application bytes,residentis jemalloc's contribution to RSS — a widening gap between the two is retained-but-unused allocator memory, the first thing to check before suspecting a leak. - Durable event outbox (ADR-0072) —
bgp_event_outbox_committed_total{category},bgp_event_outbox_dropped_total{category, reason},bgp_event_outbox_db_size_bytes,bgp_event_outbox_latest_event_id,bgp_event_outbox_degraded,bgp_event_outbox_cursor_gap_total. The degraded gauge is the alert-on-this signal — flips to1on a durability-impacting drop, decode/codec failure, or open failure since process start; does not auto-clear in v1, so any non-zero value warrants investigation. The cursor-gap counter alerts on collectors whose persistedlast_seen_event_idfell below the daemon's retention floor — typically a sign the collector was offline longer than[event_history].max_events/max_bytesplanned for. External collectors stream the cursor via theSubscribeFromEventgRPC RPC;examples/event-bridge/is the reference skeleton — copy and replace the stdout writer with your Kafka / NATS / Vector / journald sink, persistinglast_seen_event_idafter the sink confirms durable receipt. See[event_history]inCONFIGURATION.mdfor tuning and recovery semantics, and the "Durable Event Cursor" section inOPERATIONS.mdfor the alert + sizing playbook.
Policy filtering visibility — Prometheus. bgp_policy_routes_total {peer, policy, direction, action} attributes each import and export
policy evaluation to the terminal-decision policy in the chain with
policy="…" (the configured name) or policy="inline" for inline
statements and permit-all peers without an explicit chain. Initial
table dumps, route refreshes, dirty resyncs, and forced outbound
refreshes can increment export counters because they re-evaluate export
policy. The Prometheus counter is monotonic for the lifetime of the
daemon process — use Prometheus rate() / increase() to read it.
# Routes denied by a named filter on each peer's import side:
rate(bgp_policy_routes_total{direction="import", action="deny"}[5m])
Policy evaluation errors — Prometheus. bgp_policy_eval_errors_total {direction, kind} counts routes denied by the fail-closed evaluation-
error rail (ADR-0103 Decision 4: checked-arithmetic failure, absent
operand, fuel/loop-cap exhaustion, ...). These denies also appear in
bgp_policy_routes_total as action="deny"; this counter separates
"the policy said no" from "the policy is broken". Any nonzero rate
deserves a look — rbgp policy stats names the failing chain, policy,
and term (eval_errors count + last_error per chain), and the
rate-limited daemon WARN carries the same blame line.
# Any policy erroring anywhere is alert-worthy:
sum by (kind) (rate(bgp_policy_eval_errors_total[5m])) > 0
Policy filtering visibility — gRPC scalar aggregates. NeighborState
carries four per-peer running totals to give operators a cheap
sanity-check on the labelled Prometheus counter:
| Field | Direction | Scope |
|---|---|---|
import_policy_routes_permitted |
import | Per session — resets on session-down (lives on PeerSessionState). |
import_policy_routes_denied |
import | Per session — same. |
export_policy_routes_permitted |
export | Per RIB peer-attach — resets on handle_peer_down, i.e. session-down. |
export_policy_routes_denied |
export | Per RIB peer-attach — same. |
Both directions reset together on the next session establishment, so
"how many routes did this session permit / deny in total" is straight
subtraction; "how many across reconnects" requires Prometheus history.
The CLI surfaces these in rbgp neighbor <peer> as a Policy Stats
block:
Policy Stats:
Import — permitted: 1,247 denied: 31
Export — permitted: 892 denied: 0
JSON output (rbgp --json neighbor <peer>) carries the same
fields under import_policy_routes_permitted / import_policy_routes_denied
/ export_policy_routes_permitted / export_policy_routes_denied,
elided when zero.
If [bmp] is configured, rustbgpd opens a TCP session to the BMP
collector and exports per-peer Adj-RIB-In + Peer Up / Peer Down
events. See CONFIGURATION.md for the schema.
The gNMI adapter (ADR-0070) exposes Capabilities / Get / Subscribe
telemetry plus the static numbered-neighbor Set subset over the same
sockets as the native API. Native gNMI is registered on
[global.telemetry.grpc_tcp] only when native mTLS is configured; the
[global.telemetry.grpc_uds] listener serves gNMI unconditionally as a
local-only extension. RFC 7951 JSON encoding. See GNMI.md
for the path namespace and supported mutation leaves.
rbgp is the operator interface for read queries.
rbgp neighbor # list all neighbors
rbgp neighbor 10.0.0.2 # detail
rbgp neighbor 10.0.0.2 --compare 10.0.0.3 # live update-group relationship
rbgp rib # browse Loc-RIB
rbgp bfd # BFD sessions (ADR-0067)
rbgp evpn # EVPN instances + Type 2/3 RIB
rbgp top # live TUI dashboardMost data-oriented read commands support --json for scripting. Commands with
fixed formats, such as metrics, completions, and top, keep their
command-specific output.
For an Established neighbor, rbgp neighbor <addr> reports whether the peer
negotiated Route Refresh, Enhanced Route Refresh, and Extended Messages, plus
the directional maximum BGP message size rustbgpd may send (4096 or 65535
bytes). The JSON fields preserve false as an explicit negotiated result and
omit values an older daemon did not expose; a missing live-session snapshot is
reported separately by negotiation_available.
sudo systemctl stop rustbgpd
sudo install -m 0755 /tmp/rustbgpd-vX.Y.Z /usr/local/bin/rustbgpd
sudo systemctl start rustbgpdGR is on by default; after a coordinated shutdown rustbgpd advertises R=1
on the next OPEN. A GR-aware peer may retain eligible routes while sessions
rebuild, but rustbgpd advertises forwarding_preserved = false: this is not a
guarantee of forwarding continuity, and the peer may withdraw or replace routes
until normal convergence.
TOML format is not frozen between minor versions while rustbgpd is public alpha. Before upgrading across a minor version:
- Read the CHANGELOG entry for the target version. Breaking config
changes are called out under
### Changed/### Removed. - Build the new binary or pull the new container image.
- Run
rustbgpd --check /etc/rustbgpd/config.tomlagainst the new binary. Fix any errors before swapping. - Swap and start.
There is no in-place schema migration tool. If the CHANGELOG describes a field rename, edit the config file by hand.
Everything in runtime_state_dir (default /var/lib/rustbgpd):
Assign a distinct runtime_state_dir to every concurrently running rustbgpd
daemon. The directory is single-writer runtime state; sharing it between live
processes is unsupported even when their configuration files differ.
| File | Purpose | Survives restart |
|---|---|---|
gr-restart.toml |
Graceful Restart coordination marker. Written on clean shutdown, read on startup to set the R-bit in OPEN. | Yes |
warm-bundle-v1/ |
Optional owner-private shutdown checkpoint (manifest.json plus a content-addressed MRT artifact). Published only when warm_cache_checkpoint_on_shutdown = true; not restored on startup. |
Yes |
commit-confirm-journal.json |
Owner-private pre-transaction config snapshot for crash-safe confirmed commits; consumed by confirm/abort/timeout or boot revert. | Until the transaction is terminal |
config-history/*.toml |
Last 20 distinct applied configs for rbgp config history / rollback N; owner-private and secret-bearing. |
Yes |
fib-owned.json |
FIB ownership receipt — which kernel routes the daemon installed (ADR-0061). Used to drain orphan installs on next start. | Yes |
grpc.sock |
gRPC UDS endpoint (if [global.telemetry.grpc_uds] configured). |
Recreated on start |
Routing state is not restored. The optional shutdown checkpoint contains only eligible post-import-policy Adj-RIB-In views for future use; Loc-RIB, Adj-RIB-Out, and policy evaluation state are not checkpointed, and the current startup path loads none of it. Routing state rebuilds from peer routes after restart. GR and checkpoint state support bounded control-plane restart handling, but do not guarantee forwarding continuity; peers may withdraw or replace routes, and therefore forwarding, until normal convergence.
The config file itself is mutable across runs: neighbor add/delete
operations via gRPC persist back to the config file (see
CONFIGURATION.md → "Config Persistence").
The repo ships nine config profiles under
examples/ covering the standard deployment
shapes. Pick the closest match, copy, edit:
| Profile | File | Use case |
|---|---|---|
| Minimal | examples/minimal/config.toml |
Single eBGP peer; dev-friendly, state in /tmp. |
| IX route server | examples/route-server/config.toml |
RPKI, Add-Path, dual-stack, per-member policy chains. |
| Fabric edge / Linux FIB | examples/linux-edge-fib/config.toml |
FIB integration on configured unicast tables; ECMP, weighted multipath. |
| EVPN VTEP leaf | examples/evpn-vtep-leaf/config.toml |
Bidirectional VTEP: kernel FDB → Type 2 origination. |
| EVPN RR fabric | examples/rr-evpn-fabric/config.toml |
Route Reflector mode, stateless EVI (no [[evpn_instances]]). |
| Route collector | examples/route-collector/config.toml |
Passive listener, MRT dumps, BMP export. |
| DDoS mitigation | examples/ddos-mitigation/config.toml |
FlowSpec injection + RTBH (RFC 7999 BLACKHOLE). |
| Hosting provider | examples/hosting-provider/config.toml |
iBGP injector for customer-prefix automation. |
| Docker Compose | examples/docker-compose/ |
Quick-start with an FRR peer (gRPC on TCP). |
Each profile validates with rustbgpd --check out of the box; edit
the AS, addresses, and TLS material before deploying.
See SECURITY.md for the full posture document. The
short version for first deployment:
- Bind addresses.
prometheus_addr,grpc_tcp.address,grpc_uds.pathdefault to listening on what the config says. Don't expose the Prometheus or gRPC endpoint to untrusted networks without authentication. - gRPC. Use mTLS in production. Set
[security.grpc] enforcement = "tier"and map principals to roles under[security.grpc.roles](observer/automation/operator) perSECURITY.md; the per-method tier matrix itself is compiled into the daemon (crates/api/src/authz.rs), not configured per-tier in TOML. UDS with a restrictivemodeis fine for single-host operator access. - BGP authentication. TCP-MD5 (RFC 2385) and TCP-AO (RFC 5925) are both supported. TCP-AO is preferred; see ADR-0062.
- Firewall. Allow inbound TCP/179 only from configured peer
addresses (or address ranges if using
[[dynamic_neighbors]]). Block everything else.
| Symptom | Where to look |
|---|---|
| Session won't establish | rbgp neighbor <addr> + journalctl -u rustbgpd -p warning |
| Routes received but not installed | rbgp rib, then rbgp rib received <peer> --rejected; for the statement-level trace, rbgp policy explain --neighbor <peer> --prefix <CIDR> (opt-in: needs [policy.explain] enabled = true) |
| Reload didn't change behavior | rustbgpd --diff <file> + cross-reference reload matrix |
| FIB programming failures | bgp_fib_kernel_failures_total Prometheus counter + journalctl for kernel-dataplane lines |
| EVPN-specific issues | evpn-vtep-troubleshooting.md |
| Anything not above | OPERATIONS.md "Debugging" section |
For deeper investigation, raise the daemon log level globally via
[global.telemetry] log_format = "json" and per-peer via
[[neighbors]] log_level = "debug". Restart required for the global
setting; per-peer log_level is live — re-applied on SIGHUP via a
tracing reload handle that rebuilds the filter (base level plus every
per-peer directive), so a level edit takes effect without a restart (see
the reload matrix).
OPERATIONS.md— start/stop, reload, upgrade, debugging.CONFIGURATION.md— config reference.reload-matrix.md— per-field reload classification.SECURITY.md— security posture + firewall guidance.INTEROP.md— M-series interop matrix.adr/— architectural decisions per protocol and design choice.