An active call sends a DMRD/DMRE packet roughly every 90-100ms; with
OBP_PROXY.DEBUG on, that logged one RX line per packet, flooding the
log for the whole call. Log once when a stream_id first appears instead.
Non-voice frames (control) are unaffected, they were never the flood.
Review follow-up on the fan-in demux fixes.
1. apply_obp_proxy_config_reload() rebuilds the bridge registry, and a rebuild
creates fresh ObpIngressReplyTransport objects. Nothing re-pinned them to the
legacy sockets that stay bound, so the first `systemctl reload` put egress
back on the shared fan-in port and reintroduced the bug the pin fixes.
Bound legacy transports are now kept in the service state and re-pinned after
every registry rebuild (pin_legacy_egress()). For the same reason the rebuild
now describes the ports that are actually bound (the running LISTEN_PORT /
BIND_LEGACY_PORTS) instead of the incoming ones, which are only logged as
"restart required".
2. Control-frame demux ranked bridges by the live SYSTEMS.<name>["TARGET_SOCK"],
the very dict RELAX_CHECKS rewrites in place when a bridge accepts traffic
from an unexpected source. One misattributed frame could therefore move a
bridge's target and then keep matching it there. Each entry now carries
peer_hint, the peer as configured, snapshotted when the registry is built
(startup and every reload, both of which read a freshly normalized config).
Order is: configured IP:port, live IP:port, configured IP, live IP, then
registration order — so a peer that legitimately moved is still recognised,
but never at the expense of the bridge that has that address in the YAML.
Tests
- test_reload_keeps_egress_pinned_to_the_bridge_socket: starts the service on a
fake reactor, reloads twice, asserts egress still leaves through each bridge's
own socket and nothing goes out of the fan-in port.
- test_control_frame_prefers_configured_peer_over_relaxed_target: a bridge
dragged onto another bridge's address does not steal that peer's keepalives.
- test_control_frame_unknown_peer_falls_back_to_passphrase now asserts the
deterministic outcome (first registered bridge) instead of "either one".
Both new tests fail without this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three issues seen on a production ADN Systems master (2.5.6) with five
OpenBridge links, all rooted in the shared fan-in port.
1. YamlConfigLoader rebuilt the config from a fixed list of top-level keys
that did not include OBP_PROXY, so the whole block in adn-server.yaml was
silently ignored: ENABLED, LISTEN_PORT, LISTEN_IP and BIND_LEGACY_PORTS
always fell back to their defaults (--doctor reported 62032 for a config
that asked for 62031, and ENABLED: false still bound the fan-in).
2. Control frames (BCKA/BCSQ/BCST/BCVE) carry no NETWORK_ID, so the fan-in
could only tell bridges apart by passphrase. Every ADN Systems bridge
shares one passphrase, so all of them resolved to whichever bridge was
registered first: keepalives from the Spanish and Portuguese peers were
attributed to the French bridge and, with RELAX_CHECKS, rewrote its
TARGET_SOCK ("Source IP and Port has changed ... updating", in a loop).
Voice was unaffected because DMRD/DMRE do carry NETWORK_ID.
Control frames now prefer the bridge whose configured peer matches the
datagram source (exact IP:port, then same IP, then the passphrase scan).
3. ObpIngressReplyTransport fell back to the shared fan-in socket, so egress
left from LISTEN_PORT; remote peers learn that source and answer there,
funnelling everything onto the one port where control frames cannot be
told apart. Egress is now pinned to the bridge's own legacy socket when
one is bound.
Tests: five regression tests; all five fail without these changes.