#82 logged the RX debug line once per stream_id per bridge, but a bridge
with BOTH_SLOTS/several concurrent talkgroups interleaves packets from
multiple calls, so the "last seen" stream flips almost every packet and
the line logs nearly as often as before the fix — confirmed against
production output (three interleaved streams on OBP-CL).
DMRD/DMRE demux by NETWORK_ID, which is never ambiguous, and
*CALL START*/*CALL END* (application/routing_use_cases.py,
application/routing/obp_forward.py) already log once per real call with
better detail (SUB, TGID, TS, duration). Drop the fan-in RX line for voice
frames entirely instead of trying to approximate "once per call" with no
session state; control frames (BCKA/BCSQ/BCST/BCVE) keep logging every
time, since they are rare and that visibility mattered for the #79 fix.
An active call sends a DMRD/DMRE packet roughly every 90-100ms; with
OBP_PROXY.DEBUG on, that logged one RX line per packet, flooding the
log for the whole call. Log once when a stream_id first appears instead.
Non-voice frames (control) are unaffected, they were never the flood.
Validated against the production ADN 213 master: the calls it logged appear in
the replay, on the right bridge, with the right subscriber and talkgroup. Three
things that only show up against real traffic are now in the page, in both
languages.
Capture a port range, not a list: a peer answering from an unexpected port is
the case worth looking at and a narrow filter hides it.
A call legitimately arrives on several bridges at once on a mesh — the master
logs one because loop control, later in routing, keeps one leg and drops the
rest. The replay stops at ingress, so it shows all of them.
And `outbound` frames (what this server sent) are reported, not judged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two things a real capture showed that synthetic frames could not. An
unfiltered tcpdump carries both directions, and our own egress was being run
through the ingress rules, where it fails the NETWORK_ID check by definition:
10698 frames on one bridge reported as network-id-mismatch that were simply
ours. Direction is now decided by the peer address (both ends of a link
normally share the port number, so the port alone cannot tell), with the
listening port as a fallback for a peer behind NAT; --replay-both-directions
keeps the old behaviour.
And VALIDATE_SERVER_IDS was firing offline against an empty list, because the
server-id table is loaded at runtime and not from the YAML: 10107 frames on
another bridge blamed on source-server-unknown. Offline the check is skipped
unless the table is actually there.
Both were reported against the live ADN 213 master (5 bridges, 30669 real
datagrams in 120s); with them fixed the report says what the master does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The engine from phase 3 takes plain values and answers with effects, so it does
not need the server to run. ``adn-server --replay capture.pcap`` uses that: the
frames from a tcpdump capture go through the same ingress, against the
operator's own adn-server.yaml, and each one comes back with the bridge it
belongs to and either a delivery or the reason it was refused.
12:04:31 82.65.127.86:62201 OBP-FR DMRD v1 2130001 -> 214 delivered
12:04:31 85.241.222.7:62268 OBP-PT DMRE v5 2680015 -> 9 dropped (tg-filter-server) +BCSQ
12:04:32 203.0.113.9:50000 - DMRD v1 unmatched
Nothing is sent and no port is bound, so it runs beside a live master. Which
bridge a frame belongs to is decided on the evidence the server has — the port
it arrived on, then the configured peer, then whoever can verify it — which is
also the answer to "whose keepalive is this?" on a mesh where every bridge
shares one passphrase.
``--system`` narrows it to one link, ``--replay-limit`` stops early and
``--replay-summary`` prints the tally alone.
``infrastructure/pcap.py`` reads classic pcap with no dependencies: both
endiannesses, microsecond and nanosecond timestamps, Ethernet (VLAN tags
included), Linux cooked v1 and v2 (``tcpdump -i any`` writes SLL2, found while
running this against a real capture), raw IP and loopback, IPv4 and IPv6 UDP.
pcapng says which command converts it.
Docs: the OBP proxy page gains a "why did that call not cross" section in both
languages, and its RELAX_CHECKS note now says what phase 2 made true — what the
wire teaches lives in the session, TARGET_IP stays as written.
Tests: 27 new, 97% of the replay module and 87% of the pcap reader, plus an
end-to-end run of the real command against a real YAML. Full suite 1025 passed,
2 skipped (the 2 failures are this machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 3, and the end of the hblink shape in this path. Phase 1 lifted the
admission rules out of the adapter, phase 2 took the state they read; what was
left in ``_obp_datagram_received`` was the plumbing between them — two long
branches that decoded, decided, logged, quenched and called into routing, all
inside a Twisted ``DatagramProtocol`` where none of it could be run on its own.
``domain/mesh_engine.py`` now takes a decoded frame, the link's session and a
``BridgePolicy``, and answers with a list of effects: ``Reject``, ``Log``,
``NoteStream``, ``StoreTalkerAlias``, ``Deliver``, ``RequestVersion``. It reads
no configuration, opens no socket and calls no logger. The adapter keeps the
three things that are genuinely I/O — verify the MAC, build the policy, carry
out the effects in order — and the handler goes from ~180 lines of nested
branches to ~45 of dispatch.
Two things this buys beyond the shape. A datagram can be replayed through the
engine anywhere: a test, a laptop, a capture from a sysop, with no reactor in
sight. And every refusal carries a ``reason``, now tallied per bridge in the
session and printed by the keepalive loop at debug level — "my call does not
cross" is answered by ``tg-filter-mcc=12, system-sub-acl=3`` instead of by
grepping the log.
No behaviour change intended, and this time checked two ways. The differential
harness ran 9594 frames through this tree and through upstream ``develop``:
identical delivery, quench, egress address and log lines. And the corpus in
``tests/fixtures/obp_ingress_effects.jsonl`` — 96 recorded cases covering the
talkgroup filters, both ACL scopes, the bits byte, the DMRE envelope (age,
hops, source server) and BCKA/BCSQ/BCST from three addresses — was recorded
from the pre-engine handler, verified frame by frame against develop, and still
passes untouched. It stays as the contract for whatever comes next; regenerate
with CAPTURE=1 and read the diff.
Tests: 24 new unit tests for the engine (100% of the module), the recorded
corpus as a regression net, 998 passed, 2 skipped (the 2 failures are this
machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 2 of separating the OpenBridge path from its hblink ancestry. Phase 1
lifted the admission rules out of the adapter; this one takes the state they
were reading.
Legacy hblink kept what a bridge learns at runtime inside its own SYSTEMS
block: ``_bcka`` (last keepalive), ``_bcsq`` (the peer's quench table),
``_STUN`` and, worst of the four, ``TARGET_IP``/``TARGET_PORT``/``TARGET_SOCK``,
rewritten in place every time RELAX_CHECKS accepted a datagram from an address
the operator never wrote. Configuration and session state shared one mutable
dict, so a single unexpected packet could move a bridge's target for good, and
no reader could tell what came from the YAML and what came from the wire.
``domain/mesh_session.py`` now holds an ``ObpBridgeSession`` per link:
configured_peer (from the YAML, never moves) beside learned_peer (from the
wire), the last keepalive, the quench table and the BCST stun flag, each behind
the question a caller actually asks — ``peer``, ``keepalive_seen``,
``keepalive_stale(now)``, ``quenches(tg, stream)``. The store lives under a
private top-level config key next to ``_SUB_MAP`` and ``_PEER_IDS``, so every
layer that already receives the config reaches the same instance; like
``_SUB_MAP`` it is shared, not deep-copied, across a SIGHUP, and
``sync(config)`` then refreshes the configured peers and drops the sessions of
links that are gone. Editing TARGET_IP in the YAML and reloading is now a
documented way to undo a bad learned address.
Migrated readers: the keepalive gate in routing (to_target and unit data), the
60s keepalive report loop, the monitor/MQTT dashboard blocks and the quench
check and purge in the routing timers. The SYSTEMS blocks of an OPENBRIDGE
system are no longer written to at runtime.
Also removed: ``_config.pop("_no_target_log_time")``, a key nothing has written
since the port from hblink.
No behaviour change intended. The differential harness from phase 1, extended
to compare where egress actually goes (a keepalive and a voice frame sent after
every case) and to cover BCKA/BCSQ/BCST from three source addresses with
RELAX_CHECKS on and off, ran 9594 frames through this commit and through
develop: identical delivery, quench, egress address and log lines.
Tests: 24 new unit tests for the session, 98% coverage of the new module; the
RELAX_CHECKS sync test now asserts what the refactor is for — the session
follows the peer, the configured address stays put. Full suite 876 passed,
2 skipped (the 2 failures are this machine's, and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
``_obp_datagram_received`` decided admission inline: ~110 lines of nested ifs
repeating the same four-step shape (evaluate, log once per stream_id, quench,
return) eight times for the ACLs alone, reading GLOBAL and SYSTEMS state at six
different depths, and doing it twice over — once for DMRD v1 and once for
DMRE v5. Nothing in it could be exercised without a reactor and a socket.
The decisions now live in ``domain/mesh_admission.py`` as pure functions. Each
rule takes values and returns a ``Rejection`` (reason, log line, whether to
quench, whether to log once) or ``None`` to admit; ``admit_dmrd_v1`` and
``admit_dmre_v5`` chain them in the order the legacy handler applied them. The
adapter keeps what is genuinely I/O: one ``_obp_admission_context()`` that reads
config once per frame, and one ``_obp_reject()`` that applies the outcome.
No behaviour change intended. Beyond the suite, a differential harness fed
9576 cryptographically valid frames (DMRD v1 and DMRE v5, sweeping destination
TG, subscriber, bits byte, STUN, both ACL scopes, hop count, packet age, source
server and VALIDATE_SERVER_IDS) through a protocol built from this commit and
one built from develop, comparing delivery, quench and every log line: no
behavioural difference.
One deliberate deviation: the four DMRE "GLOBAL TG FILTER (local to ...)" log
calls on develop pass three arguments to a message with two placeholders, so
Python logs a formatting error instead of the drop (648 of the 9576 frames hit
this). The message now carries the talkgroup. ``test_every_rejection_can_be_
formatted`` keeps the whole family honest.
Also fixed by construction: ``reason`` gives each drop a stable handle, so a
later engine can count or trace drops without parsing log text.
Tests: 59 new unit tests, 100% coverage of the new module; full suite 849
passed, 2 skipped (the 2 failures are this machine's: kernel rmem_max and the
version installed in the venv, both failing on develop as well).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review follow-up on the fan-in demux fixes.
1. apply_obp_proxy_config_reload() rebuilds the bridge registry, and a rebuild
creates fresh ObpIngressReplyTransport objects. Nothing re-pinned them to the
legacy sockets that stay bound, so the first `systemctl reload` put egress
back on the shared fan-in port and reintroduced the bug the pin fixes.
Bound legacy transports are now kept in the service state and re-pinned after
every registry rebuild (pin_legacy_egress()). For the same reason the rebuild
now describes the ports that are actually bound (the running LISTEN_PORT /
BIND_LEGACY_PORTS) instead of the incoming ones, which are only logged as
"restart required".
2. Control-frame demux ranked bridges by the live SYSTEMS.<name>["TARGET_SOCK"],
the very dict RELAX_CHECKS rewrites in place when a bridge accepts traffic
from an unexpected source. One misattributed frame could therefore move a
bridge's target and then keep matching it there. Each entry now carries
peer_hint, the peer as configured, snapshotted when the registry is built
(startup and every reload, both of which read a freshly normalized config).
Order is: configured IP:port, live IP:port, configured IP, live IP, then
registration order — so a peer that legitimately moved is still recognised,
but never at the expense of the bridge that has that address in the YAML.
Tests
- test_reload_keeps_egress_pinned_to_the_bridge_socket: starts the service on a
fake reactor, reloads twice, asserts egress still leaves through each bridge's
own socket and nothing goes out of the fan-in port.
- test_control_frame_prefers_configured_peer_over_relaxed_target: a bridge
dragged onto another bridge's address does not steal that peer's keepalives.
- test_control_frame_unknown_peer_falls_back_to_passphrase now asserts the
deterministic outcome (first registered bridge) instead of "either one".
Both new tests fail without this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both listeners bind independent UDP sockets; if both are configured to
the same port only the first one to start would succeed, and the
server would fail at bind time with an unclear error instead of a
config validation message.
Three issues seen on a production ADN Systems master (2.5.6) with five
OpenBridge links, all rooted in the shared fan-in port.
1. YamlConfigLoader rebuilt the config from a fixed list of top-level keys
that did not include OBP_PROXY, so the whole block in adn-server.yaml was
silently ignored: ENABLED, LISTEN_PORT, LISTEN_IP and BIND_LEGACY_PORTS
always fell back to their defaults (--doctor reported 62032 for a config
that asked for 62031, and ENABLED: false still bound the fan-in).
2. Control frames (BCKA/BCSQ/BCST/BCVE) carry no NETWORK_ID, so the fan-in
could only tell bridges apart by passphrase. Every ADN Systems bridge
shares one passphrase, so all of them resolved to whichever bridge was
registered first: keepalives from the Spanish and Portuguese peers were
attributed to the French bridge and, with RELAX_CHECKS, rewrote its
TARGET_SOCK ("Source IP and Port has changed ... updating", in a loop).
Voice was unaffected because DMRD/DMRE do carry NETWORK_ID.
Control frames now prefer the bridge whose configured peer matches the
datagram source (exact IP:port, then same IP, then the passphrase scan).
3. ObpIngressReplyTransport fell back to the shared fan-in socket, so egress
left from LISTEN_PORT; remote peers learn that source and answer there,
funnelling everything onto the one port where control frames cannot be
told apart. Egress is now pinned to the bridge's own legacy socket when
one is bound.
Tests: five regression tests; all five fail without these changes.
Unit data to an individual ID (>= 1000000) was copied to every OpenBridge
(VER > 1) even when the destination was just heard on a local system and the
SUB_MAP lookup already delivers it there. A D-APRS gateway's ARS/LRRP answers
(one per minute per radio) went out on every configured bridge, multiplying
traffic and causing remote masters that mis-classify those frames to create a
dynamic talkgroup named after the radio id.
Skip the OBP fan-out when the SUB_MAP entry for the destination is on a
non-OPENBRIDGE system and younger than UNIT_DATA_LOCAL_SUB_MAX_AGE (900s).
Subscribers learned via an OpenBridge, or stale entries, keep the current
behavior, so a radio that moved to another master stays reachable.
Rescued from #72 (opened and self-closed by pyopower without a merge),
adapted to the current file layout.
TG/ID 4000 is the dynamic-TG reset control code and can never appear in
_SUB_MAP, so the unknown-destination fallback in _pvt_repeat_targets treated
it as a private call to an unlocated radio and blasted it to every connected
peer. That fallback runs before dmrd_received's own dst_id == 4000 guard, so
the control code has to be filtered at the targeting step too.
The periodic security download ran blocking HTTP on the reactor thread, so a
dead selfcare server stalled RPTPING handling well past PING_TIME * MAX_MISSED
and every peer was timed out and forced to reconnect. Move the security and
alias downloads to the thread pool, bound the DNS and HTTP timeouts below that
budget, and skip a cycle when the previous one is still in flight.
Also drop the redundant interval guard that silently stretched the real retry
period to a multiple of the loop interval, restore the process-wide socket
timeout after a failed DNS lookup, keep existing files and cached passwords
when a download returns an empty or invalid payload, and remove the inherited
50Mb cap on subscriber_ids that production was already close to tripping.
PASS_SECURITY is no longer written to the log.
Adds self-echo (a hotspot with the same TG on both slots, static or dynamic, hears its own TX on the other slot) and fixes four occurrences of the same slot-resolution bug that blocked or truncated it: peer_downlink_voice_slot always preferred an unambiguous static/SINGLE=1-locked slot over the caller's own wire slot, corrupting the busy-check and the parrot/echo anti-loopback guard alike. Also makes the "Downlink dropped" log reason specific instead of generic, and throttles repeated identical drop logs.
* fix: deliver TG on both slots when static on one, dynamic on the other
register_peer_ua_multi_tg silently dropped dynamic tracking on a slot
whenever the same TG was already static on the other slot, and the
REPEAT path never expanded to both slots at all.
* fix: strip NUL-padded CALLSIGN instead of showing raw bytes
Some peers NUL-pad instead of space-pad; reuse the existing
normalize_fixed_width_ascii helper instead of a plain .strip().
PR #53 (7ff3010c) collapsed iter_downlink_voice_slots to one delivery
slot for any peer, but peer_listen_slots only ever returns more than one
slot for a peer already confirmed duplex via peer_is_simplex -- so that
collapse could only suppress real duplex hotspots, not the unpaced
software bridges (ysf2dmr/adn-bridge) it was meant to fix, since those
already collapse correctly via their own simplex classification.
Trust peer_listen_slots directly again. Confirmed live: a duplex hotspot
with a TG on both static OPTIONS slots now receives it on both timeslots
regardless of which slot the source transmitted on.
Use SUB_MAP's stored peer id for delivery (repeat, unit-data, and
pvt_call_received), add Talker Alias support for private calls, and
report the receiving hotspot to the monitor.