- A new forward leg gets its header, terminator and embedded LC encoded
(BPTC 196,96 + RS, about 250us). Every leg of a call to the same TG gets
the same LC, so a call bridged to six OpenBridges and a MASTER spent
1.75ms on its first frame doing the same work seven times. The encoded
set is now kept per LC (bounded) and shared by the legs, which only ever
slice it.
- In-band signalling runs on each voice header and terminator and walked
every subscription of the server to find those of the source system. The
store now indexes legs by system, in the order a full scan returns them.
- The remaining full scans used to test or drop a relay table, and the copy
of a whole table just to test it is not empty, use the table index.
Short calls (50 frames), one call bridged to 6 OpenBridges and a MASTER,
200 other bridged TGs in the store, per frame, after the previous commit:
HBP ingress 136.4us -> 97.9us
OBP ingress 98.1us -> 60.7us
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every group voice frame, from HBP or OpenBridge, asked store_has_table()
whether its TG has a relay table. It answered by building a tuple of every
subscription and calling table_key() on each, so the cost of a frame grew
with the size of the network. On a live master (ADN 213, six OpenBridges)
this was 26% of the CPU spent on OpenBridge ingress.
- has_table() on the store answers from the table index it already keeps;
the port gets a default that scans, for stores without one.
- The forward plan of a group call (relay tables, resolved legs, the
MASTER/PEER dedupe and the target entries) depends only on the store and
on the source and target SYSTEMS blocks. It is now kept per
(system, slot, TG) and rebuilt when the store's new revision counter moves
or a reload replaces one of those blocks. ENABLED, quench, keepalive and
contention are still checked per frame in the loop.
- One OBP session lookup per target instead of two.
- store_has_table is imported once, not on every frame.
Per frame, one call bridged to 6 OpenBridges and a MASTER, with N other
bridged TGs in the store:
N=0 N=200
HBP ingress 81.8us -> 52.1us 447.4us -> 55.3us
OBP ingress 95.9us -> 66.8us 469.3us -> 65.7us
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
It runs once per datagram, for OBP and local traffic alike, and built a tuple
of every subscription to answer a yes/no question. Profiling a production
server at 198 datagrams/s put 67.8% of all CPU work inside it.
legs_in_table already reads the _by_table index and is on the port. Both were
added in the same commit as the scan, which never used them.
Constant time now instead of growing with the mesh: 6.8us to 0.30us at 120
subscriptions, 26.2us to 0.29us at 2400. On the server, 67.8% of work down to
0.84% and 42% less CPU at equal load.
Rule timers, in-band signalling and resets change a subscription in place
(sub.state.phase = IDLE) and then upsert the same object. upsert unindexed
"the old one" by reading its state, but the old one is that same object, now
IDLE, so the active indexes were never cleared: relay_tables_with_active_source
kept listing the leg as an active source and has_active_target_leg stayed true.
The visible effect: when a timed-out rule deactivates a hotspot's leg in the
middle of a call, the rest of the call is still forwarded, where the legacy
router checked ACTIVE on the source row for every frame and stopped. The
downlink hang logic also kept treating the system as having that leg.
The store now records what it indexed each leg under and removes exactly
that. Tests cover a leg deactivated in place, one activated in place next to
another active leg on the same key, a remove after an in-place change, and
the timed-out source end to end. All four fail on develop.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
_master_datagram_received and _peer_datagram_received each carried the same
32 lines: global SUB_ACL, the TG list for the frame's slot, then the system's
own three, every rejection logging once per stream through _laststrid. The two
copies were identical down to the wording of the six log messages, so a fix in
one could silently miss the other.
They now call _acl_rejects_dmrd, which keeps the legacy order, the same
messages and the same once-per-stream guard, and reads the TG list as
TG{slot}_ACL instead of spelling out slot 1 and slot 2. The frame's slot is
always 1 or 2 (it comes from bit 0x80), so no branch is lost.
tests/hbp/test_acl_gate.py covers both paths, the per-slot TG list and the
on-demand unit-service exemption; three of its four tests pass unchanged
against the previous code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The MASTER ingress path asks the same two questions of every voice frame.
peer_single_mode parses the peer's OPTIONS blob through
parse_peer_options_fields, and peer_rf_mode falls back to derive_peer_rf_mode
because nothing ever wrote the RF_MODE key its docstring anticipates. Around a
static-TG hotspot the downlink helpers add five more parse_peer_options_static
calls on the same blob, and each of those parses it twice
(peer_options_static_valid, then _parse_options_kv). A hotspot with
TS2_1=214;TS1_1=91;SINGLE=1;TIMER=15; had that string re-parsed more than ten
times per frame.
None of it moves at frame rate: OPTIONS only changes on RPTO, which already
drops the memo added for cached_peer_static_tgs, and SLOTS with the two
frequencies only arrive in RPTC. So peer_options_fields memoizes against the
blob the same way, the five remaining parse_peer_options_static call sites go
through cached_peer_static_tgs, and RPTC classifies the RF mode once at login.
Measured on the MASTER ingress path (tests/support/hbp_repeat_stack, 20k group
voice frames from a static-TG hotspot, ACLs on): 138.8us -> 39.1us per frame,
or 42.0us for a peer that has not sent RPTC yet and still derives its mode.
Emitted packets are byte-identical across 72 combinations of OPTIONS blobs
(including invalid ones and PASS=) and simplex/duplex peers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
STATUS is keyed by stream_id alone and the trimmer only drops a row after
180s idle, so when a peer reuses an id for another destination the second
call lands on the first one's row, fails the TGID check in loop control and
is discarded frame by frame. to_target already evicts a stale row on a
forward leg; ingress did not.
Ingress now evicts it too, once the old row has been idle for 2s so two
genuinely interleaved streams keep their own rows.
Also drops the bogus TS from the loop-control warning (it printed a slot
index from an unrelated loop), gives that branch the once-per-stream guard
its siblings have, and says what the condition is.
Seen on a production bridge: one talkgroup took 116 frames and completed no
calls at all.
_load_server_tsv_with_backup was a near-copy of _load_id_dict_with_backup:
same verify, same fallback, same backup, differing only in the parser and in
having "server_ids" written into its log messages. #94 added the skip-if-
unchanged path to one of them, so server_ids.tsv stayed the one alias file
re-read on every 15 minute tick.
They are one function now, taking the parser as an argument. Removes 51
duplicated lines and leaves no second path to drift.
The reload loop ticks every 900s so a failed download is retried soon, but
STALE_DAYS replaces the files about once a day. Every tick in between
re-read them: blake2b over 50MB, a JSON parse, a 50MB copy to .bak, and a
300k-entry profile rebuild, all to produce the same dicts.
Parsed results are now kept against each file's (mtime_ns, size, inode) and
the backup is written only when the primary has been verified, which is the
point where it is worth keeping as a fallback. A primary that fails drops
its remembered parse so the next tick looks at the file again.
The profile build also ran in merge_reload_into_config, on the reactor
thread, at over a second per cycle. It is built in the thread pool now and
handed in.
A tick with nothing new: 2568ms -> 0.2ms.
- Build BridgePolicy/AdmissionContext and PeerMeshConfig once, not per frame.
Dropped on reload, and on the _SERVER_IDS swap an alias refresh does.
- Carry the DMRE timestamp on MeshIngress: the trailer was parsed twice.
- Let the engine be the only one vetting the source; it already answers None.
- Resolve the session and its peer once per datagram instead of 3-5 times.
- Dispatch effects most-frequent-first.
47.1us -> 28.6us per frame on the v5 path. Recorded effects corpus unchanged.
MeshSessionStore.session refreshed dns_host from the config on every call,
and deciding whether TARGET_IP is a name or an address costs a thrown and
caught ValueError for every bridge configured by name. The OBP ingress
reads self._session three to four times per datagram, so a mesh of eight
named bridges paid that exception several times per frame, per bridge.
Nothing needed it there: the config is read-only at runtime, and sync
already follows a reload. It refreshed configured_peer but not dns_host,
which is the only reason the per-lookup refresh existed, so sync now owns
both and session derives dns_host only for a session it creates.
Measured on the v5 ingress path with a bridge configured by name:
66.2us to 47.9us per frame.
Anchoring compared the full socket, so a peer answering from a source port
other than the one we send to was refused: its name resolved to the right
host, but the port differed and every frame was discarded. NAT rewrites that
port, and a peer needs not bind the port it is reached on.
A name now pins the host only. The wire may still refine the port within
that host, and a re-resolution that moves the peer elsewhere drops a port
learned for the host it just left.
A name that has never resolved anchors nothing. normalize_obp_config leaves
TARGET_SOCK as (None, port) when startup resolution fails, and anchoring on
that refused every source forever, taking the link off the air until the
process restarted. Those bridges fall back to RELAX_CHECKS instead.
Control frames carry no NETWORK_ID, so three things can tell two OPENBRIDGE
bridges apart: a legacy port of their own, a passphrase of their own, or a
source address that matches what one of them is configured with. Any one is
enough. Sharing the fan-in port and a passphrase leaves only the address, and
a frame from an address none of them knows then goes to whichever bridge was
registered first — the misattribution behind #79 and #87.
Nothing said so. The validator checks duplicate NETWORK_IDs and duplicate
legacy ports, but treats PASSPHRASE as just another string.
openbridge_passphrase_collisions() groups the enabled OPENBRIDGE systems that
share one, and returns the names only, never the secret. It surfaces as a warn
finding in --doctor and as a line at startup next to the fan-in summary.
Kept advisory on purpose: a shared passphrase works as long as every peer
address is distinct and current, so refusing to start would stop servers that
are fine today.
The refusal log is capped to once per source address so a clone pinging
every 10s cannot flood the log. That left an operator with no way to tell
whether it was still happening: the first line scrolls away and nothing
replaces it, while the fan-in still prints "RX b'BCKA' from <clone> ->
OBP-USA", which reads as if the frame had been delivered.
Refusing skips the engine, and the engine is what calls count_drop() for
every other reason, so these frames were tallied nowhere either — not in
the periodic "(ROUTER) system X refused frames: ..." line, not in the
report. Count them as source-not-peer, so the tally answers "is the clone
still knocking?" once a minute without repeating the explanation.
A peer on a dynamic IP forces RELAX_CHECKS on, and RELAX_CHECKS meant
"accept from any address on earth". On a shared-passphrase mesh that is
enough for a second host to be taken for the peer: production showed
OBP-USA with two live instances (74.132.44.239, the configured peer, and
129.80.176.29, a stale clone), both authenticating, both sending voice,
keepalives and quenches. The session address flapped between them every
few seconds, so half of what we transmitted went to the wrong host, loss
climbed to 30%, and hop counts escalated until MAX HOPS dropped frames.
The network already had the answer: TARGET_IP was written as a name,
3103.adn.systems, which tracks the dynamic IP by DNS. normalize_obp_config
resolved it once and then overwrote TARGET_IP with the address, losing the
name, so nothing could ever ask again — while PEER systems have kept
_MASTER_IP and re-resolved through reactor.resolve() all along.
OPENBRIDGE now gets the same treatment:
- normalize_obp_config keeps the name in _TARGET_IP, as PEER keeps
_MASTER_IP.
- A session whose TARGET_IP was a name is DNS-anchored: learn_peer()
refuses to move it, and only adopt_resolved() can.
- accepts_source() stops widening to "anywhere" for an anchored bridge.
RELAX_CHECKS keeps its meaning for bridges configured with an address.
- Control frames get that same check. They had none at all, which is how a
foreign BCSQ could quench a live stream and a foreign BCST could STUN the
bridge outright.
- A frame from elsewhere is refused and schedules a re-resolution, rate
limited so unknown traffic cannot drive a lookup per packet. If the name
now answers with that address, the peer migrates and the next frame is
accepted; a periodic loop keeps it fresh while the link is idle. A
resolver failure keeps the address we have.
Verified on the 213 master: the clone's frames are refused and logged once,
the flapping is gone, and a real call bridged to eight systems at 0.37%
loss and 6 hops.
BCKA carries no NETWORK_ID, so on a shared-passphrase mesh anyone's
keepalive verifies against any bridge. It was moving session.peer with no
gate at all — not even RELAX_CHECKS, which the recorded corpus shows:
"bcka from 9.9.9.9:62201 relax=False" moved egress to 9.9.9.9. Since
session.peer is where voice and control are sent, a second instance of a
peer (seen in production on OBP-USA: two hosts, two 10s keepalive timers)
took the traffic over every few seconds.
A keepalive now only confirms liveness. It may still bootstrap a bridge
that has no address yet (inbound-only, no TARGET_IP), since there is
nothing to steal there and it is the only way such a bridge learns where
to answer. Relocation is left to DMRD/DMRE, which identify themselves.
Also: the fan-in demux ranked bridges by sys_cfg's TARGET_SOCK, frozen
since #81 moved runtime state into the session, so the "live" ranks were
dead code and a peer that really moved was no longer recognised. It now
reads learned_peer from the session store, closing the integration #81
left pending.
#84 added the fields to _TOPOLOGY_PEER_FIELDS with no test covering either
direction: that they appear (validated against schemas/report-v2.json)
when the peer sends coordinates, and that they're omitted, not emitted as
null/empty, when it doesn't.
#82 logged the RX debug line once per stream_id per bridge, but a bridge
with BOTH_SLOTS/several concurrent talkgroups interleaves packets from
multiple calls, so the "last seen" stream flips almost every packet and
the line logs nearly as often as before the fix — confirmed against
production output (three interleaved streams on OBP-CL).
DMRD/DMRE demux by NETWORK_ID, which is never ambiguous, and
*CALL START*/*CALL END* (application/routing_use_cases.py,
application/routing/obp_forward.py) already log once per real call with
better detail (SUB, TGID, TS, duration). Drop the fan-in RX line for voice
frames entirely instead of trying to approximate "once per call" with no
session state; control frames (BCKA/BCSQ/BCST/BCVE) keep logging every
time, since they are rare and that visibility mattered for the #79 fix.
An active call sends a DMRD/DMRE packet roughly every 90-100ms; with
OBP_PROXY.DEBUG on, that logged one RX line per packet, flooding the
log for the whole call. Log once when a stream_id first appears instead.
Non-voice frames (control) are unaffected, they were never the flood.
Two things a real capture showed that synthetic frames could not. An
unfiltered tcpdump carries both directions, and our own egress was being run
through the ingress rules, where it fails the NETWORK_ID check by definition:
10698 frames on one bridge reported as network-id-mismatch that were simply
ours. Direction is now decided by the peer address (both ends of a link
normally share the port number, so the port alone cannot tell), with the
listening port as a fallback for a peer behind NAT; --replay-both-directions
keeps the old behaviour.
And VALIDATE_SERVER_IDS was firing offline against an empty list, because the
server-id table is loaded at runtime and not from the YAML: 10107 frames on
another bridge blamed on source-server-unknown. Offline the check is skipped
unless the table is actually there.
Both were reported against the live ADN 213 master (5 bridges, 30669 real
datagrams in 120s); with them fixed the report says what the master does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The engine from phase 3 takes plain values and answers with effects, so it does
not need the server to run. ``adn-server --replay capture.pcap`` uses that: the
frames from a tcpdump capture go through the same ingress, against the
operator's own adn-server.yaml, and each one comes back with the bridge it
belongs to and either a delivery or the reason it was refused.
12:04:31 82.65.127.86:62201 OBP-FR DMRD v1 2130001 -> 214 delivered
12:04:31 85.241.222.7:62268 OBP-PT DMRE v5 2680015 -> 9 dropped (tg-filter-server) +BCSQ
12:04:32 203.0.113.9:50000 - DMRD v1 unmatched
Nothing is sent and no port is bound, so it runs beside a live master. Which
bridge a frame belongs to is decided on the evidence the server has — the port
it arrived on, then the configured peer, then whoever can verify it — which is
also the answer to "whose keepalive is this?" on a mesh where every bridge
shares one passphrase.
``--system`` narrows it to one link, ``--replay-limit`` stops early and
``--replay-summary`` prints the tally alone.
``infrastructure/pcap.py`` reads classic pcap with no dependencies: both
endiannesses, microsecond and nanosecond timestamps, Ethernet (VLAN tags
included), Linux cooked v1 and v2 (``tcpdump -i any`` writes SLL2, found while
running this against a real capture), raw IP and loopback, IPv4 and IPv6 UDP.
pcapng says which command converts it.
Docs: the OBP proxy page gains a "why did that call not cross" section in both
languages, and its RELAX_CHECKS note now says what phase 2 made true — what the
wire teaches lives in the session, TARGET_IP stays as written.
Tests: 27 new, 97% of the replay module and 87% of the pcap reader, plus an
end-to-end run of the real command against a real YAML. Full suite 1025 passed,
2 skipped (the 2 failures are this machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 3, and the end of the hblink shape in this path. Phase 1 lifted the
admission rules out of the adapter, phase 2 took the state they read; what was
left in ``_obp_datagram_received`` was the plumbing between them — two long
branches that decoded, decided, logged, quenched and called into routing, all
inside a Twisted ``DatagramProtocol`` where none of it could be run on its own.
``domain/mesh_engine.py`` now takes a decoded frame, the link's session and a
``BridgePolicy``, and answers with a list of effects: ``Reject``, ``Log``,
``NoteStream``, ``StoreTalkerAlias``, ``Deliver``, ``RequestVersion``. It reads
no configuration, opens no socket and calls no logger. The adapter keeps the
three things that are genuinely I/O — verify the MAC, build the policy, carry
out the effects in order — and the handler goes from ~180 lines of nested
branches to ~45 of dispatch.
Two things this buys beyond the shape. A datagram can be replayed through the
engine anywhere: a test, a laptop, a capture from a sysop, with no reactor in
sight. And every refusal carries a ``reason``, now tallied per bridge in the
session and printed by the keepalive loop at debug level — "my call does not
cross" is answered by ``tg-filter-mcc=12, system-sub-acl=3`` instead of by
grepping the log.
No behaviour change intended, and this time checked two ways. The differential
harness ran 9594 frames through this tree and through upstream ``develop``:
identical delivery, quench, egress address and log lines. And the corpus in
``tests/fixtures/obp_ingress_effects.jsonl`` — 96 recorded cases covering the
talkgroup filters, both ACL scopes, the bits byte, the DMRE envelope (age,
hops, source server) and BCKA/BCSQ/BCST from three addresses — was recorded
from the pre-engine handler, verified frame by frame against develop, and still
passes untouched. It stays as the contract for whatever comes next; regenerate
with CAPTURE=1 and read the diff.
Tests: 24 new unit tests for the engine (100% of the module), the recorded
corpus as a regression net, 998 passed, 2 skipped (the 2 failures are this
machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 2 of separating the OpenBridge path from its hblink ancestry. Phase 1
lifted the admission rules out of the adapter; this one takes the state they
were reading.
Legacy hblink kept what a bridge learns at runtime inside its own SYSTEMS
block: ``_bcka`` (last keepalive), ``_bcsq`` (the peer's quench table),
``_STUN`` and, worst of the four, ``TARGET_IP``/``TARGET_PORT``/``TARGET_SOCK``,
rewritten in place every time RELAX_CHECKS accepted a datagram from an address
the operator never wrote. Configuration and session state shared one mutable
dict, so a single unexpected packet could move a bridge's target for good, and
no reader could tell what came from the YAML and what came from the wire.
``domain/mesh_session.py`` now holds an ``ObpBridgeSession`` per link:
configured_peer (from the YAML, never moves) beside learned_peer (from the
wire), the last keepalive, the quench table and the BCST stun flag, each behind
the question a caller actually asks — ``peer``, ``keepalive_seen``,
``keepalive_stale(now)``, ``quenches(tg, stream)``. The store lives under a
private top-level config key next to ``_SUB_MAP`` and ``_PEER_IDS``, so every
layer that already receives the config reaches the same instance; like
``_SUB_MAP`` it is shared, not deep-copied, across a SIGHUP, and
``sync(config)`` then refreshes the configured peers and drops the sessions of
links that are gone. Editing TARGET_IP in the YAML and reloading is now a
documented way to undo a bad learned address.
Migrated readers: the keepalive gate in routing (to_target and unit data), the
60s keepalive report loop, the monitor/MQTT dashboard blocks and the quench
check and purge in the routing timers. The SYSTEMS blocks of an OPENBRIDGE
system are no longer written to at runtime.
Also removed: ``_config.pop("_no_target_log_time")``, a key nothing has written
since the port from hblink.
No behaviour change intended. The differential harness from phase 1, extended
to compare where egress actually goes (a keepalive and a voice frame sent after
every case) and to cover BCKA/BCSQ/BCST from three source addresses with
RELAX_CHECKS on and off, ran 9594 frames through this commit and through
develop: identical delivery, quench, egress address and log lines.
Tests: 24 new unit tests for the session, 98% coverage of the new module; the
RELAX_CHECKS sync test now asserts what the refactor is for — the session
follows the peer, the configured address stays put. Full suite 876 passed,
2 skipped (the 2 failures are this machine's, and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
``_obp_datagram_received`` decided admission inline: ~110 lines of nested ifs
repeating the same four-step shape (evaluate, log once per stream_id, quench,
return) eight times for the ACLs alone, reading GLOBAL and SYSTEMS state at six
different depths, and doing it twice over — once for DMRD v1 and once for
DMRE v5. Nothing in it could be exercised without a reactor and a socket.
The decisions now live in ``domain/mesh_admission.py`` as pure functions. Each
rule takes values and returns a ``Rejection`` (reason, log line, whether to
quench, whether to log once) or ``None`` to admit; ``admit_dmrd_v1`` and
``admit_dmre_v5`` chain them in the order the legacy handler applied them. The
adapter keeps what is genuinely I/O: one ``_obp_admission_context()`` that reads
config once per frame, and one ``_obp_reject()`` that applies the outcome.
No behaviour change intended. Beyond the suite, a differential harness fed
9576 cryptographically valid frames (DMRD v1 and DMRE v5, sweeping destination
TG, subscriber, bits byte, STUN, both ACL scopes, hop count, packet age, source
server and VALIDATE_SERVER_IDS) through a protocol built from this commit and
one built from develop, comparing delivery, quench and every log line: no
behavioural difference.
One deliberate deviation: the four DMRE "GLOBAL TG FILTER (local to ...)" log
calls on develop pass three arguments to a message with two placeholders, so
Python logs a formatting error instead of the drop (648 of the 9576 frames hit
this). The message now carries the talkgroup. ``test_every_rejection_can_be_
formatted`` keeps the whole family honest.
Also fixed by construction: ``reason`` gives each drop a stable handle, so a
later engine can count or trace drops without parsing log text.
Tests: 59 new unit tests, 100% coverage of the new module; full suite 849
passed, 2 skipped (the 2 failures are this machine's: kernel rmem_max and the
version installed in the venv, both failing on develop as well).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review follow-up on the fan-in demux fixes.
1. apply_obp_proxy_config_reload() rebuilds the bridge registry, and a rebuild
creates fresh ObpIngressReplyTransport objects. Nothing re-pinned them to the
legacy sockets that stay bound, so the first `systemctl reload` put egress
back on the shared fan-in port and reintroduced the bug the pin fixes.
Bound legacy transports are now kept in the service state and re-pinned after
every registry rebuild (pin_legacy_egress()). For the same reason the rebuild
now describes the ports that are actually bound (the running LISTEN_PORT /
BIND_LEGACY_PORTS) instead of the incoming ones, which are only logged as
"restart required".
2. Control-frame demux ranked bridges by the live SYSTEMS.<name>["TARGET_SOCK"],
the very dict RELAX_CHECKS rewrites in place when a bridge accepts traffic
from an unexpected source. One misattributed frame could therefore move a
bridge's target and then keep matching it there. Each entry now carries
peer_hint, the peer as configured, snapshotted when the registry is built
(startup and every reload, both of which read a freshly normalized config).
Order is: configured IP:port, live IP:port, configured IP, live IP, then
registration order — so a peer that legitimately moved is still recognised,
but never at the expense of the bridge that has that address in the YAML.
Tests
- test_reload_keeps_egress_pinned_to_the_bridge_socket: starts the service on a
fake reactor, reloads twice, asserts egress still leaves through each bridge's
own socket and nothing goes out of the fan-in port.
- test_control_frame_prefers_configured_peer_over_relaxed_target: a bridge
dragged onto another bridge's address does not steal that peer's keepalives.
- test_control_frame_unknown_peer_falls_back_to_passphrase now asserts the
deterministic outcome (first registered bridge) instead of "either one".
Both new tests fail without this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both listeners bind independent UDP sockets; if both are configured to
the same port only the first one to start would succeed, and the
server would fail at bind time with an unclear error instead of a
config validation message.
Three issues seen on a production ADN Systems master (2.5.6) with five
OpenBridge links, all rooted in the shared fan-in port.
1. YamlConfigLoader rebuilt the config from a fixed list of top-level keys
that did not include OBP_PROXY, so the whole block in adn-server.yaml was
silently ignored: ENABLED, LISTEN_PORT, LISTEN_IP and BIND_LEGACY_PORTS
always fell back to their defaults (--doctor reported 62032 for a config
that asked for 62031, and ENABLED: false still bound the fan-in).
2. Control frames (BCKA/BCSQ/BCST/BCVE) carry no NETWORK_ID, so the fan-in
could only tell bridges apart by passphrase. Every ADN Systems bridge
shares one passphrase, so all of them resolved to whichever bridge was
registered first: keepalives from the Spanish and Portuguese peers were
attributed to the French bridge and, with RELAX_CHECKS, rewrote its
TARGET_SOCK ("Source IP and Port has changed ... updating", in a loop).
Voice was unaffected because DMRD/DMRE do carry NETWORK_ID.
Control frames now prefer the bridge whose configured peer matches the
datagram source (exact IP:port, then same IP, then the passphrase scan).
3. ObpIngressReplyTransport fell back to the shared fan-in socket, so egress
left from LISTEN_PORT; remote peers learn that source and answer there,
funnelling everything onto the one port where control frames cannot be
told apart. Egress is now pinned to the bridge's own legacy socket when
one is bound.
Tests: five regression tests; all five fail without these changes.
Unit data to an individual ID (>= 1000000) was copied to every OpenBridge
(VER > 1) even when the destination was just heard on a local system and the
SUB_MAP lookup already delivers it there. A D-APRS gateway's ARS/LRRP answers
(one per minute per radio) went out on every configured bridge, multiplying
traffic and causing remote masters that mis-classify those frames to create a
dynamic talkgroup named after the radio id.
Skip the OBP fan-out when the SUB_MAP entry for the destination is on a
non-OPENBRIDGE system and younger than UNIT_DATA_LOCAL_SUB_MAX_AGE (900s).
Subscribers learned via an OpenBridge, or stale entries, keep the current
behavior, so a radio that moved to another master stays reachable.
Rescued from #72 (opened and self-closed by pyopower without a merge),
adapted to the current file layout.
TG/ID 4000 is the dynamic-TG reset control code and can never appear in
_SUB_MAP, so the unknown-destination fallback in _pvt_repeat_targets treated
it as a private call to an unlocated radio and blasted it to every connected
peer. That fallback runs before dmrd_received's own dst_id == 4000 guard, so
the control code has to be filtered at the targeting step too.
The periodic security download ran blocking HTTP on the reactor thread, so a
dead selfcare server stalled RPTPING handling well past PING_TIME * MAX_MISSED
and every peer was timed out and forced to reconnect. Move the security and
alias downloads to the thread pool, bound the DNS and HTTP timeouts below that
budget, and skip a cycle when the previous one is still in flight.
Also drop the redundant interval guard that silently stretched the real retry
period to a multiple of the loop interval, restore the process-wide socket
timeout after a failed DNS lookup, keep existing files and cached passwords
when a download returns an empty or invalid payload, and remove the inherited
50Mb cap on subscriber_ids that production was already close to tripping.
PASS_SECURITY is no longer written to the log.
Adds self-echo (a hotspot with the same TG on both slots, static or dynamic, hears its own TX on the other slot) and fixes four occurrences of the same slot-resolution bug that blocked or truncated it: peer_downlink_voice_slot always preferred an unambiguous static/SINGLE=1-locked slot over the caller's own wire slot, corrupting the busy-check and the parrot/echo anti-loopback guard alike. Also makes the "Downlink dropped" log reason specific instead of generic, and throttles repeated identical drop logs.
* fix: deliver TG on both slots when static on one, dynamic on the other
register_peer_ua_multi_tg silently dropped dynamic tracking on a slot
whenever the same TG was already static on the other slot, and the
REPEAT path never expanded to both slots at all.
* fix: strip NUL-padded CALLSIGN instead of showing raw bytes
Some peers NUL-pad instead of space-pad; reuse the existing
normalize_fixed_width_ascii helper instead of a plain .strip().
PR #53 (7ff3010c) collapsed iter_downlink_voice_slots to one delivery
slot for any peer, but peer_listen_slots only ever returns more than one
slot for a peer already confirmed duplex via peer_is_simplex -- so that
collapse could only suppress real duplex hotspots, not the unpaced
software bridges (ysf2dmr/adn-bridge) it was meant to fix, since those
already collapse correctly via their own simplex classification.
Trust peer_listen_slots directly again. Confirmed live: a duplex hotspot
with a TG on both static OPTIONS slots now receives it on both timeslots
regardless of which slot the source transmitted on.
Use SUB_MAP's stored peer id for delivery (repeat, unit-data, and
pvt_call_received), add Talker Alias support for private calls, and
report the receiving hotspot to the monitor.
Announcement/TTS injection reuses a real peer's DMR id as rf_src, which made
resolve_voice_peer_id's rf_src fallback (and the inject-only proxy's fuzzy
peer match) misattribute the call to that peer: monitor showed it TX/red and
learned bogus dynamic TGs. Guard both resolvers with a new
synthetic_announcement flag, and add a trailing is_announcement field to
every GROUP VOICE report (CSV and JSON voice_event) for monitor-side use.
Dedupe the INFO line when a peer keys into a busy TG during an active QSO,
bind RX_STREAM_ID on silent activation to avoid per-frame re-entry, and
clear the dedupe marker on suppressed-stream VTERM.
A BRIDGES scan can legitimately list the same MASTER on both TS1 and
TS2 for one TG (e.g. an inject-only proxy whose connected hotspots
collectively use both slots). SubscriptionRouter.resolve() kept both
as distinct legs, so dmrd_received called send_to_system twice for
that MASTER per source frame.
Legacy's dumb send_peers() broadcast made that harmless: each
hotspot's own radio silently dropped the slot it didn't want. This
server's send_peers() instead resolves every peer's real listen slot
from its own OPTIONS regardless of the wire slot
(iter_downlink_voice_slots), so the second leg delivered the same
audio a second time to every subscribed peer - doubling the downlink
rate and making unpaced bridges (e.g. ysf2dmr) sound slow. Only
OBP-sourced calls hit this: a MASTER never forwards to itself, so
locally-originated (HBP) calls never produced a redundant leg.
Collapse forward legs to a MASTER/PEER target down to one per
(target_system, target_tgid), keeping distinct legs for OpenBridge
targets and for legitimate TG-translation legs to the same target on
a different TGID.