It runs once per datagram, for OBP and local traffic alike, and built a tuple
of every subscription to answer a yes/no question. Profiling a production
server at 198 datagrams/s put 67.8% of all CPU work inside it.
legs_in_table already reads the _by_table index and is on the port. Both were
added in the same commit as the scan, which never used them.
Constant time now instead of growing with the mesh: 6.8us to 0.30us at 120
subscriptions, 26.2us to 0.29us at 2400. On the server, 67.8% of work down to
0.84% and 42% less CPU at equal load.
STATUS is keyed by stream_id alone and the trimmer only drops a row after
180s idle, so when a peer reuses an id for another destination the second
call lands on the first one's row, fails the TGID check in loop control and
is discarded frame by frame. to_target already evicts a stale row on a
forward leg; ingress did not.
Ingress now evicts it too, once the old row has been idle for 2s so two
genuinely interleaved streams keep their own rows.
Also drops the bogus TS from the loop-control warning (it printed a slot
index from an unrelated loop), gives that branch the once-per-stream guard
its siblings have, and says what the condition is.
Seen on a production bridge: one talkgroup took 116 frames and completed no
calls at all.
_load_server_tsv_with_backup was a near-copy of _load_id_dict_with_backup:
same verify, same fallback, same backup, differing only in the parser and in
having "server_ids" written into its log messages. #94 added the skip-if-
unchanged path to one of them, so server_ids.tsv stayed the one alias file
re-read on every 15 minute tick.
They are one function now, taking the parser as an argument. Removes 51
duplicated lines and leaves no second path to drift.
The reload loop ticks every 900s so a failed download is retried soon, but
STALE_DAYS replaces the files about once a day. Every tick in between
re-read them: blake2b over 50MB, a JSON parse, a 50MB copy to .bak, and a
300k-entry profile rebuild, all to produce the same dicts.
Parsed results are now kept against each file's (mtime_ns, size, inode) and
the backup is written only when the primary has been verified, which is the
point where it is worth keeping as a fallback. A primary that fails drops
its remembered parse so the next tick looks at the file again.
The profile build also ran in merge_reload_into_config, on the reactor
thread, at over a second per cycle. It is built in the thread pool now and
handed in.
A tick with nothing new: 2568ms -> 0.2ms.
- Build BridgePolicy/AdmissionContext and PeerMeshConfig once, not per frame.
Dropped on reload, and on the _SERVER_IDS swap an alias refresh does.
- Carry the DMRE timestamp on MeshIngress: the trailer was parsed twice.
- Let the engine be the only one vetting the source; it already answers None.
- Resolve the session and its peer once per datagram instead of 3-5 times.
- Dispatch effects most-frequent-first.
47.1us -> 28.6us per frame on the v5 path. Recorded effects corpus unchanged.
MeshSessionStore.session refreshed dns_host from the config on every call,
and deciding whether TARGET_IP is a name or an address costs a thrown and
caught ValueError for every bridge configured by name. The OBP ingress
reads self._session three to four times per datagram, so a mesh of eight
named bridges paid that exception several times per frame, per bridge.
Nothing needed it there: the config is read-only at runtime, and sync
already follows a reload. It refreshed configured_peer but not dns_host,
which is the only reason the per-lookup refresh existed, so sync now owns
both and session derives dns_host only for a session it creates.
Measured on the v5 ingress path with a bridge configured by name:
66.2us to 47.9us per frame.
The password table was re-read and re-decrypted on a 10s timer, while the
file it reads only ever changes when the security downloader replaces it,
every 300s. The downloader already reloads the table on a successful
download, so 29 of every 30 decrypt cycles produced an identical result.
Each cycle also read the encryption key from disk once per entry, because
decrypt_password builds its own Fernet: 781 entries meant 781 stats, 781
reads of the same key file and 781 Fernet constructions, about 78 key
reads a second sustained.
Dropping the timer and decrypting a table on one Fernet takes this from
~23,400 key reads per five minutes to one, measured at 85.5ms per cycle
on a server carrying 781 entries.
The loader now keeps no config of its own: it was held only to let the
expired timer reload itself.
Anchoring compared the full socket, so a peer answering from a source port
other than the one we send to was refused: its name resolved to the right
host, but the port differed and every frame was discarded. NAT rewrites that
port, and a peer needs not bind the port it is reached on.
A name now pins the host only. The wire may still refine the port within
that host, and a re-resolution that moves the peer elsewhere drops a port
learned for the host it just left.
A name that has never resolved anchors nothing. normalize_obp_config leaves
TARGET_SOCK as (None, port) when startup resolution fails, and anchoring on
that refused every source forever, taking the link off the air until the
process restarted. Those bridges fall back to RELAX_CHECKS instead.
Control frames carry no NETWORK_ID, so three things can tell two OPENBRIDGE
bridges apart: a legacy port of their own, a passphrase of their own, or a
source address that matches what one of them is configured with. Any one is
enough. Sharing the fan-in port and a passphrase leaves only the address, and
a frame from an address none of them knows then goes to whichever bridge was
registered first — the misattribution behind #79 and #87.
Nothing said so. The validator checks duplicate NETWORK_IDs and duplicate
legacy ports, but treats PASSPHRASE as just another string.
openbridge_passphrase_collisions() groups the enabled OPENBRIDGE systems that
share one, and returns the names only, never the secret. It surfaces as a warn
finding in --doctor and as a line at startup next to the fan-in summary.
Kept advisory on purpose: a shared passphrase works as long as every peer
address is distinct and current, so refusing to start would stop servers that
are fine today.
The refusal log is capped to once per source address so a clone pinging
every 10s cannot flood the log. That left an operator with no way to tell
whether it was still happening: the first line scrolls away and nothing
replaces it, while the fan-in still prints "RX b'BCKA' from <clone> ->
OBP-USA", which reads as if the frame had been delivered.
Refusing skips the engine, and the engine is what calls count_drop() for
every other reason, so these frames were tallied nowhere either — not in
the periodic "(ROUTER) system X refused frames: ..." line, not in the
report. Count them as source-not-peer, so the tally answers "is the clone
still knocking?" once a minute without repeating the explanation.
A peer on a dynamic IP forces RELAX_CHECKS on, and RELAX_CHECKS meant
"accept from any address on earth". On a shared-passphrase mesh that is
enough for a second host to be taken for the peer: production showed
OBP-USA with two live instances (74.132.44.239, the configured peer, and
129.80.176.29, a stale clone), both authenticating, both sending voice,
keepalives and quenches. The session address flapped between them every
few seconds, so half of what we transmitted went to the wrong host, loss
climbed to 30%, and hop counts escalated until MAX HOPS dropped frames.
The network already had the answer: TARGET_IP was written as a name,
3103.adn.systems, which tracks the dynamic IP by DNS. normalize_obp_config
resolved it once and then overwrote TARGET_IP with the address, losing the
name, so nothing could ever ask again — while PEER systems have kept
_MASTER_IP and re-resolved through reactor.resolve() all along.
OPENBRIDGE now gets the same treatment:
- normalize_obp_config keeps the name in _TARGET_IP, as PEER keeps
_MASTER_IP.
- A session whose TARGET_IP was a name is DNS-anchored: learn_peer()
refuses to move it, and only adopt_resolved() can.
- accepts_source() stops widening to "anywhere" for an anchored bridge.
RELAX_CHECKS keeps its meaning for bridges configured with an address.
- Control frames get that same check. They had none at all, which is how a
foreign BCSQ could quench a live stream and a foreign BCST could STUN the
bridge outright.
- A frame from elsewhere is refused and schedules a re-resolution, rate
limited so unknown traffic cannot drive a lookup per packet. If the name
now answers with that address, the peer migrates and the next frame is
accepted; a periodic loop keeps it fresh while the link is idle. A
resolver failure keeps the address we have.
Verified on the 213 master: the clone's frames are refused and logged once,
the flapping is gone, and a real call bridged to eight systems at 0.37%
loss and 6 hops.
BCKA carries no NETWORK_ID, so on a shared-passphrase mesh anyone's
keepalive verifies against any bridge. It was moving session.peer with no
gate at all — not even RELAX_CHECKS, which the recorded corpus shows:
"bcka from 9.9.9.9:62201 relax=False" moved egress to 9.9.9.9. Since
session.peer is where voice and control are sent, a second instance of a
peer (seen in production on OBP-USA: two hosts, two 10s keepalive timers)
took the traffic over every few seconds.
A keepalive now only confirms liveness. It may still bootstrap a bridge
that has no address yet (inbound-only, no TARGET_IP), since there is
nothing to steal there and it is the only way such a bridge learns where
to answer. Relocation is left to DMRD/DMRE, which identify themselves.
Also: the fan-in demux ranked bridges by sys_cfg's TARGET_SOCK, frozen
since #81 moved runtime state into the session, so the "live" ranks were
dead code and a peer that really moved was no longer recognised. It now
reads learned_peer from the session store, closing the integration #81
left pending.
README.md's pairing note still said adn-monitor 2.0.0-rc.4 (server was on
2.5.6) — both projects' next release is 2.6, so set both sides now.
README_es.md never had this line, or any of the rest of README.md's
content beyond Acknowledgments (9 lines vs 92): translated the full file
so it matches structure and content, with docs/es/ links where a Spanish
page exists (server/development/testing.md doesn't yet, links to the
English page).
#84 added the fields to _TOPOLOGY_PEER_FIELDS with no test covering either
direction: that they appear (validated against schemas/report-v2.json)
when the peer sends coordinates, and that they're omitted, not emitted as
null/empty, when it doesn't.
A peer already tells the server where it is in its RPTC login, but the
report never passed those fields on, so a monitor could only show the
free-text LOCATION. Add latitude, longitude and height to the peer rows
of the report payloads (topology and dashboard_state), next to the
other RPTC display fields.
This is what a dashboard needs to place repeaters and hotspots on a
map; the legacy adn-dmr-server already had them in CONFIG. Peers that
do not send coordinates are unaffected: the fields are simply absent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
#82 logged the RX debug line once per stream_id per bridge, but a bridge
with BOTH_SLOTS/several concurrent talkgroups interleaves packets from
multiple calls, so the "last seen" stream flips almost every packet and
the line logs nearly as often as before the fix — confirmed against
production output (three interleaved streams on OBP-CL).
DMRD/DMRE demux by NETWORK_ID, which is never ambiguous, and
*CALL START*/*CALL END* (application/routing_use_cases.py,
application/routing/obp_forward.py) already log once per real call with
better detail (SUB, TGID, TS, duration). Drop the fan-in RX line for voice
frames entirely instead of trying to approximate "once per call" with no
session state; control frames (BCKA/BCSQ/BCST/BCVE) keep logging every
time, since they are rare and that visibility mattered for the #79 fix.
An active call sends a DMRD/DMRE packet roughly every 90-100ms; with
OBP_PROXY.DEBUG on, that logged one RX line per packet, flooding the
log for the whole call. Log once when a stream_id first appears instead.
Non-voice frames (control) are unaffected, they were never the flood.
Validated against the production ADN 213 master: the calls it logged appear in
the replay, on the right bridge, with the right subscriber and talkgroup. Three
things that only show up against real traffic are now in the page, in both
languages.
Capture a port range, not a list: a peer answering from an unexpected port is
the case worth looking at and a narrow filter hides it.
A call legitimately arrives on several bridges at once on a mesh — the master
logs one because loop control, later in routing, keeps one leg and drops the
rest. The replay stops at ingress, so it shows all of them.
And `outbound` frames (what this server sent) are reported, not judged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two things a real capture showed that synthetic frames could not. An
unfiltered tcpdump carries both directions, and our own egress was being run
through the ingress rules, where it fails the NETWORK_ID check by definition:
10698 frames on one bridge reported as network-id-mismatch that were simply
ours. Direction is now decided by the peer address (both ends of a link
normally share the port number, so the port alone cannot tell), with the
listening port as a fallback for a peer behind NAT; --replay-both-directions
keeps the old behaviour.
And VALIDATE_SERVER_IDS was firing offline against an empty list, because the
server-id table is loaded at runtime and not from the YAML: 10107 frames on
another bridge blamed on source-server-unknown. Offline the check is skipped
unless the table is actually there.
Both were reported against the live ADN 213 master (5 bridges, 30669 real
datagrams in 120s); with them fixed the report says what the master does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The engine from phase 3 takes plain values and answers with effects, so it does
not need the server to run. ``adn-server --replay capture.pcap`` uses that: the
frames from a tcpdump capture go through the same ingress, against the
operator's own adn-server.yaml, and each one comes back with the bridge it
belongs to and either a delivery or the reason it was refused.
12:04:31 82.65.127.86:62201 OBP-FR DMRD v1 2130001 -> 214 delivered
12:04:31 85.241.222.7:62268 OBP-PT DMRE v5 2680015 -> 9 dropped (tg-filter-server) +BCSQ
12:04:32 203.0.113.9:50000 - DMRD v1 unmatched
Nothing is sent and no port is bound, so it runs beside a live master. Which
bridge a frame belongs to is decided on the evidence the server has — the port
it arrived on, then the configured peer, then whoever can verify it — which is
also the answer to "whose keepalive is this?" on a mesh where every bridge
shares one passphrase.
``--system`` narrows it to one link, ``--replay-limit`` stops early and
``--replay-summary`` prints the tally alone.
``infrastructure/pcap.py`` reads classic pcap with no dependencies: both
endiannesses, microsecond and nanosecond timestamps, Ethernet (VLAN tags
included), Linux cooked v1 and v2 (``tcpdump -i any`` writes SLL2, found while
running this against a real capture), raw IP and loopback, IPv4 and IPv6 UDP.
pcapng says which command converts it.
Docs: the OBP proxy page gains a "why did that call not cross" section in both
languages, and its RELAX_CHECKS note now says what phase 2 made true — what the
wire teaches lives in the session, TARGET_IP stays as written.
Tests: 27 new, 97% of the replay module and 87% of the pcap reader, plus an
end-to-end run of the real command against a real YAML. Full suite 1025 passed,
2 skipped (the 2 failures are this machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 3, and the end of the hblink shape in this path. Phase 1 lifted the
admission rules out of the adapter, phase 2 took the state they read; what was
left in ``_obp_datagram_received`` was the plumbing between them — two long
branches that decoded, decided, logged, quenched and called into routing, all
inside a Twisted ``DatagramProtocol`` where none of it could be run on its own.
``domain/mesh_engine.py`` now takes a decoded frame, the link's session and a
``BridgePolicy``, and answers with a list of effects: ``Reject``, ``Log``,
``NoteStream``, ``StoreTalkerAlias``, ``Deliver``, ``RequestVersion``. It reads
no configuration, opens no socket and calls no logger. The adapter keeps the
three things that are genuinely I/O — verify the MAC, build the policy, carry
out the effects in order — and the handler goes from ~180 lines of nested
branches to ~45 of dispatch.
Two things this buys beyond the shape. A datagram can be replayed through the
engine anywhere: a test, a laptop, a capture from a sysop, with no reactor in
sight. And every refusal carries a ``reason``, now tallied per bridge in the
session and printed by the keepalive loop at debug level — "my call does not
cross" is answered by ``tg-filter-mcc=12, system-sub-acl=3`` instead of by
grepping the log.
No behaviour change intended, and this time checked two ways. The differential
harness ran 9594 frames through this tree and through upstream ``develop``:
identical delivery, quench, egress address and log lines. And the corpus in
``tests/fixtures/obp_ingress_effects.jsonl`` — 96 recorded cases covering the
talkgroup filters, both ACL scopes, the bits byte, the DMRE envelope (age,
hops, source server) and BCKA/BCSQ/BCST from three addresses — was recorded
from the pre-engine handler, verified frame by frame against develop, and still
passes untouched. It stays as the contract for whatever comes next; regenerate
with CAPTURE=1 and read the diff.
Tests: 24 new unit tests for the engine (100% of the module), the recorded
corpus as a regression net, 998 passed, 2 skipped (the 2 failures are this
machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 2 of separating the OpenBridge path from its hblink ancestry. Phase 1
lifted the admission rules out of the adapter; this one takes the state they
were reading.
Legacy hblink kept what a bridge learns at runtime inside its own SYSTEMS
block: ``_bcka`` (last keepalive), ``_bcsq`` (the peer's quench table),
``_STUN`` and, worst of the four, ``TARGET_IP``/``TARGET_PORT``/``TARGET_SOCK``,
rewritten in place every time RELAX_CHECKS accepted a datagram from an address
the operator never wrote. Configuration and session state shared one mutable
dict, so a single unexpected packet could move a bridge's target for good, and
no reader could tell what came from the YAML and what came from the wire.
``domain/mesh_session.py`` now holds an ``ObpBridgeSession`` per link:
configured_peer (from the YAML, never moves) beside learned_peer (from the
wire), the last keepalive, the quench table and the BCST stun flag, each behind
the question a caller actually asks — ``peer``, ``keepalive_seen``,
``keepalive_stale(now)``, ``quenches(tg, stream)``. The store lives under a
private top-level config key next to ``_SUB_MAP`` and ``_PEER_IDS``, so every
layer that already receives the config reaches the same instance; like
``_SUB_MAP`` it is shared, not deep-copied, across a SIGHUP, and
``sync(config)`` then refreshes the configured peers and drops the sessions of
links that are gone. Editing TARGET_IP in the YAML and reloading is now a
documented way to undo a bad learned address.
Migrated readers: the keepalive gate in routing (to_target and unit data), the
60s keepalive report loop, the monitor/MQTT dashboard blocks and the quench
check and purge in the routing timers. The SYSTEMS blocks of an OPENBRIDGE
system are no longer written to at runtime.
Also removed: ``_config.pop("_no_target_log_time")``, a key nothing has written
since the port from hblink.
No behaviour change intended. The differential harness from phase 1, extended
to compare where egress actually goes (a keepalive and a voice frame sent after
every case) and to cover BCKA/BCSQ/BCST from three source addresses with
RELAX_CHECKS on and off, ran 9594 frames through this commit and through
develop: identical delivery, quench, egress address and log lines.
Tests: 24 new unit tests for the session, 98% coverage of the new module; the
RELAX_CHECKS sync test now asserts what the refactor is for — the session
follows the peer, the configured address stays put. Full suite 876 passed,
2 skipped (the 2 failures are this machine's, and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
``_obp_datagram_received`` decided admission inline: ~110 lines of nested ifs
repeating the same four-step shape (evaluate, log once per stream_id, quench,
return) eight times for the ACLs alone, reading GLOBAL and SYSTEMS state at six
different depths, and doing it twice over — once for DMRD v1 and once for
DMRE v5. Nothing in it could be exercised without a reactor and a socket.
The decisions now live in ``domain/mesh_admission.py`` as pure functions. Each
rule takes values and returns a ``Rejection`` (reason, log line, whether to
quench, whether to log once) or ``None`` to admit; ``admit_dmrd_v1`` and
``admit_dmre_v5`` chain them in the order the legacy handler applied them. The
adapter keeps what is genuinely I/O: one ``_obp_admission_context()`` that reads
config once per frame, and one ``_obp_reject()`` that applies the outcome.
No behaviour change intended. Beyond the suite, a differential harness fed
9576 cryptographically valid frames (DMRD v1 and DMRE v5, sweeping destination
TG, subscriber, bits byte, STUN, both ACL scopes, hop count, packet age, source
server and VALIDATE_SERVER_IDS) through a protocol built from this commit and
one built from develop, comparing delivery, quench and every log line: no
behavioural difference.
One deliberate deviation: the four DMRE "GLOBAL TG FILTER (local to ...)" log
calls on develop pass three arguments to a message with two placeholders, so
Python logs a formatting error instead of the drop (648 of the 9576 frames hit
this). The message now carries the talkgroup. ``test_every_rejection_can_be_
formatted`` keeps the whole family honest.
Also fixed by construction: ``reason`` gives each drop a stable handle, so a
later engine can count or trace drops without parsing log text.
Tests: 59 new unit tests, 100% coverage of the new module; full suite 849
passed, 2 skipped (the 2 failures are this machine's: kernel rmem_max and the
version installed in the venv, both failing on develop as well).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review follow-up on the fan-in demux fixes.
1. apply_obp_proxy_config_reload() rebuilds the bridge registry, and a rebuild
creates fresh ObpIngressReplyTransport objects. Nothing re-pinned them to the
legacy sockets that stay bound, so the first `systemctl reload` put egress
back on the shared fan-in port and reintroduced the bug the pin fixes.
Bound legacy transports are now kept in the service state and re-pinned after
every registry rebuild (pin_legacy_egress()). For the same reason the rebuild
now describes the ports that are actually bound (the running LISTEN_PORT /
BIND_LEGACY_PORTS) instead of the incoming ones, which are only logged as
"restart required".
2. Control-frame demux ranked bridges by the live SYSTEMS.<name>["TARGET_SOCK"],
the very dict RELAX_CHECKS rewrites in place when a bridge accepts traffic
from an unexpected source. One misattributed frame could therefore move a
bridge's target and then keep matching it there. Each entry now carries
peer_hint, the peer as configured, snapshotted when the registry is built
(startup and every reload, both of which read a freshly normalized config).
Order is: configured IP:port, live IP:port, configured IP, live IP, then
registration order — so a peer that legitimately moved is still recognised,
but never at the expense of the bridge that has that address in the YAML.
Tests
- test_reload_keeps_egress_pinned_to_the_bridge_socket: starts the service on a
fake reactor, reloads twice, asserts egress still leaves through each bridge's
own socket and nothing goes out of the fan-in port.
- test_control_frame_prefers_configured_peer_over_relaxed_target: a bridge
dragged onto another bridge's address does not steal that peer's keepalives.
- test_control_frame_unknown_peer_falls_back_to_passphrase now asserts the
deterministic outcome (first registered bridge) instead of "either one".
Both new tests fail without this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both listeners bind independent UDP sockets; if both are configured to
the same port only the first one to start would succeed, and the
server would fail at bind time with an unclear error instead of a
config validation message.
Three issues seen on a production ADN Systems master (2.5.6) with five
OpenBridge links, all rooted in the shared fan-in port.
1. YamlConfigLoader rebuilt the config from a fixed list of top-level keys
that did not include OBP_PROXY, so the whole block in adn-server.yaml was
silently ignored: ENABLED, LISTEN_PORT, LISTEN_IP and BIND_LEGACY_PORTS
always fell back to their defaults (--doctor reported 62032 for a config
that asked for 62031, and ENABLED: false still bound the fan-in).
2. Control frames (BCKA/BCSQ/BCST/BCVE) carry no NETWORK_ID, so the fan-in
could only tell bridges apart by passphrase. Every ADN Systems bridge
shares one passphrase, so all of them resolved to whichever bridge was
registered first: keepalives from the Spanish and Portuguese peers were
attributed to the French bridge and, with RELAX_CHECKS, rewrote its
TARGET_SOCK ("Source IP and Port has changed ... updating", in a loop).
Voice was unaffected because DMRD/DMRE do carry NETWORK_ID.
Control frames now prefer the bridge whose configured peer matches the
datagram source (exact IP:port, then same IP, then the passphrase scan).
3. ObpIngressReplyTransport fell back to the shared fan-in socket, so egress
left from LISTEN_PORT; remote peers learn that source and answer there,
funnelling everything onto the one port where control frames cannot be
told apart. Egress is now pinned to the bridge's own legacy socket when
one is bound.
Tests: five regression tests; all five fail without these changes.
Unit data to an individual ID (>= 1000000) was copied to every OpenBridge
(VER > 1) even when the destination was just heard on a local system and the
SUB_MAP lookup already delivers it there. A D-APRS gateway's ARS/LRRP answers
(one per minute per radio) went out on every configured bridge, multiplying
traffic and causing remote masters that mis-classify those frames to create a
dynamic talkgroup named after the radio id.
Skip the OBP fan-out when the SUB_MAP entry for the destination is on a
non-OPENBRIDGE system and younger than UNIT_DATA_LOCAL_SUB_MAX_AGE (900s).
Subscribers learned via an OpenBridge, or stale entries, keep the current
behavior, so a radio that moved to another master stays reachable.
Rescued from #72 (opened and self-closed by pyopower without a merge),
adapted to the current file layout.