MeshSessionStore.session refreshed dns_host from the config on every call,
and deciding whether TARGET_IP is a name or an address costs a thrown and
caught ValueError for every bridge configured by name. The OBP ingress
reads self._session three to four times per datagram, so a mesh of eight
named bridges paid that exception several times per frame, per bridge.
Nothing needed it there: the config is read-only at runtime, and sync
already follows a reload. It refreshed configured_peer but not dns_host,
which is the only reason the per-lookup refresh existed, so sync now owns
both and session derives dns_host only for a session it creates.
Measured on the v5 ingress path with a bridge configured by name:
66.2us to 47.9us per frame.
Anchoring compared the full socket, so a peer answering from a source port
other than the one we send to was refused: its name resolved to the right
host, but the port differed and every frame was discarded. NAT rewrites that
port, and a peer needs not bind the port it is reached on.
A name now pins the host only. The wire may still refine the port within
that host, and a re-resolution that moves the peer elsewhere drops a port
learned for the host it just left.
A name that has never resolved anchors nothing. normalize_obp_config leaves
TARGET_SOCK as (None, port) when startup resolution fails, and anchoring on
that refused every source forever, taking the link off the air until the
process restarted. Those bridges fall back to RELAX_CHECKS instead.
A peer on a dynamic IP forces RELAX_CHECKS on, and RELAX_CHECKS meant
"accept from any address on earth". On a shared-passphrase mesh that is
enough for a second host to be taken for the peer: production showed
OBP-USA with two live instances (74.132.44.239, the configured peer, and
129.80.176.29, a stale clone), both authenticating, both sending voice,
keepalives and quenches. The session address flapped between them every
few seconds, so half of what we transmitted went to the wrong host, loss
climbed to 30%, and hop counts escalated until MAX HOPS dropped frames.
The network already had the answer: TARGET_IP was written as a name,
3103.adn.systems, which tracks the dynamic IP by DNS. normalize_obp_config
resolved it once and then overwrote TARGET_IP with the address, losing the
name, so nothing could ever ask again — while PEER systems have kept
_MASTER_IP and re-resolved through reactor.resolve() all along.
OPENBRIDGE now gets the same treatment:
- normalize_obp_config keeps the name in _TARGET_IP, as PEER keeps
_MASTER_IP.
- A session whose TARGET_IP was a name is DNS-anchored: learn_peer()
refuses to move it, and only adopt_resolved() can.
- accepts_source() stops widening to "anywhere" for an anchored bridge.
RELAX_CHECKS keeps its meaning for bridges configured with an address.
- Control frames get that same check. They had none at all, which is how a
foreign BCSQ could quench a live stream and a foreign BCST could STUN the
bridge outright.
- A frame from elsewhere is refused and schedules a re-resolution, rate
limited so unknown traffic cannot drive a lookup per packet. If the name
now answers with that address, the peer migrates and the next frame is
accepted; a periodic loop keeps it fresh while the link is idle. A
resolver failure keeps the address we have.
Verified on the 213 master: the clone's frames are refused and logged once,
the flapping is gone, and a real call bridged to eight systems at 0.37%
loss and 6 hops.
Phase 3, and the end of the hblink shape in this path. Phase 1 lifted the
admission rules out of the adapter, phase 2 took the state they read; what was
left in ``_obp_datagram_received`` was the plumbing between them — two long
branches that decoded, decided, logged, quenched and called into routing, all
inside a Twisted ``DatagramProtocol`` where none of it could be run on its own.
``domain/mesh_engine.py`` now takes a decoded frame, the link's session and a
``BridgePolicy``, and answers with a list of effects: ``Reject``, ``Log``,
``NoteStream``, ``StoreTalkerAlias``, ``Deliver``, ``RequestVersion``. It reads
no configuration, opens no socket and calls no logger. The adapter keeps the
three things that are genuinely I/O — verify the MAC, build the policy, carry
out the effects in order — and the handler goes from ~180 lines of nested
branches to ~45 of dispatch.
Two things this buys beyond the shape. A datagram can be replayed through the
engine anywhere: a test, a laptop, a capture from a sysop, with no reactor in
sight. And every refusal carries a ``reason``, now tallied per bridge in the
session and printed by the keepalive loop at debug level — "my call does not
cross" is answered by ``tg-filter-mcc=12, system-sub-acl=3`` instead of by
grepping the log.
No behaviour change intended, and this time checked two ways. The differential
harness ran 9594 frames through this tree and through upstream ``develop``:
identical delivery, quench, egress address and log lines. And the corpus in
``tests/fixtures/obp_ingress_effects.jsonl`` — 96 recorded cases covering the
talkgroup filters, both ACL scopes, the bits byte, the DMRE envelope (age,
hops, source server) and BCKA/BCSQ/BCST from three addresses — was recorded
from the pre-engine handler, verified frame by frame against develop, and still
passes untouched. It stays as the contract for whatever comes next; regenerate
with CAPTURE=1 and read the diff.
Tests: 24 new unit tests for the engine (100% of the module), the recorded
corpus as a regression net, 998 passed, 2 skipped (the 2 failures are this
machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 2 of separating the OpenBridge path from its hblink ancestry. Phase 1
lifted the admission rules out of the adapter; this one takes the state they
were reading.
Legacy hblink kept what a bridge learns at runtime inside its own SYSTEMS
block: ``_bcka`` (last keepalive), ``_bcsq`` (the peer's quench table),
``_STUN`` and, worst of the four, ``TARGET_IP``/``TARGET_PORT``/``TARGET_SOCK``,
rewritten in place every time RELAX_CHECKS accepted a datagram from an address
the operator never wrote. Configuration and session state shared one mutable
dict, so a single unexpected packet could move a bridge's target for good, and
no reader could tell what came from the YAML and what came from the wire.
``domain/mesh_session.py`` now holds an ``ObpBridgeSession`` per link:
configured_peer (from the YAML, never moves) beside learned_peer (from the
wire), the last keepalive, the quench table and the BCST stun flag, each behind
the question a caller actually asks — ``peer``, ``keepalive_seen``,
``keepalive_stale(now)``, ``quenches(tg, stream)``. The store lives under a
private top-level config key next to ``_SUB_MAP`` and ``_PEER_IDS``, so every
layer that already receives the config reaches the same instance; like
``_SUB_MAP`` it is shared, not deep-copied, across a SIGHUP, and
``sync(config)`` then refreshes the configured peers and drops the sessions of
links that are gone. Editing TARGET_IP in the YAML and reloading is now a
documented way to undo a bad learned address.
Migrated readers: the keepalive gate in routing (to_target and unit data), the
60s keepalive report loop, the monitor/MQTT dashboard blocks and the quench
check and purge in the routing timers. The SYSTEMS blocks of an OPENBRIDGE
system are no longer written to at runtime.
Also removed: ``_config.pop("_no_target_log_time")``, a key nothing has written
since the port from hblink.
No behaviour change intended. The differential harness from phase 1, extended
to compare where egress actually goes (a keepalive and a voice frame sent after
every case) and to cover BCKA/BCSQ/BCST from three source addresses with
RELAX_CHECKS on and off, ran 9594 frames through this commit and through
develop: identical delivery, quench, egress address and log lines.
Tests: 24 new unit tests for the session, 98% coverage of the new module; the
RELAX_CHECKS sync test now asserts what the refactor is for — the session
follows the peer, the configured address stays put. Full suite 876 passed,
2 skipped (the 2 failures are this machine's, and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
``_obp_datagram_received`` decided admission inline: ~110 lines of nested ifs
repeating the same four-step shape (evaluate, log once per stream_id, quench,
return) eight times for the ACLs alone, reading GLOBAL and SYSTEMS state at six
different depths, and doing it twice over — once for DMRD v1 and once for
DMRE v5. Nothing in it could be exercised without a reactor and a socket.
The decisions now live in ``domain/mesh_admission.py`` as pure functions. Each
rule takes values and returns a ``Rejection`` (reason, log line, whether to
quench, whether to log once) or ``None`` to admit; ``admit_dmrd_v1`` and
``admit_dmre_v5`` chain them in the order the legacy handler applied them. The
adapter keeps what is genuinely I/O: one ``_obp_admission_context()`` that reads
config once per frame, and one ``_obp_reject()`` that applies the outcome.
No behaviour change intended. Beyond the suite, a differential harness fed
9576 cryptographically valid frames (DMRD v1 and DMRE v5, sweeping destination
TG, subscriber, bits byte, STUN, both ACL scopes, hop count, packet age, source
server and VALIDATE_SERVER_IDS) through a protocol built from this commit and
one built from develop, comparing delivery, quench and every log line: no
behavioural difference.
One deliberate deviation: the four DMRE "GLOBAL TG FILTER (local to ...)" log
calls on develop pass three arguments to a message with two placeholders, so
Python logs a formatting error instead of the drop (648 of the 9576 frames hit
this). The message now carries the talkgroup. ``test_every_rejection_can_be_
formatted`` keeps the whole family honest.
Also fixed by construction: ``reason`` gives each drop a stable handle, so a
later engine can count or trace drops without parsing log text.
Tests: 59 new unit tests, 100% coverage of the new module; full suite 849
passed, 2 skipped (the 2 failures are this machine's: kernel rmem_max and the
version installed in the venv, both failing on develop as well).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: accept NUL-padded RPTC callsigns in login check
Fixed-width RPTC fields may be padded with NUL (ipsc2hbp) or spaces
(MMDVM). str.rstrip() left trailing NULs so matching DB callsigns were
rejected with MSTNAK after a successful passphrase exchange.
* fix: strip NUL padding from RPTO OPTIONS in logs and storage
ipsc2hbp pads the fixed-width RPTO body with NULs; logs showed ^@ noise
and peer CALLSIGN as raw bytes on the options line. Normalize on ingest
and in redact_pass_in_options.
* fix: strip NUL padding from monitor export and shared HBP fields
Centralize fixed-width NUL/space trimming in domain helpers and use them
for dashboard_state peer fields, OPTIONS parsing, proxy self-service, and
remaining CALLSIGN log lines.
* feat: poll peer_dynamic_tgs.need_reload for dynamic TG purge
Proxy send_opts and fallback loop apply TG-4000-equivalent reset when the
monitor sets need_reload, with migration 006 and restore filter updates.
* chore: fix import order in voice subscription plan test
Introduce downlink.py as the single authority for hotspot eligibility (OPTIONS,
slot busy, post-TX GROUP_HANGTIME). send_peer and BRDG fan-out share the same
gate; monitor events carry stream id from BRDG so new QSOs after hangtime
display without replaying calls blocked during the hangtime window.
Planned release: 2.2.0