- Parse the subscriber file once per reload, turning each record into a
profile as the decoder finishes it, instead of parsing it twice and
building the whole object tree first.
- Profiles are plain tuples with interned names: a dict each doubled the
memory, and a NamedTuple stays tracked by the garbage collector.
- One copy of subscriber_ids, and the default local subscriber file (the
subscriber file itself) is not parsed again.
As soon as any plugin was loaded, every voice frame built a full
CallLegContext and event (~13 us) and scheduled its own callLater
(~2-7 us), even for plugins that never look at voice.
- A plugin may declare `events` (the event classes it handles). The bus
keeps the union and `wants()` / `wants_any()`; the voice and data
bridges check it before building anything, and emit() only calls
plugins that take that event. Undeclared plugins and internal
handlers still get everything.
- emit_deferred batches: one call_later per reactor tick, same order.
- The group context's per-stream constant part (IDs, mode, proxy flag,
aliases) is resolved once per stream; each event still gets its own
`extra` dict.
- Fix: _plugin_started / _plugin_ended grew by one entry per call
forever; END now drops the START key and ended keys are capped.
Per voice frame, group voice to 6 OBP + a MASTER (48 us with no plugin):
any plugin 66.0 -> 59.4 us; a data-only plugin 49.4 us.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Found on ADN 2131: a MASTER added to adn-server.yaml and loaded with
SIGHUP logged "added system ... listening" but answered no login.
reload_server_config builds the new systems into the config it is
reloading, which the server swaps in only at the end; create_protocol(name)
built the protocol from the live config, where the system was not there
yet. HBPProtocol copied an empty system block: no MODE, so none of the
MASTER state was set up and every packet was ignored. Systems that already
existed are fine: they get apply_system_config(<reloaded config>).
create_protocol now receives the config being reloaded, and
peer_server._create_hbp_protocol builds from it (startup is unchanged).
Regression test with the real HBPProtocol wired as the server wires it
(ConfigProxy, prepare_reload_config, swap): the added MASTER has its MODE
and answers RPTL with RPTACK. It fails without the fix.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Found on a live master (2131): a plugin beacon on TG 213 reached the
OpenBridges with 2 frames out of 115. PluginIngress marked the ingress
MASTER slot as TX of the stream, while its RX fields still held the last
radio's call. From the second frame on, group_voice_tg_ingress_collision
saw TG 213 active on that slot (our own TX mark) with a stream and a
source that were not ours (the stale RX ones), took it for a QSO and
suppressed the uplink: only the MASTER's own hotspots heard the rest.
A hotspot frame doesn't have this: udp_hbp records the slot's RX state
(RX_STREAM_ID, RX_LC, RX_RFS, RX_TGID, RX_TYPE, RX_TIME…) once routing
accepts it, so the next frame is the same stream, not a new one.
PluginIngress now does exactly that instead of the TX mark. The slot
still reads busy for routed voice while the stream plays (its RX leg),
a radio or another stream still makes the next frame fail, and the
terminator or the 1 s silence sweep sets RX_TYPE back to VTERM.
Regression test: stale RX state on the slot, a whole beacon forwarded,
header and embedded LC both the beacon's (the LC itself was never
affected: routing builds each leg's LC from the target TG and rf_src).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
They are now the voice-announcements plugin (previous commit). Removed
from the core:
- VoiceUseCases: scheduled_announcement, scheduled_tts_announcement, the
broadcast queue, target and slot selection, slot marking, legacy
monitor reports and apply_voice_config (~900 lines). What stays:
voice ident, on-demand 999x files and the disconnected prompt, which
answer a radio rather than follow a schedule.
- inject_announcement_ptt and its bootstrap wiring; plugin frames enter
through inject_plugin_dmrd.
- The TTS engine moves into the plugin (plugin/infrastructure), with its
tests; VoiceProvider.ensure_tts_ambe is gone; tts_ambe.py was a dead
stub.
Nothing to change for sysops:
- the plugin is versioned and enabled by default, and still reads
VOICE.ANNOUNCEMENTS / TTS_ANNOUNCEMENTS from adn-voice.yaml;
- without a PLUGINS.send entry the server grants voice-announcements
exactly the TGs and DMR IDs those items use (disabled ones included,
so enabling one needs no restart); an explicit entry wins;
- the manager now always hands voice_slot_for_tg with send_dmrd; both
check the talkgroup grant on every call.
Tests: the announcement tests of the core are replaced by the plugin's
(schedule, per-TG queue, busy slots, QSO cut, hourly, TTS and its
failures, default DMR ID, derived grant) and by routing tests of plugin
voice ported from test_announcement_ptt_inject (OBP fan-out, no
misattribution to a real peer, dynamic/bridged slot choice, no
pre-armed bridge needed).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The scheduled announcements and TTS announcements of VoiceUseCases, as
an official plugin that speaks through send_dmrd. Same configuration
(VOICE.ANNOUNCEMENTS / TTS_ANNOUNCEMENTS in adn-voice.yaml, followed
every 15 s) and the same behaviour: interval or hourly, one playback
per TG at a time 1.5 s apart, retries while the TG or every slot is
busy, 58 ms per frame, stop on a refused frame, TTS encoded in a
thread. Versioned next to plugins/example.
Core side, found by an end-to-end run of the plugin:
- PluginIngress ends a plugin voice stream that has been silent for a
second without a terminator (frees the slot, monitor END);
- the default max_frames_per_s adds 18 frames/s per group voice TG, so
one stream per TG fits (a voice stream is ~17 frames/s).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What a voice plugin needs from the core that it can't see itself:
- ServerContext.voice_slot_for_tg(tg): the MASTER slot to speak a TG on
now, or None while every slot is busy. Same choice announcements make
today (dynamic UA session or active bridge leg of the TG on that
MASTER first, then TS2, then TS1), ported to PluginIngress so the
announcement code can leave the core. Only for granted talkgroups.
- PluginIngress reports GROUP VOICE START/END,TX on the MASTER a plugin
stream plays on, the line announcements send today: its hotspots hear
the stream but no bridge leg reports it. A stream cut by a radio, or
never terminated, still gets its END.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review of #104: master_kill and revoking PLUGINS.send did not take
effect on SIGHUP, because merge_top_level_config never re-read PLUGINS.
Going through the real path showed it goes further: YamlConfigLoader
builds the config from a fixed list of sections and dropped PLUGINS
altogether, so the documented block (directory, master_kill, overrides)
was never read, at startup either. Both predate this PR.
- YamlConfigLoader keeps PLUGINS when present (like OBP_PROXY).
- merge_top_level_config replaces PLUGINS on every reload, absent
included: removing the block removes everything in it.
- Tests through prepare_reload_config + merge_top_level_config +
swap_runtime_config and a ConfigProxy, as the server wires them:
entry removed, master_kill, whole section removed, a TG granted; and
one with the real loader on a real adn-server.yaml, boot and reload.
Also from the review:
- parse_dmrd_header uses call_attributes(); PluginIngress uses
server_id_bytes().
- A max_frames_per_s that is not a finite positive number grants nothing.
- The allowlist is documented as a guard against buggy plugins, not a
sandbox: a plugin runs in-process with the live config.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Extract a voice burst's embedded LC from the 5 bytes that hold it
instead of running the full voice() decode.
- Once a stream's alias has decoded, later bursts only refresh its
expiry instead of reassembling it again.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
First step to move scheduled announcements, TTS and voice beacons out of
the core into a plugin: a plugin granted `group_voice_tgs` in
PLUGINS.send can send group voice on those talkgroups.
- PluginIngress (application layer) enters plugin frames on the
announcement MASTER. Group voice goes through dmrd_received with
synthetic_announcement=True, exactly the path announcements use (the
TG's bridges, OpenBridge included), then to that MASTER's hotspots.
- While a plugin stream plays it holds the MASTER slot (TX_TYPE=VHEAD,
TX_STREAM_ID, TX_RFS, TX_TGID), so routed voice finds it busy; a radio
or another stream on the slot fails the frame; VTERM frees it.
- send_dmrd called on the reactor thread returns the routing result, so
a plugin knows when to stop; from a worker thread it is queued.
- Private voice stays refused. Unit data is unchanged (local only).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ServerContext.send_dmrd(pkt) hands one DMRD frame to the routing core,
through the same synthetic ingress path scheduled announcements use
(inject_plugin_dmrd -> dmrd_received on the announcement MASTER, with
the SERVER_ID as peer). Only plugins listed in PLUGINS.send get it.
Guards (PluginDmrdSender, re-read from the live config on every frame):
- allowed_src_ids: a plugin can't send as a radio; required.
- max_frames_per_s: per-plugin token bucket, starts full.
- unit data only in this version (data header, rate 1/2, 3/4, CSBK).
- master_kill or removing the entry stops sending at once.
Routing of plugin frames (dmrd_received plugin_origin):
- delivered by the unit data path only: SUB_MAP / hotspot peer ID, to
the destination's exact hotspot, also on the ingress MASTER itself;
- never through the private call path, which would learn the plugin's
source in SUB_MAP (spreading replies over every hotspot of that
MASTER) and keep call state on its shared slot;
- no OpenBridge or DATA-GATEWAY fan-out;
- plugin events for them carry is_synthetic=True.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The state a prompt shares between its thread and the reactor is a
small dataclass (_PromptRun) instead of a dict with string keys.
- "Is the slot held by another stream?" now lives in routing/helpers.py
next to slot_has_active_voice, and both read the RX/TX legs through
one helper, instead of a third copy of that check in voice_use_cases.
- Blank lines and the test header follow the rest of the repo.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Voice ident, the 9991-9999 on-demand files and the disconnected prompt
run on a worker thread and sent every frame through send_voice_packet,
which refreshes TX_TIME but never TX_TYPE. The slot stayed at VTERM for
the whole prompt, so routing saw it free: on a system with
GROUP_HANGTIME 0 routed group voice went out as a second stream on the
same slot, and a radio keying up did not stop the prompt either.
Scheduled and TTS broadcasts already get this right. The three prompts
now share play_on_slot: the thread only paces the frames, and each one
is sent on the reactor, which marks the slot the way broadcasts do
(TX_TYPE=VHEAD, TX_STREAM_ID, TX_RFS), stops the prompt when a radio is
talking or another stream holds the slot, and frees it at the end if
the prompt still holds it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- State the invariant the plan cache relies on: from a SYSTEMS block the
plan reads only MODE, which only a reload changes, so block identity is
enough; anything else it reads must join the cache key.
- Build the plan's ingress with only what resolve() routes on, instead of
passing empty peer, radio and stream ids through _voice_forward_plan.
- SubscriptionStore.has_table is abstract; the scan default is gone.
- Both caches are created in RoutingUseCases.__init__, and a full cache
drops its oldest entry instead of all of them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A new forward leg gets its header, terminator and embedded LC encoded
(BPTC 196,96 + RS, about 250us). Every leg of a call to the same TG gets
the same LC, so a call bridged to six OpenBridges and a MASTER spent
1.75ms on its first frame doing the same work seven times. The encoded
set is now kept per LC (bounded) and shared by the legs, which only ever
slice it.
- In-band signalling runs on each voice header and terminator and walked
every subscription of the server to find those of the source system. The
store now indexes legs by system, in the order a full scan returns them.
- The remaining full scans used to test or drop a relay table, and the copy
of a whole table just to test it is not empty, use the table index.
Short calls (50 frames), one call bridged to 6 OpenBridges and a MASTER,
200 other bridged TGs in the store, per frame, after the previous commit:
HBP ingress 136.4us -> 97.9us
OBP ingress 98.1us -> 60.7us
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every group voice frame, from HBP or OpenBridge, asked store_has_table()
whether its TG has a relay table. It answered by building a tuple of every
subscription and calling table_key() on each, so the cost of a frame grew
with the size of the network. On a live master (ADN 213, six OpenBridges)
this was 26% of the CPU spent on OpenBridge ingress.
- has_table() on the store answers from the table index it already keeps;
the port gets a default that scans, for stores without one.
- The forward plan of a group call (relay tables, resolved legs, the
MASTER/PEER dedupe and the target entries) depends only on the store and
on the source and target SYSTEMS blocks. It is now kept per
(system, slot, TG) and rebuilt when the store's new revision counter moves
or a reload replaces one of those blocks. ENABLED, quench, keepalive and
contention are still checked per frame in the loop.
- One OBP session lookup per target instead of two.
- store_has_table is imported once, not on every frame.
Per frame, one call bridged to 6 OpenBridges and a MASTER, with N other
bridged TGs in the store:
N=0 N=200
HBP ingress 81.8us -> 52.1us 447.4us -> 55.3us
OBP ingress 95.9us -> 66.8us 469.3us -> 65.7us
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follow the .tmp.{pid}.{thread} convention of alias_loader.py plus a
per-process counter: the SIGHUP handler runs on the main thread, the same
one the trimmer saves from, so pid and thread alone could still collide.
A failed pickle.dump no longer leaves an orphan temp file behind.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
It runs once per datagram, for OBP and local traffic alike, and built a tuple
of every subscription to answer a yes/no question. Profiling a production
server at 198 datagrams/s put 67.8% of all CPU work inside it.
legs_in_table already reads the _by_table index and is on the port. Both were
added in the same commit as the scan, which never used them.
Constant time now instead of growing with the mesh: 6.8us to 0.30us at 120
subscriptions, 26.2us to 0.29us at 2400. On the server, 67.8% of work down to
0.84% and 42% less CPU at equal load.
Rule timers, in-band signalling and resets change a subscription in place
(sub.state.phase = IDLE) and then upsert the same object. upsert unindexed
"the old one" by reading its state, but the old one is that same object, now
IDLE, so the active indexes were never cleared: relay_tables_with_active_source
kept listing the leg as an active source and has_active_target_leg stayed true.
The visible effect: when a timed-out rule deactivates a hotspot's leg in the
middle of a call, the rest of the call is still forwarded, where the legacy
router checked ACTIVE on the source row for every frame and stopped. The
downlink hang logic also kept treating the system as having that leg.
The store now records what it indexed each leg under and removes exactly
that. Tests cover a leg deactivated in place, one activated in place next to
another active leg on the same key, a remove after an in-place change, and
the timed-out source end to end. All four fail on develop.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SUB_MAP was written once an hour. The shutdown save only runs if SIGTERM
reaches the server, which it does not under a Docker entrypoint that stays
PID 1, so a restart dropped up to an hour of routes. Private calls to those
units then go nowhere (FORWARD: []) until each radio transmits again.
The loop now ticks every 300s and writes only when a route (system, slot,
peer) changed or an entry's timestamp moved to a new hour, so the 24h trim
still sees fresh times after a restart. An idle master does not write at all.
The pickle is written to a .tmp and renamed. A crash mid-dump used to leave a
truncated file, which load() reads as an empty map.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
_master_datagram_received and _peer_datagram_received each carried the same
32 lines: global SUB_ACL, the TG list for the frame's slot, then the system's
own three, every rejection logging once per stream through _laststrid. The two
copies were identical down to the wording of the six log messages, so a fix in
one could silently miss the other.
They now call _acl_rejects_dmrd, which keeps the legacy order, the same
messages and the same once-per-stream guard, and reads the TG list as
TG{slot}_ACL instead of spelling out slot 1 and slot 2. The frame's slot is
always 1 or 2 (it comes from bit 0x80), so no branch is lost.
tests/hbp/test_acl_gate.py covers both paths, the per-slot TG list and the
on-demand unit-service exemption; three of its four tests pass unchanged
against the previous code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The MASTER ingress path asks the same two questions of every voice frame.
peer_single_mode parses the peer's OPTIONS blob through
parse_peer_options_fields, and peer_rf_mode falls back to derive_peer_rf_mode
because nothing ever wrote the RF_MODE key its docstring anticipates. Around a
static-TG hotspot the downlink helpers add five more parse_peer_options_static
calls on the same blob, and each of those parses it twice
(peer_options_static_valid, then _parse_options_kv). A hotspot with
TS2_1=214;TS1_1=91;SINGLE=1;TIMER=15; had that string re-parsed more than ten
times per frame.
None of it moves at frame rate: OPTIONS only changes on RPTO, which already
drops the memo added for cached_peer_static_tgs, and SLOTS with the two
frequencies only arrive in RPTC. So peer_options_fields memoizes against the
blob the same way, the five remaining parse_peer_options_static call sites go
through cached_peer_static_tgs, and RPTC classifies the RF mode once at login.
Measured on the MASTER ingress path (tests/support/hbp_repeat_stack, 20k group
voice frames from a static-TG hotspot, ACLs on): 138.8us -> 39.1us per frame,
or 42.0us for a peer that has not sent RPTC yet and still derives its mode.
Emitted packets are byte-identical across 72 combinations of OPTIONS blobs
(including invalid ones and PASS=) and simplex/duplex peers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
STATUS is keyed by stream_id alone and the trimmer only drops a row after
180s idle, so when a peer reuses an id for another destination the second
call lands on the first one's row, fails the TGID check in loop control and
is discarded frame by frame. to_target already evicts a stale row on a
forward leg; ingress did not.
Ingress now evicts it too, once the old row has been idle for 2s so two
genuinely interleaved streams keep their own rows.
Also drops the bogus TS from the loop-control warning (it printed a slot
index from an unrelated loop), gives that branch the once-per-stream guard
its siblings have, and says what the condition is.
Seen on a production bridge: one talkgroup took 116 frames and completed no
calls at all.
_load_server_tsv_with_backup was a near-copy of _load_id_dict_with_backup:
same verify, same fallback, same backup, differing only in the parser and in
having "server_ids" written into its log messages. #94 added the skip-if-
unchanged path to one of them, so server_ids.tsv stayed the one alias file
re-read on every 15 minute tick.
They are one function now, taking the parser as an argument. Removes 51
duplicated lines and leaves no second path to drift.
The reload loop ticks every 900s so a failed download is retried soon, but
STALE_DAYS replaces the files about once a day. Every tick in between
re-read them: blake2b over 50MB, a JSON parse, a 50MB copy to .bak, and a
300k-entry profile rebuild, all to produce the same dicts.
Parsed results are now kept against each file's (mtime_ns, size, inode) and
the backup is written only when the primary has been verified, which is the
point where it is worth keeping as a fallback. A primary that fails drops
its remembered parse so the next tick looks at the file again.
The profile build also ran in merge_reload_into_config, on the reactor
thread, at over a second per cycle. It is built in the thread pool now and
handed in.
A tick with nothing new: 2568ms -> 0.2ms.
- Build BridgePolicy/AdmissionContext and PeerMeshConfig once, not per frame.
Dropped on reload, and on the _SERVER_IDS swap an alias refresh does.
- Carry the DMRE timestamp on MeshIngress: the trailer was parsed twice.
- Let the engine be the only one vetting the source; it already answers None.
- Resolve the session and its peer once per datagram instead of 3-5 times.
- Dispatch effects most-frequent-first.
47.1us -> 28.6us per frame on the v5 path. Recorded effects corpus unchanged.
MeshSessionStore.session refreshed dns_host from the config on every call,
and deciding whether TARGET_IP is a name or an address costs a thrown and
caught ValueError for every bridge configured by name. The OBP ingress
reads self._session three to four times per datagram, so a mesh of eight
named bridges paid that exception several times per frame, per bridge.
Nothing needed it there: the config is read-only at runtime, and sync
already follows a reload. It refreshed configured_peer but not dns_host,
which is the only reason the per-lookup refresh existed, so sync now owns
both and session derives dns_host only for a session it creates.
Measured on the v5 ingress path with a bridge configured by name:
66.2us to 47.9us per frame.
The password table was re-read and re-decrypted on a 10s timer, while the
file it reads only ever changes when the security downloader replaces it,
every 300s. The downloader already reloads the table on a successful
download, so 29 of every 30 decrypt cycles produced an identical result.
Each cycle also read the encryption key from disk once per entry, because
decrypt_password builds its own Fernet: 781 entries meant 781 stats, 781
reads of the same key file and 781 Fernet constructions, about 78 key
reads a second sustained.
Dropping the timer and decrypting a table on one Fernet takes this from
~23,400 key reads per five minutes to one, measured at 85.5ms per cycle
on a server carrying 781 entries.
The loader now keeps no config of its own: it was held only to let the
expired timer reload itself.
Anchoring compared the full socket, so a peer answering from a source port
other than the one we send to was refused: its name resolved to the right
host, but the port differed and every frame was discarded. NAT rewrites that
port, and a peer needs not bind the port it is reached on.
A name now pins the host only. The wire may still refine the port within
that host, and a re-resolution that moves the peer elsewhere drops a port
learned for the host it just left.
A name that has never resolved anchors nothing. normalize_obp_config leaves
TARGET_SOCK as (None, port) when startup resolution fails, and anchoring on
that refused every source forever, taking the link off the air until the
process restarted. Those bridges fall back to RELAX_CHECKS instead.
Control frames carry no NETWORK_ID, so three things can tell two OPENBRIDGE
bridges apart: a legacy port of their own, a passphrase of their own, or a
source address that matches what one of them is configured with. Any one is
enough. Sharing the fan-in port and a passphrase leaves only the address, and
a frame from an address none of them knows then goes to whichever bridge was
registered first — the misattribution behind #79 and #87.
Nothing said so. The validator checks duplicate NETWORK_IDs and duplicate
legacy ports, but treats PASSPHRASE as just another string.
openbridge_passphrase_collisions() groups the enabled OPENBRIDGE systems that
share one, and returns the names only, never the secret. It surfaces as a warn
finding in --doctor and as a line at startup next to the fan-in summary.
Kept advisory on purpose: a shared passphrase works as long as every peer
address is distinct and current, so refusing to start would stop servers that
are fine today.
The refusal log is capped to once per source address so a clone pinging
every 10s cannot flood the log. That left an operator with no way to tell
whether it was still happening: the first line scrolls away and nothing
replaces it, while the fan-in still prints "RX b'BCKA' from <clone> ->
OBP-USA", which reads as if the frame had been delivered.
Refusing skips the engine, and the engine is what calls count_drop() for
every other reason, so these frames were tallied nowhere either — not in
the periodic "(ROUTER) system X refused frames: ..." line, not in the
report. Count them as source-not-peer, so the tally answers "is the clone
still knocking?" once a minute without repeating the explanation.
A peer on a dynamic IP forces RELAX_CHECKS on, and RELAX_CHECKS meant
"accept from any address on earth". On a shared-passphrase mesh that is
enough for a second host to be taken for the peer: production showed
OBP-USA with two live instances (74.132.44.239, the configured peer, and
129.80.176.29, a stale clone), both authenticating, both sending voice,
keepalives and quenches. The session address flapped between them every
few seconds, so half of what we transmitted went to the wrong host, loss
climbed to 30%, and hop counts escalated until MAX HOPS dropped frames.
The network already had the answer: TARGET_IP was written as a name,
3103.adn.systems, which tracks the dynamic IP by DNS. normalize_obp_config
resolved it once and then overwrote TARGET_IP with the address, losing the
name, so nothing could ever ask again — while PEER systems have kept
_MASTER_IP and re-resolved through reactor.resolve() all along.
OPENBRIDGE now gets the same treatment:
- normalize_obp_config keeps the name in _TARGET_IP, as PEER keeps
_MASTER_IP.
- A session whose TARGET_IP was a name is DNS-anchored: learn_peer()
refuses to move it, and only adopt_resolved() can.
- accepts_source() stops widening to "anywhere" for an anchored bridge.
RELAX_CHECKS keeps its meaning for bridges configured with an address.
- Control frames get that same check. They had none at all, which is how a
foreign BCSQ could quench a live stream and a foreign BCST could STUN the
bridge outright.
- A frame from elsewhere is refused and schedules a re-resolution, rate
limited so unknown traffic cannot drive a lookup per packet. If the name
now answers with that address, the peer migrates and the next frame is
accepted; a periodic loop keeps it fresh while the link is idle. A
resolver failure keeps the address we have.
Verified on the 213 master: the clone's frames are refused and logged once,
the flapping is gone, and a real call bridged to eight systems at 0.37%
loss and 6 hops.
BCKA carries no NETWORK_ID, so on a shared-passphrase mesh anyone's
keepalive verifies against any bridge. It was moving session.peer with no
gate at all — not even RELAX_CHECKS, which the recorded corpus shows:
"bcka from 9.9.9.9:62201 relax=False" moved egress to 9.9.9.9. Since
session.peer is where voice and control are sent, a second instance of a
peer (seen in production on OBP-USA: two hosts, two 10s keepalive timers)
took the traffic over every few seconds.
A keepalive now only confirms liveness. It may still bootstrap a bridge
that has no address yet (inbound-only, no TARGET_IP), since there is
nothing to steal there and it is the only way such a bridge learns where
to answer. Relocation is left to DMRD/DMRE, which identify themselves.
Also: the fan-in demux ranked bridges by sys_cfg's TARGET_SOCK, frozen
since #81 moved runtime state into the session, so the "live" ranks were
dead code and a peer that really moved was no longer recognised. It now
reads learned_peer from the session store, closing the integration #81
left pending.
A peer already tells the server where it is in its RPTC login, but the
report never passed those fields on, so a monitor could only show the
free-text LOCATION. Add latitude, longitude and height to the peer rows
of the report payloads (topology and dashboard_state), next to the
other RPTC display fields.
This is what a dashboard needs to place repeaters and hotspots on a
map; the legacy adn-dmr-server already had them in CONFIG. Peers that
do not send coordinates are unaffected: the fields are simply absent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
#82 logged the RX debug line once per stream_id per bridge, but a bridge
with BOTH_SLOTS/several concurrent talkgroups interleaves packets from
multiple calls, so the "last seen" stream flips almost every packet and
the line logs nearly as often as before the fix — confirmed against
production output (three interleaved streams on OBP-CL).
DMRD/DMRE demux by NETWORK_ID, which is never ambiguous, and
*CALL START*/*CALL END* (application/routing_use_cases.py,
application/routing/obp_forward.py) already log once per real call with
better detail (SUB, TGID, TS, duration). Drop the fan-in RX line for voice
frames entirely instead of trying to approximate "once per call" with no
session state; control frames (BCKA/BCSQ/BCST/BCVE) keep logging every
time, since they are rare and that visibility mattered for the #79 fix.
An active call sends a DMRD/DMRE packet roughly every 90-100ms; with
OBP_PROXY.DEBUG on, that logged one RX line per packet, flooding the
log for the whole call. Log once when a stream_id first appears instead.
Non-voice frames (control) are unaffected, they were never the flood.
Two things a real capture showed that synthetic frames could not. An
unfiltered tcpdump carries both directions, and our own egress was being run
through the ingress rules, where it fails the NETWORK_ID check by definition:
10698 frames on one bridge reported as network-id-mismatch that were simply
ours. Direction is now decided by the peer address (both ends of a link
normally share the port number, so the port alone cannot tell), with the
listening port as a fallback for a peer behind NAT; --replay-both-directions
keeps the old behaviour.
And VALIDATE_SERVER_IDS was firing offline against an empty list, because the
server-id table is loaded at runtime and not from the YAML: 10107 frames on
another bridge blamed on source-server-unknown. Offline the check is skipped
unless the table is actually there.
Both were reported against the live ADN 213 master (5 bridges, 30669 real
datagrams in 120s); with them fixed the report says what the master does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The engine from phase 3 takes plain values and answers with effects, so it does
not need the server to run. ``adn-server --replay capture.pcap`` uses that: the
frames from a tcpdump capture go through the same ingress, against the
operator's own adn-server.yaml, and each one comes back with the bridge it
belongs to and either a delivery or the reason it was refused.
12:04:31 82.65.127.86:62201 OBP-FR DMRD v1 2130001 -> 214 delivered
12:04:31 85.241.222.7:62268 OBP-PT DMRE v5 2680015 -> 9 dropped (tg-filter-server) +BCSQ
12:04:32 203.0.113.9:50000 - DMRD v1 unmatched
Nothing is sent and no port is bound, so it runs beside a live master. Which
bridge a frame belongs to is decided on the evidence the server has — the port
it arrived on, then the configured peer, then whoever can verify it — which is
also the answer to "whose keepalive is this?" on a mesh where every bridge
shares one passphrase.
``--system`` narrows it to one link, ``--replay-limit`` stops early and
``--replay-summary`` prints the tally alone.
``infrastructure/pcap.py`` reads classic pcap with no dependencies: both
endiannesses, microsecond and nanosecond timestamps, Ethernet (VLAN tags
included), Linux cooked v1 and v2 (``tcpdump -i any`` writes SLL2, found while
running this against a real capture), raw IP and loopback, IPv4 and IPv6 UDP.
pcapng says which command converts it.
Docs: the OBP proxy page gains a "why did that call not cross" section in both
languages, and its RELAX_CHECKS note now says what phase 2 made true — what the
wire teaches lives in the session, TARGET_IP stays as written.
Tests: 27 new, 97% of the replay module and 87% of the pcap reader, plus an
end-to-end run of the real command against a real YAML. Full suite 1025 passed,
2 skipped (the 2 failures are this machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 3, and the end of the hblink shape in this path. Phase 1 lifted the
admission rules out of the adapter, phase 2 took the state they read; what was
left in ``_obp_datagram_received`` was the plumbing between them — two long
branches that decoded, decided, logged, quenched and called into routing, all
inside a Twisted ``DatagramProtocol`` where none of it could be run on its own.
``domain/mesh_engine.py`` now takes a decoded frame, the link's session and a
``BridgePolicy``, and answers with a list of effects: ``Reject``, ``Log``,
``NoteStream``, ``StoreTalkerAlias``, ``Deliver``, ``RequestVersion``. It reads
no configuration, opens no socket and calls no logger. The adapter keeps the
three things that are genuinely I/O — verify the MAC, build the policy, carry
out the effects in order — and the handler goes from ~180 lines of nested
branches to ~45 of dispatch.
Two things this buys beyond the shape. A datagram can be replayed through the
engine anywhere: a test, a laptop, a capture from a sysop, with no reactor in
sight. And every refusal carries a ``reason``, now tallied per bridge in the
session and printed by the keepalive loop at debug level — "my call does not
cross" is answered by ``tg-filter-mcc=12, system-sub-acl=3`` instead of by
grepping the log.
No behaviour change intended, and this time checked two ways. The differential
harness ran 9594 frames through this tree and through upstream ``develop``:
identical delivery, quench, egress address and log lines. And the corpus in
``tests/fixtures/obp_ingress_effects.jsonl`` — 96 recorded cases covering the
talkgroup filters, both ACL scopes, the bits byte, the DMRE envelope (age,
hops, source server) and BCKA/BCSQ/BCST from three addresses — was recorded
from the pre-engine handler, verified frame by frame against develop, and still
passes untouched. It stays as the contract for whatever comes next; regenerate
with CAPTURE=1 and read the diff.
Tests: 24 new unit tests for the engine (100% of the module), the recorded
corpus as a regression net, 998 passed, 2 skipped (the 2 failures are this
machine's and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>