SUB_MAP was written once an hour. The shutdown save only runs if SIGTERM
reaches the server, which it does not under a Docker entrypoint that stays
PID 1, so a restart dropped up to an hour of routes. Private calls to those
units then go nowhere (FORWARD: []) until each radio transmits again.
The loop now ticks every 300s and writes only when a route (system, slot,
peer) changed or an entry's timestamp moved to a new hour, so the 24h trim
still sees fresh times after a restart. An idle master does not write at all.
The pickle is written to a .tmp and renamed. A crash mid-dump used to leave a
truncated file, which load() reads as an empty map.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reload loop ticks every 900s so a failed download is retried soon, but
STALE_DAYS replaces the files about once a day. Every tick in between
re-read them: blake2b over 50MB, a JSON parse, a 50MB copy to .bak, and a
300k-entry profile rebuild, all to produce the same dicts.
Parsed results are now kept against each file's (mtime_ns, size, inode) and
the backup is written only when the primary has been verified, which is the
point where it is worth keeping as a fallback. A primary that fails drops
its remembered parse so the next tick looks at the file again.
The profile build also ran in merge_reload_into_config, on the reactor
thread, at over a second per cycle. It is built in the thread pool now and
handed in.
A tick with nothing new: 2568ms -> 0.2ms.
Phase 2 of separating the OpenBridge path from its hblink ancestry. Phase 1
lifted the admission rules out of the adapter; this one takes the state they
were reading.
Legacy hblink kept what a bridge learns at runtime inside its own SYSTEMS
block: ``_bcka`` (last keepalive), ``_bcsq`` (the peer's quench table),
``_STUN`` and, worst of the four, ``TARGET_IP``/``TARGET_PORT``/``TARGET_SOCK``,
rewritten in place every time RELAX_CHECKS accepted a datagram from an address
the operator never wrote. Configuration and session state shared one mutable
dict, so a single unexpected packet could move a bridge's target for good, and
no reader could tell what came from the YAML and what came from the wire.
``domain/mesh_session.py`` now holds an ``ObpBridgeSession`` per link:
configured_peer (from the YAML, never moves) beside learned_peer (from the
wire), the last keepalive, the quench table and the BCST stun flag, each behind
the question a caller actually asks — ``peer``, ``keepalive_seen``,
``keepalive_stale(now)``, ``quenches(tg, stream)``. The store lives under a
private top-level config key next to ``_SUB_MAP`` and ``_PEER_IDS``, so every
layer that already receives the config reaches the same instance; like
``_SUB_MAP`` it is shared, not deep-copied, across a SIGHUP, and
``sync(config)`` then refreshes the configured peers and drops the sessions of
links that are gone. Editing TARGET_IP in the YAML and reloading is now a
documented way to undo a bad learned address.
Migrated readers: the keepalive gate in routing (to_target and unit data), the
60s keepalive report loop, the monitor/MQTT dashboard blocks and the quench
check and purge in the routing timers. The SYSTEMS blocks of an OPENBRIDGE
system are no longer written to at runtime.
Also removed: ``_config.pop("_no_target_log_time")``, a key nothing has written
since the port from hblink.
No behaviour change intended. The differential harness from phase 1, extended
to compare where egress actually goes (a keepalive and a voice frame sent after
every case) and to cover BCKA/BCSQ/BCST from three source addresses with
RELAX_CHECKS on and off, ran 9594 frames through this commit and through
develop: identical delivery, quench, egress address and log lines.
Tests: 24 new unit tests for the session, 98% coverage of the new module; the
RELAX_CHECKS sync test now asserts what the refactor is for — the session
follows the peer, the configured address stays put. Full suite 876 passed,
2 skipped (the 2 failures are this machine's, and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The periodic security download ran blocking HTTP on the reactor thread, so a
dead selfcare server stalled RPTPING handling well past PING_TIME * MAX_MISSED
and every peer was timed out and forced to reconnect. Move the security and
alias downloads to the thread pool, bound the DNS and HTTP timeouts below that
budget, and skip a cycle when the previous one is still in flight.
Also drop the redundant interval guard that silently stretched the real retry
period to a multiple of the loop interval, restore the process-wide socket
timeout after a failed DNS lookup, keep existing files and cached passwords
when a download returns an empty or invalid payload, and remove the inherited
50Mb cap on subscriber_ids that production was already close to tripping.
PASS_SECURITY is no longer written to the log.
Use SUB_MAP's stored peer id for delivery (repeat, unit-data, and
pvt_call_received), add Talker Alias support for private calls, and
report the receiving hotspot to the monitor.
Route announcements/TTS through synthetic PTT on the proxy MASTER (SERVER_ID
peer, normal dmrd_received forwarding). Emit START/END TX report events for
inject so monitor fans out to SYSTEM-N; configurable server voice DMR_ID.
* feat: poll peer_dynamic_tgs.need_reload for dynamic TG purge
Proxy send_opts and fallback loop apply TG-4000-equivalent reset when the
monitor sets need_reload, with migration 006 and restore filter updates.
* chore: fix import order in voice subscription plan test
* fix: audit wave 1 server hygiene refactors
Inject call_later into PlaybackUseCases, move voice config mtime watch to
bootstrap, complete SubscriptionStore port methods, and relocate echo
routing seed to the application layer.
* fix: document dashboard_state in report-v2 schema
Add dashboard_state to report-v2.json with a two-master example fixture,
update public protocol docs for HELLO → STATE_SND connect flow, and drop
the unused TOPOLOGY_JSON HELLO feature token.
* fix: add ReportWire contract tests against report-v2 schema
Assert state_frames and bridge_event_frames output validates against
committed example fixtures and the report-v2 JSON schema.
Fix test imports for echo_seed module relocation.
* fix: align OPTIONS static validity checks across routing and report
Delegate subscription_table validation to peer_options_static_valid so empty
OPTIONS is valid and PASS-mixed strings are rejected consistently.
* fix: OBP DMRE source-server validation without ALLOW_UNREG_ID bypass
Port OPENBRIDGE.validate_id lookup for 6-7 digit source servers so OBP
ingress matches legacy production config (VALIDATE_SERVER_IDS=True).
* fix: audit items 11-13 coverage, infra tests, and warning logs
* fix: add tests/fakes shim for application test decoupling
* fix: per-stream OBP bridge TX legs for concurrent MASTER downlink
When two OBP voice streams share the same MASTER timeslot, stop
flip-flopping the flat TX row so per-peer downlink gates stay stable.
* fix: ruff lint in OBP concurrent streams downlink test
Remove dead code after return and unused start_tx_events variable.
* fix: raise UDP SO_RCVBUF on voice listeners (C-LOCAL)
Apply a 4 MB receive buffer on system and proxy UDP sockets to reduce
kernel RcvbufErrors under OBP load; size is configurable via GLOBAL.UDP_RCVBUF.
* fix: exclude byte-identical duplicates from HBP rate counter (A.2)
Check lastData before incrementing the ingress packet counter so compressed
duplicate bursts do not trigger legitimate RATE DROP on call start.
* fix: duplicate-safe TA embed phase and REPEAT VHEAD DMRA (B)
Ignore byte-identical B-E embed bursts so duplicate uplinks do not desync the
TA phase machine, and re-emit DMRA on every VHEAD on the REPEAT path.
* chore: document echo point-to-point and logged_in reconciliation (EN/ES)
Document multi-hotspot echo/service delivery via RX_PEER, the lst_seen
logged_in reconcile loop, and cross-links between user and dev guides.
Introduce downlink.py as the single authority for hotspot eligibility (OPTIONS,
slot busy, post-TX GROUP_HANGTIME). send_peer and BRDG fan-out share the same
gate; monitor events carry stream id from BRDG so new QSOs after hangtime
display without replaying calls blocked during the hangtime window.
Planned release: 2.2.0
* feat: persist peer dynamic TGs in MariaDB across reconnects
Add DATABASE config, async DynamicTgStore, and restore on RPTC so
SINGLE=0/1 dynamics survive hotspot disconnects and server restarts
without blocking the DMRD voice path.
* fix: ensure peer_dynamic_tgs table on server startup
Apply migration 004 idempotently at boot so the server does not depend
on adn-monitor db_bootstrap when peer_dynamic_tgs is missing.
* fix: import DynamicTgEntry for ruff F821 in subscription_table
* fix: log clear MariaDB startup failures to file and stderr
Validate DATABASE at config load and on connect; map common MySQL
errors to actionable messages so the server does not fail silently.
* fix: TG 4000 clears STATUS and bridge legs after dynamic reset
Clear RX slot state to stop RPTO re-seeding cleared sessions, run
in-band 4000 deactivation on inject-only paths, and mark downlink dirty.
* fix: complete TG 4000 reset for monitor and dynamic TG persistence
Emit INGRESS BRDG_EVENT so SINGLE=0 UA chips clear without stuck TX;
wipe all peer dynamic rows from memory and MariaDB on reset. Never store
TG 4000 as a UA session. Require DATABASE only for full peer-server configs.
Normalize all Python sources to the standard ADN copyright block with
complete GPLv3 notice. Add legacy attribution on routing and dmr_utils
ports; drop SemVer wording from changelog and fix an unused test import.
Replace internal bridge terminology with routing (RoutingUseCases, AclRouter,
routing_table export). SubscriptionStore remains runtime authority with O(1)
indexes for router and downlink filters. Fix STATIC TG parity on OPTIONS/RPTO,
parrot in-band edge cases, and remove per-packet routing_table export from the
hot path that caused high CPU under multi-hotspot OBP load.
Require subscription_store in BridgeUseCases and route timer, OPTIONS,
static TG, and OBP mutations through store ops with export-only BRIDGES shim.
Fix dashboard YAML static TG fallback and sole-hotspot monitor remap for
dynamic UA when a bridge leg is active.
Sync subscription store after timer and in-band mutations, add optional
store authority with BRIDGES export shim, wire MeshCodecRegistry into
OBP udp_hbp paths, and extend harness parity tests.
Enforce strict SINGLE=1 RX exclusivity, always filter inject-only downlink by
each peer's own OPTIONS, and push self-service DB options only after PASS=.
Stop merged SYSTEM static TG lists from leaking across hotspots in topology.