perf/alias-memory
develop
master
v2.5.6
v2.5.5
v2.5.4
v2.5.3
v2.5.2
v2.5.1
v2.5.0
v2.4.3
v2.4.2
v2.4.1
v2.4.0
v2.3.3
v2.3.2
v2.3.1
v2.3.0
v2.2.6
v2.2.5
v2.2.4
v2.2.3
v2.2.2
v2.2.1
v2.2.0
v2.1.1
v2.1.0
v2.0.6
v2.0.5
v2.0.4
v2.0.3
v2.0.2
v2.0.1
v2.0.0
v1.0.0
v2.0.0-rc.1
v2.0.0-rc.2
v2.0.0-rc.3
v2.0.0-rc.4
${ noResults }
4 Commits (55faf797dae1102bb2e92fbca1b11cc6b7743fe3)
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
55faf797da |
fix(obp): count frames refused for coming from the wrong address
The refusal log is capped to once per source address so a clone pinging every 10s cannot flood the log. That left an operator with no way to tell whether it was still happening: the first line scrolls away and nothing replaces it, while the fan-in still prints "RX b'BCKA' from <clone> -> OBP-USA", which reads as if the frame had been delivered. Refusing skips the engine, and the engine is what calls count_drop() for every other reason, so these frames were tallied nowhere either — not in the periodic "(ROUTER) system X refused frames: ..." line, not in the report. Count them as source-not-peer, so the tally answers "is the clone still knocking?" once a minute without repeating the explanation. |
1 week ago |
|
|
d5d44b1fde |
feat(obp): anchor a bridge's peer to DNS when TARGET_IP is a hostname
A peer on a dynamic IP forces RELAX_CHECKS on, and RELAX_CHECKS meant "accept from any address on earth". On a shared-passphrase mesh that is enough for a second host to be taken for the peer: production showed OBP-USA with two live instances (74.132.44.239, the configured peer, and 129.80.176.29, a stale clone), both authenticating, both sending voice, keepalives and quenches. The session address flapped between them every few seconds, so half of what we transmitted went to the wrong host, loss climbed to 30%, and hop counts escalated until MAX HOPS dropped frames. The network already had the answer: TARGET_IP was written as a name, 3103.adn.systems, which tracks the dynamic IP by DNS. normalize_obp_config resolved it once and then overwrote TARGET_IP with the address, losing the name, so nothing could ever ask again — while PEER systems have kept _MASTER_IP and re-resolved through reactor.resolve() all along. OPENBRIDGE now gets the same treatment: - normalize_obp_config keeps the name in _TARGET_IP, as PEER keeps _MASTER_IP. - A session whose TARGET_IP was a name is DNS-anchored: learn_peer() refuses to move it, and only adopt_resolved() can. - accepts_source() stops widening to "anywhere" for an anchored bridge. RELAX_CHECKS keeps its meaning for bridges configured with an address. - Control frames get that same check. They had none at all, which is how a foreign BCSQ could quench a live stream and a foreign BCST could STUN the bridge outright. - A frame from elsewhere is refused and schedules a re-resolution, rate limited so unknown traffic cannot drive a lookup per packet. If the name now answers with that address, the peer migrates and the next frame is accepted; a periodic loop keeps it fresh while the link is idle. A resolver failure keeps the address we have. Verified on the 213 master: the clone's frames are refused and logged once, the flapping is gone, and a real call bridged to eight systems at 0.37% loss and 6 hops. |
1 week ago |
|
|
e105d493f0 |
refactor(obp): drive OpenBridge ingress from an engine that answers with effects
Phase 3, and the end of the hblink shape in this path. Phase 1 lifted the admission rules out of the adapter, phase 2 took the state they read; what was left in ``_obp_datagram_received`` was the plumbing between them — two long branches that decoded, decided, logged, quenched and called into routing, all inside a Twisted ``DatagramProtocol`` where none of it could be run on its own. ``domain/mesh_engine.py`` now takes a decoded frame, the link's session and a ``BridgePolicy``, and answers with a list of effects: ``Reject``, ``Log``, ``NoteStream``, ``StoreTalkerAlias``, ``Deliver``, ``RequestVersion``. It reads no configuration, opens no socket and calls no logger. The adapter keeps the three things that are genuinely I/O — verify the MAC, build the policy, carry out the effects in order — and the handler goes from ~180 lines of nested branches to ~45 of dispatch. Two things this buys beyond the shape. A datagram can be replayed through the engine anywhere: a test, a laptop, a capture from a sysop, with no reactor in sight. And every refusal carries a ``reason``, now tallied per bridge in the session and printed by the keepalive loop at debug level — "my call does not cross" is answered by ``tg-filter-mcc=12, system-sub-acl=3`` instead of by grepping the log. No behaviour change intended, and this time checked two ways. The differential harness ran 9594 frames through this tree and through upstream ``develop``: identical delivery, quench, egress address and log lines. And the corpus in ``tests/fixtures/obp_ingress_effects.jsonl`` — 96 recorded cases covering the talkgroup filters, both ACL scopes, the bits byte, the DMRE envelope (age, hops, source server) and BCKA/BCSQ/BCST from three addresses — was recorded from the pre-engine handler, verified frame by frame against develop, and still passes untouched. It stays as the contract for whatever comes next; regenerate with CAPTURE=1 and read the diff. Tests: 24 new unit tests for the engine (100% of the module), the recorded corpus as a regression net, 998 passed, 2 skipped (the 2 failures are this machine's and fail on develop too). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
1 week ago |
|
|
05915a8207 |
refactor(obp): give each bridge a session instead of writing state into its config
Phase 2 of separating the OpenBridge path from its hblink ancestry. Phase 1
lifted the admission rules out of the adapter; this one takes the state they
were reading.
Legacy hblink kept what a bridge learns at runtime inside its own SYSTEMS
block: ``_bcka`` (last keepalive), ``_bcsq`` (the peer's quench table),
``_STUN`` and, worst of the four, ``TARGET_IP``/``TARGET_PORT``/``TARGET_SOCK``,
rewritten in place every time RELAX_CHECKS accepted a datagram from an address
the operator never wrote. Configuration and session state shared one mutable
dict, so a single unexpected packet could move a bridge's target for good, and
no reader could tell what came from the YAML and what came from the wire.
``domain/mesh_session.py`` now holds an ``ObpBridgeSession`` per link:
configured_peer (from the YAML, never moves) beside learned_peer (from the
wire), the last keepalive, the quench table and the BCST stun flag, each behind
the question a caller actually asks — ``peer``, ``keepalive_seen``,
``keepalive_stale(now)``, ``quenches(tg, stream)``. The store lives under a
private top-level config key next to ``_SUB_MAP`` and ``_PEER_IDS``, so every
layer that already receives the config reaches the same instance; like
``_SUB_MAP`` it is shared, not deep-copied, across a SIGHUP, and
``sync(config)`` then refreshes the configured peers and drops the sessions of
links that are gone. Editing TARGET_IP in the YAML and reloading is now a
documented way to undo a bad learned address.
Migrated readers: the keepalive gate in routing (to_target and unit data), the
60s keepalive report loop, the monitor/MQTT dashboard blocks and the quench
check and purge in the routing timers. The SYSTEMS blocks of an OPENBRIDGE
system are no longer written to at runtime.
Also removed: ``_config.pop("_no_target_log_time")``, a key nothing has written
since the port from hblink.
No behaviour change intended. The differential harness from phase 1, extended
to compare where egress actually goes (a keepalive and a voice frame sent after
every case) and to cover BCKA/BCSQ/BCST from three source addresses with
RELAX_CHECKS on and off, ran 9594 frames through this commit and through
develop: identical delivery, quench, egress address and log lines.
Tests: 24 new unit tests for the session, 98% coverage of the new module; the
RELAX_CHECKS sync test now asserts what the refactor is for — the session
follows the peer, the configured address stays put. Full suite 876 passed,
2 skipped (the 2 failures are this machine's, and fail on develop too).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
1 week ago |