docs: simplify performance page for sysops

Rewrite EN/ES performance docs with a short user-facing summary and
reference CPU/RAM gains for OpenBridge-heavy nodes on 2.3.3+.
pull/50/head
Rodrigo Pérez 3 months ago
parent 9623482f92
commit 8e8e09ac6a

@ -1,99 +1,49 @@
# Performance (2.x)
**adn-server 2.x** and **adn-monitor 2.x** include several changes that reduce CPU work and memory footprint compared with **adn-dmr-server** and the old monitor/proxy stack. This page lists **what** improves and **what causes it**.
**adn-server 2.x** and **adn-monitor 2.x** use less CPU and RAM than **adn-dmr-server**
and the old monitor/proxy stack. You do not need to tune anything — the gains come from
how routing, reporting, and the integrated proxy are built.
## At a glance
Pair **server 2.x with monitor 2.x** to get the full benefit on the dashboard side.
| Area | Typical effect | Main cause |
|------|----------------|------------|
| **Voice downlink (inject proxy)** | Lower CPU under busy group traffic | **`PeerDownlinkIndex`** — fan-out to peers that match `(slot, TG)` instead of scanning every connected hotspot per packet |
| **Bridge source lookup** | Faster “am I the ACTIVE source?” | **`SubscriptionStore`** indexes (`relay_tables_with_active_source`) — O(1) by `(system, slot, tgid)` vs scanning table rows |
| **Background CPU** | Fewer wakeups | **Event-driven OPTIONS / static TG** — removed legacy **26 s** `options_config_loop` ([Behaviour and timers](behaviour-and-timers.md)) |
| **Mass peer login** | Less redundant CONFIG traffic | **`ConfigPushThrottle`** — adaptive debounce on CONFIG push to the monitor |
| **Reporting vs voice** | Voice path less blocked by reports | **`BoundedReportQueue`** — coalesced snapshots, bounded drain per tick |
| **Server → monitor wire** | Less serialize/send work | **Report v2** JSON (`routing_table`, `topology`, `voice_event`) instead of periodic full pickle of `CONFIG`/`BRIDGES` ([Report protocol v2](../protocols/report-v2.md)) |
| **Process count (RAM)** | One Python process instead of two | **Integrated `PROXY`** in `adn-server.py` — no separate **adn-proxy** process ([Hotspot proxy](../user-guide/hotspot-proxy.md)) |
| **Monitor RAM / WS load** | Smaller in-memory dashboard state | **Slim `dashboard_state` wire**, `clean_sys_dict`, lighter WebSocket fingerprints ([Monitor architecture](../../monitor/architecture.md)) |
## What improved
## Server: inject-only downlink index
| Improvement | What you get |
|-------------|--------------|
| **Smarter voice routing** | Under busy group traffic the server does less work per packet — especially on proxies with many hotspots and on OpenBridge-heavy nodes. |
| **Integrated hotspot proxy** | One `adn-server` process instead of server + standalone **adn-proxy** — less RAM and simpler ops. See [Hotspot proxy](../user-guide/hotspot-proxy.md). |
| **No legacy 26 s timer** | Static talkgroups refresh on events (startup, config reload, peer OPTIONS), not a background loop every 26 seconds. See [Behaviour and timers](behaviour-and-timers.md). |
| **Lighter monitor link** | **Report v2** sends compact JSON instead of heavy periodic pickle dumps. See [Report protocol v2](../protocols/report-v2.md). |
| **Monitor 2.x** | Slimmer dashboard state, less memory growth on panels left open for days. See [Monitor architecture](../../monitor/architecture.md). |
The largest **CPU** win on many ADN networks is on the **MASTER inject-only** path (`PROXY` with inject-only mode).
## OpenBridge-heavy servers (2.3.3+)
**Legacy:** `send_peers` walks **every registered peer** for each downlink packet → cost grows as **O(peers × packets/s)**.
If you run **many OpenBridges** and upgrade from an earlier **2.x** build to **2.3.3 or
newer**, a production node under comparable OBP load showed roughly:
**2.x:** `PeerDownlinkIndex` precomputes candidates from each peer’s **OPTIONS** (static TGs) and **UA session** state. For each group voice frame, only peers that **might** want that `(slot, TG)` are considered; each candidate still passes `peer_should_receive_group_voice`.
| | Before (2.x) | After (2.3.3+) |
|--|--------------|----------------|
| **CPU** | baseline | **~25% of before** |
| **RAM** | baseline | **~75% of before** |
```text
Legacy: every DMRD → try all N peers
2.x: every DMRD → index lookup → try k peers (k ≪ N on busy proxies)
```
Exact numbers depend on traffic and hardware; treat these as a reference, not a guarantee.
OPTIONS parsing is **cached per peer** (`_CACHED_OPTIONS_STATIC`): if the OPTIONS blob is unchanged, already-parsed static TGs are reused instead of re-parsing on every packet.
## When you will notice it most
| Code | Role |
|------|------|
| `application/routing/peer_downlink_index.py` | Index build and `(slot, tgid) → candidates` |
| `infrastructure/twisted_adapters/udp_hbp.py` | `_iter_downlink_peers`, `send_peers` |
| `tests/infrastructure/test_peer_downlink_fanout.py` | Inject-only fan-out tests |
| Your network | Effect |
|--------------|--------|
| Small install, few peers, light traffic | Modest — the stack is simply lighter overall. |
| **Inject proxy with many hotspots** | **Clear CPU win** when group voice is busy. |
| **Many OpenBridges, steady mesh traffic** | **Clear CPU and RAM win** after **2.3.3+** (see table above). |
| **Many hotspots logging in at once** | Less load on server and monitor during login bursts. |
| **Monitor open 24/7** | Lower and more stable RAM on **adn-monitor 2.x**. |
**When it matters:** proxy with **tens to hundreds** of hotspots and steady group voice. On a small conference with few peers, the difference is minor.
## Server: routing indexes
On every group voice frame the server must find relay tables where **this system is the ACTIVE source**.
**Legacy:** scan rows inside `BRIDGES[table_key]` (and related tables).
**2.x:** `InMemorySubscriptionStore.relay_tables_with_active_source()` uses a maintained **`_source_tables`** index — lookup by `(system, slot, dst_tgid)` without walking all legs.
This lives in the subscription store implementation; it is an **algorithmic index**, not a separate feature you configure.
| Code | Role |
|------|------|
| `infrastructure/subscription_store.py` | `_source_tables`, `_by_table`, `_active_target_counts` |
| `application/subscription/router.py` | `SubscriptionRouter.resolve()` |
## Server: less periodic and login-storm work
| Change | What it avoids |
|--------|----------------|
| **No 26 s OPTIONS loop** | Timer firing every 26 s across all systems to refresh static bridges when RPTO/startup/reload already handle it |
| **`ConfigPushThrottle`** | Flooding the monitor with CONFIG snapshots when many peers connect within a few seconds (debounce widens from ~0.3 s to ~2 s during bursts) |
| **`BoundedReportQueue`** | Doing pickle/JSON encode and TCP send synchronously on the voice hot path; coalesces duplicate config/bridge snapshots |
## Server: reporting and deployment
- **Report v2** — structured JSON replaces opaque pickle snapshots for bridge/config state on the **2.x monitor** wire. See [Monitoring and reports](../user-guide/monitoring.md) and [Report protocol v2](../protocols/report-v2.md).
- **Integrated proxy** — `PROXY` runs **in-process**; dropping the standalone **adn-proxy** saves baseline **RAM** (one interpreter, shared config) and simplifies ops.
## Monitor (adn-monitor 2.x)
Pair **adn-server 2.x** with **adn-monitor 2.x** to get the reporting-side gains:
| Change | Effect |
|--------|--------|
| **Slim wire / `dashboard_state`** | Monitor ingests compact JSON state instead of holding full duplicated pickle trees from v1 |
| **`clean_sys_dict`** | Periodic eviction of stale in-memory entries (caps runaway growth on long-lived panels) |
| **Last-heard row cache, lighter WS fingerprints** | Less work per dashboard refresh |
| **Unified FastAPI stack** | Removed separate PHP API and standalone monitor **proxy** process |
Details: [Monitor architecture](../../monitor/architecture.md).
## When you will notice a difference
| Deployment | CPU | RAM |
|------------|-----|-----|
| Few masters, no inject proxy, light traffic | Small | Small |
| **Inject-only proxy, many hotspots, busy TG** | **Clear** (downlink index) | Moderate (single server process vs server+proxy) |
| Long-lived monitor + report v2 | Moderate (less serialize on wire) | **Clearer** on monitor (slim state, `clean_sys_dict`) |
Crypto, AMBE, and OpenBridge MAC work still dominate on OpenBridge-heavy paths — routing-table optimizations do not remove that cost.
Crypto, AMBE, and OpenBridge wire work still cost CPU on busy OBP paths — routing
optimizations remove redundant bridge work, not voice codec or encryption overhead.
## Related reading
- [Architecture](architecture.md) — layers and entrypoint
- [BRIDGES vs Subscriptions](bridges-vs-subscriptions.md) — routing model (not a performance feature)
- [Behaviour and timers](behaviour-and-timers.md) — event-driven OPTIONS vs legacy 26 s loop
- [Hotspot proxy](../user-guide/hotspot-proxy.md) — integrated `PROXY` / inject-only
- [Report protocol v2](../protocols/report-v2.md) — JSON wire to monitor
- Release notes: `CHANGELOG.md` at the repository root (`Performance` under **2.0.0-rc.1**).
- [Hotspot proxy](../user-guide/hotspot-proxy.md) — integrated `PROXY`
- [Monitoring and reports](../user-guide/monitoring.md) — report v2 pairing
- [Behaviour and timers](behaviour-and-timers.md) — event-driven OPTIONS
- Release notes: `CHANGELOG.md` at the repository root

@ -1,99 +1,50 @@
# Rendimiento (2.x)
**adn-server 2.x** y **adn-monitor 2.x** incluyen varios cambios que reducen trabajo de CPU y huella de memoria frente a **adn-dmr-server** y al stack antiguo de monitor/proxy. Esta página resume **qué** mejora y **qué lo provoca**.
**adn-server 2.x** y **adn-monitor 2.x** consumen menos CPU y RAM que **adn-dmr-server**
y el stack antiguo de monitor/proxy. No hace falta ajustar nada — las mejoras vienen de
cómo están hechos el enrutado, los informes y el proxy integrado.
## Resumen
Empareja **servidor 2.x con monitor 2.x** para aprovechar también el lado del panel.
| Área | Efecto típico | Causa principal |
|------|---------------|-----------------|
| **Downlink de voz (proxy inject)** | Menos CPU con tráfico de grupo intenso | **`PeerDownlinkIndex`** — fan-out solo a peers que encajan `(slot, TG)` en lugar de escanear todos los hotspots por paquete |
| **Origen ACTIVE en bridge** | Lookup más rápido | **Índices del `SubscriptionStore`** (`relay_tables_with_active_source`) — O(1) por `(system, slot, tgid)` frente a recorrer filas |
| **CPU de fondo** | Menos despertares | **OPTIONS / TG estática por eventos** — eliminado el bucle legacy cada **26 s** `options_config_loop` ([Comportamiento y temporizadores](behaviour-and-timers.md)) |
| **Ráfaga de logins** | Menos CONFIG redundante | **`ConfigPushThrottle`** — debounce adaptativo al empujar CONFIG al monitor |
| **Informes vs voz** | La voz se bloquea menos por informes | **`BoundedReportQueue`** — snapshots coalescidos, drenado acotado por tick |
| **Cable servidor → monitor** | Menos serializar/enviar | **Informe v2** JSON (`routing_table`, `topology`, `voice_event`) en lugar de pickle periódico de `CONFIG`/`BRIDGES` ([Protocolo de informes v2](../protocols/report-v2.md)) |
| **Procesos (RAM)** | Un proceso Python en lugar de dos | **`PROXY` integrado** en `adn-server.py` — sin proceso **adn-proxy** aparte ([Proxy hotspot](../user-guide/hotspot-proxy.md)) |
| **RAM / WS del monitor** | Estado de panel más compacto | **Wire slim `dashboard_state`**, `clean_sys_dict`, fingerprints WS más ligeros ([Arquitectura del monitor](../../monitor/architecture.md)) |
## Qué mejoró
## Servidor: índice de downlink inject-only
La mayor ganancia de **CPU** en muchas redes ADN está en el camino **MASTER inject-only** (`PROXY` en modo inject-only).
**Legacy:** `send_peers` recorre **todos los peers registrados** por cada paquete de downlink → coste **O(peers × paquetes/s)**.
**2.x:** `PeerDownlinkIndex` precalcula candidatos desde **OPTIONS** (TG estáticas) y estado **UA** de cada peer. Por cada frame de voz de grupo solo se consideran peers que **podrían** querer ese `(slot, TG)`; cada candidato sigue pasando `peer_should_receive_group_voice`.
```text
Legacy: cada DMRD → probar los N peers
2.x: cada DMRD → lookup en índice → probar k peers (k ≪ N en proxies cargados)
```
El parse de OPTIONS se **guarda en caché por peer** (`_CACHED_OPTIONS_STATIC`): si el blob OPTIONS no cambió, se reutilizan las TG estáticas ya parseadas en lugar de volver a interpretarlo en cada paquete.
| Código | Rol |
|--------|-----|
| `application/routing/peer_downlink_index.py` | Construcción del índice y `(slot, tgid) → candidatos` |
| `infrastructure/twisted_adapters/udp_hbp.py` | `_iter_downlink_peers`, `send_peers` |
| `tests/infrastructure/test_peer_downlink_fanout.py` | Tests de fan-out inject-only |
**Cuándo se nota:** proxy con **decenas o cientos** de hotspots y voz de grupo continua. En una conferencia pequeña con pocos peers, la diferencia es pequeña.
## Servidor: índices de enrutado
En cada frame de voz de grupo el servidor debe encontrar tablas donde **este system es origen ACTIVE**.
**Legacy:** recorrer filas dentro de `BRIDGES[clave]`.
**2.x:** `InMemorySubscriptionStore.relay_tables_with_active_source()` usa el índice **`_source_tables`** — lookup por `(system, slot, dst_tgid)` sin recorrer todas las patas.
Está en la implementación del store; es un **índice algorítmico**, no una opción de configuración aparte.
| Código | Rol |
|--------|-----|
| `infrastructure/subscription_store.py` | `_source_tables`, `_by_table`, `_active_target_counts` |
| `application/subscription/router.py` | `SubscriptionRouter.resolve()` |
## Servidor: menos trabajo periódico y en tormenta de logins
| Cambio | Qué evita |
| Mejora | Qué ganas |
|--------|-----------|
| **Sin bucle OPTIONS 26 s** | Timer cada 26 s en todos los systems cuando RPTO/arranque/reload ya refrescan bridges estáticos |
| **`ConfigPushThrottle`** | Inundar al monitor con snapshots CONFIG cuando muchos peers conectan en pocos segundos (debounce ~0,3 s → ~2 s en ráfaga) |
| **`BoundedReportQueue`** | Encode pickle/JSON y envío TCP en el hot path de voz; coalesce de snapshots config/bridge duplicados |
| **Enrutado de voz más eficiente** | Con tráfico de grupo intenso el servidor hace menos trabajo por paquete — sobre todo en proxies con muchos hotspots y en nodos con muchos OpenBridge. |
| **Proxy hotspot integrado** | Un solo proceso `adn-server` en lugar de servidor + **adn-proxy** aparte — menos RAM y operación más simple. Ver [Proxy hotspot](../user-guide/hotspot-proxy.md). |
| **Sin timer legacy de 26 s** | Las TG estáticas se refrescan por eventos (arranque, recarga de config, OPTIONS del peer), no con un bucle de fondo cada 26 segundos. Ver [Comportamiento y temporizadores](behaviour-and-timers.md). |
| **Cable al monitor más liviano** | **Informe v2** envía JSON compacto en lugar de volcados pickle pesados. Ver [Protocolo de informes v2](../protocols/report-v2.md). |
| **Monitor 2.x** | Estado de panel más compacto, menos crecimiento de memoria con el panel abierto días. Ver [Arquitectura del monitor](../../monitor/architecture.md). |
## Servidor: informes y despliegue
## Servidores con muchos OpenBridge (2.3.3+)
- **Informe v2** — JSON estructurado sustituye snapshots pickle opacos de bridge/config en el cable hacia **monitor 2.x**. Ver [Monitor e informes](../user-guide/monitoring.md) y [Protocolo de informes v2](../protocols/report-v2.md).
- **Proxy integrado** — `PROXY` **in-process**; quitar **adn-proxy** standalone ahorra **RAM** base (un intérprete, config compartida) y simplifica operación.
Si tienes **muchos OpenBridge** y actualizas desde un **2.x** anterior a **2.3.3 o
superior**, un nodo en producción con carga OBP comparable mostró aproximadamente:
## Monitor (adn-monitor 2.x)
| | Antes (2.x) | Después (2.3.3+) |
|--|-------------|------------------|
| **CPU** | línea base | **~25% de la línea base** |
| **RAM** | línea base | **~75% de la línea base** |
Empareja **adn-server 2.x** con **adn-monitor 2.x** para las mejoras del lado informes:
Las cifras exactas dependen del tráfico y del hardware; tómalo como referencia, no como garantía.
| Cambio | Efecto |
|--------|--------|
| **Wire slim / `dashboard_state`** | El monitor ingiere JSON compacto en lugar de duplicar árboles pickle v1 |
| **`clean_sys_dict`** | Expulsión periódica de entradas obsoletas en memoria (tope de crecimiento en paneles largos) |
| **Caché lastheard, fingerprints WS ligeros** | Menos trabajo por refresco del dashboard |
| **Stack FastAPI unificado** | Eliminados API PHP y proceso **proxy** standalone del monitor |
Detalle: [Arquitectura del monitor](../../monitor/architecture.md).
## Cuándo se nota más
## Cuándo se nota la diferencia
| Despliegue | CPU | RAM |
|------------|-----|-----|
| Pocos masters, sin proxy inject, tráfico bajo | Poca | Poca |
| **Proxy inject-only, muchos hotspots, TG activa** | **Clara** (índice downlink) | Moderada (un proceso servidor vs servidor+proxy) |
| Monitor largo + informe v2 | Moderada (menos serializar en cable) | **Más clara** en monitor (estado slim, `clean_sys_dict`) |
| Tu red | Efecto |
|--------|--------|
| Instalación pequeña, pocos peers, poco tráfico | Moderado — el stack es más liviano en general. |
| **Proxy inject con muchos hotspots** | **Ganancia clara de CPU** con voz de grupo activa. |
| **Muchos OpenBridge, tráfico mesh continuo** | **Ganancia clara de CPU y RAM** tras **2.3.3+** (ver tabla anterior). |
| **Muchos hotspots conectando a la vez** | Menos carga en servidor y monitor en ráfagas de login. |
| **Monitor abierto 24/7** | RAM más baja y estable en **adn-monitor 2.x**. |
Crypto, AMBE y MAC OpenBridge siguen dominando en tramos OBP cargados — optimizar la tabla de bridge no elimina ese coste.
Crypto, AMBE y el trabajo de cable OpenBridge siguen costando CPU en tramos OBP
cargados — las optimizaciones de routing quitan trabajo redundante de bridges, no el
codec de voz ni el cifrado.
## Lecturas relacionadas
- [Arquitectura](architecture.md) — capas y entrypoint
- [BRIDGES vs Subscriptions](bridges-vs-subscriptions.md) — modelo de enrutado (no es feature de rendimiento)
- [Comportamiento y temporizadores](behaviour-and-timers.md) — OPTIONS por eventos vs bucle 26 s legacy
- [Proxy hotspot](../user-guide/hotspot-proxy.md) — `PROXY` integrado / inject-only
- [Protocolo de informes v2](../protocols/report-v2.md) — cable JSON al monitor
- Notas de versión: `CHANGELOG.md` en la raíz del repositorio (`Performance` en **2.0.0-rc.1**).
- [Proxy hotspot](../user-guide/hotspot-proxy.md) — `PROXY` integrado
- [Monitor e informes](../user-guide/monitoring.md) — emparejar informe v2
- [Comportamiento y temporizadores](behaviour-and-timers.md) — OPTIONS por eventos
- Notas de versión: `CHANGELOG.md` en la raíz del repositorio

Loading…
Cancel
Save

Powered by TurnKey Linux.