diff --git a/docs/en/server/development/performance.md b/docs/en/server/development/performance.md index 5dfdc50..7bad48a 100644 --- a/docs/en/server/development/performance.md +++ b/docs/en/server/development/performance.md @@ -1,99 +1,49 @@ # Performance (2.x) -**adn-server 2.x** and **adn-monitor 2.x** include several changes that reduce CPU work and memory footprint compared with **adn-dmr-server** and the old monitor/proxy stack. This page lists **what** improves and **what causes it**. +**adn-server 2.x** and **adn-monitor 2.x** use less CPU and RAM than **adn-dmr-server** +and the old monitor/proxy stack. You do not need to tune anything — the gains come from +how routing, reporting, and the integrated proxy are built. -## At a glance +Pair **server 2.x with monitor 2.x** to get the full benefit on the dashboard side. -| Area | Typical effect | Main cause | -|------|----------------|------------| -| **Voice downlink (inject proxy)** | Lower CPU under busy group traffic | **`PeerDownlinkIndex`** — fan-out to peers that match `(slot, TG)` instead of scanning every connected hotspot per packet | -| **Bridge source lookup** | Faster “am I the ACTIVE source?” | **`SubscriptionStore`** indexes (`relay_tables_with_active_source`) — O(1) by `(system, slot, tgid)` vs scanning table rows | -| **Background CPU** | Fewer wakeups | **Event-driven OPTIONS / static TG** — removed legacy **26 s** `options_config_loop` ([Behaviour and timers](behaviour-and-timers.md)) | -| **Mass peer login** | Less redundant CONFIG traffic | **`ConfigPushThrottle`** — adaptive debounce on CONFIG push to the monitor | -| **Reporting vs voice** | Voice path less blocked by reports | **`BoundedReportQueue`** — coalesced snapshots, bounded drain per tick | -| **Server → monitor wire** | Less serialize/send work | **Report v2** JSON (`routing_table`, `topology`, `voice_event`) instead of periodic full pickle of `CONFIG`/`BRIDGES` ([Report protocol v2](../protocols/report-v2.md)) | -| **Process count (RAM)** | One Python process instead of two | **Integrated `PROXY`** in `adn-server.py` — no separate **adn-proxy** process ([Hotspot proxy](../user-guide/hotspot-proxy.md)) | -| **Monitor RAM / WS load** | Smaller in-memory dashboard state | **Slim `dashboard_state` wire**, `clean_sys_dict`, lighter WebSocket fingerprints ([Monitor architecture](../../monitor/architecture.md)) | +## What improved -## Server: inject-only downlink index +| Improvement | What you get | +|-------------|--------------| +| **Smarter voice routing** | Under busy group traffic the server does less work per packet — especially on proxies with many hotspots and on OpenBridge-heavy nodes. | +| **Integrated hotspot proxy** | One `adn-server` process instead of server + standalone **adn-proxy** — less RAM and simpler ops. See [Hotspot proxy](../user-guide/hotspot-proxy.md). | +| **No legacy 26 s timer** | Static talkgroups refresh on events (startup, config reload, peer OPTIONS), not a background loop every 26 seconds. See [Behaviour and timers](behaviour-and-timers.md). | +| **Lighter monitor link** | **Report v2** sends compact JSON instead of heavy periodic pickle dumps. See [Report protocol v2](../protocols/report-v2.md). | +| **Monitor 2.x** | Slimmer dashboard state, less memory growth on panels left open for days. See [Monitor architecture](../../monitor/architecture.md). | -The largest **CPU** win on many ADN networks is on the **MASTER inject-only** path (`PROXY` with inject-only mode). +## OpenBridge-heavy servers (2.3.3+) -**Legacy:** `send_peers` walks **every registered peer** for each downlink packet → cost grows as **O(peers × packets/s)**. +If you run **many OpenBridges** and upgrade from an earlier **2.x** build to **2.3.3 or +newer**, a production node under comparable OBP load showed roughly: -**2.x:** `PeerDownlinkIndex` precomputes candidates from each peer’s **OPTIONS** (static TGs) and **UA session** state. For each group voice frame, only peers that **might** want that `(slot, TG)` are considered; each candidate still passes `peer_should_receive_group_voice`. +| | Before (2.x) | After (2.3.3+) | +|--|--------------|----------------| +| **CPU** | baseline | **~25% of before** | +| **RAM** | baseline | **~75% of before** | -```text -Legacy: every DMRD → try all N peers -2.x: every DMRD → index lookup → try k peers (k ≪ N on busy proxies) -``` +Exact numbers depend on traffic and hardware; treat these as a reference, not a guarantee. -OPTIONS parsing is **cached per peer** (`_CACHED_OPTIONS_STATIC`): if the OPTIONS blob is unchanged, already-parsed static TGs are reused instead of re-parsing on every packet. +## When you will notice it most -| Code | Role | -|------|------| -| `application/routing/peer_downlink_index.py` | Index build and `(slot, tgid) → candidates` | -| `infrastructure/twisted_adapters/udp_hbp.py` | `_iter_downlink_peers`, `send_peers` | -| `tests/infrastructure/test_peer_downlink_fanout.py` | Inject-only fan-out tests | +| Your network | Effect | +|--------------|--------| +| Small install, few peers, light traffic | Modest — the stack is simply lighter overall. | +| **Inject proxy with many hotspots** | **Clear CPU win** when group voice is busy. | +| **Many OpenBridges, steady mesh traffic** | **Clear CPU and RAM win** after **2.3.3+** (see table above). | +| **Many hotspots logging in at once** | Less load on server and monitor during login bursts. | +| **Monitor open 24/7** | Lower and more stable RAM on **adn-monitor 2.x**. | -**When it matters:** proxy with **tens to hundreds** of hotspots and steady group voice. On a small conference with few peers, the difference is minor. - -## Server: routing indexes - -On every group voice frame the server must find relay tables where **this system is the ACTIVE source**. - -**Legacy:** scan rows inside `BRIDGES[table_key]` (and related tables). - -**2.x:** `InMemorySubscriptionStore.relay_tables_with_active_source()` uses a maintained **`_source_tables`** index — lookup by `(system, slot, dst_tgid)` without walking all legs. - -This lives in the subscription store implementation; it is an **algorithmic index**, not a separate feature you configure. - -| Code | Role | -|------|------| -| `infrastructure/subscription_store.py` | `_source_tables`, `_by_table`, `_active_target_counts` | -| `application/subscription/router.py` | `SubscriptionRouter.resolve()` | - -## Server: less periodic and login-storm work - -| Change | What it avoids | -|--------|----------------| -| **No 26 s OPTIONS loop** | Timer firing every 26 s across all systems to refresh static bridges when RPTO/startup/reload already handle it | -| **`ConfigPushThrottle`** | Flooding the monitor with CONFIG snapshots when many peers connect within a few seconds (debounce widens from ~0.3 s to ~2 s during bursts) | -| **`BoundedReportQueue`** | Doing pickle/JSON encode and TCP send synchronously on the voice hot path; coalesces duplicate config/bridge snapshots | - -## Server: reporting and deployment - -- **Report v2** — structured JSON replaces opaque pickle snapshots for bridge/config state on the **2.x monitor** wire. See [Monitoring and reports](../user-guide/monitoring.md) and [Report protocol v2](../protocols/report-v2.md). -- **Integrated proxy** — `PROXY` runs **in-process**; dropping the standalone **adn-proxy** saves baseline **RAM** (one interpreter, shared config) and simplifies ops. - -## Monitor (adn-monitor 2.x) - -Pair **adn-server 2.x** with **adn-monitor 2.x** to get the reporting-side gains: - -| Change | Effect | -|--------|--------| -| **Slim wire / `dashboard_state`** | Monitor ingests compact JSON state instead of holding full duplicated pickle trees from v1 | -| **`clean_sys_dict`** | Periodic eviction of stale in-memory entries (caps runaway growth on long-lived panels) | -| **Last-heard row cache, lighter WS fingerprints** | Less work per dashboard refresh | -| **Unified FastAPI stack** | Removed separate PHP API and standalone monitor **proxy** process | - -Details: [Monitor architecture](../../monitor/architecture.md). - -## When you will notice a difference - -| Deployment | CPU | RAM | -|------------|-----|-----| -| Few masters, no inject proxy, light traffic | Small | Small | -| **Inject-only proxy, many hotspots, busy TG** | **Clear** (downlink index) | Moderate (single server process vs server+proxy) | -| Long-lived monitor + report v2 | Moderate (less serialize on wire) | **Clearer** on monitor (slim state, `clean_sys_dict`) | - -Crypto, AMBE, and OpenBridge MAC work still dominate on OpenBridge-heavy paths — routing-table optimizations do not remove that cost. +Crypto, AMBE, and OpenBridge wire work still cost CPU on busy OBP paths — routing +optimizations remove redundant bridge work, not voice codec or encryption overhead. ## Related reading -- [Architecture](architecture.md) — layers and entrypoint -- [BRIDGES vs Subscriptions](bridges-vs-subscriptions.md) — routing model (not a performance feature) -- [Behaviour and timers](behaviour-and-timers.md) — event-driven OPTIONS vs legacy 26 s loop -- [Hotspot proxy](../user-guide/hotspot-proxy.md) — integrated `PROXY` / inject-only -- [Report protocol v2](../protocols/report-v2.md) — JSON wire to monitor -- Release notes: `CHANGELOG.md` at the repository root (`Performance` under **2.0.0-rc.1**). +- [Hotspot proxy](../user-guide/hotspot-proxy.md) — integrated `PROXY` +- [Monitoring and reports](../user-guide/monitoring.md) — report v2 pairing +- [Behaviour and timers](behaviour-and-timers.md) — event-driven OPTIONS +- Release notes: `CHANGELOG.md` at the repository root diff --git a/docs/es/server/development/performance.md b/docs/es/server/development/performance.md index 41a2179..bdfc65a 100644 --- a/docs/es/server/development/performance.md +++ b/docs/es/server/development/performance.md @@ -1,99 +1,50 @@ # Rendimiento (2.x) -**adn-server 2.x** y **adn-monitor 2.x** incluyen varios cambios que reducen trabajo de CPU y huella de memoria frente a **adn-dmr-server** y al stack antiguo de monitor/proxy. Esta página resume **qué** mejora y **qué lo provoca**. +**adn-server 2.x** y **adn-monitor 2.x** consumen menos CPU y RAM que **adn-dmr-server** +y el stack antiguo de monitor/proxy. No hace falta ajustar nada — las mejoras vienen de +cómo están hechos el enrutado, los informes y el proxy integrado. -## Resumen +Empareja **servidor 2.x con monitor 2.x** para aprovechar también el lado del panel. -| Área | Efecto típico | Causa principal | -|------|---------------|-----------------| -| **Downlink de voz (proxy inject)** | Menos CPU con tráfico de grupo intenso | **`PeerDownlinkIndex`** — fan-out solo a peers que encajan `(slot, TG)` en lugar de escanear todos los hotspots por paquete | -| **Origen ACTIVE en bridge** | Lookup más rápido | **Índices del `SubscriptionStore`** (`relay_tables_with_active_source`) — O(1) por `(system, slot, tgid)` frente a recorrer filas | -| **CPU de fondo** | Menos despertares | **OPTIONS / TG estática por eventos** — eliminado el bucle legacy cada **26 s** `options_config_loop` ([Comportamiento y temporizadores](behaviour-and-timers.md)) | -| **Ráfaga de logins** | Menos CONFIG redundante | **`ConfigPushThrottle`** — debounce adaptativo al empujar CONFIG al monitor | -| **Informes vs voz** | La voz se bloquea menos por informes | **`BoundedReportQueue`** — snapshots coalescidos, drenado acotado por tick | -| **Cable servidor → monitor** | Menos serializar/enviar | **Informe v2** JSON (`routing_table`, `topology`, `voice_event`) en lugar de pickle periódico de `CONFIG`/`BRIDGES` ([Protocolo de informes v2](../protocols/report-v2.md)) | -| **Procesos (RAM)** | Un proceso Python en lugar de dos | **`PROXY` integrado** en `adn-server.py` — sin proceso **adn-proxy** aparte ([Proxy hotspot](../user-guide/hotspot-proxy.md)) | -| **RAM / WS del monitor** | Estado de panel más compacto | **Wire slim `dashboard_state`**, `clean_sys_dict`, fingerprints WS más ligeros ([Arquitectura del monitor](../../monitor/architecture.md)) | +## Qué mejoró -## Servidor: índice de downlink inject-only - -La mayor ganancia de **CPU** en muchas redes ADN está en el camino **MASTER inject-only** (`PROXY` en modo inject-only). - -**Legacy:** `send_peers` recorre **todos los peers registrados** por cada paquete de downlink → coste **O(peers × paquetes/s)**. - -**2.x:** `PeerDownlinkIndex` precalcula candidatos desde **OPTIONS** (TG estáticas) y estado **UA** de cada peer. Por cada frame de voz de grupo solo se consideran peers que **podrían** querer ese `(slot, TG)`; cada candidato sigue pasando `peer_should_receive_group_voice`. - -```text -Legacy: cada DMRD → probar los N peers -2.x: cada DMRD → lookup en índice → probar k peers (k ≪ N en proxies cargados) -``` - -El parse de OPTIONS se **guarda en caché por peer** (`_CACHED_OPTIONS_STATIC`): si el blob OPTIONS no cambió, se reutilizan las TG estáticas ya parseadas en lugar de volver a interpretarlo en cada paquete. - -| Código | Rol | -|--------|-----| -| `application/routing/peer_downlink_index.py` | Construcción del índice y `(slot, tgid) → candidatos` | -| `infrastructure/twisted_adapters/udp_hbp.py` | `_iter_downlink_peers`, `send_peers` | -| `tests/infrastructure/test_peer_downlink_fanout.py` | Tests de fan-out inject-only | - -**Cuándo se nota:** proxy con **decenas o cientos** de hotspots y voz de grupo continua. En una conferencia pequeña con pocos peers, la diferencia es pequeña. - -## Servidor: índices de enrutado - -En cada frame de voz de grupo el servidor debe encontrar tablas donde **este system es origen ACTIVE**. - -**Legacy:** recorrer filas dentro de `BRIDGES[clave]`. - -**2.x:** `InMemorySubscriptionStore.relay_tables_with_active_source()` usa el índice **`_source_tables`** — lookup por `(system, slot, dst_tgid)` sin recorrer todas las patas. - -Está en la implementación del store; es un **índice algorítmico**, no una opción de configuración aparte. - -| Código | Rol | -|--------|-----| -| `infrastructure/subscription_store.py` | `_source_tables`, `_by_table`, `_active_target_counts` | -| `application/subscription/router.py` | `SubscriptionRouter.resolve()` | - -## Servidor: menos trabajo periódico y en tormenta de logins - -| Cambio | Qué evita | +| Mejora | Qué ganas | |--------|-----------| -| **Sin bucle OPTIONS 26 s** | Timer cada 26 s en todos los systems cuando RPTO/arranque/reload ya refrescan bridges estáticos | -| **`ConfigPushThrottle`** | Inundar al monitor con snapshots CONFIG cuando muchos peers conectan en pocos segundos (debounce ~0,3 s → ~2 s en ráfaga) | -| **`BoundedReportQueue`** | Encode pickle/JSON y envío TCP en el hot path de voz; coalesce de snapshots config/bridge duplicados | +| **Enrutado de voz más eficiente** | Con tráfico de grupo intenso el servidor hace menos trabajo por paquete — sobre todo en proxies con muchos hotspots y en nodos con muchos OpenBridge. | +| **Proxy hotspot integrado** | Un solo proceso `adn-server` en lugar de servidor + **adn-proxy** aparte — menos RAM y operación más simple. Ver [Proxy hotspot](../user-guide/hotspot-proxy.md). | +| **Sin timer legacy de 26 s** | Las TG estáticas se refrescan por eventos (arranque, recarga de config, OPTIONS del peer), no con un bucle de fondo cada 26 segundos. Ver [Comportamiento y temporizadores](behaviour-and-timers.md). | +| **Cable al monitor más liviano** | **Informe v2** envía JSON compacto en lugar de volcados pickle pesados. Ver [Protocolo de informes v2](../protocols/report-v2.md). | +| **Monitor 2.x** | Estado de panel más compacto, menos crecimiento de memoria con el panel abierto días. Ver [Arquitectura del monitor](../../monitor/architecture.md). | -## Servidor: informes y despliegue +## Servidores con muchos OpenBridge (2.3.3+) -- **Informe v2** — JSON estructurado sustituye snapshots pickle opacos de bridge/config en el cable hacia **monitor 2.x**. Ver [Monitor e informes](../user-guide/monitoring.md) y [Protocolo de informes v2](../protocols/report-v2.md). -- **Proxy integrado** — `PROXY` **in-process**; quitar **adn-proxy** standalone ahorra **RAM** base (un intérprete, config compartida) y simplifica operación. +Si tienes **muchos OpenBridge** y actualizas desde un **2.x** anterior a **2.3.3 o +superior**, un nodo en producción con carga OBP comparable mostró aproximadamente: -## Monitor (adn-monitor 2.x) +| | Antes (2.x) | Después (2.3.3+) | +|--|-------------|------------------| +| **CPU** | línea base | **~25% de la línea base** | +| **RAM** | línea base | **~75% de la línea base** | -Empareja **adn-server 2.x** con **adn-monitor 2.x** para las mejoras del lado informes: +Las cifras exactas dependen del tráfico y del hardware; tómalo como referencia, no como garantía. -| Cambio | Efecto | -|--------|--------| -| **Wire slim / `dashboard_state`** | El monitor ingiere JSON compacto en lugar de duplicar árboles pickle v1 | -| **`clean_sys_dict`** | Expulsión periódica de entradas obsoletas en memoria (tope de crecimiento en paneles largos) | -| **Caché lastheard, fingerprints WS ligeros** | Menos trabajo por refresco del dashboard | -| **Stack FastAPI unificado** | Eliminados API PHP y proceso **proxy** standalone del monitor | - -Detalle: [Arquitectura del monitor](../../monitor/architecture.md). +## Cuándo se nota más -## Cuándo se nota la diferencia - -| Despliegue | CPU | RAM | -|------------|-----|-----| -| Pocos masters, sin proxy inject, tráfico bajo | Poca | Poca | -| **Proxy inject-only, muchos hotspots, TG activa** | **Clara** (índice downlink) | Moderada (un proceso servidor vs servidor+proxy) | -| Monitor largo + informe v2 | Moderada (menos serializar en cable) | **Más clara** en monitor (estado slim, `clean_sys_dict`) | +| Tu red | Efecto | +|--------|--------| +| Instalación pequeña, pocos peers, poco tráfico | Moderado — el stack es más liviano en general. | +| **Proxy inject con muchos hotspots** | **Ganancia clara de CPU** con voz de grupo activa. | +| **Muchos OpenBridge, tráfico mesh continuo** | **Ganancia clara de CPU y RAM** tras **2.3.3+** (ver tabla anterior). | +| **Muchos hotspots conectando a la vez** | Menos carga en servidor y monitor en ráfagas de login. | +| **Monitor abierto 24/7** | RAM más baja y estable en **adn-monitor 2.x**. | -Crypto, AMBE y MAC OpenBridge siguen dominando en tramos OBP cargados — optimizar la tabla de bridge no elimina ese coste. +Crypto, AMBE y el trabajo de cable OpenBridge siguen costando CPU en tramos OBP +cargados — las optimizaciones de routing quitan trabajo redundante de bridges, no el +codec de voz ni el cifrado. ## Lecturas relacionadas -- [Arquitectura](architecture.md) — capas y entrypoint -- [BRIDGES vs Subscriptions](bridges-vs-subscriptions.md) — modelo de enrutado (no es feature de rendimiento) -- [Comportamiento y temporizadores](behaviour-and-timers.md) — OPTIONS por eventos vs bucle 26 s legacy -- [Proxy hotspot](../user-guide/hotspot-proxy.md) — `PROXY` integrado / inject-only -- [Protocolo de informes v2](../protocols/report-v2.md) — cable JSON al monitor -- Notas de versión: `CHANGELOG.md` en la raíz del repositorio (`Performance` en **2.0.0-rc.1**). +- [Proxy hotspot](../user-guide/hotspot-proxy.md) — `PROXY` integrado +- [Monitor e informes](../user-guide/monitoring.md) — emparejar informe v2 +- [Comportamiento y temporizadores](behaviour-and-timers.md) — OPTIONS por eventos +- Notas de versión: `CHANGELOG.md` en la raíz del repositorio