Adds MeshCoreConnector.addContactByKey, the device-side half of the
identity exchange. It uses the stock CMD_ADD_UPDATE_CONTACT (command 9),
whose handler creates the contact outright when the public key is
unknown. Verified byte-identical to upstream meshcore-dev/MeshCore at
0679dbef, so this needs no capability gate and works on stock firmware,
not only Offband builds.
buildUpdateContactPathFrame gains an optional lastAdvert. It defaults to
now, leaving all three existing path-update callers byte-identical, and
addContactByKey passes the epoch.
That epoch is the whole point. The firmware compares this field against
every incoming advert with 'timestamp <= last_advert_timestamp' and
silently discards the non-greater ones as replay attacks
(BaseChatMesh.cpp:142-145). Advert timestamps come from the sender's
clock, and two nodes on this mesh currently advertise with 2024 clocks,
so stamping now would leave a key-added contact permanently deaf to its
own adverts. A negative input clamps to zero rather than wrapping to a
huge value, which would be the worst case for that guard.
The frame is sent at full length with the 0xFF flood sentinel. The
firmware length guard is only 'len >= 36' but updateContactFromFrame
reads through offset 136 regardless (MyMesh.cpp:295-318), so the frame
must never be trimmed. A test pins that too.
Local state keeps lastSeen at the epoch so the contact reads as
unverified until a real advert upgrades it (#630).
Epic #619.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Channel frames carry no key, so the sender name is resolved against known
and discovered contacts. Adds resolveContactsByName (returning Contacts,
which importDiscoveredContact needs) and rebuilds resolveContactKeysByName
on top of it so one matching rule serves both.
The avatar is the shortcut; the sender name stays inert because it sits too
close to the message body to hit reliably on a phone. Long-press keeps every
action it had and gains the same contact rows. A name several nodes claim
lists all of them, and an unheard name says so instead of failing silently.
Three findings reviewed, two applied, one rejected on evidence.
APPLIED. Suppression logging is de-duplicated. Device-info arrives on
every connect, so a flapping radio would have emitted the same
suppression line on each reconnect. It now logs once per distinct
reason, and the memo clears when a gate opens or when per-radio state is
cleared, so the next genuine suppression is still reported. Real, and
exactly the flooding the project forbids.
APPLIED, though not reachable today. The scope reply logged
reply.scope!, which is safe only because parseNotifyScopeReply returns
an ERROR reply when the scope code is unknown, so a non-error reply
always carries a scope. That invariant lives in another file, so it is
now checked at the boundary instead of asserted with a bang across it.
The review rated this High on the theory that an unknown sub-code could
reach the handler; it cannot, because the parser returns null for any
sub-code it does not own and the dispatcher only forwards what it
returned. Defect not reachable, guard still worth having.
REJECTED. The review claimed the matrix reply log is unbounded because
the wire count byte allows up to 255 rows, estimating 7.5 KB per line.
The log is built from ButtonMatrix.assignments, a Map keyed by
ButtonSequence, which has exactly four values, and the parser skips
sequence codes it does not recognise rather than adding them. The map
therefore holds at most four entries and the line is bounded at roughly
160 characters no matter what the device sends. The count byte cannot
inflate it.
Full suite 740 pass, analyze clean, format clean.
Agent: CalmBay (session d14220d9)
The device-UI surface logged only refusals, so a successful read or
write produced nothing at all.
That gap cost a full diagnosis on 2026-08-02. When the notification
scope appeared not to refresh, neither the app debug log nor a firmware
serial capture could distinguish "the client never asked" from "the
client asked and got an unchanged value back". Answering it needed the
firmware counterpart and a serial capture, for a question the client
should have been able to answer alone.
- Every request logs what it sent and the frame shape.
- Every reply logs what came back: the scope value, or the matrix mask
and its decoded rows.
- A SET reply is distinguishable from a GET reply, so a confirmed write
is not mistaken for a read.
- Refusals keep their existing warn-level logging, unrecognised reason
codes still surface as raw hex.
The important one is the third case, which did not exist before: a
request that is NEVER SENT now says so and says which gate closed, with
the caps2 value that closed it. Silence was the ambiguity; suppression
is now explicit.
No behaviour change. Logging only, at info for normal traffic and warn
for refusals, tagged DeviceUI to match the existing handler logs.
No flooding risk: these fire on device-info, on pane open, and on user
action. Nothing polls, and no per-frame logging was added (SAFELANE §11
rule 10).
Full suite 740 pass, analyze clean, format clean.
Agent: CalmBay (session d14220d9)
The flood arm read as though the model were consulted:
if (pathLength < 0) {
// Flood: trust ML, only enforce firmware formula as floor
if (mlTimeout < physicsMin) return physicsMin;
}
return mlTimeout.clamp(physicsMin, physicsMax);
It never did that. _physicsMinTimeout and _physicsMaxTimeout return the
identical expression for pathLength < 0, mirroring the firmware's
calcFloodTimeoutMillisFor, so a clamp between them cannot preserve a
prediction. Whichever way control went, flood returned 500 + 16 * airtime.
The comment described behaviour the code did not have, which is exactly
the kind of load-bearing false comment that re-causes a bug later.
Replaced with an explicit early return and the constraint written down.
Behaviour is unchanged.
Equivalence is tested, not asserted. A new group attaches a real
TimeoutPredictionService, trains it on 40-56 s delivery times so any
prediction is far from 1300, and checks flood is unmoved while a direct
path is still clamped to its ceiling, which proves the clamp is live
rather than the predictor being absent. That group was run against the
pre-change branch and passes there too, so this is a simplification and
not a behaviour change.
Deliberately NOT decided here: whether flood should ever trust a
prediction. Every round trip measured to date was 0-hop direct, so there
is no flood data to decide it from. The verification needed to answer it
is written up on the issue.
Part of epic #473, stacked on #528, #529, #531 and #532.
The comment justified not modelling command execution time by claiming
`wifi on 30` does real work bringing up an interface. It does not. The
firmware handler sets a persistence deadline and sprintf's its reply
immediately (CommonCLI.cpp, "wifi on"), so execution is near-instant.
The owner caught it: he recalled reissuing the command several times
against on-screen errors, not one command taking 20 seconds.
The measurement supports him. On the 20.33 s case the reply carried
claimed=01:56:32 against a command sent at 01:56:26.938, and the RF frame
did not reach our radio until 01:56:47.271979. So roughly 5 s to reach the
repeater and be answered, then roughly 15 s in its transmit queue. The
tail is scheduling on both radios.
No behaviour change. The budget is unchanged and still has to tolerate a
20.33 s round trip; only the stated reason was wrong, and a wrong reason
in a load-bearing comment re-causes the bug later.
The CLI timeout used calculateTimeout, which mirrors the firmware's
calcDirectTimeoutMillisFor and estimates ONE-WAY delivery. A CLI command
is a request, an execution and a reply, so the budget was structurally
short.
Measured against rpt-01 on 910.525 MHz / SF7 / BW 62.5k / CR 4:5, where
the old window was 4074 ms: 27 commands sent, 12 replies matched, and 4
of those 12 arrived after the client had already given up, at 6.17 s,
8.28 s, 15.52 s and 20.33 s against a median of 2.75 s.
calculateCliTimeout sums four terms, each with a source rather than a
chosen value:
outbound leg calculateTimeout, which is what it actually models
cliReplyDelayMs 600, firmware CLI_REPLY_DELAY_MILLIS, unconditional
reply leg the reply is a second packet the ACK formula omits
retrieval budget replies are pull-based; the radio raises MSG_WAITING
and the app must ask, granting itself 5000 ms per
attempt across 3 retries
The retrieval term is derived from the sync constants rather than
restated, so the command timeout cannot drift below the layer it depends
on. That is the #530 invariant holding by construction, not by two
numbers being maintained in agreement.
Execution time is deliberately not modelled. The same verb, wifi on 30,
returned in both 2.31 s and 20.33 s, so it is not a per-command constant
that could be tabulated. The retrieval term carries that tail.
The reply leg uses physics only. The predictor is trained on
direct-message ACK latency, so asking it about a CLI reply leg would be
extrapolation; that is #534 and #535, not this change.
Tests cover the construction, the never-below-retrieval invariant, growth
with path length, coverage of the 20.33 s worst case actually observed,
and negatively that the direct-message ACK path is untouched.
Part of epic #473. Does not change the reported duration string (#531)
or the stale-prefix fallback (#532), and does nothing for commands that
draw no reply at all (#541).
extractScalarValue required ' = ' (space-equals-space) and trimRight()'d first,
so a blank field's reply 'key =' (no space after =) failed the match and the
whole 'mqtt.broker.N.field =' line leaked in as the value. Split on the first
'=' and trim instead: blank -> empty, values with '=' preserved. Regression
tests added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Badge stays observers-primary; tooltip/long-press now read 'N observers . M
observations' so they reconcile with CoreScope's 'Observations (N)' feed (the
two are distinct metrics: 15 distinct observers vs 34 total sightings). Refresh
feedback moved from a bottom snackbar to an auto-dismissed top MaterialBanner,
out of the way of the composer. Part of #524.
Agent: QuietSnow (session 31eaba02)
Owner testing: the inline badge needed pixel-accurate taps, the refresh gave no
feedback, and there was no refresh in the message action sheet. Now the badge is
a padded InkWell (real tap target), the long-press panel has a 'Refresh CoreScope
observers' action, both show a snackbar with the fetched count, and the auto-poll
is more patient (adds 90s tail steps, stabilises after 3 flat checks) so it climbs
closer to the total before you need to tap. Part of #524.
Agent: QuietSnow (session 31eaba02)
Poll now runs quick-then-slow (10/20/30/60x4s, ~5min cap) and stops early once
the count is flat for 2 checks, so it climbs to the near-final observer count
instead of stopping at a fixed cutoff. The badge is tappable to re-query on
demand (counts only ever grow), covering later re-checks. Part of #524.
Agent: QuietSnow (session 31eaba02)
Observer counts accrue over time (the packet must propagate the mesh and be
reported before CoreScope has any record), so an instant query at send+0.2s
always missed it. Now polls at ~10/30/60/120s, keeping the highest count as
observers report in; bails on disconnect. Part of #524.
Agent: QuietSnow (session 31eaba02)
Info logs at each step (0xC6 reply match, hash, CoreScope GET url + result,
stored count, message-not-found) so the app debug log shows exactly where the
badge chain breaks. Part of #524.
Agent: QuietSnow (session 31eaba02)
After a channel send, when firmware advertises 0xC6 (cap2 0x08 + ver>=22) AND
the owner enabled the feature, the connector queries the packet hash, then
CoreScope for observer_count, and stamps it on the message (background,
best-effort). Adds the coreScopeObserverCountEnabled AppSettings flag, a gated
switch in the #509 Experimental section (disabled until the radio supports the
capability), and a distinct cloud badge next to the radio-heard count. All
inert until the firmware PR (#611) lands. Completes the client side of #524.
Agent: QuietSnow (session 31eaba02)
Unifies the previously-duplicated send timestamps (frame builder vs outgoing
message each called now() separately) into one monotonic-per-channel value, so
every channel send has a unique (ts, channel_idx) key. That is the client-side
guarantee the 0xC6 correlation needs (VioletBarn caught that AES-128-ECB
determinism would otherwise make two same-second messages a wrong-hash query).
Adds transient onAirHash + coreScopeObserverCount to ChannelMessage. Part of #524.
Agent: QuietSnow (session 31eaba02)
Builder/parser for the client-issued 0xC6 CMD_OFFBAND_PKT_HASH query (per
firmware #611 contract) and firmwareSupportsPktHash (cap2 0x08 + ver >= 22).
Inert until firmware advertises the capability. Part of epic #524.
Agent: QuietSnow (session 31eaba02)
Children of #471 (parent stays open for AAB hardware validation, #507).
The block list was one global unscoped list on the phone, unioned onto
every radio on connect. A stale block for one of the owner's own radios
therefore rode onto every fresh/erased radio, and clearing a radio never
stuck: any other radio still holding it re-seeded the global list on
connect, which re-pushed it back.
- #505 block_store.dart: scope keys/names by connected device key (like
the app's other stores); dropLegacyGlobal() deletes the legacy
block_keys_v1/block_names_v1 (drop-and-start-fresh, owner-approved).
No global list.
- #505 block_service.dart: load() drops the legacy global and starts
empty; loadForDevice(deviceKey) swaps the in-memory set per radio and
runs the #250 self-heal. All mutating ops serialized through a Future
chain so loadForDevice and importKeys (different connector frame
handlers, both unawaited) cannot interleave and wipe each other.
- #506 meshcore_connector.dart: the device-key hook loads the connected
radio's list, so the offload union reconciles within that one radio.
Existing radio-side firmware blocks left untouched (no auto-CLEAR on
migration, owner-approved). Tests: per-radio isolation, clear-sticks
across reconnect, legacy-drop, disconnect-clears, load/import race.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Switching radios showed the previous radio's channel history (including
its outgoing messages) on the new radio. The in-memory caches
_channelMessages, _conversations, and _loadedConversationKeys are keyed
by channel index / contact key, not by radio, and were never cleared on
a switch; a new radio's empty store could not overwrite them
(_loadChannelMessages only writes on a non-empty read). On-disk stores
are already per-radio (device+PSK since #277), so no re-keying or
migration is needed.
Clear the three caches in _resetConnectionHandshakeState (runs at the
start of every connect). loadAllChannelMessages and _loadMessagesForContact
repopulate from the new radio's store. Adds a regression test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Three defects. One found by the review, one found while verifying it, one
the review reported at the wrong severity.
1. The device-info reconcile was in the WRONG FUNCTION. It sat inside the
`gps:1`/`gps:0` custom-var handler instead of `_handleDeviceInfo`, so
the headless-UI state was never pulled when caps byte 2 actually
landed. It worked only because the pane re-requests when opened, and
the previous commit message's claim that it re-reads on device-info
refresh was false as written. Now fires where caps2 is parsed.
2. Device-UI state was never cleared between radios. Capability bits
refresh from every device-info frame but the values behind them do
not, so a reconnect could display one radio's notification scope as
another's. Cleared explicitly before re-reading. The review did not
find this one.
3. A confirmed matrix write was silently discarded when no matrix was
held: `_buttonMatrix?.withAssignment(...)` resolves to null and drops
a change the device had already applied. Now re-reads instead.
Review rated this High on the grounds that a user could set an action
before the initial read returned. That path is not reachable: the
dropdown only renders once a matrix exists, so the user cannot fire a
write first. The defect is real, the severity was not.
Also applied the review's structural finding: the dispatcher parsed each
frame to route it and the handler parsed it again. Parsed once and passed
down, so the two copies cannot drift and drop a valid frame.
Full suite 740 pass, analyze clean, format clean.
Epic: #474, #475
Agent: CalmBay (session d14220d9)
Client UI for the button-action matrix (#474) and the device notification
scope (#475), under Settings > Node Settings.
Speaks the canonical 0xC5 contract published by firmware in
OffbandConfigProtocol.h: one command byte, sub-code selects the surface
(0x01/0x02 scope get/set, 0x03/0x04 matrix get/set, 0x7F error). Not
folded into 0xC0, which is observer-only and would make the feature
unreachable on the headless trackers it exists for.
- Notification scope All/Self/None, labelled as this radio's buzzer and
kept distinct from the per-channel app notify mode. Re-read on open and
on every device-info refresh so a scope changed by triple-pressing the
device is never shown stale.
- Button actions per press sequence, assignable only from the action set
the DEVICE reports via its supported-actions mask, so a board with no
buzzer or no GPS never offers a choice it would refuse. Single press
defaults to unassigned.
- Long press is deliberately not assignable. Firmware owns it for CLI
rescue and power off, and remapping it could leave a screenless board
unrecoverable.
- Failures show the device's own reason via the firmware-owned reason
codes, not a generic error, and the banner persists until dismissed.
An unrecognised reason is surfaced with its raw code rather than
swallowed.
- State only ever follows the device's reply, never the request, so the
UI can never show an assignment the radio rejected and nothing is
faked or stored unacknowledged.
- A radio advertising neither capability bit gets a diagnosis, its raw
caps byte 2 and an explicit statement that no command will be sent,
rather than a blank screen that is indistinguishable from a bug.
caps2 bit assignment is firmware-confirmed: 0x01 notification scope, set
only where PIN_BUZZER is defined, 0x02 button matrix. That is the reverse
of the order the epics were filed in, so it cannot be inferred from issue
numbers.
21 protocol tests: gating, frame encoding, a lying count byte, unknown
sequence/action/scope codes, truncated and foreign frames, reason-code
mapping, and that no sequence models a long press.
NOT YET EXERCISED AGAINST A DEVICE. The firmware 0xC5 handler is written
but unmerged, so until it answers, a read shows its loading row and a
write is not confirmed. Hardware validation of the pair is the owner's
gate and has not happened.
Full suite 740 pass, analyze clean, format clean.
Epic: #474, #475
Agent: CalmBay (session d14220d9)
Two defects the adversarial review found, both verified before accepting.
1. `label` trimmed the name for display while the token stayed raw, so
two contacts differing only by surrounding whitespace rendered as
identical rows with no way to tell which one was about to be
addressed. That hides the exact identity #497 exists to preserve.
Display now defaults to the raw name.
2. Candidate deduplication used String.toLowerCase() while matching uses
the ASCII fold. Verified empirically: 'É' and 'é' compare equal under
toLowerCase and unequal under foldAscii, so of two real contacts one
silently vanished from the mention list while remaining matchable.
foldAscii is now public and is the single equivalence rule used for
both dedup and matching.
Regression tests added for both.
Full suite 719 pass, analyze clean, format clean.
Agent: CalmBay (session d14220d9)
Adds parseOffbandCaps2 alongside the existing tail-byte helpers, a
_offbandCaps2 field plus getter, and byte 2 in the device-info
capability log line.
Byte 2 sits at offset 84, deliberately NOT adjacent to byte 1 at 82:
offset 83 is already the FEM LNA state byte and every tail field is read
at a fixed absolute offset, so an adjacent insert would shift the FEM
state and make shipped clients misread a bitmask as the LNA toggle.
Firmware appends it at the end of the frame for that reason.
Absence is "no byte-2 capabilities", never an error, so any radio
predating the firmware change reads null and behaves unchanged.
No bit constants yet: byte-2 bits 0 and 1 are earmarked for firmware
epics but neither is claimed, so nothing gates on them here.
Wire contract from OffbandMesh/meshcore-firmware PR #515 (branch
feat/508-caps-byte2, commit 7664c29f). That PR is open pending this
client-side validation, tracked at #481.
Epic: #474
Agent: CalmBay (session d14220d9)
Owner ruling 2026-08-01. Mention autocomplete trimmed the contact name
when building the @[...] token while the device stores and matches it
byte-for-byte, so any name with leading or trailing whitespace was
unmentionable and never beeped. Confirmed on the wire: the contact
record keeps the 0x20, the outgoing mention drops it.
@[...] is a wire token, not display text. Entry, firmware memcpy,
advert encode, advert parse and the contact record are all verbatim by
design; this trim was the only transformation applied to a node name
anywhere, and it changed the identity of the addressee.
- MentionCandidate now separates the two concerns: `name` is raw and
goes on the wire, `label` is display-only and may be tidied. The
const constructor is preserved, so existing const call sites still
compile.
- The candidate builder keeps sender and contact names raw. Emptiness is
probed on a trimmed copy; the stored value is untouched.
- _mentionsSelf drops its own trim as a direct consequence: a raw token
requires a raw comparison, or this node stops recognising mentions of
its own name.
- Contract clause written at both the insertion site and _mentionsSelf,
stating the token is byte-for-byte and that a future .trim() tidy-up
is forbidden. The contract was silent on raw vs normalised, which is
why both sides were reasonable and incompatible.
Deliberately out of scope per the ruling: no entry-side trim in
settings_screen. It fixes no deployed name and would silently alter
deliberate spacing.
Test updated to assert the new truth: ' Ben ' does NOT match @[Ben],
and DOES match @[ Ben ]. Full suite 732 pass, analyze clean.
Agent: CalmBay (session d14220d9)
_mentionsSelf folded with String.toLowerCase(), which applies full
Unicode case mapping. Firmware adopts this same match rule with a
byte-wise fold that does not, so a node name carrying any non-ASCII
character could produce one self-mention verdict on the client and the
opposite on the device for the same message. A silent wrong answer, not
a visible failure.
Owner decision 2026-07-31: both sides fold ASCII A-Z only, so non-ASCII
names compare case-sensitively.
Extracts the rule into a testable static (mentionsName) and documents it
as a cross-repo contract: widening it, whether by accepting a bare
@name, restoring Unicode folding, or anchoring the match, is a breaking
change that ships only in an aligned client and firmware build pair.
Behaviour change: notifications for non-ASCII node names go from
case-insensitive to case-sensitive. Deliberate, per the decision above.
Epic: #475
Agent: CalmBay (session d14220d9)
resolvePathSelection returned no width and, on the override branch, a byte
count in place of a hop count. preparePathForContactSend then called
setContactPath without a width, so encodePathLen packed mode bits 00 and a
2-byte route went out as twice as many 1-byte hops (path_len 0x06 for a
6-hop 2-byte route instead of 0x46). The radio routed on wrong 1-byte
prefixes, confirmed on Bandit's 2026-08-01 log and via CoreScope. Direct
and flood were immune because neither uses the path bytes.
PathSelection now carries hashWidth and a true hop count on every branch,
using the contact's pathHashWidth as the single width authority. Both
setContactPath call sites thread it. Also addresses the override timeout
inflation half of #299 (hopCount was a byte count feeding calculateTimeout).
Gemini review found two override/history width edge cases -> split to #494.
Refs #299
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Ben's style rule: no em-dash character in copy, docs, or code comments
(it reads as an AI tell), and the repo is going public. Replaced the
em-dashes in code comments and English UI strings with commas, colons,
or periods, whichever reads best. User-facing strings were hand-tuned
for natural punctuation rather than a blanket comma.
Preserved the lone "no data" glyph placeholders (a standalone dash used
as a not-available indicator in status displays); those are a design
element, not prose.
Regenerated app_localizations*.dart from app_en.arb (the English
fallback for untranslated keys propagates to every locale's generated
file). Non-English ARB translations left untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Firmware #465 now writes RX RSSI into reserved2 (byte[3]) of the v3
contact-msg-recv frame as a clamped int8 dBm, 0 when unset
(MyMesh.cpp:550-551). Read it in the parse instead of skipping: gate on
!= 0 (RSSI is always negative for a real RX, so 0 = no data), null for
device-composed outgoing messages.
reserved1 (byte[2]) is untouched — that's #429's outgoing flag; RSSI
lives in byte[3] after the res1/res2 collision fix (#464/#465).
The Message.rssi field, persistence, param-passing, and the (hidden-
while-null) RSSI row on the Packet Path screen were all staged in #438,
so this is just the wire read + 3 gate tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Smaz is a client-side text convention with no MeshCore protocol or firmware
support (findings: GH-315). When enabled it compressed ordinary prose into
`s:`+base64, which any non-lineage client renders as garbage — the same
cross-client interop failure as the old `g:<code>` GIF token.
Phase 1 of the Smaz-removal epic (GH-314): stop sending Smaz and remove the
user-facing surface, while KEEPING decode so messages from lineage peers still
on Smaz, and any legacy compressed rows, keep rendering. Decode removal is
deferred to Phase 2 (GH-421).
Removed: the Smaz branch in prepareContact/ChannelOutboundText (Cyr2Lat branch
and the structured-payload guard preserved); connector state/API (the enabled
maps, is*/set*SmazEnabled, ensureContactSmazSettingLoaded, the warm-up and
channel loaders); the per-channel and per-contact toggles plus their
mutual-exclusion; loadSmazEnabled/saveSmazEnabled and the `*_smaz_` key prefixes
(stores kept, they also hold Cyr2Lat); l10n `channels_smazCompression` and the
orphaned `chat_compressOutgoingMessages` across 18 locales (+ regenerated
app_localizations).
Retained for Phase 2: the 5 Smaz.tryDecodePrefixed decode sites and
helpers/smaz.dart.
Safety (verified): no storage migration, the app already persists plaintext
(receive decodes before store; send stores the pre-compression text), so the
`s:` form was wire-only. ACK matching is unaffected, the expected hash derives
from prepare*'s output, so dropping compression keeps both sides hashing
plaintext.
Behavior change: the composer byte-counter now reflects raw size, so anyone who
had Smaz on can type slightly fewer chars per message.
Tests: gif_url_outbound_guard_test rewritten to pin plaintext passthrough and
decode retention. Full suite green, analyze clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tapping a DM opens the Path screen; it now shows the RF details the app
was already receiving but discarding:
- SNR: captured from the v3 contact-msg-recv frame (was skipBytes(1)).
Firmware sends (int8)(snr_dB*4), so dB = byte/4.0 (MyMesh.cpp:512) —
the old commented-out code multiplied by 4, which was 16x wrong.
- Path type: firmware sends path_len 0xFF for a direct/routed frame and
the hop count for a flood (MyMesh.cpp:545). The app collapsed 0xFF->0,
colliding "direct" with "flood, 0 hops". Capture the distinction into
Message.isFloodRoute so the Path row reads "direct (routed)" vs
"flood, N hops".
- RSSI: row wired to Message.rssi but left null — the wire byte is a
hardcoded reserved 0 today; populated once firmware ships it (#439).
Message gains snr/rssi/isFloodRoute (nullable, additive, persisted like
rxTime). 10 tests cover scaling, the 0xFF discriminator, and copyWith.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Gemini review (standards#145) flagged that a burst of 0x91 pushes (e.g. a
bulk channel edit on the device) would each trigger a full getChannels
re-sync. getChannels already guards concurrent syncs (_isSyncingChannels),
so there is no request flooding, but sequential re-syncs after each completes
would repeatedly clear+repopulate the channel list (UI flicker). Coalesce the
burst with a 400ms debounce so it settles into a single re-sync. Timer is
cancelled in dispose().
Citadel: meshcore-open-p12
Client half of the coordinated firmware+client change (firmware wadamesh#50,
verified against source).
Part A: handle push code 0x91 (pushCodeChannelsChanged) by re-polling
getChannels(force: true), so channels added/removed on the device appear in the
client without a reconnect.
Part B: read reserved1 bit0 of the V3 contact and channel message frames as an
outgoing flag. A message composed on the device (DM or channel) now renders as
sent-by-me (isOutgoing: true) instead of received. The channel self-echo guard
is bypassed for outgoing so device-composed channel messages are not dropped.
Removed the dead upstream hasPath/path-bytes branch in ChannelMessage.fromFrame
(no firmware, stock or wadamesh, appends path bytes to this frame; verified).
Backward-compatible: old firmware sends 0 (received, as today); an old client
ignores the bit.
Citadel: meshcore-open-p12
Gemini adversarial review flagged the getter's clamp: the floor used the
ATT_MTU minimum (23) where it should use the writable minimum (20 = 23-3),
overstating the cap by 3 for tiny MTUs; and an unknown MTU defaulted to
maxFrameSize (172), the wrong direction for a safety cap. Floor at 20 and
default an unknown MTU conservatively. No behaviour change on radios that
grant MTU >= 175 (they hit the maxFrameSize short-circuit).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A max-length DM built a 170-byte frame but this BLE link only writes
ATT_MTU-3 (=169) bytes, so writeCharacteristic threw. The DM path swallowed
the throw and never resolved the message, leaving it "active" forever and
silently blocking every later DM to that contact until force-stop+reconnect.
- Size: maxContact/ChannelMessageBytes take an MTU-aware frame budget
(BLE = mtuNow-3, USB/TCP = maxFrameSize); composers pass
connector.effectiveMaxFrameSize. The channel cap now also subtracts the
"Name: " prefix so a small-MTU link can't overflow.
- DM wedge: _sendMessageDirect propagates the failure and the retry callback
is awaited, so a failed send marks the message failed and drains the
per-contact queue. Added a RESP_CODE_SENT safety timeout.
- Channel wedge: sendChannelMessage clears the stuck queue id and marks the
message failed on a failed write.
- Tests: MTU-aware caps + failed-send-does-not-wedge regression.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Gemini review finding (medium): _handleErrorFrame only failed the caplog
download on RESP_CODE_ERR; control ops (enable/disable/erase) also get
RESP_CODE_ERR when the device is busy (firmware rejects non-STATUS ops while a
stream is in flight), so their completers hung 5s then threw a misleading
TimeoutException. Now fail the pending ack/status completers fast with
CaplogBusyException. analyze clean; full suite 643 green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds the received CHUNK-frame count to CaplogTruncatedException and the
on-screen error ("received X of Y bytes in N chunks"). This diagnostic
pinpointed the near-full BLE truncation: 94 chunks (all frames arrived) but
each full frame 3 bytes short = BLE MTU clipping 176-byte caplog frames to
173. Firmware chunk-size cap for BLE tracked with TopazHill.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fixes the reboot-test failure (support latched "unsupported" after the
reconnect window) and reworks the boot-log UX per Ben.
- meshcore_protocol.dart: offbandCapCaplog (0x20) + firmwareSupportsOffbandCaplog
(0x20 bit AND FIRMWARE_VER_CODE >= 17), mirroring the block-cap pattern.
- meshcore_connector.dart: supportsOffbandCaplog getter.
- serial_capture_screen.dart: gate support on the static cap bit (reactive via
the connector), not a one-shot STATUS probe, so it never latches "unsupported"
after a reboot. Derive capturing state from device STATUS so an auto-resumed
capture (post firmware #428) shows STOP not START. Cancel timers on disconnect,
re-query STATUS on reconnect. New red "Start & Reboot" (no timer) boot-log flow.
- 4 cap-gate unit tests; 18 caplog tests total green; analyze clean.
Root cause confirmed with firmware (TopazHill): caplog cap bit (0x20) is
advertised statically across reboots; the client's probe raced the reconnect
window and latched. Firmware #428 (persist flag + boot capture) is the other half.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Extends the 0xC4 caplog protocol with the companion control sub-codes the
firmware exposes (the companion has no CLI; #395 CLI verbs are repeater-
only). Firmware #417/#408.
- meshcore_protocol.dart: request sub-codes ENABLE/DISABLE/ERASE/STATUS +
builders; ACK / STATUS parsers (CaplogAck, CaplogDeviceStatus).
- meshcore_connector.dart: setDeviceCaplogEnabled / eraseDeviceCaplog /
getDeviceCaplogStatus; 0xC4 response routing split so ACK (0x10) and
STATUS (0x11) dispatch to own completers, download stream unchanged.
- 5 unit tests for builders + parsers; analyze clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Slice 1 of the client half of serial-capture (#430): the protocol layer
to download the device's serial-capture buffer over the companion link.
- meshcore_protocol.dart: cmdOffbandCaplog / respCodeOffbandCaplog = 0xC4
(NOT 0xC3, which collides with cmdOffbandFemLna; see firmware #406),
START/CHUNK/END sub-codes + request builder.
- caplog_reassembler.dart: pure START/CHUNK*/END reassembly state machine
with truncation detection, unit-tested in isolation.
- meshcore_connector.dart: downloadCaplog() + 0xC4 frame dispatch + fast
busy-reject on RESP_CODE_ERR while awaiting START.
- 9 unit tests passing; flutter analyze clean.
Integration test gated on the firmware 0xC4 fix merging. Not pushed
(human-test gate).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Wadamesh keys its history watermark (last_delivered_seq) by client_id. We
sent none, so we shared the empty-string slot with every other MeshCore
client on the machine: whichever connected first drained the device history
ring and the next app got NO_MORE_MESSAGES for frames it never received.
cid_len is 6 by necessity, not preference. Stock reads cmd_frame[1..7] as
reserved with the app name at a fixed offset 8; Wadamesh reads the name at
2 + cid_len. Only 6 puts the name at 8 on both, so one frame serves both
firmwares with no firmware change. Covered by app_start_frame_test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adversarial review caught a data-loss defect in the reactor-name
fallback. For a reaction in a room from an author not in the local
contacts, _resolveContactSenderName returns null and the chain fell
through to reactionContact?.name. But reactionContact is the room server,
not the reactor, so every unknown-author reaction was attributed to the
room's display name. Two distinct unknown authors then both resolved to
that one name, so applyReaction's per-reactor dedup treated the second as
a duplicate and dropped it, count and all.
Fall back to the per-author hex (from the frame's own fourByteRoomContact
Key) instead, which is unique per reactor. In a true 1:1 that hex is
empty and the contact genuinely is the reactor, so its name stays the
correct fallback.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reactions stored only a Map<String,int> of emoji to count, so "who
reacted" was unanswerable and MeshCore One's tap-to-see-who had nothing
to show against. This captures the reactor for every reaction, both our
own r: format and the PocketMesh / MeshCore One format.
Data layer only. The tap-a-badge-to-see-who UI is a separate follow-on;
this makes the data available and does not change any screen.
Additive by design. A new reactionSenders field (Map<String,List<String>>,
emoji -> reactor names) sits alongside the existing count map, serialised
under a new JSON key. Both maps are always written. An older build reading
a newer store ignores the unknown key and still gets correct counts from
reactions; it never hits the hard `value as int` cast that would fail the
whole message-list load (the #355 data-loss shape). Records written before
this field load with an empty sender map, so their counts survive with
names simply absent.
- ChannelMessage + Message: new reactionSenders field, constructor,
copyWith
- both stores: serialise the new key; deserialise via a shared
ReactionHelper.reactionSendersFromJson that returns empty for a missing
key and skips malformed entries rather than throwing
- applyReaction: records the reactor and dedups per reactor per emoji.
This is persistent dedup (survives restart), unlike the connector's
in-memory processed-set. A pre-existing count with no sender list is
incremented from its stored value, not recomputed from the partial
list, so old counts are preserved.
- connector: threads the reactor name into all reaction paths. Channel
uses the frame sender; a room resolves the author via
fourByteRoomContactKey; an outgoing reaction is attributed to self.
- pending queue: retry now hands the stored reactor to its callback
Tests: reactor capture, two-reactor count, same-reactor no-double-count,
pre-#383 count preservation, and JSON round-trip + backward-compat +
malformed-entry handling for the new field. Full suite 604 passing.
Epic #376.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two of three Gemini findings held up against the code; the third did not.
Accepted, message-swallow risk: the emoji check looked only at the first
rune against a range list that included arrows, so a multi-line message
beginning with an arrow and ending in eight Crockford characters would
have been consumed and displayed as a reaction whose "emoji" was the
whole sentence. Ben's own capture has U+2192 mid-sentence in ordinary
channel traffic, so this was reachable, not theoretical.
- drop U+2190-U+21FF and U+2934-U+2935; arrows carry the Unicode Emoji
property but read as punctuation in prose
- cap the emoji segment at 8 runes, which is clear of the longest ZWJ
sequence and nowhere near a sentence. This is the guard that holds
regardless of how the range list evolves.
Accepted, dedup: the contact path keyed on target hash plus emoji only.
A room server is many-party over the 1:1 transport, so two members
sending the same emoji collapsed into one and the count stuck at 1, the
same defect already fixed on the channel path. The reacting author
prefix is now part of the key; in a true 1:1 it is empty and the key is
unchanged.
Rejected, room-server sender matching: the review claimed getSenderName
resolves to the room itself and that room text carries a "Sender: "
prefix. Neither is true. _resolveContactSenderName resolves the author
via fourByteRoomContactKey to the real per-sender contact name, and room
message text is stored bare with the author in its own field. One real
sub-case survives: an author who is not in our contacts resolves to null
and will not match, which degrades to queue-then-expire with a warn log.
Also rejected: switching @[ lookup from first to last occurrence. The
reference implementation uses the first, and diverging risks mismatching
payloads it accepts.
Reactions sent from MeshCore One arrived as junk text: an emoji line
followed by an 8-character token such as "dyps6yf0". Those tokens are
Crockford Base32 target-message hashes, the second line of a two-line
reaction payload we did not recognise.
Receive-side only. Offband keeps sending its own r:hhhh:ii format; their
client already parses ours, so nothing about what we transmit changes.
Wire format (confirmed against a live capture, see #378):
channel: {emoji}@[{targetSender}]\n{hash}
direct: {emoji}\n{hash}
hash: sha256(body utf8 + timestamp uint32 LE seconds)[0:5],
Crockford Base32, 8 chars, lowercase
The body is hashed without the channel "SenderName: " prefix, which is
why the sender travels in @[...] instead.
- crockford_base32.dart: encode and normalise, 12-bit accumulator so the
web target's 32-bit bitwise ops cannot truncate a 40-bit value
- pocketmesh_reaction.dart: hash and a parser mirroring the reference
implementation, with a conservative leading-emoji check so a real
message is never swallowed
- reaction_helper.dart: ReactionInfo carries a dialect; applyReaction
picks the matching hash and, for the channel form, requires an exact
sender-name match (a node name can carry emoji and variation
selectors)
- pending_reactions.dart: a reaction arriving before its target is held
and retried rather than silently dropped, bounded at 50 entries with a
15 minute TTL and a warn-level log on expiry (closes the silent-drop
path in #382)
- the channel dedup key now includes the reacting sender, so two people
sending the same emoji no longer collapse into one
- notification tray summarises the foreign format as a reaction instead
of showing the raw token
Known limitation: on channels with Smaz or Cyr2Lat enabled the wire text
differs from the text we store, so a hash computed by another client
will not match. Flagged for a decision rather than worked around.
Epic #376. Fast-follow #383 adds reactor identity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The persisted store was being overwritten with the windowed in-memory list, so
any channel or DM with more than _messageWindowSize (200) messages lost
everything older than the newest 200 on the first save after load. Observed
live: Public went 232 -> 201 in one session on a build that already had the
#333 fix, so this was not the race - it was windowing truncating the store.
This is the slow-erosion cause behind the whole 566 -> 231 -> 206 -> 201
history.
saveChannelMessages / saveMessages now MERGE into the persisted set instead of
overwriting: upsert by message identity so older persisted messages are kept,
new ones added, and the in-memory copy wins for edits/reactions/status. If the
existing history fails to decode, the save aborts loudly rather than merging
into an empty base and truncating (SAFELANE 6).
Deletion is now an explicit path - removeChannelMessage / removeMessage - since
the app deletes individual messages. Routing delete through the merging save
would resurrect them; the connector's deleteChannelMessage / deleteMessage now
call the explicit remove.
Windowing stays for display and memory; it no longer dictates what is stored.
Tests: 250-message history survives a 200-window save; new message appends;
edit is captured not duplicated; delete does not resurrect. 508 pass, analyze
and format clean.
Cost: a save now re-reads and re-encodes the channel/contact history. Cheap
with drift's per-key writes (#335); a future append-only schema removes even
that.
Measured, not guessed. Cumulative per-code frame timing named the owner:
frame handler cumulative (2002ms total sync): code=3 2002ms/105calls
code=3 is RESP_CODE_CONTACT, and it accounted for ALL of the synchronous
frame time at ~19ms per contact. That is why a per-call 100ms threshold never
fired: no single call is slow, the volume is.
The 19ms is _handleContact calling notifyListeners() for every contact. Each
notification is a synchronous full-tree rebuild, and each rebuild walks the
contact list, so the cost grows with the address book. The channel handler
directly alongside already guards its notify with _isLoadingChannels; the
contact handler had no equivalent.
Notifications are now throttled to one per 250ms while _isLoadingContacts is
set, and remain immediate outside a pull so single adverts still update live.
Throttled rather than suppressed so the sync progress bar keeps advancing. The
existing notifyListeners() when the pull completes still fires, so the final
state cannot be stale.
Scope of what this accounts for, stated honestly: measured synchronous frame
work was ~2s per report window and the pull showed ~4s total, against a ~30s
observed freeze. This removes the largest measured contributor. If the stall
persists, the remainder is outside the frame handlers and the cumulative
counters will show frame time has dropped, which narrows it further.
flutter analyze clean, dart format clean, 494 tests pass.
The conversation-load metric was misread, and nearly caused a fix to the wrong
thing. It reported store=30138ms across 300 contacts, which looked like
loadMessages costing ~100ms each. It is not:
- The device-wide fallback key `messages_<dev>` does not exist on the
reporting machine, and only 11 per-contact message keys exist in total, so
loadMessages does three in-memory map lookups and returns empty.
- storeWatch spans an `await`, so it measures WALL time including event-loop
queueing, not work. merge, which is synchronous and has no await, measures
0ms - consistent with the isolate being busy elsewhere while those awaits
wait their turn.
So the 30s is ~300 awaits each waiting for a busy isolate, not 30s of storage
work. The metric label now says so, rather than inviting the same misreading.
The per-call 100ms threshold cannot see the actual shape here: a handler that
takes 40ms and runs 234 times owns ~9s of isolate time and never trips it.
Frame handlers now accumulate synchronous time per frame code and report the
top codes by total, which names the owner regardless of per-call cost.
Diagnostics only, no behaviour change.
flutter analyze clean, dart format clean, 494 tests pass.
The startup loads are now clean (all steps ~0ms, down from 37792ms), but the
contact pull still stalls in multi-second bursts: 8.0s and 7.7s at +25s..+52s
on the last run, matching the reported freeze from ~15s to ~49s.
No frame handler exceeds 100ms, because the work is not in one.
_loadMessagesForContact is fired unawaited from the contact handler, so it runs
after the handler returns and escapes that timing entirely.
This times it directly and reports cumulative store time versus merge+notify
time every 25 contacts, so the next run says which half is responsible rather
than requiring another guess. Two candidates it will discriminate:
- store: prefs read + jsonDecode per contact
- merge+notify: the per-contact notifyListeners, i.e. one full widget tree
rebuild per contact, each of which walks the contact list
Diagnostics only, no behaviour change.
flutter analyze clean, dart format clean.
Two reported freezes: one shortly after connect, one during channel sync,
with the sync progress bar never repainting and then completing all at once.
That is the UI isolate being blocked, and the log already shows the
signature - a long silence followed by many frames sharing one timestamp,
which is queued input flushing after the loop frees up.
The previous fix (batching 600 per-contact notifications) moved the stall from
43.6s to 40.1s, which is noise. It was the wrong culprit, so this replaces
guessing with measurement:
- The post-SELF_INFO cache load is split into named steps, each timed and
logged as `startup-load <step> took Nms`. The step order is unchanged, and
cachedChannels still runs before channelMessages so PSK-keyed history
resolves.
- Every RX frame handler is timed, and any that occupies the isolate for more
than 100ms logs `slow frame handler: code=N blocked UI for Nms`. This names
the handler wherever it is, rather than requiring a guess about which one.
No behaviour change. Diagnostics only, kept as Perf-tagged logging because
this class of stall is worth catching again.
flutter analyze clean, dart format clean, 469 tests pass.
Diagnosed from the app's own log. After SELF_INFO the client went completely
silent for 43.7s, then flushed a burst of radio frames all sharing the
identical timestamp - queued while the event loop was blocked, not a slow
radio.
Cause: loadContactCache() looped over every cached contact calling
_ensureContactSmazSettingLoaded and _ensureContactCyr2LatSettingLoaded, and
each of those notified on completion. With 300 contacts that is 600
notifyListeners() calls in a burst, each rebuilding the whole widget tree, and
each rebuild itself walks the contact list. O(contacts^2) on the UI isolate.
Both loaders are now awaitable and take notify: false for bulk use.
loadContactCache awaits them all and notifies once.
Measured on the reporting device (radio 4f6585264e, TCP): 300 contacts, 1090
discovered contacts, 12 channels holding 1841 messages, 6.9 MB prefs file.
Log evidence: 43.7s stall between "Pulled battery" and the first channel
response, followed by same-millisecond frame delivery.
This is the Windows-visible freeze in #306. The platform asymmetry is
consistent with contact-list size rather than anything desktop-specific, so
whether Android is genuinely faster or simply has a smaller address book on
its paired radio is still open - #306 stays open pending a like-for-like
re-measure.
flutter analyze clean, dart format clean, 469 tests pass.
Channel history is keyed by the channel's PSK (#194), and the resolver set on
the store reads that PSK out of _channels. After SELF_INFO, loadCachedChannels()
and loadAllChannelMessages() were both fired without awaiting, so the messages
could load while _channels was still empty. The resolver then returned null,
_storageKey() silently fell back to the legacy slot-index key, and channels
read as empty even though their history was intact under the PSK key.
Silent because the fallback is a legitimate code path: there is no error, no
parse failure and no log line, just an empty list. It is indistinguishable
from data loss to the user.
Observed on a live radio: Public held 566 messages under
channel_messages_<dev>psk_8b3387e9..., and the app displayed nothing.
loadAllChannelMessages() now runs only after loadCachedChannels() completes.
The startup call in main.dart already awaits in the right order; it runs before
any radio is connected, which is the source of the "Public key hex is not set"
warnings at launch, and is harmless once this post-connect load is correct.
Was previously committed against #291 on the nav-redesign branch; split out
here under its own issue.