The timer ran timeoutMs while the message printed
(timeoutMs / 1000).ceil(), so every window in (4000, 5000] announced
"timeout after 5 seconds". The owner's 0-hop window was 4074 ms and fired
at 4.07 s while claiming 5, which is what made the behaviour look
arbitrary rather than deterministic: the number shown was never the
number used.
The service now throws a typed RepeaterCommandTimeout carrying the window
that was actually armed, formatted to one decimal. Every existing caller
already stringifies the error, so all of them inherit an honest figure
without being touched; the CLI screen additionally renders it through a
new localized string rather than the generic error wrapper.
Tests pin that 4074 reports 4.1 rather than 5, that the new 28748 ms
budget reports 28.7 rather than 29, and that two windows inside the same
second no longer collapse to the same text, which was the defect's
signature.
Part of epic #473, stacked on #528 and #529.
The comment justified not modelling command execution time by claiming
`wifi on 30` does real work bringing up an interface. It does not. The
firmware handler sets a persistence deadline and sprintf's its reply
immediately (CommonCLI.cpp, "wifi on"), so execution is near-instant.
The owner caught it: he recalled reissuing the command several times
against on-screen errors, not one command taking 20 seconds.
The measurement supports him. On the 20.33 s case the reply carried
claimed=01:56:32 against a command sent at 01:56:26.938, and the RF frame
did not reach our radio until 01:56:47.271979. So roughly 5 s to reach the
repeater and be answered, then roughly 15 s in its transmit queue. The
tail is scheduling on both radios.
No behaviour change. The budget is unchanged and still has to tolerate a
20.33 s round trip; only the stated reason was wrong, and a wrong reason
in a load-bearing comment re-causes the bug later.
The CLI timeout used calculateTimeout, which mirrors the firmware's
calcDirectTimeoutMillisFor and estimates ONE-WAY delivery. A CLI command
is a request, an execution and a reply, so the budget was structurally
short.
Measured against rpt-01 on 910.525 MHz / SF7 / BW 62.5k / CR 4:5, where
the old window was 4074 ms: 27 commands sent, 12 replies matched, and 4
of those 12 arrived after the client had already given up, at 6.17 s,
8.28 s, 15.52 s and 20.33 s against a median of 2.75 s.
calculateCliTimeout sums four terms, each with a source rather than a
chosen value:
outbound leg calculateTimeout, which is what it actually models
cliReplyDelayMs 600, firmware CLI_REPLY_DELAY_MILLIS, unconditional
reply leg the reply is a second packet the ACK formula omits
retrieval budget replies are pull-based; the radio raises MSG_WAITING
and the app must ask, granting itself 5000 ms per
attempt across 3 retries
The retrieval term is derived from the sync constants rather than
restated, so the command timeout cannot drift below the layer it depends
on. That is the #530 invariant holding by construction, not by two
numbers being maintained in agreement.
Execution time is deliberately not modelled. The same verb, wifi on 30,
returned in both 2.31 s and 20.33 s, so it is not a per-command constant
that could be tabulated. The retrieval term carries that tail.
The reply leg uses physics only. The predictor is trained on
direct-message ACK latency, so asking it about a CLI reply leg would be
extrapolation; that is #534 and #535, not this change.
Tests cover the construction, the never-below-retrieval invariant, growth
with path length, coverage of the 20.33 s worst case actually observed,
and negatively that the direct-message ACK path is untouched.
Part of epic #473. Does not change the reported duration string (#531)
or the stale-prefix fallback (#532), and does nothing for commands that
draw no reply at all (#541).
A repeater reply that arrived after its command's window closed hit
`if (commandId.isEmpty) return;` in RepeaterCommandService.handleResponse
and was dropped with no log and no UI. Since the command timeout can be
shorter than the app's own message-retrieval budget, that made "the
command ran but the response was never reported" the normal outcome
rather than an edge case, and it affected every repeater screen, not
just the CLI one.
RepeaterCommandService now remembers a timed-out command's prefix for two
minutes, so a reply arriving afterwards can be attributed to the request
it answers. Replies that reach no waiting command are handed to a new
onUnmatchedResponse sink and logged through appLogger. No path through
handleResponse returns without either completing a command, surfacing the
payload, or logging why it could not.
The CLI screen renders these as a distinct history entry naming the
original command and how late it was. The settings screen applies the
value if it is a `get` reply and tells the user it arrived late.
repeater_status_screen already parsed responses independently of the
service, so it had no silent-loss path to fix.
Part of epic #473. Does not change the timeout window itself (#529),
the reported duration (#531), or the stale-prefix fallback (#532).
extractScalarValue required ' = ' (space-equals-space) and trimRight()'d first,
so a blank field's reply 'key =' (no space after =) failed the match and the
whole 'mqtt.broker.N.field =' line leaked in as the value. Split on the first
'=' and trim instead: blank -> empty, values with '=' preserved. Regression
tests added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Badge stays observers-primary; tooltip/long-press now read 'N observers . M
observations' so they reconcile with CoreScope's 'Observations (N)' feed (the
two are distinct metrics: 15 distinct observers vs 34 total sightings). Refresh
feedback moved from a bottom snackbar to an auto-dismissed top MaterialBanner,
out of the way of the composer. Part of #524.
Agent: QuietSnow (session 31eaba02)
Unifies the previously-duplicated send timestamps (frame builder vs outgoing
message each called now() separately) into one monotonic-per-channel value, so
every channel send has a unique (ts, channel_idx) key. That is the client-side
guarantee the 0xC6 correlation needs (VioletBarn caught that AES-128-ECB
determinism would otherwise make two same-second messages a wrong-hash query).
Adds transient onAirHash + coreScopeObserverCount to ChannelMessage. Part of #524.
Agent: QuietSnow (session 31eaba02)
Builder/parser for the client-issued 0xC6 CMD_OFFBAND_PKT_HASH query (per
firmware #611 contract) and firmwareSupportsPktHash (cap2 0x08 + ver >= 22).
Inert until firmware advertises the capability. Part of epic #524.
Agent: QuietSnow (session 31eaba02)
Best-effort GET /api/packets?hash=H&groupByHash=true against a CoreScope
instance (default map.okimesh.org); returns observer_count or null. Never
throws, so the chat UI degrades silently to radio-only when offline or the
hash is unknown. Part of epic #524.
Agent: QuietSnow (session 31eaba02)
7 taps on the About version row (rolling ~3s window) reveal a neutral
'Experimental' settings category; persisted on AppSettings, survives
restart, re-hidden from a switch inside the section. Empty of toggles
for now (Fast Sync #118, CoreScope #524 land into it later).
Countdown snackbars after tap 4; already-unlocked feedback. Version-text
tap is absorbed so it counts without opening the About dialog.
Agent: QuietSnow (session 31eaba02)
Children of #471 (parent stays open for AAB hardware validation, #507).
The block list was one global unscoped list on the phone, unioned onto
every radio on connect. A stale block for one of the owner's own radios
therefore rode onto every fresh/erased radio, and clearing a radio never
stuck: any other radio still holding it re-seeded the global list on
connect, which re-pushed it back.
- #505 block_store.dart: scope keys/names by connected device key (like
the app's other stores); dropLegacyGlobal() deletes the legacy
block_keys_v1/block_names_v1 (drop-and-start-fresh, owner-approved).
No global list.
- #505 block_service.dart: load() drops the legacy global and starts
empty; loadForDevice(deviceKey) swaps the in-memory set per radio and
runs the #250 self-heal. All mutating ops serialized through a Future
chain so loadForDevice and importKeys (different connector frame
handlers, both unawaited) cannot interleave and wipe each other.
- #506 meshcore_connector.dart: the device-key hook loads the connected
radio's list, so the offload union reconciles within that one radio.
Existing radio-side firmware blocks left untouched (no auto-CLEAR on
migration, owner-approved). Tests: per-radio isolation, clear-sticks
across reconnect, legacy-drop, disconnect-clears, load/import race.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A transient upstream 503 (e.g. Hugging Face) hard-failed the model download and
surfaced it as an error (#229). The download requests (HEAD, single GET, range
GET) now go through sendModelDownloadWithRetry: 5xx/429 responses and network
exceptions retry with bounded exponential backoff (1..30s, <=5 attempts, honors
Retry-After, cancellable); terminal 4xx (e.g. 404) fail immediately. The retry
logic is a pure injectable top-level function, unit-tested (503-then-200,
persistent-503, 404-no-retry, network-exception, cancel).
Manual retry after a sustained failure is the existing "Download model" button
(re-invokes the download); no new UI added. Build of epic #422.
Switching radios showed the previous radio's channel history (including
its outgoing messages) on the new radio. The in-memory caches
_channelMessages, _conversations, and _loadedConversationKeys are keyed
by channel index / contact key, not by radio, and were never cleared on
a switch; a new radio's empty store could not overwrite them
(_loadChannelMessages only writes on a non-empty read). On-disk stores
are already per-radio (device+PSK since #277), so no re-keying or
migration is needed.
Clear the three caches in _resetConnectionHandshakeState (runs at the
start of every connect). loadAllChannelMessages and _loadMessagesForContact
repopulate from the new radio's store. Adds a regression test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Client UI for the button-action matrix (#474) and the device notification
scope (#475), under Settings > Node Settings.
Speaks the canonical 0xC5 contract published by firmware in
OffbandConfigProtocol.h: one command byte, sub-code selects the surface
(0x01/0x02 scope get/set, 0x03/0x04 matrix get/set, 0x7F error). Not
folded into 0xC0, which is observer-only and would make the feature
unreachable on the headless trackers it exists for.
- Notification scope All/Self/None, labelled as this radio's buzzer and
kept distinct from the per-channel app notify mode. Re-read on open and
on every device-info refresh so a scope changed by triple-pressing the
device is never shown stale.
- Button actions per press sequence, assignable only from the action set
the DEVICE reports via its supported-actions mask, so a board with no
buzzer or no GPS never offers a choice it would refuse. Single press
defaults to unassigned.
- Long press is deliberately not assignable. Firmware owns it for CLI
rescue and power off, and remapping it could leave a screenless board
unrecoverable.
- Failures show the device's own reason via the firmware-owned reason
codes, not a generic error, and the banner persists until dismissed.
An unrecognised reason is surfaced with its raw code rather than
swallowed.
- State only ever follows the device's reply, never the request, so the
UI can never show an assignment the radio rejected and nothing is
faked or stored unacknowledged.
- A radio advertising neither capability bit gets a diagnosis, its raw
caps byte 2 and an explicit statement that no command will be sent,
rather than a blank screen that is indistinguishable from a bug.
caps2 bit assignment is firmware-confirmed: 0x01 notification scope, set
only where PIN_BUZZER is defined, 0x02 button matrix. That is the reverse
of the order the epics were filed in, so it cannot be inferred from issue
numbers.
21 protocol tests: gating, frame encoding, a lying count byte, unknown
sequence/action/scope codes, truncated and foreign frames, reason-code
mapping, and that no sequence models a long press.
NOT YET EXERCISED AGAINST A DEVICE. The firmware 0xC5 handler is written
but unmerged, so until it answers, a read shows its loading row and a
write is not confirmed. Hardware validation of the pair is the owner's
gate and has not happened.
Full suite 740 pass, analyze clean, format clean.
Epic: #474, #475
Agent: CalmBay (session d14220d9)
Two defects the adversarial review found, both verified before accepting.
1. `label` trimmed the name for display while the token stayed raw, so
two contacts differing only by surrounding whitespace rendered as
identical rows with no way to tell which one was about to be
addressed. That hides the exact identity #497 exists to preserve.
Display now defaults to the raw name.
2. Candidate deduplication used String.toLowerCase() while matching uses
the ASCII fold. Verified empirically: 'É' and 'é' compare equal under
toLowerCase and unequal under foldAscii, so of two real contacts one
silently vanished from the mention list while remaining matchable.
foldAscii is now public and is the single equivalence rule used for
both dedup and matching.
Regression tests added for both.
Full suite 719 pass, analyze clean, format clean.
Agent: CalmBay (session d14220d9)
Adds parseOffbandCaps2 alongside the existing tail-byte helpers, a
_offbandCaps2 field plus getter, and byte 2 in the device-info
capability log line.
Byte 2 sits at offset 84, deliberately NOT adjacent to byte 1 at 82:
offset 83 is already the FEM LNA state byte and every tail field is read
at a fixed absolute offset, so an adjacent insert would shift the FEM
state and make shipped clients misread a bitmask as the LNA toggle.
Firmware appends it at the end of the frame for that reason.
Absence is "no byte-2 capabilities", never an error, so any radio
predating the firmware change reads null and behaves unchanged.
No bit constants yet: byte-2 bits 0 and 1 are earmarked for firmware
epics but neither is claimed, so nothing gates on them here.
Wire contract from OffbandMesh/meshcore-firmware PR #515 (branch
feat/508-caps-byte2, commit 7664c29f). That PR is open pending this
client-side validation, tracked at #481.
Epic: #474
Agent: CalmBay (session d14220d9)
Owner ruling 2026-08-01. Mention autocomplete trimmed the contact name
when building the @[...] token while the device stores and matches it
byte-for-byte, so any name with leading or trailing whitespace was
unmentionable and never beeped. Confirmed on the wire: the contact
record keeps the 0x20, the outgoing mention drops it.
@[...] is a wire token, not display text. Entry, firmware memcpy,
advert encode, advert parse and the contact record are all verbatim by
design; this trim was the only transformation applied to a node name
anywhere, and it changed the identity of the addressee.
- MentionCandidate now separates the two concerns: `name` is raw and
goes on the wire, `label` is display-only and may be tidied. The
const constructor is preserved, so existing const call sites still
compile.
- The candidate builder keeps sender and contact names raw. Emptiness is
probed on a trimmed copy; the stored value is untouched.
- _mentionsSelf drops its own trim as a direct consequence: a raw token
requires a raw comparison, or this node stops recognising mentions of
its own name.
- Contract clause written at both the insertion site and _mentionsSelf,
stating the token is byte-for-byte and that a future .trim() tidy-up
is forbidden. The contract was silent on raw vs normalised, which is
why both sides were reasonable and incompatible.
Deliberately out of scope per the ruling: no entry-side trim in
settings_screen. It fixes no deployed name and would silently alter
deliberate spacing.
Test updated to assert the new truth: ' Ben ' does NOT match @[Ben],
and DOES match @[ Ben ]. Full suite 732 pass, analyze clean.
Agent: CalmBay (session d14220d9)
_mentionsSelf folded with String.toLowerCase(), which applies full
Unicode case mapping. Firmware adopts this same match rule with a
byte-wise fold that does not, so a node name carrying any non-ASCII
character could produce one self-mention verdict on the client and the
opposite on the device for the same message. A silent wrong answer, not
a visible failure.
Owner decision 2026-07-31: both sides fold ASCII A-Z only, so non-ASCII
names compare case-sensitively.
Extracts the rule into a testable static (mentionsName) and documents it
as a cross-repo contract: widening it, whether by accepting a bare
@name, restoring Unicode folding, or anchoring the match, is a breaking
change that ships only in an aligned client and firmware build pair.
Behaviour change: notifications for non-ASCII node names go from
case-insensitive to case-sensitive. Deliberate, per the decision above.
Epic: #475
Agent: CalmBay (session d14220d9)
resolvePathSelection returned no width and, on the override branch, a byte
count in place of a hop count. preparePathForContactSend then called
setContactPath without a width, so encodePathLen packed mode bits 00 and a
2-byte route went out as twice as many 1-byte hops (path_len 0x06 for a
6-hop 2-byte route instead of 0x46). The radio routed on wrong 1-byte
prefixes, confirmed on Bandit's 2026-08-01 log and via CoreScope. Direct
and flood were immune because neither uses the path bytes.
PathSelection now carries hashWidth and a true hop count on every branch,
using the contact's pathHashWidth as the single width authority. Both
setContactPath call sites thread it. Also addresses the override timeout
inflation half of #299 (hopCount was a byte count feeding calculateTimeout).
Gemini review found two override/history width edge cases -> split to #494.
Refs #299
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Follow-up to #457. Swept the em-dash character (U+2014) out of the test
tree (test descriptions and comments) and cleaned 8 em-dashes that
landed in lib comments via the #456 refactor after #457 merged, so the
tree is back to zero. Same rules: replaced with commas/colons/periods,
preserved the lone "no data" glyph placeholders, left non-English ARB
untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Profile is now a container of capability-scoped sections (schema_version 2):
- wifi (WifiConfig) — any wifi device
- mqtt (MqttSection = region + status_interval + brokers) — observer capability
so MQTT is defined once and shared by observer / observer-repeater /
observer-companion, never redone per device. Future radio/repeater/companion/
display sections slot in alongside.
Parser reads the sectioned YAML and rejects the old flat v1 layout with a clear
message. Enumerator reads from sections but emits the SAME firmware keys, so
apply/diff/screens are unchanged. 45 config-profile tests updated + green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Ben's style rule: no em-dash character in copy, docs, or code comments
(it reads as an AI tell), and the repo is going public. Replaced the
em-dashes in code comments and English UI strings with commas, colons,
or periods, whichever reads best. User-facing strings were hand-tuned
for natural punctuation rather than a blanket comma.
Preserved the lone "no data" glyph placeholders (a standalone dash used
as a not-available indicator in status displays); those are a design
element, not prose.
Regenerated app_localizations*.dart from app_en.arb (the English
fallback for untranslated keys propagates to every locale's generated
file). Non-English ARB translations left untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Firmware #465 now writes RX RSSI into reserved2 (byte[3]) of the v3
contact-msg-recv frame as a clamped int8 dBm, 0 when unset
(MyMesh.cpp:550-551). Read it in the parse instead of skipping: gate on
!= 0 (RSSI is always negative for a real RX, so 0 = no data), null for
device-composed outgoing messages.
reserved1 (byte[2]) is untouched — that's #429's outgoing flag; RSSI
lives in byte[3] after the res1/res2 collision fix (#464/#465).
The Message.rssi field, persistence, param-passing, and the (hidden-
while-null) RSSI row on the Packet Path screen were all staged in #438,
so this is just the wire read + 3 gate tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Smaz is a client-side text convention with no MeshCore protocol or firmware
support (findings: GH-315). When enabled it compressed ordinary prose into
`s:`+base64, which any non-lineage client renders as garbage — the same
cross-client interop failure as the old `g:<code>` GIF token.
Phase 1 of the Smaz-removal epic (GH-314): stop sending Smaz and remove the
user-facing surface, while KEEPING decode so messages from lineage peers still
on Smaz, and any legacy compressed rows, keep rendering. Decode removal is
deferred to Phase 2 (GH-421).
Removed: the Smaz branch in prepareContact/ChannelOutboundText (Cyr2Lat branch
and the structured-payload guard preserved); connector state/API (the enabled
maps, is*/set*SmazEnabled, ensureContactSmazSettingLoaded, the warm-up and
channel loaders); the per-channel and per-contact toggles plus their
mutual-exclusion; loadSmazEnabled/saveSmazEnabled and the `*_smaz_` key prefixes
(stores kept, they also hold Cyr2Lat); l10n `channels_smazCompression` and the
orphaned `chat_compressOutgoingMessages` across 18 locales (+ regenerated
app_localizations).
Retained for Phase 2: the 5 Smaz.tryDecodePrefixed decode sites and
helpers/smaz.dart.
Safety (verified): no storage migration, the app already persists plaintext
(receive decodes before store; send stores the pre-compression text), so the
`s:` form was wire-only. ACK matching is unaffected, the expected hash derives
from prepare*'s output, so dropping compression keeps both sides hashing
plaintext.
Behavior change: the composer byte-counter now reflects raw size, so anyone who
had Smaz on can type slightly fewer chars per message.
Tests: gif_url_outbound_guard_test rewritten to pin plaintext passthrough and
decode retention. Full suite green, analyze clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- config_profile_diff.dart: pure current->new diff builder (add/change,
drops no-ops, flags danger + secret rows). 4 tests.
- splitProfileWrites: partitions writes into safe vs credential/identity for
the two-tier apply; a mixed broker splits, enabled rides the safe half only.
- ConfigProfilePreviewScreen: full sub-screen — reads current state, renders
the diff (amber overwrites, red danger section), normal Apply for plain
config + a separate red gate (with confirm dialog listing exactly which
credential/identity values change) for the danger set. Secrets masked.
Re-diffs after apply for partial-save recovery.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Device-agnostic write enumerator (config_profile_writes.dart): a profile ->
ordered flat writes + per-broker field maps. Skips null/empty (never clobbers),
skips jwt_token (live-minted), holds enabled out for last-write, flags the
danger set (username/password/jwt_owner/jwt_email + wifi.pwd) for #406's gate.
Observer executor (observer_apply_service.dart): flats via setFlat, brokers via
the existing saveBroker (disable-first, fields, enabled LAST, stop-on-error
partial-safe #80); reads current enabled to preserve it when a profile omits it.
Result labels name keys/slots only, never values (no secret leak).
8 enumerator tests. Executor is thin orchestration over the tested service;
end-to-end covered by #408 hardware.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tapping a DM opens the Path screen; it now shows the RF details the app
was already receiving but discarding:
- SNR: captured from the v3 contact-msg-recv frame (was skipBytes(1)).
Firmware sends (int8)(snr_dB*4), so dB = byte/4.0 (MyMesh.cpp:512) —
the old commented-out code multiplied by 4, which was 16x wrong.
- Path type: firmware sends path_len 0xFF for a direct/routed frame and
the hop count for a flood (MyMesh.cpp:545). The app collapsed 0xFF->0,
colliding "direct" with "flood, 0 hops". Capture the distinction into
Message.isFloodRoute so the Path row reads "direct (routed)" vs
"flood, N hops".
- RSSI: row wired to Message.rssi but left null — the wire byte is a
hardcoded reserved 0 today; populated once firmware ships it (#439).
Message gains snr/rssi/isFloodRoute (nullable, additive, persisted like
rxTime). 10 tests cover scaling, the 0xFF discriminator, and copyWith.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A max-length DM built a 170-byte frame but this BLE link only writes
ATT_MTU-3 (=169) bytes, so writeCharacteristic threw. The DM path swallowed
the throw and never resolved the message, leaving it "active" forever and
silently blocking every later DM to that contact until force-stop+reconnect.
- Size: maxContact/ChannelMessageBytes take an MTU-aware frame budget
(BLE = mtuNow-3, USB/TCP = maxFrameSize); composers pass
connector.effectiveMaxFrameSize. The channel cap now also subtracts the
"Name: " prefix so a small-MTU link can't overflow.
- DM wedge: _sendMessageDirect propagates the failure and the retry callback
is awaited, so a failed send marks the message failed and drains the
per-contact queue. Added a RESP_CODE_SENT safety timeout.
- Channel wedge: sendChannelMessage clears the stuck queue id and marks the
message failed on a failed write.
- Tests: MTU-aware caps + failed-send-does-not-wedge regression.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The 172->176 ceiling raise made the existing "drops oversized frames and
resyncs" test's 173-byte frame a valid length, so it buffered instead of
dropping (it was red). Derive the over-max length from the ceiling so it stays
a genuine oversized frame and still exercises drop-and-resync, confirming the
higher ceiling doesn't weaken corrupt-frame recovery. Full suite: 643 green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The frame decoder capped companion frames at 172, but the firmware's real
MAX_FRAME_SIZE is 176 (BaseSerialInterface.h, "+4 for transport codes").
Full-size caplog CHUNK frames (176 B) were silently rejected; only the final
sub-172 partial chunk survived, so a download reassembled just the last chunk
(e.g. "received 95 of 5141"). Caplog is the first feature to use full frames,
so nothing exposed this before.
- usb_serial_frame_codec.dart: usbSerialMaxPayloadLength 172 -> 176.
- serial_capture_screen.dart: erase-on-start (Start / Start&Reboot) for a clean
session ("didn't start at 0").
- 3 decoder regression tests; 113 tests total green; analyze clean.
Root cause confirmed with firmware (TopazHill): firmware streams the full
buffer correctly (wire trace 5175/5175); the client decoder dropped oversized
chunks. My earlier firmware-ring hypothesis was wrong.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fixes the reboot-test failure (support latched "unsupported" after the
reconnect window) and reworks the boot-log UX per Ben.
- meshcore_protocol.dart: offbandCapCaplog (0x20) + firmwareSupportsOffbandCaplog
(0x20 bit AND FIRMWARE_VER_CODE >= 17), mirroring the block-cap pattern.
- meshcore_connector.dart: supportsOffbandCaplog getter.
- serial_capture_screen.dart: gate support on the static cap bit (reactive via
the connector), not a one-shot STATUS probe, so it never latches "unsupported"
after a reboot. Derive capturing state from device STATUS so an auto-resumed
capture (post firmware #428) shows STOP not START. Cancel timers on disconnect,
re-query STATUS on reconnect. New red "Start & Reboot" (no timer) boot-log flow.
- 4 cap-gate unit tests; 18 caplog tests total green; analyze clean.
Root cause confirmed with firmware (TopazHill): caplog cap bit (0x20) is
advertised statically across reboots; the client's probe raced the reconnect
window and latched. Firmware #428 (persist flag + boot capture) is the other half.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Extends the 0xC4 caplog protocol with the companion control sub-codes the
firmware exposes (the companion has no CLI; #395 CLI verbs are repeater-
only). Firmware #417/#408.
- meshcore_protocol.dart: request sub-codes ENABLE/DISABLE/ERASE/STATUS +
builders; ACK / STATUS parsers (CaplogAck, CaplogDeviceStatus).
- meshcore_connector.dart: setDeviceCaplogEnabled / eraseDeviceCaplog /
getDeviceCaplogStatus; 0xC4 response routing split so ACK (0x10) and
STATUS (0x11) dispatch to own completers, download stream unchanged.
- 5 unit tests for builders + parsers; analyze clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Slice 1 of the client half of serial-capture (#430): the protocol layer
to download the device's serial-capture buffer over the companion link.
- meshcore_protocol.dart: cmdOffbandCaplog / respCodeOffbandCaplog = 0xC4
(NOT 0xC3, which collides with cmdOffbandFemLna; see firmware #406),
START/CHUNK/END sub-codes + request builder.
- caplog_reassembler.dart: pure START/CHUNK*/END reassembly state machine
with truncation detection, unit-tested in isolation.
- meshcore_connector.dart: downloadCaplog() + 0xC4 frame dispatch + fast
busy-reject on RESP_CODE_ERR while awaiting START.
- 9 unit tests passing; flutter analyze clean.
Integration test gated on the firmware 0xC4 fix merging. Not pushed
(human-test gate).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Address 3 review findings (all confirmed real): thread section context through
the optional-field accessors so nested errors name the broker (brokers[i]."port"),
reject broker ports outside 1..65535, and reject negative integer fields
(status_interval, jwt_refresh) via a shared _optUint. +3 tests (14 total).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds the yaml dependency and parseConfigProfile(): YAML -> ConfigProfile
(#402). Strict by design since profiles are untrusted input (#139) — unknown
keys, wrong types, out-of-range/duplicate broker slots, and unknown
transport/auth values all throw ConfigProfileFormatException with a
user-facing message. Only keys present populate the model. 11 unit tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Wadamesh keys its history watermark (last_delivered_seq) by client_id. We
sent none, so we shared the empty-string slot with every other MeshCore
client on the machine: whichever connected first drained the device history
ring and the next app got NO_MORE_MESSAGES for frames it never received.
cid_len is 6 by necessity, not preference. Stock reads cmd_frame[1..7] as
reserved with the app name at a fixed offset 8; Wadamesh reads the name at
2 + cid_len. Only 6 puts the name at 8 on both, so one frame serves both
firmwares with no firmware change. Covered by app_start_frame_test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reactions stored only a Map<String,int> of emoji to count, so "who
reacted" was unanswerable and MeshCore One's tap-to-see-who had nothing
to show against. This captures the reactor for every reaction, both our
own r: format and the PocketMesh / MeshCore One format.
Data layer only. The tap-a-badge-to-see-who UI is a separate follow-on;
this makes the data available and does not change any screen.
Additive by design. A new reactionSenders field (Map<String,List<String>>,
emoji -> reactor names) sits alongside the existing count map, serialised
under a new JSON key. Both maps are always written. An older build reading
a newer store ignores the unknown key and still gets correct counts from
reactions; it never hits the hard `value as int` cast that would fail the
whole message-list load (the #355 data-loss shape). Records written before
this field load with an empty sender map, so their counts survive with
names simply absent.
- ChannelMessage + Message: new reactionSenders field, constructor,
copyWith
- both stores: serialise the new key; deserialise via a shared
ReactionHelper.reactionSendersFromJson that returns empty for a missing
key and skips malformed entries rather than throwing
- applyReaction: records the reactor and dedups per reactor per emoji.
This is persistent dedup (survives restart), unlike the connector's
in-memory processed-set. A pre-existing count with no sender list is
incremented from its stored value, not recomputed from the partial
list, so old counts are preserved.
- connector: threads the reactor name into all reaction paths. Channel
uses the frame sender; a room resolves the author via
fourByteRoomContactKey; an outgoing reaction is attributed to self.
- pending queue: retry now hands the stored reactor to its callback
Tests: reactor capture, two-reactor count, same-reactor no-double-count,
pre-#383 count preservation, and JSON round-trip + backward-compat +
malformed-entry handling for the new field. Full suite 604 passing.
Epic #376.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two of three Gemini findings held up against the code; the third did not.
Accepted, message-swallow risk: the emoji check looked only at the first
rune against a range list that included arrows, so a multi-line message
beginning with an arrow and ending in eight Crockford characters would
have been consumed and displayed as a reaction whose "emoji" was the
whole sentence. Ben's own capture has U+2192 mid-sentence in ordinary
channel traffic, so this was reachable, not theoretical.
- drop U+2190-U+21FF and U+2934-U+2935; arrows carry the Unicode Emoji
property but read as punctuation in prose
- cap the emoji segment at 8 runes, which is clear of the longest ZWJ
sequence and nowhere near a sentence. This is the guard that holds
regardless of how the range list evolves.
Accepted, dedup: the contact path keyed on target hash plus emoji only.
A room server is many-party over the 1:1 transport, so two members
sending the same emoji collapsed into one and the count stuck at 1, the
same defect already fixed on the channel path. The reacting author
prefix is now part of the key; in a true 1:1 it is empty and the key is
unchanged.
Rejected, room-server sender matching: the review claimed getSenderName
resolves to the room itself and that room text carries a "Sender: "
prefix. Neither is true. _resolveContactSenderName resolves the author
via fourByteRoomContactKey to the real per-sender contact name, and room
message text is stored bare with the author in its own field. One real
sub-case survives: an author who is not in our contacts resolves to null
and will not match, which degrades to queue-then-expire with a warn log.
Also rejected: switching @[ lookup from first to last occurrence. The
reference implementation uses the first, and diverging risks mismatching
payloads it accepts.
Reactions sent from MeshCore One arrived as junk text: an emoji line
followed by an 8-character token such as "dyps6yf0". Those tokens are
Crockford Base32 target-message hashes, the second line of a two-line
reaction payload we did not recognise.
Receive-side only. Offband keeps sending its own r:hhhh:ii format; their
client already parses ours, so nothing about what we transmit changes.
Wire format (confirmed against a live capture, see #378):
channel: {emoji}@[{targetSender}]\n{hash}
direct: {emoji}\n{hash}
hash: sha256(body utf8 + timestamp uint32 LE seconds)[0:5],
Crockford Base32, 8 chars, lowercase
The body is hashed without the channel "SenderName: " prefix, which is
why the sender travels in @[...] instead.
- crockford_base32.dart: encode and normalise, 12-bit accumulator so the
web target's 32-bit bitwise ops cannot truncate a 40-bit value
- pocketmesh_reaction.dart: hash and a parser mirroring the reference
implementation, with a conservative leading-emoji check so a real
message is never swallowed
- reaction_helper.dart: ReactionInfo carries a dialect; applyReaction
picks the matching hash and, for the channel form, requires an exact
sender-name match (a node name can carry emoji and variation
selectors)
- pending_reactions.dart: a reaction arriving before its target is held
and retried rather than silently dropped, bounded at 50 entries with a
15 minute TTL and a warn-level log on expiry (closes the silent-drop
path in #382)
- the channel dedup key now includes the reacting sender, so two people
sending the same emoji no longer collapse into one
- notification tray summarises the foreign format as a reaction instead
of showing the raw token
Known limitation: on channels with Smaz or Cyr2Lat enabled the wire text
differs from the text we store, so a hash computed by another client
will not match. Flagged for a decision rather than worked around.
Epic #376. Fast-follow #383 adds reactor identity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
AppShell decided top-level vs detail by `selectedIndex != null`, but pushed
detail screens (a channel chat, the LOS map) also set selectedIndex to keep the
bottom bar visible. So Back from inside a channel backgrounded/left the app
instead of popping to the channel list.
Decouple the two concerns: add `isTopLevel` (default true), separate from
`selectedIndex`. Detail screens pass `isTopLevel: false` (keep the bar, but Back
pops). The decision is extracted into a pure `AppShell.backAction`
(drawer -> close; detail with a route below -> pop; else -> background), with an
assert that a detail is actually poppable. Unit tests cover the matrix; widget
tests exercise the real system-Back -> PopScope -> pop wiring.
Follow-up #390 tracks the separate AppBar back-arrow path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
When the database can't open (e.g. the native sqlite library fails to load),
every read/write silently failed and the app opened to an empty, normal-looking
screen — the user thinks their history was wiped (SAFELANE §6 violation).
Probe the storage layer in main() before any store reads
(BlobStore.verifyReadWrite). On failure, a StorageHealthService one-way latch
records it, and a persistent, non-dismissable banner is shown above the whole
app: history is not lost, storage is unavailable, messages are NOT being saved,
restart after fixing. Cross-platform (probe goes through drift, covers web too).
Tests: service state/latch + banner show/hide widget test. Gemini-reviewed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#363 pins the DB to one directory and migrates ONE prior store in, but a user
who ran differently-built copies can have data split across several stores.
On startup (native only, after the prefs->drift migration) this discovers
every store on the machine and unions its bulk blobs (messages by id, contacts
by public key) into the current one, so nothing shows as a gap. Sources are
read-only and never deleted.
Not a one-shot: instead of a permanent flag it tracks each store's signature
(mtime+size), so a store that a stray older build later creates or grows is
re-merged rather than stranded. A store over a 200 MB guard, or one that
cannot be read, is skipped and NOT recorded as done, so it retries later.
Identity dedup is key-order independent. sqlite3 (dart:ffi) stays out of the
web build via a conditional import.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The DB directory came from getApplicationSupportDirectory(), which on Windows
is derived from the executable's Company/Product version metadata. Two builds
with different metadata (or a future rebrand) resolved DIFFERENT %APPDATA%
folders and read DIFFERENT databases, so a user's history appeared to vanish
when they ran a different build.
Pin the desktop DB to a constant path (Windows: %APPDATA%\Offband MeshCore;
Linux pinned too, its support dir derives from the exe name). On first run at
the pinned location, migrate an existing DB in via SQLite VACUUM INTO (a
consistent snapshot that is safe even under a concurrent writer), choosing the
DB from the known canonical locations first and falling back to a bounded
scan of the app-data roots (no name blocklist) so a DB under an unknown
folder is still found. On any snapshot failure it aborts cleanly, leaving the
source intact and opening a fresh DB - never a raw file copy of a live WAL DB.
Adds sqlite3 as a direct dependency (pinned to drift 2.34.2's resolved 3.5.0).
Tests cover source selection, the folder-agnostic scan, and the snapshot
happy + abort paths.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The migration's already-present branch called prefs.remove(key) and threw
the prefs copy away when drift already held the key. A build that writes to
SharedPreferences (a non-drift build, or any in-between test build) collects
new messages there, so the next drift run silently gapped that history out.
It cost 574 real channel messages, recovered from backups.
Union the prefs copy into drift by element identity (messageId, else
publicKey, else canonical JSON), keeping drift's live copy on a collision and
appending prefs-only elements. Verify the merged write before removing the
source. New `merged` counter in the report.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adversarial review (standards#145) of the storage/migration surface found four
real defects, all data-integrity. All confirmed against the code and fixed; no
false positives.
BLOCKER - concurrent save race. saveChannelMessages / saveMessages are
read-modify-write with an await gap, so two saves to the same key raced and the
second clobbered the first, silently losing messages (e.g. a message and its
delivery ack arriving together). BlobStore now provides synchronized(key, ...),
a per-key operation chain; every RMW - merge-save, remove, and the load-path
legacy migration - runs through it. Different keys stay concurrent. New test
fires two concurrent saves and asserts the union survives.
BLOCKER - non-atomic DB relocation. The documents->support move copied and
deleted each file (.sqlite/-wal/-shm) in turn, so a failure after the main file
moved stranded the DB across two locations and corrupted it. Now copies all
files, verifies each by size, and only then deletes the sources; on any failure
it rolls back the destination and leaves the original intact.
MAJOR - load-path legacy migration raced save. The #194 index->PSK adoption did
a blind write to the PSK key that could clobber a save that landed first. It now
runs under the key lock and MERGES (union) instead of overwriting, so both the
adopted history and any fresh message survive.
MINOR - channel merge key lacked a sender. Two senders posting identical text at
the same timestamp without a messageId would collide and lose one. The key now
includes the sender, matching MessageStore.
flutter analyze clean, dart format clean, 509 tests pass. Full Gemini log in
docs/llm-consultations/.
The persisted store was being overwritten with the windowed in-memory list, so
any channel or DM with more than _messageWindowSize (200) messages lost
everything older than the newest 200 on the first save after load. Observed
live: Public went 232 -> 201 in one session on a build that already had the
#333 fix, so this was not the race - it was windowing truncating the store.
This is the slow-erosion cause behind the whole 566 -> 231 -> 206 -> 201
history.
saveChannelMessages / saveMessages now MERGE into the persisted set instead of
overwriting: upsert by message identity so older persisted messages are kept,
new ones added, and the in-memory copy wins for edits/reactions/status. If the
existing history fails to decode, the save aborts loudly rather than merging
into an empty base and truncating (SAFELANE 6).
Deletion is now an explicit path - removeChannelMessage / removeMessage - since
the app deletes individual messages. Routing delete through the merging save
would resurrect them; the connector's deleteChannelMessage / deleteMessage now
call the explicit remove.
Windowing stays for display and memory; it no longer dictates what is stored.
Tests: 250-message history survives a 200-window save; new message appends;
edit is captured not duplicated; delete does not resurrect. 508 pass, analyze
and format clean.
Cost: a save now re-reads and re-encodes the channel/contact history. Cheap
with drift's per-key writes (#335); a future append-only schema removes even
that.
The guard read pubspec.lock with a newline-sensitive regex, so it passed on
LF (CI) and failed on CRLF (Windows) for the same, correct assets. A guard
that is itself platform-fragile is worse than none. Normalise CRLF->LF before
matching.
Surfaced by the 290+306+335 integration merge, where the lockfile came through
with CRLF endings.
Moves message history, contacts and discovered contacts out of the settings
store. Settings stay in SharedPreferences, which is what it is for.
Ordering is the whole safety argument: WRITE, VERIFY BY READING BACK, and only
then remove the source. #333 was a storage path that chose a key silently and
made 566 real messages read as empty; deleting before verifying would make
that class of mistake permanent instead of cosmetic. On any failure the source
is left intact and the error is logged - never a silent drop (SAFELANE 6).
Idempotent by construction: a key already present in drift is not overwritten,
so re-running is a no-op. If an older build re-writes a migrated key into
prefs, the migrated copy wins and the stale prefs copy is discarded rather
than promoted.
Rehearsed against a COPY of a real 7 MB store, as the plan required before
touching live data:
REHEARSAL: 69 migrated, 0 failed, 5.21 MB, 69 bulk keys expected
Every key checked for exact length, confirmed removed from prefs, and every
settings key confirmed untouched. The live store was never opened.
Two things the tests caught that review would not have:
1. getString THROWS on a non-string value rather than returning null, so a
non-string under a bulk prefix was counted as a migration FAILURE. It now
type-checks with prefs.get() and skips. Alarming falsely is its own bug.
2. The `contacts` prefix was checked against the real store rather than
assumed: it matches only the 6 bulk contact blobs, and correctly does NOT
match contact_unread_count*.
Not yet wired into app startup - that is the switchover, and it is deliberately
a separate commit so this can be reviewed on its own.
flutter analyze clean, dart format clean, 504 tests pass.