connect() on non-Linux platforms now treats Android GATT status 133
(ANDROID_SPECIFIC_ERROR) as the transient failure it usually is: clean
close, 800 ms delay, one bounded retry, then surface. The scanner's
connect-failure snackbar no longer shows raw exception text: it maps
133, timeouts and everything else to localized plain-language messages,
persists until dismissed, and sends the raw detail to the app log
(#522 presentation acceptance).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
disconnect() now caps the platform-confirm wait at 10 s (the vendored
Android plugin's backstop resolves within ~5 s), retries once after
500 ms, and when both attempts fail sets bleReleaseUnconfirmed instead
of pretending the release worked. The scanner screen shows a
persistent, dismissible banner naming the recovery (toggle Bluetooth /
force-stop), per the error-visibility standard. Cleared on the next
connect or disconnect attempt.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#528 surfaced unmatched and late replies in the UI because they were
previously destroyed in silence. Once #529 widened the window, late
arrivals stopped being the exception and the on-screen output became
noise during ordinary use. Owner's call: log it, keep it off the screen.
Removed: the CLI screen's late-response history entry and its renderer
branch, the settings screen's snackbar, and the four l10n strings that
only those two used.
Also removed the settings screen's silent late-apply. A field changing
value seconds after the fact with no explanation is the same problem as
the snackbar, only quieter, and applying while saying nothing would be
worse than either showing or not showing. One line to re-enable if that
turns out to be wanted.
Kept: the service still detects every unmatched reply and writes it to
the app log via appLogger, now including the payload as well as the
repeater, the originating command when known, and how late it arrived.
That is the diagnostics channel and is what keeps this from being a
silent drop.
onUnmatchedResponse and UnmatchedRepeaterResponse stay on the service.
They are the seam the no-silent-drop guarantee is tested through; the
#528 and #532 tests exercise them directly and pass unchanged, which is
the evidence the service contract did not move.
807 tests pass, analyze clean, format clean.
Adversarial review raised three findings. Two survived verification.
Rejected: a claimed RangeError when abbreviating a short public key. The
code already guards with a length check before either substring, so the
crash it describes cannot occur.
Rejected as described, fixed as found: contact position (0, 0) was said to
be dropped on import. It is not. The frame builder writes the position block
whenever lastModified is present, which it always is here, so a suppressed
position still writes 0/0 and the bytes are identical either way. The
hasPosition conditional was therefore doing nothing except misleading a
reader, which is exactly what happened. Removed, and the behavior is pinned
with a test.
Accepted: the service added sections to 'applied' after sendFrame returned,
which only means the frame left our side. A lost or refused frame was
indistinguishable from success, so the user could be told an import worked
when nothing changed on the device.
A correct fix needs an awaited per-command acknowledgement in the connector,
keyed on command code so concurrent commands cannot steal each other's OK.
That is real connector work and is filed as #584. What is fixed here is the
overclaim: 'applied' and the two counts are documented as sent rather than
confirmed, and the result string now reads 'Sent N contacts and M channels
to the device'. The code no longer states something it cannot know.
Mirrors stock's Import Config screen, including the detail that nothing is
checked by default. Export opts out, import opts in; that asymmetry is
stock's own and is the safe default in each direction.
Only the sections the chosen file actually carries are offered, since a stock
export omits whatever the exporting user deselected. Channel and contact
counts on this screen come from the file, not the device.
Per-section wording is stock's, because the behavior is stock's: contacts are
updated, existing channels are left unchanged. Replacing the node identity is
irreversible, so it gets its own explicit confirmation instead of riding
along with the rest of the selection.
The result dialog names every channel skipped and every section not applied,
each with its reason. An import that quietly did less than the user asked for
is the silent failure SAFELANE section 6 forbids, and stock's own screen does
not tell you.
A parse failure reports the offending JSON path from the model's exception,
so the user learns which part of the file is wrong rather than 'invalid file'.
Both screens are reachable from Settings, above the GPX exports, which are
map data only and a different thing.
Mirrors stock's Export Config screen: section list, Select All and Deselect
All, everything checked by default. The asymmetry with import, which will
start with nothing checked, is stock's own and is deliberate: opt out on the
way out, opt in on the way in.
A section the user ticked that could not be gathered is shown before the file
is written, with its specific reason, and the user chooses whether to export
anyway. Firmware-built-without-it reads differently from device-did-not-answer
because only one of them is worth retrying.
The screen also states plainly that this is the stock format and cannot carry
Offband-only data, so importing the file back will not restore it.
File naming follows stock, <name>_meshcore_config_<stamp>.json, and is
sanitized for the filesystem without touching the name written inside the
file, which has to round-trip exactly. Real device names in the reference
corpus contain emoji and a trailing space, both covered by tests.
Writing reuses LogExport for the per-platform mechanism (share sheet on
mobile, Save As on desktop, download on web) rather than adding a second
export path that could drift. The temp file can contain the node private key,
so it is written under the app's own temp directory.
The public key is abbreviated on screen; the private key is never rendered.
6 tests on the file name.
Applies a stock config file to the device with stock's own merge semantics,
quoted from its Import Config screen:
- contacts are upserted on public key, and nothing is ever deleted
- channels are additive only, so a rotated PSK for a channel the user
already has does not apply
- importing an identity overwrites the node outright
We match that behavior but not its silence. Every skipped channel comes back
named, with a reason, so the UI can show it; a channel that quietly did not
import is the silent no-op SAFELANE section 6 forbids.
The channel merge rule is extracted as planChannelImport, a pure function, so
the surprising part is testable without a radio. It treats a channel as
already present when either its PSK or its name matches, which is the
non-destructive reading of a rule stock states without defining. Pinning down
what stock actually does is a T1 question, recorded in the doc comment.
New channels take the lowest free slot, since the file carries no index.
other_settings is reported as not fully writable rather than claimed as
applied: buildSetOtherParamsFrame deliberately pins the auto-add-contacts
byte, so manual_add_contacts cannot be written without changing app-wide
behavior. Advert location policy is applied.
10 tests over the planner and the result type.
Gathers live device state into a StockConfig, section by section, matching
stock's Export Config screen where each section is independently selectable
and the file simply omits what was not ticked.
A requested section that cannot be gathered is reported in the result rather
than dropped, with the reason kept specific: unsupported (firmware built
without identity export) is distinct from noReply and rejected, so the
export screen can say the radio cannot do it instead of offering a pointless
retry.
Two unit traps handled explicitly:
- currentFreqHz is misnamed; the value is kHz, which is also what stock's
frequency field carries, so it passes through unconverted. currentBwHz
really is Hz. The mismatch exists in both our state and the file.
- other_settings.manual_add_contacts must be the raw device byte. The
connector's _manualAddContacts is an inverted derived view of bit 0
(firmware treats a clear bit as auto-add enabled), so exporting it would
have written the wrong value. The raw byte is now retained and exposed,
with the inversion documented at the parse site. No behavior change.
Path mapping keeps our hash count and per-hop width (#309) intact: stock
encodes one hash per comma-separated element, so the width survives as
element length. A flood route, an over-long declared hop count, or a width
stock cannot express all yield no path rather than a guessed one.
13 tests over the pure mapping functions.
Prerequisite for the export/import epic (#568): two device values the
connector could not read or write, both already exposed by firmware.
Identity (CMD_EXPORT_PRIVATE_KEY 23 / CMD_IMPORT_PRIVATE_KEY 24):
- adds the commands, RESP_CODE_PRIVATE_KEY 14 and the previously unhandled
RESP_CODE_DISABLED 15
- both firmware commands sit behind build flags, so a radio can legitimately
answer DISABLED. IdentityTransfer keeps unsupported distinct from rejected
so the UI can say the radio cannot do it rather than offering a retry that
can never succeed
- a short PRIVATE_KEY frame reports rejected instead of handing back a
truncated key that could be written to an export file
- import is destructive and documented as such; the key is never logged
Auto-add hop limit:
- GET already returned max_hops as a third byte and we discarded it; it is
now parsed and exposed
- SET learns an optional maxHops. Omitting it keeps the frame two bytes, so
firmware leaves the device value alone and existing callers are unchanged
- clamped to 64 on our side to match the firmware clamp
Adds handleFrameForTest so parse paths can be exercised without a radio,
following the file's existing visibleForTesting convention. 12 new tests.
Models the MeshCore stock companion config export JSON so files interchange
in both directions (epic #568). Shape reverse-engineered from ten real
exports across five radios; evidence recorded on #569.
Format rules the parser enforces, none of them obvious:
- no version field exists, so validation is by shape and every top-level
section is independently optional
- public_key and private_key travel as one unit
- radio_settings mixes units: frequency kHz, bandwidth Hz, coding_rate a
bare denominator
- coordinates are JSON strings, never numbers
- channels carry no index; array position is the index
- out_path_list is one comma-separated hop hash per element, so hop count
and hash width are both recovered rather than inferred (#309)
- contact timestamps are untrusted and never sanitized
- null and empty-string out_path_list stay distinct on re-encode; both mean
no route, but real exports contain both and the difference is unexplained
Unknown top-level keys are ignored so a future stock release cannot break
the reader. Errors carry the offending JSON path for the import UI.
Test fixtures are synthetic with obviously fake keys. Validated separately
against the owner's ten real exports, which round-trip key-for-key; that
harness reads a path outside the repo and was deliberately not committed.
The radio is authoritative for the expected-ACK hash it reports, but the
client looked that hash up in a map keyed by a hash it recomputed locally.
When the two disagreed the lookup missed silently (debugPrint only), the
8000ms watchdog from #395 marked the message failed, and the genuine ACK
later matched nothing because every downstream map is populated only on the
match path. Delivery worked the whole time.
Captured on a Wadamesh HV4 TFT in #449: five DMs, all shown as errors, two
with a confirmed ACK whose hash equalled the radio's RESP_CODE_SENT value
exactly. Wadamesh builds the payload differently, so the client cannot
predict its digest, and Offband is expected to work against stock MeshCore,
Wadamesh and other forks.
RESP_CODE_SENT is the reply to our own CMD_SEND_TXT_MSG, so adopt the radio's
value when exactly one send awaits confirmation. Zero or several candidates
keep the previous behaviour rather than guessing. Channel sends draw the same
frame, so adoption is refused while one is outstanding.
The lookup miss is now a warn on debugLogService instead of a bare
debugPrint (SAFELANE section 6).
A send-order correlation queue was developed alongside this and has been
split out: its ordering premise needs a transport send mutex that does not
exist yet, and three review rounds each found a fresh defect in it. This
commit deliberately carries only the adoption fallback, which is what the
owner's hardware test validated.
Refs #449
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A truncated caplog download used to completeError and discard every received
byte, so the user got nothing from a buffer still intact on the device (a tester
lost 8461/8608 bytes, 98.3%, deterministically per firmware#711 so retry can't
help). Now truncation is a soft outcome:
- downloadCaplog() returns a CaplogDownload (bytes + received/expected/chunks +
truncated) instead of throwing; busy/timeout still throw.
- serial_capture_screen still writes + LogExport.shareFile the partial bytes,
marks the file "# PARTIAL ..." with the counts, and shows a non-fatal warning
instead of a red error.
- Button label is now platform-aware (Download & save on desktop, & share on
mobile) via LogExport.actionVerb, matching the App/BLE log screens.
Replaces CaplogTruncatedException with CaplogDownload. Reassembler unchanged
(already keeps the bytes). Adds CaplogDownload contract tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two accepted findings from the standards#145 adversarial review.
Finding 4, the sharp one: the app emitted the compact <key:type:name>
card but could not parse it. A user copying a card out of a channel and
pasting it into Add by public key would have been rejected by the app
that produced it. Adds Contact.fromChannelShare and tries it after
fromShareUri in the dialog, so both real formats are accepted.
The parser splits on the FIRST TWO colons and takes the remainder as the
name, because names carry colons, spaces, emoji and CJK. It also finds a
card embedded in a longer message, which is how one actually arrives,
since people caption them.
Rendering a received card as a tappable affordance remains #610; this
only covers text pasted or scanned into the add flow. #610 is
correspondingly smaller now.
Finding 2, performance: resolveContactVerification ran an O(N) message
scan from a widget build inside a ListView. Now scans newest-first,
since a delivered message is overwhelmingly likely to be recent, and
advert-verified contacts still return on a single comparison without
touching the message list.
A lastMessageAt == epoch shortcut was written, then removed after
checking _setContactLastMessageAt: it maintains that field only for
advTypeChat, so a key-added repeater that had been messaged would have
shown the wrong badge. A cheap wrong answer is worse than a slightly
slower right one, and the rejected approach is documented in place.
Epic #619.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Owner feedback from testing: a scan button living only inside the public
key field is awkward placement, because scanning is how most people will
actually add someone and it should not be reachable only from inside a
field they have to open first.
Now in both places, routed to the right thing:
- Contacts overflow menu, Scan contact QR, opens the scanner and hands
the result to the add dialog already populated. Nothing to paste.
- The key field keeps its scan button, which fills in place, for when
the dialog is already open.
Same scanner and same parser either way; only the entry point differs.
The dialog gains an optional initialKeyText, and seeds it through the
same handler as a paste, so a scanned link populates name and type
rather than sitting there as raw text.
Platform availability moves to a shared contactQrScanAvailable getter on
the scanner screen, since it now has two consumers. Both the menu entry
and the field button hide on Windows and Linux, where mobile_scanner has
no support and the render half is the desktop path.
Epic #619.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The composer GIF button becomes a +, opening a short picker with GIF and
My contact card. Owner decision after seeing the build: one affordance
beside the text entry rather than a button per attachable thing, because
the sendable set stays small when a channel message shares a 160-byte
payload with the Sender: prefix.
Sends the COMPACT format, not the meshcore:// URI:
<{64-hex key}:{type}:{name}>
That shape was observed on the live mesh, twice, from different senders
in #test and #hamradio four weeks apart. It costs about 75 bytes against
117 for the equivalent URI, so on a 160-byte budget the difference is
airtime rather than tidiness. The URI form stays correct for a QR, a DM,
or an out-of-band paste; this is the channel idiom.
Angle brackets are stripped from an emitted name because they are the
delimiters, and a name carrying one would truncate the payload for every
parser reading it. A colon is left alone: the name is the final field,
so a correct parser splits on the first two colons and takes the rest.
Inserts into the composer rather than sending, matching the GIF picker,
so the card can be captioned and reviewed first. It appends, so a
caption already typed is not destroyed.
Parsing a received share is #610 and is not in this change, so an
inbound card still renders as raw text for now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds a calm three-state indicator beside each contact name, per the
owner steer: a green check for advert-verified, a neutral outline check
for key-confirmed, a muted key for key-only. No amber and no hazard
glyph, because nothing is wrong with a key-added contact. The scale
reads as how much we know, never as how risky.
The states are not cosmetic:
- advertVerified: a signed advert arrived, so the node itself asserted
its name, type and position, and the raw packet is stored, which is
also what makes the contact re-shareable.
- keyConfirmed: a message with this contact went through. Direct
messages are encrypted with an ECDH secret derived from the contact
key, and the ACK is computed over the decrypted plaintext
(BaseChatMesh.cpp:442,451), so this is proof the holder of the
matching private key is live. It does NOT prove the person is who the
name claims.
- keyOnly: someone supplied a key and nothing has confirmed it on air.
Costs almost nothing to compute. Contact.isAdvertVerified falls out of
the epoch last_advert_timestamp that #627 already writes, and it clears
itself when a real advert rewrites the field. The advert check runs
first so only an unconfirmed contact pays for a message scan, which
keeps a long contact list cheap.
Also fixes a defect the epoch sentinel introduced: _formatLastSeen ran
the epoch through the relative formatter and claimed a key-added contact
was last seen tens of thousands of days ago. It now reads Not heard yet.
Epic #619.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both halves of the visual exchange, sharing one format and one parser.
Render: showMyContactQrDialog puts this device's own identity on screen
as a QR plus the same link as selectable, copyable text. It refuses to
render when not connected, since an empty key would encode a QR nobody
can add.
Scan: ContactQrScannerScreen validates with Contact.isValidShareUri, the
same check the paste path uses, and pops the raw string. The add dialog
routes a scan through the same handler as a paste, so a QR gets no
separate code path.
The scan affordance is gated to platforms mobile_scanner supports
(Android, iOS, macOS, web). On Windows and Linux the QR half is the
render side, which is the better desktop flow anyway: put your code on
the big screen and let the other person scan it with a phone.
Fixes a real defect found by the new test, not a test artifact:
QrCodeDisplay built its QrImageView through a LayoutBuilder, which
cannot answer intrinsic dimension queries, so any intrinsic-measuring
parent threw. AlertDialog measures its content's max intrinsic height,
so the dialog crashed on open. The QR is now bounded by a tight
SizedBox, which answers the intrinsic itself. This was the widget's
first real call site, so the bug had never been exercised.
Entry point is provisional alongside Add by key; #632 reorganises.
Epic #619.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
First user-reachable piece of the identity exchange. A dialog takes a
public key, a name and a contact type, and hands the stub to
addContactByKey.
The key field also accepts a whole meshcore://contact/add link and
absorbs every field from it, because that is what someone actually
pastes, and a QR is only that link rendered visually. Whitespace is
stripped so a key copied across a line break still works.
The stub is built by round-tripping through buildShareUri and
fromShareUri rather than constructing a Contact directly, so manual
entry and a scanned QR cannot drift apart. One code path, one set of
invariants.
An informational note states plainly that the contact is not confirmed
on air and that the name is whatever the user typed until the node
adverts. Deliberately informational rather than a warning: nothing is
wrong with a key-added contact (#630).
Entry point is provisional, sitting in the contacts overflow menu so the
feature is reachable. The proper add-contact surface, split away from the
advert affordance, is #632 under epic #623.
Five widget tests pin what reaches the connector: key, typed name,
chosen type, the flood sentinel, and the epoch lastSeen that keeps the
firmware replay guard from muting the contact. Also covered: pasted-link
prefill, wrapped-key whitespace, an invalid key sending nothing, and the
missing-name fallback.
Epic #619.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds MeshCoreConnector.addContactByKey, the device-side half of the
identity exchange. It uses the stock CMD_ADD_UPDATE_CONTACT (command 9),
whose handler creates the contact outright when the public key is
unknown. Verified byte-identical to upstream meshcore-dev/MeshCore at
0679dbef, so this needs no capability gate and works on stock firmware,
not only Offband builds.
buildUpdateContactPathFrame gains an optional lastAdvert. It defaults to
now, leaving all three existing path-update callers byte-identical, and
addContactByKey passes the epoch.
That epoch is the whole point. The firmware compares this field against
every incoming advert with 'timestamp <= last_advert_timestamp' and
silently discards the non-greater ones as replay attacks
(BaseChatMesh.cpp:142-145). Advert timestamps come from the sender's
clock, and two nodes on this mesh currently advertise with 2024 clocks,
so stamping now would leave a key-added contact permanently deaf to its
own adverts. A negative input clamps to zero rather than wrapping to a
huge value, which would be the worst case for that guard.
The frame is sent at full length with the 0xFF flood sentinel. The
firmware length guard is only 'len >= 36' but updateContactFromFrame
reads through offset 136 regardless (MyMesh.cpp:295-318), so the frame
must never be trimmed. A test pins that too.
Local state keeps lastSeen at the epoch so the contact reads as
unverified until a real advert upgrades it (#630).
Epic #619.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds Contact.toShareUri and the static Contact.buildShareUri, the inverse
of #625's parser. Emitting the reference-app format is what makes an
Offband card or QR importable by a stock, non-Offband user.
buildShareUri takes raw parts rather than a Contact so the local device
can share its OWN identity, which is a public key plus a node name and
never a Contact instance.
Spaces are percent-encoded rather than emitted as '+'. Both decode to a
space, and this matches Channel.toShareUri, which the codebase already
documents as round-tripping with the reference app's QR (#161).
Tests cover the emitted parameter shape, the companion default, and
round-tripping through fromShareUri for names with spaces, emoji and
reserved characters. A raw & or = in a name would otherwise truncate or
forge query parameters, so that case is pinned explicitly.
Epic #619.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds Contact.fromShareUri / isValidShareUri for the reference-app format
documented in firmware docs/qr_codes.md:
meshcore://contact/add?name=<url-encoded>&public_key=<64 hex>&type=<1-4>
The payload is a bare public key, not a signed advert, so the parsed
Contact is an identity stub: no path, no position, no rawPacket. A QR is
this same URI rendered visually, so scanning will share this parser.
lastSeen is deliberately the epoch rather than DateTime.now(). It maps to
the firmware's last_advert_timestamp, which the advert handler compares
with 'timestamp <= last_advert_timestamp' and treats as a replay attack
(BaseChatMesh.cpp:142-145). Stamping now would leave the contact
permanently deaf to its own adverts, since advert timestamps come from
the sender's clock and two nodes on this mesh currently advertise with
2024 clocks. A regression test pins this.
The fork's older meshcore://<raw advert hex> form returns null here, so
callers keep routing it to the existing advert import path.
Epic #619. Not yet wired to the UI: the import path needs the send half
(#627) before there is anything to add.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The unresolved state claimed "Haven't heard this node's advert yet". The
condition behind it is only that the claimed name matched nothing, which is
not the same thing: a node renamed since its last advert is known to us, just
not under the name it is posting with.
Now reads "No node known by that name" with a hint that adding needs an
advert from the node, so it states what we know and what unblocks it without
asserting something we cannot check. Tracked as issue 579.
Gemini pre-PR review finding. importDiscoveredContact no-ops when the radio
is gone, and sendFrame throws on a mid-write disconnect or an unwritable BLE
characteristic. Both reached the user as a tap that appeared to work.
Checks the connection first, catches the write failure, logs it, and shows a
persistent error snackbar in either case (SAFELANE section 6: no silent
failures).
Channel frames carry no key, so the sender name is resolved against known
and discovered contacts. Adds resolveContactsByName (returning Contacts,
which importDiscoveredContact needs) and rebuilds resolveContactKeysByName
on top of it so one matching rule serves both.
The avatar is the shortcut; the sender name stays inert because it sits too
close to the message body to hit reliably on a phone. Long-press keeps every
action it had and gains the same contact rows. A name several nodes claim
lists all of them, and an unheard name says so instead of failing silently.
Three findings reviewed, two applied, one rejected on evidence.
APPLIED. Suppression logging is de-duplicated. Device-info arrives on
every connect, so a flapping radio would have emitted the same
suppression line on each reconnect. It now logs once per distinct
reason, and the memo clears when a gate opens or when per-radio state is
cleared, so the next genuine suppression is still reported. Real, and
exactly the flooding the project forbids.
APPLIED, though not reachable today. The scope reply logged
reply.scope!, which is safe only because parseNotifyScopeReply returns
an ERROR reply when the scope code is unknown, so a non-error reply
always carries a scope. That invariant lives in another file, so it is
now checked at the boundary instead of asserted with a bang across it.
The review rated this High on the theory that an unknown sub-code could
reach the handler; it cannot, because the parser returns null for any
sub-code it does not own and the dispatcher only forwards what it
returned. Defect not reachable, guard still worth having.
REJECTED. The review claimed the matrix reply log is unbounded because
the wire count byte allows up to 255 rows, estimating 7.5 KB per line.
The log is built from ButtonMatrix.assignments, a Map keyed by
ButtonSequence, which has exactly four values, and the parser skips
sequence codes it does not recognise rather than adding them. The map
therefore holds at most four entries and the line is bounded at roughly
160 characters no matter what the device sends. The count byte cannot
inflate it.
Full suite 740 pass, analyze clean, format clean.
Agent: CalmBay (session d14220d9)
The device-UI surface logged only refusals, so a successful read or
write produced nothing at all.
That gap cost a full diagnosis on 2026-08-02. When the notification
scope appeared not to refresh, neither the app debug log nor a firmware
serial capture could distinguish "the client never asked" from "the
client asked and got an unchanged value back". Answering it needed the
firmware counterpart and a serial capture, for a question the client
should have been able to answer alone.
- Every request logs what it sent and the frame shape.
- Every reply logs what came back: the scope value, or the matrix mask
and its decoded rows.
- A SET reply is distinguishable from a GET reply, so a confirmed write
is not mistaken for a read.
- Refusals keep their existing warn-level logging, unrecognised reason
codes still surface as raw hex.
The important one is the third case, which did not exist before: a
request that is NEVER SENT now says so and says which gate closed, with
the caps2 value that closed it. Silence was the ambiguity; suppression
is now explicit.
No behaviour change. Logging only, at info for normal traffic and warn
for refusals, tagged DeviceUI to match the existing handler logs.
No flooding risk: these fire on device-info, on pane open, and on user
action. Nothing polls, and no per-frame logging was added (SAFELANE §11
rule 10).
Full suite 740 pass, analyze clean, format clean.
Agent: CalmBay (session d14220d9)
1. An unprefixed reply no longer resolves an arbitrary pending command.
#532 constrained prefixed replies only; the unprefixed fallback still
used firstWhere, which with two or more in flight is map-iteration
order, so a caller could receive another command's output. The
fallback cannot just be deleted: firmware only echoes the prefix when
the command exceeds four characters including it
(simple_repeater/MyMesh.cpp, strlen(command) > 4 && command[2] == '|'),
so a very short command legitimately answers without one. It now
resolves only when exactly one command is pending, and surfaces the
ambiguity otherwise.
2. A late reply no longer overwrites unsaved edits. The settings screen
applied a late `get` payload unconditionally, so a reply arriving after
its timeout could revert a field the user had typed into while waiting.
It now applies only when nothing is dirty, and says which happened.
3. _expiredCommands is pruned on timeout as well as on receive. It was
pruned only in handleResponse, so a run of commands that all time out
with no traffic coming back accumulated records until disposal.
The reviewer described 3 as unbounded growth. It is bounded at 256 by the
prefix token space, and stale records could never be mis-attributed
because handleResponse prunes before matching, so the severity was
overstated. Fixed anyway; it is nearly free.
A fourth finding, that the new l10n strings are untranslated in the other
17 locales, is rejected. That is this project's established pipeline:
strings are authored in app_en.arb, gen-l10n emits English fallbacks, and
untranslated.json tracks the gap, which it now does for these keys. Every
string in the app arrived this way, so it is not a defect this branch
introduces.
The ambiguity test was verified to fail against the pre-fix code.
The flood arm read as though the model were consulted:
if (pathLength < 0) {
// Flood: trust ML, only enforce firmware formula as floor
if (mlTimeout < physicsMin) return physicsMin;
}
return mlTimeout.clamp(physicsMin, physicsMax);
It never did that. _physicsMinTimeout and _physicsMaxTimeout return the
identical expression for pathLength < 0, mirroring the firmware's
calcFloodTimeoutMillisFor, so a clamp between them cannot preserve a
prediction. Whichever way control went, flood returned 500 + 16 * airtime.
The comment described behaviour the code did not have, which is exactly
the kind of load-bearing false comment that re-causes a bug later.
Replaced with an explicit early return and the constraint written down.
Behaviour is unchanged.
Equivalence is tested, not asserted. A new group attaches a real
TimeoutPredictionService, trains it on 40-56 s delivery times so any
prediction is far from 1300, and checks flood is unmoved while a direct
path is still clamped to its ceiling, which proves the clamp is live
rather than the predictor being absent. That group was run against the
pre-change branch and passes there too, so this is a simplification and
not a behaviour change.
Deliberately NOT decided here: whether flood should ever trust a
prediction. Every round trip measured to date was 0-hop direct, so there
is no flood data to decide it from. The verification needed to answer it
is written up on the issue.
Part of epic #473, stacked on #528, #529, #531 and #532.
handleResponse looked up the reply's prefix, and when that matched nothing
pending it fell through to "first pending command for this repeater". So a
straggler from an expired command could complete an unrelated in-flight
one, reporting one command's output as another command's result.
The service already carries a correlation token in every command and every
reply echoes it. The fallback discarded that. Now a reply that carries a
prefix is only ever matched to the prefix's owner; if there is no owner it
goes to the unmatched handling added in #528, where an expired command is
still named. Only a reply with no prefix at all, which has no correlation
token to honour, may fall back to matching by repeater.
Not reachable from today's callers: all four pass retries: 1 and the
settings refresh awaits each command, so two are never in flight to one
repeater. It becomes live the moment anything issues concurrent commands,
which any retry work would.
The two negative tests were verified to fail against the old logic and
pass against the new, so they are a real guard rather than a restatement.
The two preservation tests pass either way by design.
_registerPending is extracted so the test seam and the real send path
cannot drift apart.
Part of epic #473, stacked on #528, #529 and #531.
The timer ran timeoutMs while the message printed
(timeoutMs / 1000).ceil(), so every window in (4000, 5000] announced
"timeout after 5 seconds". The owner's 0-hop window was 4074 ms and fired
at 4.07 s while claiming 5, which is what made the behaviour look
arbitrary rather than deterministic: the number shown was never the
number used.
The service now throws a typed RepeaterCommandTimeout carrying the window
that was actually armed, formatted to one decimal. Every existing caller
already stringifies the error, so all of them inherit an honest figure
without being touched; the CLI screen additionally renders it through a
new localized string rather than the generic error wrapper.
Tests pin that 4074 reports 4.1 rather than 5, that the new 28748 ms
budget reports 28.7 rather than 29, and that two windows inside the same
second no longer collapse to the same text, which was the defect's
signature.
Part of epic #473, stacked on #528 and #529.
The comment justified not modelling command execution time by claiming
`wifi on 30` does real work bringing up an interface. It does not. The
firmware handler sets a persistence deadline and sprintf's its reply
immediately (CommonCLI.cpp, "wifi on"), so execution is near-instant.
The owner caught it: he recalled reissuing the command several times
against on-screen errors, not one command taking 20 seconds.
The measurement supports him. On the 20.33 s case the reply carried
claimed=01:56:32 against a command sent at 01:56:26.938, and the RF frame
did not reach our radio until 01:56:47.271979. So roughly 5 s to reach the
repeater and be answered, then roughly 15 s in its transmit queue. The
tail is scheduling on both radios.
No behaviour change. The budget is unchanged and still has to tolerate a
20.33 s round trip; only the stated reason was wrong, and a wrong reason
in a load-bearing comment re-causes the bug later.
The CLI timeout used calculateTimeout, which mirrors the firmware's
calcDirectTimeoutMillisFor and estimates ONE-WAY delivery. A CLI command
is a request, an execution and a reply, so the budget was structurally
short.
Measured against rpt-01 on 910.525 MHz / SF7 / BW 62.5k / CR 4:5, where
the old window was 4074 ms: 27 commands sent, 12 replies matched, and 4
of those 12 arrived after the client had already given up, at 6.17 s,
8.28 s, 15.52 s and 20.33 s against a median of 2.75 s.
calculateCliTimeout sums four terms, each with a source rather than a
chosen value:
outbound leg calculateTimeout, which is what it actually models
cliReplyDelayMs 600, firmware CLI_REPLY_DELAY_MILLIS, unconditional
reply leg the reply is a second packet the ACK formula omits
retrieval budget replies are pull-based; the radio raises MSG_WAITING
and the app must ask, granting itself 5000 ms per
attempt across 3 retries
The retrieval term is derived from the sync constants rather than
restated, so the command timeout cannot drift below the layer it depends
on. That is the #530 invariant holding by construction, not by two
numbers being maintained in agreement.
Execution time is deliberately not modelled. The same verb, wifi on 30,
returned in both 2.31 s and 20.33 s, so it is not a per-command constant
that could be tabulated. The retrieval term carries that tail.
The reply leg uses physics only. The predictor is trained on
direct-message ACK latency, so asking it about a CLI reply leg would be
extrapolation; that is #534 and #535, not this change.
Tests cover the construction, the never-below-retrieval invariant, growth
with path length, coverage of the 20.33 s worst case actually observed,
and negatively that the direct-message ACK path is untouched.
Part of epic #473. Does not change the reported duration string (#531)
or the stale-prefix fallback (#532), and does nothing for commands that
draw no reply at all (#541).
A repeater reply that arrived after its command's window closed hit
`if (commandId.isEmpty) return;` in RepeaterCommandService.handleResponse
and was dropped with no log and no UI. Since the command timeout can be
shorter than the app's own message-retrieval budget, that made "the
command ran but the response was never reported" the normal outcome
rather than an edge case, and it affected every repeater screen, not
just the CLI one.
RepeaterCommandService now remembers a timed-out command's prefix for two
minutes, so a reply arriving afterwards can be attributed to the request
it answers. Replies that reach no waiting command are handed to a new
onUnmatchedResponse sink and logged through appLogger. No path through
handleResponse returns without either completing a command, surfacing the
payload, or logging why it could not.
The CLI screen renders these as a distinct history entry naming the
original command and how late it was. The settings screen applies the
value if it is a `get` reply and tells the user it arrived late.
repeater_status_screen already parsed responses independently of the
service, so it had no silent-loss path to fix.
Part of epic #473. Does not change the timeout window itself (#529),
the reported duration (#531), or the stale-prefix fallback (#532).
Adds three tappable rows to the About page opening in the external browser via
url_launcher: offband.org, the Google Play listing (app.offband.meshcore), and
offband.org/donate (Ko-fi + GitHub Sponsors options, live-verified). Failure
surfaces a snackbar. Completes epic #525.
Agent: QuietSnow (session 31eaba02)
About moves out of the Debug pane into a top-level 'About' category (its own
pane, not the stock dialog). Keeps the marketing version display, adds a
description + refreshed legalese (Offband, zjs81/MeshCore MIT attribution kept),
and a View licenses row (showLicensePage). Build identity stays in Device Info
(#553). #527 adds the outbound links here next. Part of epic #525.
Agent: QuietSnow (session 31eaba02)
#456 removed enabled from the profile template but the apply path still
re-enabled via wasLive, carrying over the prior enabled state on import,
against the explicit directive that import brings in NEITHER enabled nor
disabled. Firmware already force-disables a slot on any field write (#53);
apply now passes enable=false and never re-enables. The slot is left disabled
and the operator re-enables intentionally.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
extractScalarValue required ' = ' (space-equals-space) and trimRight()'d first,
so a blank field's reply 'key =' (no space after =) failed the match and the
whole 'mqtt.broker.N.field =' line leaked in as the value. Split on the first
'=' and trim instead: blank -> empty, values with '=' preserved. Regression
tests added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The #256 channel Send Again gated on status != sent, which hid it at the
radio-ack checkmark (status=sent) -- exactly the "transmitted but no repeat-back"
case it was for. It only surfaced on the waiting clock (pending). Gate on
repeatCount == 0 instead, so it shows until a repeater repeats the message back
(repeatCount > 0), then closes off. Matches the original design intent. Follow-up
to #256; DM gate (failed-only) unchanged.
The build stamp rendered as SelectableText, which ate taps for selection, so
only the plain 'Build' label triggered the unlock. When a row is tappable it
now uses a plain Text so the whole row (label + value) is one InkWell target;
copy stays on the button. Part of #509/#553.
Agent: QuietSnow (session 31eaba02)
Hiding the section left enabled features running invisibly. The action is now
'Disable experimental features': disableExperimental() turns off every
experimental flag (CoreScope observer counts, future ones) AND re-locks the
section in one atomic write. Also fixes the subtitle to say 'Build row' (the
#553 anchor), not 'version row'. Part of the #509/#553 experimental section.
Agent: QuietSnow (session 31eaba02)
Follow-up to #509. Settings is post-connect, so About had no availability edge,
and wrapping the About version subtitle fought the About dialog's onTap. The
7-tap unlock now anchors on the Build (BuildInfo.stamp) row in Device Info,
mirroring Android's build-number gesture; About reverts to a plain row. The
countdown/unlock hint renders inline under the Build row.
Co-located on the #524 test branch so the owner tests one APK; will be split
to its own PR off dev at land time. Agent: QuietSnow (session 31eaba02)
Badge stays observers-primary; tooltip/long-press now read 'N observers . M
observations' so they reconcile with CoreScope's 'Observations (N)' feed (the
two are distinct metrics: 15 distinct observers vs 34 total sightings). Refresh
feedback moved from a bottom snackbar to an auto-dismissed top MaterialBanner,
out of the way of the composer. Part of #524.
Agent: QuietSnow (session 31eaba02)
Owner testing: the inline badge needed pixel-accurate taps, the refresh gave no
feedback, and there was no refresh in the message action sheet. Now the badge is
a padded InkWell (real tap target), the long-press panel has a 'Refresh CoreScope
observers' action, both show a snackbar with the fetched count, and the auto-poll
is more patient (adds 90s tail steps, stabilises after 3 flat checks) so it climbs
closer to the total before you need to tap. Part of #524.
Agent: QuietSnow (session 31eaba02)
Poll now runs quick-then-slow (10/20/30/60x4s, ~5min cap) and stops early once
the count is flat for 2 checks, so it climbs to the near-final observer count
instead of stopping at a fixed cutoff. The badge is tappable to re-query on
demand (counts only ever grow), covering later re-checks. Part of #524.
Agent: QuietSnow (session 31eaba02)
Observer counts accrue over time (the packet must propagate the mesh and be
reported before CoreScope has any record), so an instant query at send+0.2s
always missed it. Now polls at ~10/30/60/120s, keeping the highest count as
observers report in; bails on disconnect. Part of #524.
Agent: QuietSnow (session 31eaba02)
Info logs at each step (0xC6 reply match, hash, CoreScope GET url + result,
stored count, message-not-found) so the app debug log shows exactly where the
badge chain breaks. Part of #524.
Agent: QuietSnow (session 31eaba02)
After a channel send, when firmware advertises 0xC6 (cap2 0x08 + ver>=22) AND
the owner enabled the feature, the connector queries the packet hash, then
CoreScope for observer_count, and stamps it on the message (background,
best-effort). Adds the coreScopeObserverCountEnabled AppSettings flag, a gated
switch in the #509 Experimental section (disabled until the radio supports the
capability), and a distinct cloud badge next to the radio-heard count. All
inert until the firmware PR (#611) lands. Completes the client side of #524.
Agent: QuietSnow (session 31eaba02)
Unifies the previously-duplicated send timestamps (frame builder vs outgoing
message each called now() separately) into one monotonic-per-channel value, so
every channel send has a unique (ts, channel_idx) key. That is the client-side
guarantee the 0xC6 correlation needs (VioletBarn caught that AES-128-ECB
determinism would otherwise make two same-second messages a wrong-hash query).
Adds transient onAirHash + coreScopeObserverCount to ChannelMessage. Part of #524.
Agent: QuietSnow (session 31eaba02)