Files
gronod ca723d15cf Publish sanitized protocol captures and documentation
Replace device credentials and identifiers consistently across packet
captures, documentation, and tests so the protocol evidence can be shared
publicly without exposing private device or network identities.
2026-09-25 13:49:11 +01:00

36 KiB
Raw Permalink Blame History

Ecovacs legacy XMPP protocol — full specification (robot-facing side)

Privacy notice: Device serials, MAC addresses, authentication values, controller identifiers, and other identifying values shown here have been consistently replaced with synthetic values.

Derived from four captures of a wukong-platform Ecovacs robot, device class 155:

Capture Window Duration Pkts Notes
packetcapture-ix1.12-20260923211202.pcap 21:12:21–21:14:24 ~2 min 326 boot + basic commands
packetcapture-ix1.12-20260923214817.pcap 21:48:43–21:59:15 ~10.5 min 841 clean state; schedules, stair-protection, errors, dock charge
packetcapture-ix1.12-20260923223825.pcap 22:38:44–22:46:01 ~7.3 min 100 forced wifi reconnects, cached endpoint reconnect, black-holed ping
packetcapture-ix1.12-20260923230041.pcap 23:01:02.724–23:05:08.209 245.5 s 37 half-open timeout through complete automatic recovery (35 IP frames + 2 ARP)

Device identity

  • Model platform: wukong, device class 155 (appears as XMPP domain 155.ecorobot.net and in the firmware URL)
  • Serial / DID / XMPP uid: E2998877665544332211
  • MAC: 02:00:00:00:95:01; DHCP hostname request: deebot
  • SASL credential: a4f19c2e7b603d58e1c947ab25d8063f — constant across all captured handshakes, therefore a persistent device credential rather than a session token
  • Robot IP: 172.20.20.12 via DHCP; gateway/DNS 172.20.0.1

Scope of this document: everything needed to implement the robot-facing side of a replacement server — i.e. the services a provisioned robot contacts from power-on until it is ready to accept commands. The app-facing side (portal REST API, user auth) is not required for the robot to connect and is covered only where the robot's behavior depends on it.


1. Boot sequence — shared order (captures 1 and 2)

The two power-on captures follow the same order. They are not packet-identical: capture 1 was offered a DHCP lease of 6761 s and sent one IGMPv2 report; capture 2 was offered 7200 s and sent two reports about 0.21 s apart. DNS retransmit gaps were 4.2 ms and 7.4 ms. Absolute times below are capture 1; capture 2 matches to within a few tens of milliseconds.

t+0.00  DHCP Offer + ACK          (hostname "deebot"; Discover/Request not in the capture)
t+0.01  ARP probe, t+0.49 second probe, t+1.08 announcement
t+2.00  ARP who-has gateway       → gateway MAC 02:00:00:00:00:01
t+2.00  IGMPv2 report → 226.1.1.1 (×1 in capture 1, ×2 in capture 2)
t+2.00  DNS A? lbo.ecouser.net    (2 queries, src port 4096, 4–7 ms apart, before the answer)
t+3.00  TCP → {lbo-ip}:8007 ×2 parallel   POST /lookup.do  (EcoMsgNew & EcoUpdate)
t+3.19  TCP → {update-ip}:8005            GET …/firmware/latest.json
t+3.53  TCP → {xmpp-ip}:5223              XMPP stream open
t+3.83  SASL PLAIN → success → re-stream → bind 'atom' → session → presence
t+4.03  READY (hello-world presence). No controller traffic until the app connects.

Boot→ready ≈ 4.0 s (presence at 4.026 s and 3.984 s). A stale session for the same JID is closed with </stream:stream> + FIN after the new session IQ, about 20 ms after that IQ and before hello world (capture 1: old port 16128, FIN at t=4.021, session result in the same millisecond; bind result was 24 ms earlier). The robot answers that FIN with two RSTs. Closing at bind is early relative to the captured server but is safe: the kick is done before presence.

Dependencies: only DNS and the lookup.do host are strictly needed to redirect the robot — everything downstream is learned from lookup.do responses. Firmware check may fail/404 without consequence.


2. DNS

  • Resolver used: the first server from DHCP option 6 (here 172.20.0.1; DHCP offered 172.20.0.1, 192.168.0.5, 192.168.0.4).
  • Query: A? lbo.ecouser.net — sent twice, 4.2 ms apart in capture 1 and 7.4 ms apart in capture 2, both before the response arrives.
  • Source port: 4096 on both queries in both boots.
  • Response observed: CNAME slb-oldweb-iot-eu.ww.ecouser.net → A 8.211.16.120. That address is also the EcoUpdate host returned by lookup.do. The XMPP host is a different literal address (47.87.130.1).
  • To redirect: resolve lbo.ecouser.net (and lbo.ecovacs.net for safety — see bumper docs) to the replacement server's IP. A wildcard address=/ecouser.net/{ip} + /ecovacs.net/{ip} + /ecovacs.com/{ip} covers everything, including {class}.ecorobot.net should the robot ever resolve it (it does not — it connects to the IP from lookup.do).

3. Service discovery — POST /lookup.do (plaintext HTTP, port 8007)

The robot opens one TCP connection per service, in parallel (EcoMsgNew's SYN leads EcoUpdate by under 1 ms; either response may arrive first). HTTP/1.0, no Host header, tiny JSON body (newline + tab indent, and a tab between : and the value), one request per connection. A replacement listener must accept HTTP/1.0 without Host. The captured server answered HTTP/1.1 with Connection: close and FIN; the robot then RSTs. One of the two lookup sockets was reset by the robot without a captured server FIN. Send the response and close; tolerate an immediate RST.

POST /lookup.do HTTP/1.0
Content-Length: 48
Accept: Application/json
Content-Type: application/json

{
	"todo":	"FindBest",
	"service":	"EcoMsgNew"
}
HTTP/1.1 200 OK
Content-Type: application/json; charset=utf-8
Content-Length: 46

{"result":"ok","ip":"47.87.130.1","port":5223}
service Returns Used for
EcoMsgNew {"result":"ok","ip":"<xmpp-ip>","port":5223} XMPP control channel
EcoUpdate {"result":"ok","ip":"<ota-ip>","port":8005} firmware check host

Critical formatting: body must be compact JSON with no spaces and port as a JSON number (not string). (Bumper comment: bot is "very picky" about this.)

For a local server: answer EcoMsgNew → {your-ip}, 5223. EcoUpdate may point anywhere — the robot tolerates a failed/absent update check, but pointing it at the local server and answering §4 keeps the boot path fully on-LAN.

4. Firmware check — GET on the EcoUpdate host (port 8005)

GET /products/{product}/class/{class}/firmware/latest.json HTTP/1.0
Connection: Close
Accept: Application/json
Content-Type: application/json

→ here /products/wukong/class/155/firmware/latest.json → 404, accepted gracefully; robot proceeds to XMPP about 180 ms later without retry. The captured 404 was HTTP/1.1, Content-Type: text/plain; charset=utf-8, Content-Length: 9, body exactly Not Found (no trailing newline). A JSON body or a successful manifest was not observed; do not invent one. XMPP starts only after this response.

5. XMPP service — TCP 5223, plaintext

Despite 5223 being the legacy SSL port, no TLS is negotiated: the server advertises <starttls><required/></starttls> and the robot ignores it, proceeding straight to SASL PLAIN. A replacement server does not need TLS at all for this robot; advertising starttls is optional cosmetic fidelity.

5.1 Handshake — verbatim

C→S <?xml version='1.0'?><stream:stream xmlns:stream='http://etherx.jabber.org/streams'
        xmlns='jabber:client' to='155.ecorobot.net' version='1.0'>
S→C <stream:stream xmlns:stream="http://etherx.jabber.org/streams"
        xmlns="jabber:client" version="1.0" id="{32hex}" from="155.ecorobot.net">
S→C <stream:features>
      <auth xmlns="http://jabber.org/features/iq-auth"/>
      <starttls xmlns="urn:ietf:params:xml:ns:xmpp-tls"><required/></starttls>
      <mechanisms xmlns="urn:ietf:params:xml:ns:xmpp-sasl"><mechanism>PLAIN</mechanism></mechanisms>
    </stream:features>
C→S <auth xmlns='urn:ietf:params:xml:ns:xmpp-sasl' mechanism='PLAIN'>{base64}</auth>
S→C <success xmlns="urn:ietf:params:xml:ns:xmpp-sasl"/>
C→S <?xml version='1.0'?><stream:stream … to='155.ecorobot.net' version='1.0'>   <!-- reopen -->
S→C <stream:stream … id="{same-32hex}" from="155.ecorobot.net">
S→C <stream:features><bind xmlns="urn:ietf:params:xml:ns:xmpp-bind"/>
      <session xmlns="urn:ietf:params:xml:ns:xmpp-session"/></stream:features>
C→S <iq type='set' id='0'><bind xmlns='urn:ietf:params:xml:ns:xmpp-bind'>
      <resource>atom</resource></bind></iq>
S→C <iq type="result" id="0"><bind xmlns="urn:ietf:params:xml:ns:xmpp-bind">
      <jid>E2998877665544332211@155.ecorobot.net/atom</jid></bind></iq>
C→S <iq type='set' id='1'><session xmlns='urn:ietf:params:xml:ns:xmpp-session'/></iq>
S→C <iq type="result" id="1"/>
C→S <presence><status>hello world</status></presence>
S→C <presence to="E2998877665544332211@155.ecorobot.net/atom"> dummy </presence>

5.2 Details a server must get right

  • Stream to= carries {class}.ecorobot.net — this is how the server learns the device class. Parse it out of the (unclosed) stream tag.
  • SASL PLAIN payload = base64("\0" + serial + "\0" + token); here \0E2998877665544332211\0a4f19c2e7b603d58e1c947ab25d8063f. Authcid = serial = uid. The password is a persistent factory/account credential. A local server accepts the presented password unconditionally.
  • JID: {serial}@{class}.ecorobot.net/atom. Resource is always atom.
  • Stream id: 32 hex digits. The post-SASL <stream:stream> repeats the same id. The next TCP connection gets a new id. The robot does not check it.
  • iq id: stanzas the robot originates (bind, session, pings, pushes, ctl results) use a counter that is monotonic for the whole boot, with no gaps, not per session. Result acks echo the controller's iq id and are not part of that counter. Bind=0/session=1 only on the first connection after power-on. Capture 1 ran 0..51, capture 2 ran 0..172, capture 3 continued 201..213 across two reconnects (bind 208/210, session 209/211), and capture 4 continued 222..225. The jump 213→222 is the eight uncaptured 120 s pings between those files, not a reset. Don't assume small numbers. The app's own iq ids are decimal strings of 3–8 digits and were reused (6677 after 16.9 s, 7689 after 315 s), once per later command, each with one ack. There was no duplicate ack of a single request.
  • After session → READY. The robot emits <presence><status>hello world </status></presence>; answer with a <presence> dummy </presence> addressed to its full JID.
  • <auth> (iq-auth) is advertised but never used by the robot.
  • Session uniqueness: on a new session for the same JID, close the previous connection (</stream:stream> + FIN). Observed ~20 ms after the robot's session IQ and before presence. The robot double-RSTs if that FIN arrives. If the old path is black-holed the FIN is never delivered; the new bind must still be accepted.
  • Framing: stanzas are a continuous non-document XML stream — implement a streaming parser: the robot may send several stanzas in one TCP segment, and a stanza may straddle segments. Treat <?xml …?> + <stream:stream> (never closed by the client) and a lone </stream:stream> (session end) specially.

6. The com:ctl command layer

6.1 Addressing

  • Bot JID: {serial}@{class}.ecorobot.net/atom
  • Controller JID: {uid}@ecouser.net/{resource} — e.g. the real app used demouser01234567@ecouser.net/LABclient001. The server relays stanzas between these JIDs; the robot directs all its responses/reports to the controller JID learned from incoming from= attributes. With one controller it is unambiguous; with several, route reports to the controller that last commanded the bot (or broadcast — observed data can't distinguish).
  • Controller announcement: the app announces itself with <iq type='get' to='{bot-jid}' from='{ctl-jid}'><ping xmlns='urn:xmpp:ping'/></iq>. That is the first post-ready traffic, but the delay is when the user opened the app, not a robot timer: 12.877 s in capture 2 and 38.751 s in capture 1. The app then sent a second ping (0.13–0.27 s later) and SetTime about 1 s after the first ping. Steady-state controller pings are ~90 s (89.96–92.21 s in capture 2). The bot answers type='result' to that JID. No report was emitted on a session that never received a controller from= (capture 3, both reconnects, including several minutes on the second). The first report in both boots was an empty Sched2, ~100 ms after the SetTime result and not in the gap between the announce pings and SetTime. A bridge must send its own from= on every session (ping, then SetTime, matching the app). The learned JID does not survive a re-bind.

6.2 Command (controller → bot)

<iq id="{sid}" to="{bot-jid}" from="{ctl-jid}" type="set">
  <query xmlns="com:ctl"><ctl td="{Command}" id="{cid}">…payload…</ctl></query>
</iq>
  • sid = stanza id (controller-chosen, 3–8 digits observed). cid = ctl correlation id. Every captured cid is a zero-padded 8-digit decimal (01410553, 00027119). The robot echoes it verbatim. The app reused a cid on an immediate status retry before the first response (GetCleanState 46393039 and 57306986); one ctl result then arrived. A bridge should keep cids unique among outstanding commands and still accept one result for a cid that was issued twice.
  • Several complete iq stanzas were written in one TCP segment (two Get*s, and once those two plus a ping). Segment boundaries are not stanza boundaries.
  • Move commands are sent without id on <ctl> and produce only the stanza ack (no ctl response). All others carry id. A new Move was sent without a preceding stop (TurnAround then SpinLeft, forward then forward). Observed move bursts lasted 0.24–3.1 s; no firmware auto-stop was seen inside that window.

6.3 Response — asymmetric two-stanza pattern

  1. Stanza-level ack: <iq type='result' … id='{sid}'/>
  2. Payload response as a new <iq type='set'> bot → controller:
<iq to='{ctl-jid}' type='set' id='{bot-seq}'>
  <query xmlns='com:ctl'><ctl id='{cid}' ret='ok' errno=''>…payload…</ctl></query>
</iq>

Correlation is by ctl/@id — not the iq id. ret was only ok in these captures (fail was not seen). errno='' is present on Get* and SetCleanSpeed results and absent on SetTime, Clean, Charge, PlaySound, AddSched, ModSched, and DelSched (<ctl id='…' ret='ok'/>). Treat a missing errno as no error. Do not require the attribute.

6.4 Reports (bot → controller, unsolicited)

<iq to='{ctl-jid}' type='set' id='{bot-seq}'>
  <query xmlns='com:ctl'><ctl td='{Report}'>…</ctl></query></iq>

td ∈ Sched2, CleanReport, ChargeState, BatteryInfo, error — detailed in §8/§9.

Push stanzas carry to= but no from= (don't require one when parsing), and — contrary to normal XMPP iq semantics — the controller never acks them: zero type='result' replies to push ids appear in ~10.5 min of capture. The bot neither expects nor notices. A bridge must not wait for acks on pushes and need not emit them.

6.5 Pings

  • Controller→bot: <iq type='get' to='{bot-jid}' from='{ctl-jid}'><ping xmlns='urn:xmpp:ping'/></iq> — the app's announce/keepalive, sent every ~90 s (also the first post-session stanza, §6.1); bot answers type='result' to the controller JID.
  • Bot→server: <iq from='{bot-jid}' to='155.ecorobot.net' type='get'><ping xmlns='urn:xmpp:ping'/></iq> every ~120 s; server must answer <iq type='result' from='155.ecorobot.net' to='{bot-jid}' id='{n}'/>.
  • The server itself never pings the bot (all bot-ward pings carry the controller from=).

7. Command reference (all td values seen on the wire)

Request/response XML is verbatim. {cid} = ctl id echoed in response.

7.1 SetTime — clock sync (always the first command after connect)

<ctl td="SetTime" id="{cid}"><time t="1790194386" tz="1" tzm="0"/></ctl>
→ <ctl id='{cid}' ret='ok'/>

t=epoch seconds, tz=tz hours offset, tzm=tz minutes offset (UTC+1 here). A standalone server should emit this itself once a bot session is ready.

7.2 GetBatteryInfo

<ctl id="{cid}" td="GetBatteryInfo"/>
→ <ctl id='{cid}' ret='ok' errno=''><battery power='076'/></ctl>

power = 0–100, zero-padded to 3 digits.

7.3 GetCleanState

<ctl id="{cid}" td="GetCleanState"/>
→ <ctl id='{cid}' ret='ok' errno=''><clean type='stop' speed='standard' st='h' t='' a=''/></ctl>

type = current/last clean mode. Observed snapshots: type='stop' st='h' and type='auto' st='s'. p/r were not in a GetCleanState reply. t and a were empty strings in every reply (not the single space used by CleanReport's st/rsn).

7.4 GetChargeState

<ctl id="{cid}" td="GetChargeState"/>
→ <ctl id='{cid}' ret='ok' errno=''><charge type='Idle'/></ctl>

type ∈ Idle (not docked/not charging), going (returning to dock), SlotCharging (on dock, charging).

7.5 GetCleanSpeed / SetCleanSpeed

<ctl id="{cid}" td="GetCleanSpeed"/>  → <ctl id='{cid}' ret='ok' errno='' speed='standard'/>
<ctl id="{cid}" td="SetCleanSpeed" speed="strong"/> → <ctl id='{cid}' ret='ok' errno=''/>

speed ∈ standard, strong. The ctl result is ret='ok' errno='' and does not echo speed. During an active auto clean, SetCleanSpeed was followed by a CleanReport at the new speed. While type='stop', a speed change was not consistently followed by a report. Apply the new speed when ret='ok', and also accept a later CleanReport or GetCleanSpeed.

7.6 GetLifeSpan — consumable life

<ctl id="{cid}" td="GetLifeSpan" type="SideBrush"/>
→ <ctl id='{cid}' ret='ok' errno='' type='SideBrush' val='068' total='365'/>

type ∈ SideBrush, Brush, DustCaseHeap (filter). val = % remaining (zero-padded, 068 = 68 %). total was 365 for all three types. The unit was not established; keep the integer and do not label it hours.

7.7 Move — manual driving

<ctl td="Move"><move action="forward"/></ctl>

action ∈ forward, backward (sucks-known, not exercised), SpinLeft, SpinRight, TurnAround, stop. No id on ctl → stanza-ack only.

7.8 Clean — start/stop a clean

<ctl id="{cid}" td="Clean"><clean type="auto" speed="strong" act="s"/></ctl>
→ <ctl id='{cid}' ret='ok'/>          (+ CleanReport push)

type ∈ auto, border (edge), spot, singleRoom (camelCase on the wire — sucks maps singleroom, a real vocab discrepancy), also stop for the stop-command itself. act = s start / h halt; p,r (pause/resume) known from sucks. speed ∈ standard|strong.

7.9 Charge — dock control

<ctl id="{cid}" td="Charge"><charge type="go"/></ctl>
→ <ctl id='{cid}' ret='ok'/>  (+ CleanReport stop + ChargeState 'going')

type = go (return to dock) / stopGo (cancel return). go is followed within ~50 ms by CleanReport stop and ChargeState going, and by SlotCharging when the robot is on the dock (21.4 s later in capture 2; again, with errno 100 and CleanReport stop, at the start of capture 3). stopGo is followed by CleanReport stop and ChargeState Idle.

7.10 PlaySound — find-me beep

<ctl id="{cid}" td="PlaySound" sid="0"/> → <ctl id='{cid}' ret='ok'/>

8. Schedule subsystem — fully captured

8.1 AddSched

<ctl id="{cid}" td="AddSched">
  <sched name="17901966021514" on="1" time="19:30" repeat="0001000">
    <ctl td="Clean"><clean type="auto"/></ctl>
  </sched></ctl>
→ <ctl id='{cid}' ret='ok'/>          (+ Sched2 push)

8.2 ModSched

<ctl id="{cid}" td="ModSched">
  <ModSched name="17901966021514">
    <sched name="17901966021514" on="0" time="19:30" repeat="0001000">
      <ctl td="Clean"><clean type="auto"/></ctl>
    </sched></ModSched></ctl>
→ <ctl id='{cid}' ret='ok'/>          (+ Sched2 push)

Wrapper <ModSched name="{existing-name}"> selects the entry to replace. Captured edits changed on, time, and repeat. The inner <sched name> was always equal to the wrapper name; a rename was not tested. The inner clean type was always auto.

8.3 DelSched

<ctl id="{cid}" td="DelSched"><DelSched name="17901966423286"/></ctl>
→ <ctl id='{cid}' ret='ok'/>          (+ Sched2 push)

8.4 GetSched and the Sched2 report — <s> element format

<ctl id='{cid}' ret='ok' errno=''>
  <s n='17901966302986' o='1' t='01:50' r='1101011' f='p'> <ctl td='clean' type='auto'/> </s>
  <s n='17901966021514' o='1' t='21:30' r='0111000' f='p'> <ctl td='clean' type='auto'/> </s>
</ctl>
attr meaning
n schedule name/id — app-generated unique string ≈ epoch_seconds*10^4 + suffix; opaque to the robot, echoed verbatim
o on/enabled 0/1
t local time HH:MM
r 7-char repeat bitmask index 0 = Sunday … index 6 = Saturday. Verified: 0001000 (index 3, Wednesday) fired on Wednesday 23 Sep 2026. Also observed: 0111000, 1101011, 1111111
f flag, constant 'p' in every <s> (meaning unknown; echo as stored)
  • The inner action is <ctl td='clean' type='auto'/> — lowercase clean, type directly on ctl. Only auto was stored. The command dialect uses a <clean type=…/> child and td="Clean". Each <s> contains a space before the inner <ctl> and a space before </s>. Adjacent <s> elements are concatenated with no text between them (</s><s).
  • Sched2 is pushed: (a) once a controller is known, immediately after the first SetTime (empty <ctl td='Sched2'/> when the table is empty — this is not at XMPP ready, which was 13–39 s earlier), (b) after every Add/Mod/Del, (c) when a schedule fires. At 21:58:59 local the bot pushed Sched2 for the 21:59 / 0001000 entry and, 73 ms later, CleanReport auto. That is under a second before the scheduled minute.
  • GetSched returns the same <s> children inside the ctl response. An empty table is a ctl result with ret='ok' errno='' and no <s> children, which is a different element from an empty Sched2 push.
  • DelSched of the last entry → Sched2 pushed with no <s> children (<ctl td='Sched2'/>), observed twice (after the first batch and after the 21:59 entry was disabled and deleted).

9. Robot-initiated reports (pushes)

All are <iq type='set' to='{ctl-jid}'><query xmlns='com:ctl'><ctl td=…/>:

td Payload When emitted
Sched2 <s …/> children or empty after the controller is known (first one ~100 ms after SetTime), after sched mutations, when a schedule fires
CleanReport <clean type='{mode}' speed='{spd}' st=' ' rsn=' '/> after Clean and Charge, and on autonomous transitions. st and rsn were a single space in every report, including while running. h/s appear only in GetCleanState. Observed report types: auto, border, spot, singleRoom, stop, speeds standard and strong. SetCleanSpeed updated a following report during an active clean only
ChargeState <charge type='Idle'/'going'/'SlotCharging'/> go → going within ~50 ms of CleanReport stop, then SlotCharging on arrival (21.4 s later in one run). stopGo → CleanReport stop + Idle. Idle is also pushed on leaving the dock (35 s after SlotCharging, with CleanReport stop). Idle while off the dock is also the GetChargeState answer during a clean
BatteryInfo <battery power='NNN'/> about every 25.00 s both while cleaning and while SlotCharging. Capture 2 fell from 77 to 68 over the clean, with one +1 step (70→71, which is the sample after docking). Capture 3 rose 67→68 on the dock. Slots were sometimes skipped (gaps of 50 s and 75 s)
error <ctl td='error' errno='NNN'/> error events — see §10
(ping) <ping xmlns='urn:xmpp:ping'/> to {class}.ecorobot.net every ~120 s

10. Error reporting — observed

<ctl td='error' errno='N'/> is a push, not a command response:

errno Context observed Behavior
103 mid auto-clean; CleanReport stop 24 ms later clean aborted — stair/cliff protection halt (per capture context). No command ctl-result carried errno 103
100 15.771 s after 103, then CleanReport auto 24 ms later (capture 2). Separately, at the start of capture 3: CleanReport stop 23 ms later and SlotCharging 56 ms later all-clear / error-cleared beacon, not a new fault. The reports that follow carry the new motion state. The stop 4.4 s after the capture-2 beacon was a later Clean act='h', not part of the beacon

errno='100' means "error cleared", not "error". In capture 2 it preceded a resumed CleanReport auto; in capture 3 it preceded CleanReport stop + SlotCharging. Consumers should treat it as clearing a prior fault and derive state from the CleanReport/ChargeState that follow, never as an error itself.

Command-level errors also exist (sucks/bumper knowledge, not exercised here): ret='fail' + errno on ctl responses (3,5,8 per sucks charge handling; 103 = permission denied on command responses — different from the td=error push!). Do not conflate: td='error' pushes are device fault reports.

11. Session lifecycle & timing behavior

  • Boot→ready ~4 s; XMPP session is long-lived.
  • Bot→server ping every ~120 s (to='{class}.ecorobot.net').
  • Controller pings relayed on demand; robot always answers result.
  • One session per JID: the new session kicks the old connection (</stream:stream>+FIN) about 20 ms after the session IQ and before presence — observed in captures 1 and 3. The robot double-RSTs when that FIN arrives.
  • Session end: </stream:stream> from either side; robot just drops TCP on power-off (no graceful close observed at shutdown).
  • Robot iq id sequence is strictly incrementing per boot, continuing across reconnects (§5.2).

11.1 Reconnect behavior (capture 3 — forced wifi drops + router state clears)

  • Reconnects skip the entire bootstrap. After a wifi drop the robot does DHCP renew (leases 7200 s then 7183 s) + ARP probe + one IGMPv2 report, then opens a fresh TCP connection directly to the cached EcoMsgNew IP:5223 — no DNS query, no lookup.do, no firmware check. The first reconnect reported 226.1.1.1; the second reported 224.0.0.1. The lookup.do result is cached for the life of the boot. The DNS/8007/8005 services are only needed at power-on. If the bridge IP changes, the robot cannot rediscover it without a reboot.
  • Every reconnect is a full re-handshake: stream → SASL PLAIN (same factory token) → re-stream (same stream id as that connection's first open) → bind atom → session → hello world. SYN to dummy presence was 0.42 s, 0.48 s, and 0.45 s. No credential or endpoint renegotiation exists.
  • Stale-session kick on each new session — </stream:stream>+FIN about 20 ms after the new session IQ, before presence. Not at the bind result.
  • No reports on a session with no announced controller — conns B and C in capture 3 produced zero pushes (the app was gone and never re-announced). Confirms §6.1: the learned controller JID is per-session and must be re-established after every reconnect.
  • Dead-path detection is pure TCP, and the RTO is adaptive. Capture 4 is the complete timeout. The last healthy bot ping (id=222) was acknowledged. Exactly 120.000 s later the bot sent ping id=223. With the return path black-holed, that identical 130-byte segment (same sequence, same XMPP id) was retransmitted at +0.668, +2.342, +5.368, +11.426, +23.468, +47.666, +95.863 s. No second XMPP stanza was generated. Capture 3 had already started this for ping id=213 and was still retransmitting when the file ended, on a longer RTO: +0.918, +2.999, +7.016, +15.043, +31.116, +63.289 s. Do not hard-code either series.
  • The robot abandons the half-open socket after 120.002 s (capture 4, measured from the original ping). It sent TCP FIN, did not wait for FIN-ACK, and did not send </stream:stream>.
  • Fresh connect after 4.999 s: new TCP connection to the same cached endpoint. SYN to dummy presence was 0.455 s; bind/session ids continued as 224/225. No DNS, lookup.do, DHCP, firmware lookup, or XEP-0198 resume. The first bot ping of a new session is not immediate (about 95 s after presence on capture 3's second session); on a stable session the period is 120.000–120.002 s.
  • The old server-side socket may remain half-open because neither its close nor the robot's FIN can traverse the cleared state. The new bind must atomically replace the JID→connection mapping and close/discard the old local socket; never reject the new bind merely because that JID appears connected.

11.2 Required reconnect state machine

ESTABLISHED
  bot ping every 120 s
  ping write/ack failure → kernel TCP retransmission
  120 s without delivery → robot sends FIN, abandons socket
  wait ~5 s
  TCP connect cached EcoMsgNew IP:port
  full XMPP authentication/bind/session (not XEP-0198 stream resumption)
  READY

A server cannot shorten the robot firmware's client-side 120 s timeout once packets are black-holed. It can improve Home Assistant accuracy independently by detecting its own failed controller pings/TCP keepalive and publishing offline before the robot reconnects.

Bridge-side improvements over what the real server demonstrated:

  1. Detect zombie sessions before the robot's own ~125 s ping-timeout/reconnect cycle: send controller pings every ~60 s and mark offline when one misses a 10–15 s result deadline. Optional TCP keepalive is an additional signal.
  2. On every new READY: re-announce + SetTime + status fan-out (mandatory — the per-session learned JID is gone).
  3. Availability will flap offline→online across reconnects; the retained state topics keep HA's entity populated throughout.

12. Value enumerations (complete observed set plus marked library values)

clean.type   auto | border | spot | singleRoom | stop      [SpotArea library-known]
clean.act    s (start) | h (halt)                          [p,r library-known]
clean.st     s (running, GetCleanState) | h (halted, GetCleanState) | ' ' (every CleanReport)
clean.speed / SetCleanSpeed.speed / GetCleanSpeed   standard | strong
move.action  forward | SpinLeft | SpinRight | TurnAround | stop  [backward library-known]
charge.type  go | stopGo                        (command)
charge state Idle | going | SlotCharging        (reports/queries)
lifespan.type SideBrush | Brush | DustCaseHeap
ctl.ret      ok | fail ;  ctl.errno '' | <numeric>
s.o          0|1 ;  s.r  7 chars, index 0 = Sunday … index 6 = Saturday ;  s.f 'p'

13. Firmware quirks to tolerate

  1. STARTTLS ignored despite <required/> — never wait for it.
  2. Two identical DNS queries 4–7 ms apart, both before the answer; two parallel lookup.do connections. No Host header.
  3. <query> may contain a bare <battery power='…'/> with no <ctl>. Seen twice, both times a full <iq type='set' id='{bot-seq}'> (capture 1 id 46 power 076; capture 2 id 60 power 077), 70–90 ms after a normal GetBatteryInfo result with the same power. The app sent no iq result for those ids. Parse the battery and do not ack.
  4. The app reused iq ids and ctl ids (see §5.2 and §6.2). Each request still got one result. There was no duplicate ack of one stanza ~600 ms apart. Complete a cid once; a later request may legally reuse it after the first result, and the captured app sometimes reused a ctl id before the result.
  5. Move ctl has no id; never expect a ctl response for it.
  6. Sched <s> elements contain literal-space text nodes and the inner action uses lowercase td='clean' with type on ctl.
  7. hello world presence has no type; answer with dummy presence.
  8. HTTP/1.0 requests, Accept: Application/json (capital A), tiny bodies; responses must be space-free JSON with numeric port.
  9. Stream from=/id= values are not validated by the robot (real server: from="{class}.ecorobot.net", random hex id — mimic for fidelity).
  10. Stanza boundaries ≠ TCP segment boundaries in both directions.

14. Minimum server checklist (robot-facing)

# Service Required behavior
1 DNS lbo.ecouser.net (and lbo.ecovacs.net) → server IP
2 TCP 8007 POST /lookup.do FindBest: EcoMsgNew→{ip,5223}, EcoUpdate→{ip,8005}; compact JSON
3 TCP 8005 GET /products/*/class/*/firmware/latest.json → 404 (or canned manifest)
4 TCP 5223 XMPP stream: features (+iq-auth, optional starttls advert, PLAIN), SASL accept-all, bind→{serial}@{class}.ecorobot.net/atom, session result, dummy presence
5 XMPP Parse to='{class}.ecorobot.net' for devclass; keep uid/JID table
6 XMPP urn:xmpp:ping: answer server-domain pings; relay controller pings
7 com:ctl Route iq/query/ctl between controller JIDs and bot JIDs verbatim (no schema validation); both type=set and result
8 com:ctl Optionally inject own commands from a virtual controller JID (app-free control) — the bot answers to from=; first from= seen per session registers the report-push destination (announce with a urn:xmpp:ping, repeat ~90 s); pushes need no ack
9 XMPP Atomically replace the JID→connection mapping and close/discard the stale local socket on same-JID re-bind; always accept the new connection even if the old half-open socket cannot receive its close
10 — Track last-reporting state (CleanReport/ChargeState/BatteryInfo/Sched2/error) for a status API

Nothing else is required: no TLS, no HTTP 443, no MQTT, no app auth — the robot is fully served by the above.

15. Bumper coverage map (updated after capture 2)

Already covered: lookup.do+FindBest/EcoMsgNew (confserver.py:422), plaintext XMPP handshake, devclass extraction from to=, SASL-accept for bots, atom bind → correct JID form, session/presence, ping handling both ways, transparent com:ctl relay both directions (including type='set' responses and all push reports — Sched2/error/CleanReport relay fine), errno=103 command-response repair flow (AddUser/SetAC/GetUserInfo).

Gaps/risks:

  1. EcoUpdate lookup → hardcoded real Ecovacs 47.88.66.164:8005; no local 8005 listener or latest.json route.
  2. lbo.ecouser.net absent from DNS docs (wildcard covers it).
  3. errno='103' substring match in _handle_result matches td='error' fault pushes. Those pushes have neither an error nor an admin attribute, so adminuser is never set and the handler raises. It should ignore td='error' pushes and only run the AddUser path for a ctl command response that actually carries error or admin.
  4. Stream from= domain (ecouser.net vs real {class}.ecorobot.net) and static stream id "1" — cosmetic.
  5. No stale-session kick on same-JID rebind.
  6. _handle_ctl crashes on to-less stanzas (all observed stanzas have to).
  7. Bumper sends GetDeviceInfo post-presence — not in the captured server behavior, and this N95 never sent or answered that command in these files. The response schema is unknown. Do not depend on it.
  8. sucks vocab: singleroom vs wire singleRoom; SetTime missing tzm; sched commands (AddSched/ModSched/DelSched/Sched2) absent from sucks entirely — documented here for the first time.

Bottom line: the protocol is now documented completely enough to implement the robot-facing server from scratch — bootstrap, discovery, handshake, command/response correlation, full td vocabulary including the schedule subsystem, all push reports, error semantics, keepalive, and firmware quirks.