Files
gronod ca723d15cf Publish sanitized protocol captures and documentation
Replace device credentials and identifiers consistently across packet
captures, documentation, and tests so the protocol evidence can be shared
publicly without exposing private device or network identities.
2026-09-25 13:49:11 +01:00

686 lines
36 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Ecovacs legacy XMPP protocol — full specification (robot-facing side)
> **Privacy notice:** Device serials, MAC addresses, authentication values, controller identifiers, and other identifying values shown here have been consistently replaced with synthetic values.
Derived from four captures of a **wukong**-platform Ecovacs robot, device **class 155**:
| Capture | Window | Duration | Pkts | Notes |
|---|---|---|---|---|
| `packetcapture-ix1.12-20260923211202.pcap` | 21:12:21–21:14:24 | ~2 min | 326 | boot + basic commands |
| `packetcapture-ix1.12-20260923214817.pcap` | 21:48:43–21:59:15 | ~10.5 min | 841 | clean state; schedules, stair-protection, errors, dock charge |
| `packetcapture-ix1.12-20260923223825.pcap` | 22:38:44–22:46:01 | ~7.3 min | 100 | forced wifi reconnects, cached endpoint reconnect, black-holed ping |
| `packetcapture-ix1.12-20260923230041.pcap` | 23:01:02.724–23:05:08.209 | 245.5 s | 37 | half-open timeout through complete automatic recovery (35 IP frames + 2 ARP) |
**Device identity**
- Model platform: `wukong`, device **class `155`** (appears as XMPP domain `155.ecorobot.net` and in the firmware URL)
- Serial / DID / XMPP uid: `E2998877665544332211`
- MAC: `02:00:00:00:95:01`; DHCP hostname request: **`deebot`**
- SASL credential: `a4f19c2e7b603d58e1c947ab25d8063f` — constant across all captured handshakes, therefore a persistent device credential rather than a session token
- Robot IP: `172.20.20.12` via DHCP; gateway/DNS `172.20.0.1`
**Scope of this document:** everything needed to implement the *robot-facing* side
of a replacement server — i.e. the services a **provisioned** robot contacts from
power-on until it is ready to accept commands. The app-facing side (portal REST
API, user auth) is *not* required for the robot to connect and is covered only
where the robot's behavior depends on it.
---
## 1. Boot sequence — shared order (captures 1 and 2)
The two power-on captures follow the same order. They are not packet-identical:
capture 1 was offered a DHCP lease of **6761 s** and sent **one** IGMPv2 report;
capture 2 was offered **7200 s** and sent **two** reports about 0.21 s apart.
DNS retransmit gaps were 4.2 ms and 7.4 ms. Absolute times below are capture 1;
capture 2 matches to within a few tens of milliseconds.
```
t+0.00 DHCP Offer + ACK (hostname "deebot"; Discover/Request not in the capture)
t+0.01 ARP probe, t+0.49 second probe, t+1.08 announcement
t+2.00 ARP who-has gateway → gateway MAC 02:00:00:00:00:01
t+2.00 IGMPv2 report → 226.1.1.1 (×1 in capture 1, ×2 in capture 2)
t+2.00 DNS A? lbo.ecouser.net (2 queries, src port 4096, 4–7 ms apart, before the answer)
t+3.00 TCP → {lbo-ip}:8007 ×2 parallel POST /lookup.do (EcoMsgNew & EcoUpdate)
t+3.19 TCP → {update-ip}:8005 GET …/firmware/latest.json
t+3.53 TCP → {xmpp-ip}:5223 XMPP stream open
t+3.83 SASL PLAIN → success → re-stream → bind 'atom' → session → presence
t+4.03 READY (hello-world presence). No controller traffic until the app connects.
```
Boot→ready ≈ **4.0 s** (presence at 4.026 s and 3.984 s). A stale session for the
same JID is closed with `</stream:stream>` + FIN **after the new session IQ**,
about 20 ms after that IQ and before `hello world` (capture 1: old port 16128,
FIN at t=4.021, session result in the same millisecond; bind result was 24 ms
earlier). The robot answers that FIN with two RSTs. Closing at bind is early
relative to the captured server but is safe: the kick is done before presence.
**Dependencies:** only **DNS** and the **`lookup.do` host** are strictly needed
to redirect the robot — everything downstream is learned from `lookup.do`
responses. Firmware check may fail/404 without consequence.
---
## 2. DNS
- Resolver used: the **first** server from DHCP option 6 (here `172.20.0.1`;
DHCP offered `172.20.0.1, 192.168.0.5, 192.168.0.4`).
- Query: `A? lbo.ecouser.net` — sent twice, 4.2 ms apart in capture 1 and 7.4 ms apart in capture 2, both before the response arrives.
- Source port: `4096` on both queries in both boots.
- Response observed: `CNAME slb-oldweb-iot-eu.ww.ecouser.net` → `A 8.211.16.120`. That address is also the `EcoUpdate` host returned by `lookup.do`. The XMPP host is a different literal address (`47.87.130.1`).
- To redirect: resolve `lbo.ecouser.net` (and `lbo.ecovacs.net` for safety — see
bumper docs) to the replacement server's IP. A wildcard
`address=/ecouser.net/{ip}` + `/ecovacs.net/{ip}` + `/ecovacs.com/{ip}`
covers everything, including `{class}.ecorobot.net` should the robot ever
resolve it (it does **not** — it connects to the IP from `lookup.do`).
## 3. Service discovery — `POST /lookup.do` (plaintext HTTP, port 8007)
The robot opens **one TCP connection per service, in parallel** (EcoMsgNew's
SYN leads EcoUpdate by under 1 ms; either response may arrive first). HTTP/1.0,
no `Host` header, tiny JSON body (newline + tab indent, and a tab between
`:` and the value), one request per connection. A replacement listener must
accept HTTP/1.0 without `Host`. The captured server answered HTTP/1.1 with
`Connection: close` and FIN; the robot then RSTs. One of the two lookup
sockets was reset by the robot without a captured server FIN. Send the
response and close; tolerate an immediate RST.
```http
POST /lookup.do HTTP/1.0
Content-Length: 48
Accept: Application/json
Content-Type: application/json
{
"todo": "FindBest",
"service": "EcoMsgNew"
}
```
```http
HTTP/1.1 200 OK
Content-Type: application/json; charset=utf-8
Content-Length: 46
{"result":"ok","ip":"47.87.130.1","port":5223}
```
| `service` | Returns | Used for |
|---|---|---|
| `EcoMsgNew` | `{"result":"ok","ip":"<xmpp-ip>","port":5223}` | XMPP control channel |
| `EcoUpdate` | `{"result":"ok","ip":"<ota-ip>","port":8005}` | firmware check host |
**Critical formatting:** body must be compact JSON with **no spaces** and
`port` as a JSON **number** (not string). (Bumper comment: bot is "very picky"
about this.)
For a local server: answer `EcoMsgNew → {your-ip}, 5223`. `EcoUpdate` may point
anywhere — the robot tolerates a failed/absent update check, but pointing it at
the local server and answering §4 keeps the boot path fully on-LAN.
## 4. Firmware check — `GET` on the `EcoUpdate` host (port 8005)
```http
GET /products/{product}/class/{class}/firmware/latest.json HTTP/1.0
Connection: Close
Accept: Application/json
Content-Type: application/json
```
→ here `/products/wukong/class/155/firmware/latest.json` → **`404`**, accepted
gracefully; robot proceeds to XMPP about 180 ms later without retry. The
captured 404 was `HTTP/1.1`, `Content-Type: text/plain; charset=utf-8`,
`Content-Length: 9`, body exactly `Not Found` (no trailing newline). A JSON
body or a successful manifest was not observed; do not invent one. XMPP starts
only after this response.
## 5. XMPP service — TCP 5223, **plaintext**
Despite 5223 being the legacy SSL port, **no TLS is negotiated**: the server
advertises `<starttls><required/></starttls>` and the robot ignores it,
proceeding straight to SASL PLAIN. A replacement server does not need TLS at
all for this robot; advertising `starttls` is optional cosmetic fidelity.
### 5.1 Handshake — verbatim
```xml
C→S <?xml version='1.0'?><stream:stream xmlns:stream='http://etherx.jabber.org/streams'
xmlns='jabber:client' to='155.ecorobot.net' version='1.0'>
S→C <stream:stream xmlns:stream="http://etherx.jabber.org/streams"
xmlns="jabber:client" version="1.0" id="{32hex}" from="155.ecorobot.net">
S→C <stream:features>
<auth xmlns="http://jabber.org/features/iq-auth"/>
<starttls xmlns="urn:ietf:params:xml:ns:xmpp-tls"><required/></starttls>
<mechanisms xmlns="urn:ietf:params:xml:ns:xmpp-sasl"><mechanism>PLAIN</mechanism></mechanisms>
</stream:features>
C→S <auth xmlns='urn:ietf:params:xml:ns:xmpp-sasl' mechanism='PLAIN'>{base64}</auth>
S→C <success xmlns="urn:ietf:params:xml:ns:xmpp-sasl"/>
C→S <?xml version='1.0'?><stream:stream … to='155.ecorobot.net' version='1.0'> <!-- reopen -->
S→C <stream:stream … id="{same-32hex}" from="155.ecorobot.net">
S→C <stream:features><bind xmlns="urn:ietf:params:xml:ns:xmpp-bind"/>
<session xmlns="urn:ietf:params:xml:ns:xmpp-session"/></stream:features>
C→S <iq type='set' id='0'><bind xmlns='urn:ietf:params:xml:ns:xmpp-bind'>
<resource>atom</resource></bind></iq>
S→C <iq type="result" id="0"><bind xmlns="urn:ietf:params:xml:ns:xmpp-bind">
<jid>E2998877665544332211@155.ecorobot.net/atom</jid></bind></iq>
C→S <iq type='set' id='1'><session xmlns='urn:ietf:params:xml:ns:xmpp-session'/></iq>
S→C <iq type="result" id="1"/>
C→S <presence><status>hello world</status></presence>
S→C <presence to="E2998877665544332211@155.ecorobot.net/atom"> dummy </presence>
```
### 5.2 Details a server must get right
- **Stream `to=`** carries `{class}.ecorobot.net` — this is how the server
learns the device class. Parse it out of the (unclosed) stream tag.
- **SASL PLAIN** payload = `base64("\0" + serial + "\0" + token)`; here
`\0E2998877665544332211\0a4f19c2e7b603d58e1c947ab25d8063f`. Authcid = serial =
uid. The password is a persistent factory/account credential. A local server
accepts the presented password unconditionally.
- **JID:** `{serial}@{class}.ecorobot.net/atom`. Resource is always `atom`.
- **Stream `id`:** 32 hex digits. The post-SASL `<stream:stream>` repeats the
same id. The next TCP connection gets a new id. The robot does not check it.
- **iq `id`:** stanzas the robot originates (bind, session, pings, pushes,
ctl results) use a counter that is **monotonic for the whole boot, with no
gaps**, not per session. Result acks echo the controller's iq id and are not
part of that counter. Bind=0/session=1 only on the first connection after
power-on. Capture 1 ran 0..51, capture 2 ran 0..172, capture 3 continued
201..213 across two reconnects (bind `208`/`210`, session `209`/`211`), and
capture 4 continued `222`..`225`. The jump 213→222 is the eight uncaptured
120 s pings between those files, not a reset. Don't assume small numbers.
The app's own iq ids are decimal strings of 3–8 digits and **were reused**
(`6677` after 16.9 s, `7689` after 315 s), once per later command, each with
one ack. There was no duplicate ack of a single request.
- After `session` → READY. The robot emits `<presence><status>hello world
</status></presence>`; answer with a `<presence> dummy </presence>` addressed
to its full JID.
- `<auth>` (iq-auth) is advertised but never used by the robot.
- **Session uniqueness:** on a new session for the same JID, close the previous
connection (`</stream:stream>` + FIN). Observed ~20 ms after the robot's
session IQ and before presence. The robot double-RSTs if that FIN arrives.
If the old path is black-holed the FIN is never delivered; the new bind must
still be accepted.
- **Framing:** stanzas are a continuous non-document XML stream — implement a
streaming parser: the robot may send several stanzas in one TCP segment, and
a stanza may straddle segments. Treat `<?xml …?>` + `<stream:stream>` (never
closed by the client) and a lone `</stream:stream>` (session end) specially.
---
## 6. The `com:ctl` command layer
### 6.1 Addressing
- **Bot JID:** `{serial}@{class}.ecorobot.net/atom`
- **Controller JID:** `{uid}@ecouser.net/{resource}` — e.g. the real app used
`demouser01234567@ecouser.net/LABclient001`. The server relays stanzas
between these JIDs; the robot directs all its responses/reports to the
controller JID learned from incoming `from=` attributes. With one controller
it is unambiguous; with several, route reports to the controller that last
commanded the bot (or broadcast — observed data can't distinguish).
- **Controller announcement:** the app announces itself with
`<iq type='get' to='{bot-jid}' from='{ctl-jid}'><ping xmlns='urn:xmpp:ping'/></iq>`.
That is the first post-ready traffic, but the delay is when the user opened
the app, not a robot timer: **12.877 s** in capture 2 and **38.751 s** in
capture 1. The app then sent a second ping (0.13–0.27 s later) and `SetTime`
about 1 s after the first ping. Steady-state controller pings are **~90 s**
(89.96–92.21 s in capture 2). The bot answers `type='result'` to that JID.
**No report was emitted on a session that never received a controller
`from=`** (capture 3, both reconnects, including several minutes on the
second). The first report in both boots was an empty `Sched2`, ~100 ms after
the `SetTime` result and not in the gap between the announce pings and
`SetTime`. A bridge must send its own `from=` on every session (ping, then
`SetTime`, matching the app). The learned JID does not survive a re-bind.
### 6.2 Command (controller → bot)
```xml
<iq id="{sid}" to="{bot-jid}" from="{ctl-jid}" type="set">
<query xmlns="com:ctl"><ctl td="{Command}" id="{cid}">…payload…</ctl></query>
</iq>
```
- `sid` = stanza id (controller-chosen, 3–8 digits observed). `cid` = **ctl
correlation id**. Every captured cid is a zero-padded 8-digit decimal
(`01410553`, `00027119`). The robot echoes it verbatim. The app reused a cid
on an immediate status retry before the first response (`GetCleanState`
`46393039` and `57306986`); one ctl result then arrived. A bridge should keep
cids unique among outstanding commands and still accept one result for a cid
that was issued twice.
- Several complete iq stanzas were written in one TCP segment (two `Get*`s, and
once those two plus a ping). Segment boundaries are not stanza boundaries.
- `Move` commands are sent **without** `id` on `<ctl>` and produce only the
stanza ack (no ctl response). All others carry `id`. A new `Move` was sent
without a preceding `stop` (`TurnAround` then `SpinLeft`, `forward` then
`forward`). Observed move bursts lasted 0.24–3.1 s; no firmware auto-stop
was seen inside that window.
### 6.3 Response — asymmetric **two-stanza** pattern
1. Stanza-level ack: `<iq type='result' … id='{sid}'/>`
2. Payload response as a **new `<iq type='set'>`** bot → controller:
```xml
<iq to='{ctl-jid}' type='set' id='{bot-seq}'>
<query xmlns='com:ctl'><ctl id='{cid}' ret='ok' errno=''>…payload…</ctl></query>
</iq>
```
Correlation is by **`ctl/@id`** — not the iq id. `ret` was only `ok` in these
captures (`fail` was not seen). `errno=''` is present on `Get*` and
`SetCleanSpeed` results and **absent** on `SetTime`, `Clean`, `Charge`,
`PlaySound`, `AddSched`, `ModSched`, and `DelSched` (`<ctl id='…' ret='ok'/>`).
Treat a missing `errno` as no error. Do not require the attribute.
### 6.4 Reports (bot → controller, unsolicited)
```xml
<iq to='{ctl-jid}' type='set' id='{bot-seq}'>
<query xmlns='com:ctl'><ctl td='{Report}'>…</ctl></query></iq>
```
`td` ∈ `Sched2`, `CleanReport`, `ChargeState`, `BatteryInfo`, `error` —
detailed in §8/§9.
**Push stanzas carry `to=` but no `from=`** (don't require one when parsing),
and — contrary to normal XMPP iq semantics — **the controller never acks
them**: zero `type='result'` replies to push ids appear in ~10.5 min of
capture. The bot neither expects nor notices. A bridge must not wait for acks
on pushes and need not emit them.
### 6.5 Pings
- Controller→bot: `<iq type='get' to='{bot-jid}' from='{ctl-jid}'><ping
xmlns='urn:xmpp:ping'/></iq>` — the app's announce/keepalive, sent **every
~90 s** (also the first post-session stanza, §6.1); bot answers
`type='result'` to the controller JID.
- Bot→server: `<iq from='{bot-jid}' to='155.ecorobot.net' type='get'><ping
xmlns='urn:xmpp:ping'/></iq>` every **~120 s**; server must answer
`<iq type='result' from='155.ecorobot.net' to='{bot-jid}' id='{n}'/>`.
- The server itself never pings the bot (all bot-ward pings carry the
controller `from=`).
---
## 7. Command reference (all `td` values seen on the wire)
Request/response XML is verbatim. `{cid}` = ctl id echoed in response.
### 7.1 `SetTime` — clock sync (always the first command after connect)
```xml
<ctl td="SetTime" id="{cid}"><time t="1790194386" tz="1" tzm="0"/></ctl>
→ <ctl id='{cid}' ret='ok'/>
```
`t`=epoch seconds, `tz`=tz hours offset, `tzm`=tz minutes offset (UTC+1 here).
**A standalone server should emit this itself** once a bot session is ready.
### 7.2 `GetBatteryInfo`
```xml
<ctl id="{cid}" td="GetBatteryInfo"/>
→ <ctl id='{cid}' ret='ok' errno=''><battery power='076'/></ctl>
```
`power` = 0–100, zero-padded to 3 digits.
### 7.3 `GetCleanState`
```xml
<ctl id="{cid}" td="GetCleanState"/>
→ <ctl id='{cid}' ret='ok' errno=''><clean type='stop' speed='standard' st='h' t='' a=''/></ctl>
```
`type` = current/last clean mode. Observed snapshots: `type='stop' st='h'`
and `type='auto' st='s'`. `p`/`r` were not in a `GetCleanState` reply.
`t` and `a` were empty strings in every reply (not the single space used by
`CleanReport`'s `st`/`rsn`).
### 7.4 `GetChargeState`
```xml
<ctl id="{cid}" td="GetChargeState"/>
→ <ctl id='{cid}' ret='ok' errno=''><charge type='Idle'/></ctl>
```
`type` ∈ `Idle` (not docked/not charging), `going` (returning to dock),
`SlotCharging` (on dock, charging).
### 7.5 `GetCleanSpeed` / `SetCleanSpeed`
```xml
<ctl id="{cid}" td="GetCleanSpeed"/> → <ctl id='{cid}' ret='ok' errno='' speed='standard'/>
<ctl id="{cid}" td="SetCleanSpeed" speed="strong"/> → <ctl id='{cid}' ret='ok' errno=''/>
```
`speed` ∈ `standard`, `strong`. The ctl result is `ret='ok' errno=''` and does
**not** echo `speed`. During an active auto clean, `SetCleanSpeed` was followed
by a `CleanReport` at the new speed. While `type='stop'`, a speed change was
not consistently followed by a report. Apply the new speed when `ret='ok'`,
and also accept a later `CleanReport` or `GetCleanSpeed`.
### 7.6 `GetLifeSpan` — consumable life
```xml
<ctl id="{cid}" td="GetLifeSpan" type="SideBrush"/>
→ <ctl id='{cid}' ret='ok' errno='' type='SideBrush' val='068' total='365'/>
```
`type` ∈ `SideBrush`, `Brush`, `DustCaseHeap` (filter). `val` = % remaining
(zero-padded, `068` = 68 %). `total` was `365` for all three types. The unit
was not established; keep the integer and do not label it hours.
### 7.7 `Move` — manual driving
```xml
<ctl td="Move"><move action="forward"/></ctl>
```
`action` ∈ `forward`, `backward` (sucks-known, not exercised), `SpinLeft`,
`SpinRight`, `TurnAround`, `stop`. **No `id` on ctl → stanza-ack only.**
### 7.8 `Clean` — start/stop a clean
```xml
<ctl id="{cid}" td="Clean"><clean type="auto" speed="strong" act="s"/></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ CleanReport push)
```
`type` ∈ `auto`, `border` (edge), `spot`, `singleRoom` (camelCase on the wire —
sucks maps `singleroom`, a real vocab discrepancy), also `stop` for the
stop-command itself. `act` = `s` start / `h` halt; `p`,`r` (pause/resume) known
from sucks. `speed` ∈ `standard|strong`.
### 7.9 `Charge` — dock control
```xml
<ctl id="{cid}" td="Charge"><charge type="go"/></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ CleanReport stop + ChargeState 'going')
```
`type` = `go` (return to dock) / `stopGo` (cancel return). `go` is followed
within ~50 ms by `CleanReport stop` and `ChargeState going`, and by
`SlotCharging` when the robot is on the dock (21.4 s later in capture 2;
again, with errno 100 and `CleanReport stop`, at the start of capture 3).
`stopGo` is followed by `CleanReport stop` and `ChargeState Idle`.
### 7.10 `PlaySound` — find-me beep
```xml
<ctl id="{cid}" td="PlaySound" sid="0"/> → <ctl id='{cid}' ret='ok'/>
```
---
## 8. Schedule subsystem — fully captured
### 8.1 `AddSched`
```xml
<ctl id="{cid}" td="AddSched">
<sched name="17901966021514" on="1" time="19:30" repeat="0001000">
<ctl td="Clean"><clean type="auto"/></ctl>
</sched></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ Sched2 push)
```
### 8.2 `ModSched`
```xml
<ctl id="{cid}" td="ModSched">
<ModSched name="17901966021514">
<sched name="17901966021514" on="0" time="19:30" repeat="0001000">
<ctl td="Clean"><clean type="auto"/></ctl>
</sched></ModSched></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ Sched2 push)
```
Wrapper `<ModSched name="{existing-name}">` selects the entry to replace.
Captured edits changed `on`, `time`, and `repeat`. The inner `<sched name>`
was always equal to the wrapper name; a rename was not tested. The inner
clean type was always `auto`.
### 8.3 `DelSched`
```xml
<ctl id="{cid}" td="DelSched"><DelSched name="17901966423286"/></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ Sched2 push)
```
### 8.4 `GetSched` and the `Sched2` report — `<s>` element format
```xml
<ctl id='{cid}' ret='ok' errno=''>
<s n='17901966302986' o='1' t='01:50' r='1101011' f='p'> <ctl td='clean' type='auto'/> </s>
<s n='17901966021514' o='1' t='21:30' r='0111000' f='p'> <ctl td='clean' type='auto'/> </s>
</ctl>
```
| attr | meaning |
|---|---|
| `n` | schedule name/id — app-generated unique string ≈ `epoch_seconds*10^4 + suffix`; opaque to the robot, echoed verbatim |
| `o` | on/enabled `0`/`1` |
| `t` | local time `HH:MM` |
| `r` | 7-char repeat bitmask **index 0 = Sunday … index 6 = Saturday**. Verified: `0001000` (index 3, Wednesday) fired on Wednesday 23 Sep 2026. Also observed: `0111000`, `1101011`, `1111111` |
| `f` | flag, constant `'p'` in every `<s>` (meaning unknown; echo as stored) |
- The inner action is `<ctl td='clean' type='auto'/>` — **lowercase `clean`,
`type` directly on `ctl`**. Only `auto` was stored. The command dialect uses
a `<clean type=…/>` child and `td="Clean"`. Each `<s>` contains a space
before the inner `<ctl>` and a space before `</s>`. Adjacent `<s>` elements
are concatenated with no text between them (`</s><s`).
- `Sched2` is pushed: **(a)** once a controller is known, immediately after the
first `SetTime` (empty `<ctl td='Sched2'/>` when the table is empty — this is
not at XMPP ready, which was 13–39 s earlier), **(b)** after every
Add/Mod/Del, **(c)** when a schedule fires. At 21:58:59 local the bot pushed
`Sched2` for the `21:59` / `0001000` entry and, 73 ms later, `CleanReport`
`auto`. That is under a second before the scheduled minute.
- `GetSched` returns the same `<s>` children inside the ctl response. An empty
table is a ctl result with `ret='ok' errno=''` and no `<s>` children, which
is a different element from an empty `Sched2` push.
- `DelSched` of the last entry → `Sched2` pushed with **no** `<s>` children
(`<ctl td='Sched2'/>`), observed twice (after the first batch and after the
21:59 entry was disabled and deleted).
---
## 9. Robot-initiated reports (pushes)
All are `<iq type='set' to='{ctl-jid}'><query xmlns='com:ctl'><ctl td=…/>`:
| `td` | Payload | When emitted |
|---|---|---|
| `Sched2` | `<s …/>` children or empty | after the controller is known (first one ~100 ms after SetTime), after sched mutations, when a schedule fires |
| `CleanReport` | `<clean type='{mode}' speed='{spd}' st=' ' rsn=' '/>` | after Clean and Charge, and on autonomous transitions. `st` and `rsn` were a single space in every report, including while running. `h`/`s` appear only in `GetCleanState`. Observed report types: `auto`, `border`, `spot`, `singleRoom`, `stop`, speeds `standard` and `strong`. `SetCleanSpeed` updated a following report during an active clean only |
| `ChargeState` | `<charge type='Idle'/'going'/'SlotCharging'/>` | `go` → `going` within ~50 ms of `CleanReport stop`, then `SlotCharging` on arrival (21.4 s later in one run). `stopGo` → `CleanReport stop` + `Idle`. `Idle` is also pushed on leaving the dock (35 s after `SlotCharging`, with `CleanReport stop`). `Idle` while off the dock is also the `GetChargeState` answer during a clean |
| `BatteryInfo` | `<battery power='NNN'/>` | about every **25.00 s** both while cleaning and while `SlotCharging`. Capture 2 fell from 77 to 68 over the clean, with one +1 step (70→71, which is the sample after docking). Capture 3 rose 67→68 on the dock. Slots were sometimes skipped (gaps of 50 s and 75 s) |
| `error` | `<ctl td='error' errno='NNN'/>` | **error events** — see §10 |
| *(ping)* | `<ping xmlns='urn:xmpp:ping'/>` to `{class}.ecorobot.net` | every ~120 s |
## 10. Error reporting — observed
`<ctl td='error' errno='N'/>` is a **push**, not a command response:
| errno | Context observed | Behavior |
|---|---|---|
| `103` | mid auto-clean; `CleanReport stop` 24 ms later | clean aborted — stair/cliff protection halt (per capture context). No command ctl-result carried errno 103 |
| `100` | 15.771 s after 103, then `CleanReport auto` 24 ms later (capture 2). Separately, at the start of capture 3: `CleanReport stop` 23 ms later and `SlotCharging` 56 ms later | **all-clear / error-cleared beacon**, not a new fault. The reports that follow carry the new motion state. The stop 4.4 s after the capture-2 beacon was a later `Clean` `act='h'`, not part of the beacon |
**`errno='100'` means "error cleared", not "error".** In capture 2 it preceded
a resumed `CleanReport auto`; in capture 3 it preceded `CleanReport stop` +
`SlotCharging`. Consumers should treat it as clearing a prior fault and derive
state from the `CleanReport`/`ChargeState` that follow, never as an error
itself.
Command-level errors also exist (sucks/bumper knowledge, not exercised here):
`ret='fail'` + `errno` on ctl responses (`3`,`5`,`8` per sucks charge handling;
`103` = permission denied on *command responses* — different from the `td=error`
push!). **Do not conflate:** `td='error'` pushes are device fault reports.
## 11. Session lifecycle & timing behavior
- **Boot→ready ~4 s**; XMPP session is long-lived.
- Bot→server ping every ~120 s (`to='{class}.ecorobot.net'`).
- Controller pings relayed on demand; robot always answers `result`.
- **One session per JID:** the new session kicks the old connection
(`</stream:stream>`+FIN) about 20 ms after the session IQ and before
presence — observed in captures 1 and 3. The robot double-RSTs when that
FIN arrives.
- Session end: `</stream:stream>` from either side; robot just drops TCP on
power-off (no graceful close observed at shutdown).
- Robot iq `id` sequence is strictly incrementing **per boot**, continuing
across reconnects (§5.2).
### 11.1 Reconnect behavior (capture 3 — forced wifi drops + router state clears)
- **Reconnects skip the entire bootstrap.** After a wifi drop the robot does
DHCP renew (leases 7200 s then 7183 s) + ARP probe + one IGMPv2 report, then
opens a fresh TCP connection **directly to the cached `EcoMsgNew` IP:5223**
— no DNS query, no `lookup.do`, no firmware check. The first reconnect
reported `226.1.1.1`; the second reported `224.0.0.1`. The `lookup.do`
result is cached for the life of the boot. The DNS/`8007`/`8005` services
are only needed at power-on. **If the bridge IP changes, the robot cannot
rediscover it without a reboot.**
- **Every reconnect is a full re-handshake**: stream → SASL PLAIN (same
factory token) → re-stream (same stream id as that connection's first open)
→ bind `atom` → session → `hello world`. SYN to dummy presence was 0.42 s,
0.48 s, and 0.45 s. No credential or endpoint renegotiation exists.
- **Stale-session kick on each new session** — `</stream:stream>`+FIN about
20 ms after the new session IQ, before presence. Not at the bind result.
- **No reports on a session with no announced controller** — conns B and C in
capture 3 produced *zero* pushes (the app was gone and never re-announced).
Confirms §6.1: the learned controller JID is per-session and must be
re-established after every reconnect.
- **Dead-path detection is pure TCP, and the RTO is adaptive.** Capture 4 is
the complete timeout. The last healthy bot ping (`id=222`) was acknowledged.
Exactly 120.000 s later the bot sent ping `id=223`. With the return path
black-holed, that identical 130-byte segment (same sequence, same XMPP id)
was retransmitted at `+0.668, +2.342, +5.368, +11.426, +23.468, +47.666,
+95.863 s`. No second XMPP stanza was generated. Capture 3 had already
started this for ping `id=213` and was still retransmitting when the file
ended, on a longer RTO: `+0.918, +2.999, +7.016, +15.043, +31.116, +63.289 s`.
Do not hard-code either series.
- **The robot abandons the half-open socket after 120.002 s** (capture 4,
measured from the original ping). It sent TCP FIN, did not wait for FIN-ACK,
and did not send `</stream:stream>`.
- **Fresh connect after 4.999 s:** new TCP connection to the same cached
endpoint. SYN to dummy presence was 0.455 s; bind/session ids continued as
`224`/`225`. No DNS, `lookup.do`, DHCP, firmware lookup, or XEP-0198 resume.
The first bot ping of a new session is not immediate (about 95 s after
presence on capture 3's second session); on a stable session the period is
120.000–120.002 s.
- The old server-side socket may remain half-open because neither its close nor
the robot's FIN can traverse the cleared state. The new bind must atomically
replace the JID→connection mapping and close/discard the old local socket;
never reject the new bind merely because that JID appears connected.
### 11.2 Required reconnect state machine
```text
ESTABLISHED
bot ping every 120 s
ping write/ack failure → kernel TCP retransmission
120 s without delivery → robot sends FIN, abandons socket
wait ~5 s
TCP connect cached EcoMsgNew IP:port
full XMPP authentication/bind/session (not XEP-0198 stream resumption)
READY
```
A server cannot shorten the robot firmware's client-side 120 s timeout once
packets are black-holed. It can improve Home Assistant accuracy independently
by detecting its own failed controller pings/TCP keepalive and publishing
`offline` before the robot reconnects.
**Bridge-side improvements over what the real server demonstrated:**
1. Detect zombie sessions before the robot's own ~125 s ping-timeout/reconnect
cycle: send controller pings every ~60 s and mark offline when one misses a
10–15 s result deadline. Optional TCP keepalive is an additional signal.
2. On every new READY: re-announce + `SetTime` + status fan-out (mandatory —
the per-session learned JID is gone).
3. Availability will flap offline→online across reconnects; the retained
state topics keep HA's entity populated throughout.
## 12. Value enumerations (complete observed set plus marked library values)
```
clean.type auto | border | spot | singleRoom | stop [SpotArea library-known]
clean.act s (start) | h (halt) [p,r library-known]
clean.st s (running, GetCleanState) | h (halted, GetCleanState) | ' ' (every CleanReport)
clean.speed / SetCleanSpeed.speed / GetCleanSpeed standard | strong
move.action forward | SpinLeft | SpinRight | TurnAround | stop [backward library-known]
charge.type go | stopGo (command)
charge state Idle | going | SlotCharging (reports/queries)
lifespan.type SideBrush | Brush | DustCaseHeap
ctl.ret ok | fail ; ctl.errno '' | <numeric>
s.o 0|1 ; s.r 7 chars, index 0 = Sunday … index 6 = Saturday ; s.f 'p'
```
## 13. Firmware quirks to tolerate
1. **STARTTLS ignored** despite `<required/>` — never wait for it.
2. Two identical DNS queries 4–7 ms apart, both before the answer; two parallel `lookup.do` connections. No `Host` header.
3. `<query>` may contain a **bare `<battery power='…'/>` with no `<ctl>`**.
Seen twice, both times a full `<iq type='set' id='{bot-seq}'>` (capture 1
id `46` power `076`; capture 2 id `60` power `077`), 70–90 ms after a normal
`GetBatteryInfo` result with the same power. The app sent no iq result for
those ids. Parse the battery and do not ack.
4. The app reused iq ids and ctl ids (see §5.2 and §6.2). Each request still
got one result. There was **no** duplicate ack of one stanza ~600 ms apart.
Complete a cid once; a later request may legally reuse it after the first
result, and the captured app sometimes reused a ctl id before the result.
5. `Move` ctl has no `id`; never expect a ctl response for it.
6. Sched `<s>` elements contain literal-space text nodes and the inner action
uses lowercase `td='clean'` with `type` on `ctl`.
7. `hello world` presence has no `type`; answer with ` dummy ` presence.
8. HTTP/1.0 requests, `Accept: Application/json` (capital A), tiny bodies;
responses must be space-free JSON with numeric `port`.
9. Stream `from=`/`id=` values are not validated by the robot (real server:
`from="{class}.ecorobot.net"`, random hex id — mimic for fidelity).
10. Stanza boundaries ≠ TCP segment boundaries in both directions.
## 14. Minimum server checklist (robot-facing)
| # | Service | Required behavior |
|---|---|---|
| 1 | DNS | `lbo.ecouser.net` (and `lbo.ecovacs.net`) → server IP |
| 2 | TCP 8007 | `POST /lookup.do` `FindBest`: `EcoMsgNew`→`{ip,5223}`, `EcoUpdate`→`{ip,8005}`; compact JSON |
| 3 | TCP 8005 | `GET /products/*/class/*/firmware/latest.json` → 404 (or canned manifest) |
| 4 | TCP 5223 | XMPP stream: features (+iq-auth, optional starttls advert, PLAIN), SASL accept-all, bind→`{serial}@{class}.ecorobot.net/atom`, session result, dummy presence |
| 5 | XMPP | Parse `to='{class}.ecorobot.net'` for devclass; keep uid/JID table |
| 6 | XMPP | `urn:xmpp:ping`: answer server-domain pings; relay controller pings |
| 7 | com:ctl | Route `iq/query/ctl` between controller JIDs and bot JIDs **verbatim** (no schema validation); both `type=set` and `result` |
| 8 | com:ctl | Optionally inject own commands from a virtual controller JID (app-free control) — the bot answers to `from=`; first `from=` seen per session registers the report-push destination (announce with a `urn:xmpp:ping`, repeat ~90 s); pushes need no ack |
| 9 | XMPP | Atomically replace the JID→connection mapping and close/discard the stale local socket on same-JID re-bind; always accept the new connection even if the old half-open socket cannot receive its close |
| 10 | — | Track last-reporting state (`CleanReport`/`ChargeState`/`BatteryInfo`/`Sched2`/`error`) for a status API |
Nothing else is required: no TLS, no HTTP 443, no MQTT, no app auth — the robot
is fully served by the above.
## 15. Bumper coverage map (updated after capture 2)
**Already covered:** `lookup.do`+`FindBest`/`EcoMsgNew` (confserver.py:422),
plaintext XMPP handshake, devclass extraction from `to=`, SASL-accept for bots,
`atom` bind → correct JID form, session/presence, ping handling both ways,
transparent `com:ctl` relay both directions (including `type='set'` responses
and all push reports — Sched2/error/CleanReport relay fine), errno=103
*command-response* repair flow (AddUser/SetAC/GetUserInfo).
**Gaps/risks:**
1. `EcoUpdate` lookup → hardcoded real Ecovacs `47.88.66.164:8005`; no local
8005 listener or `latest.json` route.
2. `lbo.ecouser.net` absent from DNS docs (wildcard covers it).
3. **`errno='103'` substring match in `_handle_result` matches `td='error'`
fault pushes.** Those pushes have neither an `error` nor an `admin`
attribute, so `adminuser` is never set and the handler raises. It should
ignore `td='error'` pushes and only run the AddUser path for a ctl command
response that actually carries `error` or `admin`.
4. Stream `from=` domain (`ecouser.net` vs real `{class}.ecorobot.net`) and
static stream id `"1"` — cosmetic.
5. No stale-session kick on same-JID rebind.
6. `_handle_ctl` crashes on `to`-less stanzas (all observed stanzas have `to`).
7. Bumper sends `GetDeviceInfo` post-presence — not in the captured server
behavior, and this N95 never sent or answered that command in these files.
The response schema is unknown. Do not depend on it.
8. sucks vocab: `singleroom` vs wire `singleRoom`; `SetTime` missing `tzm`;
sched commands (`AddSched`/`ModSched`/`DelSched`/`Sched2`) **absent from
sucks entirely** — documented here for the first time.
**Bottom line:** the protocol is now documented completely enough to implement
the robot-facing server from scratch — bootstrap, discovery, handshake,
command/response correlation, full `td` vocabulary including the schedule
subsystem, all push reports, error semantics, keepalive, and firmware quirks.