2 Commits
Author SHA1 Message Date
gronod 3ab6795be6 Merge remote-tracking branch 'origin/main' 2026-09-25 13:52:53 +01:00
gronod ca723d15cf Publish sanitized protocol captures and documentation
Replace device credentials and identifiers consistently across packet
captures, documentation, and tests so the protocol evidence can be shared
publicly without exposing private device or network identities.
2026-09-25 13:49:11 +01:00
10 changed files with 2004 additions and 33 deletions
-7
View File
@@ -1,13 +1,6 @@
.env
/n95bridge
packetcapture-ix1.12-20260923211202.pcap
packetcapture-ix1.12-20260923214817.pcap
packetcapture-ix1.12-20260923223825.pcap
packetcapture-ix1.12-20260923230041.pcap
docs/M1-MEGAPLAN.md
docs/MQTT-BRIDGE.md
docs/N95-FULL-SPECIFICATION.md
docs/PCAP-ANALYSIS.md
docs/megaplans/M1/01-config-http-bootstrap.md
docs/megaplans/M1/02-xmpp-session.md
docs/megaplans/M1/03-ctl-correlation-and-state.md
+3 -1
View File
@@ -1,5 +1,7 @@
# Deebot N95 local MQTT bridge
> **Privacy notice:** Device serials, MAC addresses, authentication values, controller identifiers, and other identifying values shown in the included captures and documentation have been consistently replaced with synthetic values.
This service redirects an **already provisioned Ecovacs Deebot N95** from its legacy cloud bootstrap and XMPP endpoint to a local bridge. It exposes the robot to Home Assistant through MQTT discovery as a native MQTT vacuum. The bridge supports multiple robots, with one MQTT client and one topic tree per robot serial number. It does not provision a robot or run a DNS server.
```text
@@ -171,6 +173,6 @@ The bridge logs to stderr. `LOG_LEVEL=debug` adds redacted XMPP stanza logging.
## Development and project documents
Run `go test ./...` from the repository root. The capture replay tests expect four `packetcapture-ix1.12-*.pcap` files at that root; these are local, untracked fixtures and may be absent from a fresh checkout. Without them, the capture tests fail even though the runtime build does not need them. A broker integration test is available with `go test -tags mqttintegration ./cmd/n95bridge` and `MQTT_TEST_URL` set to a test broker URL; it skips when that variable is unset. `MQTT_TEST_URL` is test-only, not a service setting.
Run `go test ./...` from the repository root. The capture replay tests use the four sanitized `packetcapture-ix1.12-*.pcap` fixtures committed under `docs/pcaps/`; the runtime build does not need them. A broker integration test is available with `go test -tags mqttintegration ./cmd/n95bridge` and `MQTT_TEST_URL` set to a test broker URL; it skips when that variable is unset. `MQTT_TEST_URL` is test-only, not a service setting.
Protocol and design details are in [the full N95 specification](docs/N95-FULL-SPECIFICATION.md), [packet capture analysis](docs/PCAP-ANALYSIS.md), [MQTT bridge mapping](docs/MQTT-BRIDGE.md), and [the milestone plan](docs/M1-MEGAPLAN.md). Those documents describe some intended or observed protocol behavior beyond the currently implemented runtime; the configuration and limitations above reflect the code in this repository.
+25 -25
View File
@@ -32,10 +32,10 @@ import (
const pcapMagicLittleEndian uint32 = 0xa1b2c3d4
var captureFilenames = []string{
"packetcapture-ix1.12-20260923211202.pcap",
"packetcapture-ix1.12-20260923214817.pcap",
"packetcapture-ix1.12-20260923223825.pcap",
"packetcapture-ix1.12-20260923230041.pcap",
"docs/pcaps/packetcapture-ix1.12-20260923211202.pcap",
"docs/pcaps/packetcapture-ix1.12-20260923214817.pcap",
"docs/pcaps/packetcapture-ix1.12-20260923223825.pcap",
"docs/pcaps/packetcapture-ix1.12-20260923230041.pcap",
}
type tcpSegment struct {
@@ -126,7 +126,7 @@ func readPCAPSegments(path string) ([]tcpSegment, error) {
seq := binary.BigEndian.Uint32(tcpData[4:8])
ack := binary.BigEndian.Uint32(tcpData[8:12])
offsetFlags := binary.BigEndian.Uint16(tcpData[12:14])
tcpLen := int((offsetFlags >> 12) & 0x0f) * 4
tcpLen := int((offsetFlags>>12)&0x0f) * 4
if len(tcpData) < tcpLen {
continue
}
@@ -206,7 +206,7 @@ func TestCaptureFilesPresent(t *testing.T) {
}
func TestCapture1LookupPair(t *testing.T) {
pcapPath, err := findRepoFile("packetcapture-ix1.12-20260923211202.pcap")
pcapPath, err := findRepoFile("docs/pcaps/packetcapture-ix1.12-20260923211202.pcap")
if err != nil {
t.Fatalf("pcap file not found: %v", err)
}
@@ -300,7 +300,7 @@ func TestCapture1LookupPair(t *testing.T) {
}
func TestCapture1Firmware404(t *testing.T) {
pcapPath, err := findRepoFile("packetcapture-ix1.12-20260923211202.pcap")
pcapPath, err := findRepoFile("docs/pcaps/packetcapture-ix1.12-20260923211202.pcap")
if err != nil {
t.Fatalf("pcap file not found: %v", err)
}
@@ -398,7 +398,7 @@ func (c *pcapXMPPClient) readUntil(delim string) string {
}
func TestCapture1Handshake(t *testing.T) {
pcapPath, err := findRepoFile("packetcapture-ix1.12-20260923211202.pcap")
pcapPath, err := findRepoFile("docs/pcaps/packetcapture-ix1.12-20260923211202.pcap")
if err != nil {
t.Fatalf("pcap file not found: %v", err)
}
@@ -544,7 +544,7 @@ func TestCapture1Handshake(t *testing.T) {
}
func TestCapture1BareBattery(t *testing.T) {
pcapPath, err := findRepoFile("packetcapture-ix1.12-20260923211202.pcap")
pcapPath, err := findRepoFile("docs/pcaps/packetcapture-ix1.12-20260923211202.pcap")
if err != nil {
t.Fatalf("pcap file not found: %v", err)
}
@@ -599,18 +599,18 @@ func TestCapture1BareBattery(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
testBotJID := "E2001229192001911354@155.ecorobot.net/atom"
actor := robot.NewActor(ctx, testBotJID, "E2001229192001911354", "controller@ecouser.net/ha", pubFn, nil, nil, nil)
testBotJID := "E2998877665544332211@155.ecorobot.net/atom"
actor := robot.NewActor(ctx, testBotJID, "E2998877665544332211", "controller@ecouser.net/ha", pubFn, nil, nil, nil)
actor.SessionReady(session.ReadyEvent{
Generation: 1,
JID: testBotJID,
Serial: "E2001229192001911354",
Serial: "E2998877665544332211",
Send: sendFn,
})
actor.Stanza(session.StanzaEvent{
Generation: 1,
JID: testBotJID,
Serial: "E2001229192001911354",
Serial: "E2998877665544332211",
Stanza: match,
})
time.Sleep(50 * time.Millisecond)
@@ -626,7 +626,7 @@ func TestCapture1BareBattery(t *testing.T) {
}
func TestCapture2Sched2(t *testing.T) {
pcapPath, err := findRepoFile("packetcapture-ix1.12-20260923214817.pcap")
pcapPath, err := findRepoFile("docs/pcaps/packetcapture-ix1.12-20260923214817.pcap")
if err != nil {
t.Fatalf("pcap file not found: %v", err)
}
@@ -694,15 +694,15 @@ func TestCapture2Sched2(t *testing.T) {
// PCAP-ANALYSIS.md §11.1 and N95-FULL-SPECIFICATION.md §7:
//
// Timeline from packetcapture-ix1.12-20260923230041.pcap:
// - The last healthy bot ping was answered.
// - 120 seconds later the next bot ping was black-holed. The same TCP segment was
// retransmitted at +0.668, +2.342, +5.368, +11.426, +23.468, +47.666, and +95.863 seconds.
// - An earlier incomplete series in packetcapture-ix1.12-20260923223825.pcap used
// +0.918, +2.999, +7.016, +15.043, +31.116, and +63.289 seconds.
// - The robot sent FIN at +120.002 seconds from the original ping, with no XMPP stream
// close and without waiting for FIN-ACK.
// - A new TCP connection opened 4.999 seconds after that FIN. SYN to dummy presence took
// 0.455 seconds. There was no DNS, lookup, firmware check, or XEP-0198 resume.
// - The last healthy bot ping was answered.
// - 120 seconds later the next bot ping was black-holed. The same TCP segment was
// retransmitted at +0.668, +2.342, +5.368, +11.426, +23.468, +47.666, and +95.863 seconds.
// - An earlier incomplete series in packetcapture-ix1.12-20260923223825.pcap used
// +0.918, +2.999, +7.016, +15.043, +31.116, and +63.289 seconds.
// - The robot sent FIN at +120.002 seconds from the original ping, with no XMPP stream
// close and without waiting for FIN-ACK.
// - A new TCP connection opened 4.999 seconds after that FIN. SYN to dummy presence took
// 0.455 seconds. There was no DNS, lookup, firmware check, or XEP-0198 resume.
//
// The bridge's deadline is 12 seconds (xmpp.PingResultTimeout). Reconnect handling
// accepts the down-then-ready transition without sleeping or delay.
@@ -716,8 +716,8 @@ func TestCapture4ReconnectIsNotASleep(t *testing.T) {
registry := session.NewRegistry()
bus := session.NewBus()
testJID := "E2001229192001911354@155.ecorobot.net/atom"
serial := "E2001229192001911354"
testJID := "E2998877665544332211@155.ecorobot.net/atom"
serial := "E2998877665544332211"
// Register initial generation.
gen1, replaced := registry.Bind(testJID, nil, nil)
+314
View File
@@ -0,0 +1,314 @@
# MQTT Bridge — `com:ctl` (XMPP) ↔ MQTT ↔ Home Assistant
> **Privacy notice:** Device serials, MAC addresses, authentication values, controller identifiers, and other identifying values shown here have been consistently replaced with synthetic values.
Maps the robot-facing `com:ctl` protocol (see `PCAP-ANALYSIS.md`) onto an
**existing, supported MQTT control schema**: **Home Assistant's native
`mqtt.vacuum` integration**. No custom HA component, no YAML for the entity —
the bridge publishes a discovery config and HA does the rest.
```
[Robot N95] ──XMPP com:ctl──▶ [BRIDGE: local XMPP server + MQTT client]
│ ▲
▼ │
[MQTT broker] ◀──▶ [Home Assistant mqtt.vacuum]
```
Device under bridge: Ecovacs **Deebot N95** (`wukong` platform, class `155`,
serial `E2998877665544332211`, JID `E2998877665544332211@155.ecorobot.net/atom`).
**Why this schema:** `mqtt.vacuum` is built into HA (used by ~half of MQTT
installs, e.g. Valetudo). Its vocabulary — start/pause/stop/dock/spot/locate/
fan-speed/state — maps almost 1:1 onto the `com:ctl` dialect of this robot.
Valetudo's own schema was considered and rejected: it assumes a rooted,
mapping-capable robot (segments, maps, layers) that the N95 doesn't have.
**Prerequisites for a working deployment**
- DNS override: `lbo.ecouser.net` → the bridge host (PCAP-ANALYSIS §2).
- Bridge listeners on that host: HTTP `:8007` (`/lookup.do`), HTTP `:8005`
(`/products/…/latest.json` → 404), XMPP `:5223` (§14 checklist of
PCAP-ANALYSIS covers all three; `lookup.do` should return the bridge's own
IP for both `EcoMsgNew` and `EcoUpdate`).
- An existing MQTT broker; Home Assistant's MQTT integration pointed at it
with discovery enabled (default prefix `homeassistant`).
- Bridge config: broker host/credentials, `{base}` prefix, per-robot `{did}`.
---
## 1. Topic layout
Prefix per robot so the bridge can serve several: `{base}` defaults to
`ecovacs`, `{did}` = robot serial.
| Topic (relative to `{base}/{did}/`) | Dir | Payload | Purpose |
|---|---|---|---|
| `command` | HA→bridge | string | `start`, `return_to_base`, `stop`, `clean_spot`, `locate`. `pause` is not advertised until `act="p"` is verified |
| `set_fan_speed` | HA→bridge | string | `standard`, `strong` |
| `send_command` | HA→bridge | string or JSON | bridge extension commands (§5) |
| `state` | bridge→HA | JSON | `{"state": …, "fan_speed": …}` |
| `json_attributes` | bridge→HA | JSON | battery, consumables, schedules, last error |
| `availability` | bridge→HA | `online`/`offline` | XMPP session liveness (retained) |
| `raw` | bridge→diagnostics | JSON containing direction, timestamp, raw XML and parse status | unknown/unparsed protocol diagnostics |
| `command_result` | bridge→diagnostics | JSON | `sid`/`cid`, ack/result phase, `ret` and `errno` |
Bridge-internal command surface (not HA-facing): any MQTT client can inject
stanzas via the `send_command` `raw` extension (§5).
HA discovers the entity from a **retained** message on
`homeassistant/vacuum/{unique_id}/config` — §7.
## 2. Commands — MQTT `command` topic → `com:ctl`
Each row shows the `<ctl>` element to wrap in
`<iq type="set" to="{bot-jid}" from="{bridge-jid}" id="{sid}">
<query xmlns="com:ctl">…</query></iq>` (§6 of PCAP-ANALYSIS).
| MQTT payload | `<ctl>` to send | Effect / bot response |
|---|---|---|
| `start` | `<ctl td="Clean" id="{cid}"><clean type="auto" speed="{fan}" act="s"/></ctl>` | auto clean at current fan speed → `ret='ok'` + `CleanReport auto` |
| `pause` | `<ctl td="Clean" id="{cid}"><clean type="{current}" speed="{fan}" act="p"/></ctl>` | ⚠ `act="p"` is sucks-documented, **not seen in captures** — verify on this firmware before enabling `pause` in `supported_features`; fallback = don't advertise it |
| `stop` | `<ctl td="Clean" id="{cid}"><clean type="stop" speed="{fan}" act="h"/></ctl>` | halt → `ret='ok'` + `CleanReport stop` |
| `return_to_base` | `<ctl td="Charge" id="{cid}"><charge type="go"/></ctl>` | `ret='ok'` (no errno) + `CleanReport stop` + `ChargeState going` within ~50 ms; `SlotCharging` when it arrives (21 s later in one capture) |
| `clean_spot` | `<ctl td="Clean" id="{cid}"><clean type="spot" speed="{fan}" act="s"/></ctl>` | spot clean |
| `locate` | `<ctl td="PlaySound" sid="0" id="{cid}"/>` | find-me beep → `ret='ok'` |
| `standard` / `strong` on `set_fan_speed` | `<ctl td="SetCleanSpeed" id="{cid}" speed="{v}"/>` | `ret='ok' errno=''` (speed is not echoed). During an active clean a `CleanReport` follows at the new speed. While stopped, apply the speed on `ret='ok'` even if no report arrives |
`{fan}` = last known `speed` (`standard` default). `{cid}` = fresh zero-padded
8-digit id, unique among commands that have not yet returned. The stanza ack
matches `iq/@id`; the ctl result echoes `ctl/@id`. Publish `command_result`
for both phases.
## 3. Reports — `com:ctl` pushes → MQTT state
The bot pushes `<iq type="set"><query xmlns="com:ctl"><ctl td="R" …/></query>`
to the controller JID it learned from `from=` — i.e. the bridge's own virtual
controller JID (`{bridge-uid}@ecouser.net/{resource}`). Push stanzas carry
`to=` but no `from=` and are **never acked** by the controller (verified in
capture) — the bridge must not wait for or send `iq result` on them.
| `td` | Wire payload | MQTT output |
|---|---|---|
| `CleanReport` | `<clean type='T' speed='S' st=' ' rsn=' '/>` | update `fan_speed` and `clean_type`. `auto`/`border`/`spot`/`singleRoom` → cleaning; `stop` → docked only if charge state is still `SlotCharging`, otherwise idle. `st`/`rsn` were a single space in every report |
| `ChargeState` | `<charge type='SlotCharging'/>` | `state` ← docked, `charge_state` ← `SlotCharging` |
| `ChargeState` | `<charge type='going'/>` | `state` ← returning, `charge_state` ← `going` |
| `ChargeState` | `<charge type='Idle'/>` | set `charge_state` to `Idle` and re-derive. This leaves `docked`. It does not by itself mean cleaning |
| `BatteryInfo` | `<battery power='076'/>` | `battery_level` integer 76. The bare `<query><battery power='NNN'/></query>` iq (seen twice, own bot iq id, ~70–90 ms after a GetBatteryInfo result) maps the same way and is not acked |
| `error` | `<ctl td='error' errno='103'/>` | `state` stays `error` and `last_error` stays `"103"` until errno `100`. Later CleanReport/ChargeState still update `clean_type`, `charge_state`, and fan speed |
| `error` | `<ctl td='error' errno='100'/>` | clear `last_error`. The next CleanReport/ChargeState (24–56 ms later in the captures) sets the motion state |
| `Sched2` | `<s …/>` children, or none | replace the whole `schedules` array (§4) |
| *(response)* | `ret='ok'`, `errno` only on some commands | one result per cid. `Get*` and `SetCleanSpeed` include `errno=''`; `Clean`/`Charge`/`PlaySound`/`SetTime`/schedule mutations omit `errno`. A missing `errno` with `ret='ok'` is success. `ret='fail'` was not captured. **Result payloads use the same derivation as pushes** |
### 3.1 State derivation rules (HA `state` key)
HA states: `cleaning`, `docked`, `paused`, `idle`, `returning`, `error`.
```
SlotCharging → docked
going → returning
error push (errno≠100) → error (stays error until errno 100, even if a CleanReport follows 24 ms later)
error push errno=100 → clear error (all-clear; the reports that follow set state)
clean type=stop → docked if charge_state is SlotCharging, else idle
clean type=auto|border|spot|singleRoom → cleaning
act="p" accepted → paused (not captured; do not advertise)
no state yet → idle
```
Priority when several facts are current: `error` > `docked` > `returning` >
`paused` > `cleaning` > `idle`. Errno 103 is followed immediately by
`CleanReport stop`; if the report cleared the error, HA would never show it.
Publish the full state JSON (both keys) on every change.
### 3.2 Attribute feed → `json_attributes` (retained)
```json
{
"battery_level": 76,
"side_brush": 68,
"main_brush": 90,
"filter": 70,
"lifespan_total": {"side_brush": 365, "main_brush": 365, "filter": 365},
"clean_type": "auto",
"charge_state": "Idle",
"last_error": "103",
"last_command_error": null,
"schedules": [
{"name": "17901970980846", "on": true, "time": "21:59",
"repeat": "0001000", "flag": "p", "action": {"td": "clean", "type": "auto"}}
]
}
```
Every publish is the **complete** object. A partial object would replace the
retained document. `last_error` is the robot fault beacon (cleared only by
errno 100). `last_command_error` is a ctl `ret='fail'` or a command rejected
while offline, and must not overwrite `last_error`.
- `side_brush` / `main_brush` / `filter` ← `GetLifeSpan` `val` for
`SideBrush` / `Brush` / `DustCaseHeap`. `lifespan_total.*` ← `total` (unit
unknown; all three were 365). Poll on every READY and when
`send_command` `get_lifespan` is used.
- `schedules` ← replace from every `Sched2` or `GetSched` result (n→name,
o→on, t→time, r→repeat with index 0 = Sunday, f→flag, inner ctl→action).
Spaces inside `<s>` are text nodes, not fields. Only `type='auto'` was
stored in a schedule.
- HA's MQTT vacuum state payload has `state` and optional `fan_speed` only.
`battery` is not a legal `supported_features` value. Battery and consumables
show up as attributes. Optional retained sensor discovery, same device
identifiers, can publish `{base}/{did}/battery` (and the three consumable
percents) with `device_class` `battery` for the battery sensor.
## 4. Schedule CRUD over MQTT (bridge extension)
`send_command` JSON payloads (`{"command": "…", "k": "v"}` — HA flattens params):
```json
{"command":"add_sched","name":"{id}","on":"1","time":"21:30","repeat":"0111000","clean_type":"auto"}
→ <ctl td="AddSched" id="{cid}"><sched name="{name}" on="{on}" time="{time}"
repeat="{repeat}"><ctl td="Clean"><clean type="{clean_type}"/></ctl></sched></ctl>
{"command":"mod_sched","name":"{id}","on":"0","time":"21:30","repeat":"0111000","clean_type":"auto"}
→ <ctl td="ModSched" id="{cid}"><ModSched name="{name}"><sched …/></ModSched></ctl>
{"command":"del_sched","name":"{id}"}
→ <ctl td="DelSched" id="{cid}"><DelSched name="{name}"/></ctl>
{"command":"get_sched"}
→ <ctl td="GetSched" id="{cid}"/> (answer lands in json_attributes.schedules)
```
`repeat` = 7-char bitmask, index 0 = Sunday … index 6 = Saturday (verified).
## 5. Other extension commands (`send_command`)
```json
{"command":"move","action":"forward|SpinLeft|SpinRight|TurnAround|stop"}
→ <ctl td="Move"><move action="{action}"/></ctl> (no ctl id — stanza ack only)
`backward` is sucks vocabulary and was not captured. A new Move may follow
another Move with no stop in between. Do not auto-stop; captured bursts were
under about 3 s, so a long-run firmware timeout is untested.
{"command":"clean","clean_type":"auto|border|spot|singleRoom"}
→ <ctl td="Clean" id="{cid}"><clean type="{clean_type}" speed="{fan}" act="s"/></ctl>
{"command":"cancel_return"} → <ctl td="Charge" id="{cid}"><charge type="stopGo"/></ctl>
{"command":"resume"} → <ctl td="Clean" id="{cid}"><clean type="{current}" speed="{fan}" act="r"/></ctl> ⚠ unverified
{"command":"set_time"} → <ctl td="SetTime" id="{cid}"><time t="{epoch}" tz="{h}" tzm="{m}"/></ctl>
{"command":"get_status"} → fan out GetBatteryInfo+GetCleanState+GetChargeState+GetCleanSpeed+GetSched
{"command":"get_lifespan"} → GetLifeSpan ×3 (SideBrush, Brush, DustCaseHeap)
{"command":"raw","xml":"<ctl …/>"} → passthrough, off unless explicitly enabled
```
## 6. Session behavior — register before you can monitor
**Critical:** the bot only pushes reports to the controller JID it has seen in
an inbound `from=` — and that learning is **per-session** (PCAP-ANALYSIS §6.1).
A bridge that only listens will receive *nothing* — not even `Sched2`. So on
**every** robot session reaching READY (`hello world` presence — including
each re-bind after a reconnect):
1. **Announce:** `<iq type="get" to="{bot-jid}" from="{bridge-jid}"><ping
xmlns="urn:xmpp:ping"/></iq>`. The app sent this first (sometimes twice,
a few hundred milliseconds apart). The bot answers `result`. No push was
seen in the second before the next step.
2. Send `SetTime` (current epoch + tz). In both boots the first push was an
empty or populated `Sched2` about 100 ms after this result.
3. Fan out `GetBatteryInfo`, `GetCleanState`, `GetChargeState`,
`GetCleanSpeed`, `GetSched`, and `GetLifeSpan` for `SideBrush`, `Brush`,
and `DustCaseHeap`. Give each a distinct ctl id.
4. Publish `availability = online` (retained) after the announce result.
The status fan-out may still be in flight.
Thereafter: ping the bot `from="{bridge-jid}"` and correlate its iq result.
The real app used ~90 s; **60 s is recommended for the bridge**, with a
10–15 s response deadline. This both keeps the learned JID warm and detects a
black-holed connection independently of the bot's 120 s ping/retry cycle.
On XMPP disconnect/timeout — TCP drop, `</stream:stream>`, or one bridge ping
missing its response deadline — publish `availability = offline` (retained).
The fourth capture proves the bot itself takes ~120 s after its unanswered
ping, then waits ~5 s before reconnecting; don't leave HA falsely online for
that interval. Optional TCP keepalive on the XMPP listener is an additional
signal, not a substitute for iq-result correlation. Also set the bridge's
MQTT **LWT to `…/availability = offline` (retained)** so a dead bridge process
marks the entity unavailable, and publish `offline` on graceful shutdown.
Reconnects come straight to the cached endpoint — no bootstrap. The bridge
sees a fresh TCP connect + full SASL/bind/session handshake and re-runs the
READY sequence above. On a new bind, **atomically replace the old JID session**
even if its half-open TCP socket cannot be closed over the network; close the
old local socket and ignore any later callbacks/data from that obsolete
session generation. Keep the listener IP stable: the robot cannot rediscover
a moved endpoint without rebooting.
MQTT commands arriving while the bot session is down can't be delivered —
drop them and record that on `last_command_error`, not on `last_error`.
Do not queue them.
## 7. HA discovery config (publish retained)
Topic: `homeassistant/vacuum/ecovacs_E2998877665544332211/config`
The component is already named by the topic. Do not put `platform` in this
payload; that key is for device-discovery (`homeassistant/device/...`) only.
Use the serial in the object id so a second robot does not collide.
Resend the retained payload when the bridge's MQTT session reconnects and
when Home Assistant publishes `online` to `homeassistant/status`.
```json
{
"name": "Deebot N95",
"unique_id": "ecovacs_E2998877665544332211",
"command_topic": "ecovacs/E2998877665544332211/command",
"set_fan_speed_topic": "ecovacs/E2998877665544332211/set_fan_speed",
"send_command_topic": "ecovacs/E2998877665544332211/send_command",
"state_topic": "ecovacs/E2998877665544332211/state",
"json_attributes_topic": "ecovacs/E2998877665544332211/json_attributes",
"availability_topic": "ecovacs/E2998877665544332211/availability",
"payload_available": "online",
"payload_not_available": "offline",
"fan_speed_list": ["standard", "strong"],
"supported_features": ["start","stop","return_home","status","locate",
"clean_spot","fan_speed","send_command"],
"device": {
"identifiers": ["E2998877665544332211"],
"manufacturer": "Ecovacs",
"model": "Deebot N95 (wukong/155)",
"serial_number": "E2998877665544332211",
"connections": [["mac", "02:00:00:00:95:01"]]
}
}
```
(`pause` stays omitted until `act="p"` is verified. `clean_segments` stays
omitted — non-mapping robot. Do not list `battery`; current Home Assistant
rejects it on this integration. State JSON is `state` plus optional
`fan_speed` only.)
## 8. QoS / retain policy
| Topic | QoS | Retain | Why |
|---|---|---|---|
| `state`, `json_attributes`, `availability` (incl. LWT) | 0 | **yes** | HA restarts must see last state |
| discovery `…/config` | 0 | **yes** | required for discovery |
| `command`, `set_fan_speed`, `send_command` | 0–1 | no | commands are momentary |
| `raw`, `command_result` | 0 | no | diagnostic/event streams |
## 9. Verified vs assumed
| Mapping | Basis |
|---|---|
| start/stop/spot/dock/cancel-dock/locate/fan-speed/SetTime/Get*/schedules | **captured** verbatim |
| `paused` via `act="p"`, resume via `act="r"` | sucks vocabulary only — test before enabling |
| error→`error` state | captured (errno 103/100 pushes) |
| `singleRoom` exposed via `send_command.clean` | captured; HA has no native single-room concept |
| `backward` move | sucks vocabulary only |
| errno attribute omitted on Clean/Charge/PlaySound/SetTime/schedules | captured; `Get*` and `SetCleanSpeed` include `errno=''` |
| bare `<battery>` iq, twice | captured; same mapping as `BatteryInfo`, no ack |
| `going` then `SlotCharging` 21 s later; `Idle` on `stopGo` and on leaving the dock | captured |
| schedule inner type other than `auto`, ModSched rename, `act=p`/`act=r` | not captured |
Every HA command in the advertised feature list maps to a stanza in the four
captures, and every observed bot push has an MQTT output above. `pause`,
resume, and `backward` stay out of that list until they are captured on this
firmware.
+977
View File
@@ -0,0 +1,977 @@
# Deebot N95 Local XMPP-to-Home Assistant MQTT Bridge — Full Specification
> **Privacy notice:** Device serials, MAC addresses, authentication values, controller identifiers, and other identifying values shown here have been consistently replaced with synthetic values.
## 1. Purpose and conformance
This is the complete, implementation-agnostic specification for a local
solution that:
1. boots and accepts an already-provisioned Ecovacs Deebot N95;
2. replaces the Ecovacs bootstrap, firmware-check and legacy XMPP endpoints;
3. monitors and controls the robot through its legacy `com:ctl` XMPP dialect;
4. exposes a Home Assistant entity using the native **MQTT vacuum** schema;
5. exposes capabilities outside that schema through documented auxiliary MQTT
commands and attributes; and
6. reports unparsed protocol messages without losing the connection.
An implementation conforms when a provisioned N95 can be redirected by the
single DNS record in §3, complete the exchanges in §§4–6, reconnect according
to §7, and use every mapping in §§9–13.
This specification consolidates findings from four packet captures. Values
marked **observed** appeared on the wire. Values marked **library-known** come
from the `sucks` client but were not exercised by this N95 capture set.
No captured credential is required by a replacement server: accept whatever
SASL PLAIN authcid/password the provisioned robot presents. Never log the
password or decoded SASL payload.
---
## 2. System boundaries and deployment prerequisites
```text
N95 --DNS--> local resolver
N95 --HTTP:8007--> bridge /lookup.do
N95 --HTTP:8005--> bridge firmware 404
N95 --TCP:5223 plaintext XMPP--> bridge
bridge --MQTT--> broker <--MQTT--> Home Assistant
```
Required:
- the N95 receives a DNS resolver through DHCP option 6;
- that resolver returns the bridge listener address for `lbo.ecouser.net`;
- the address remains stable for the robot's powered-on lifetime;
- TCP 8007, 8005 and 5223 are reachable at that address;
- the bridge can connect to an MQTT broker used by Home Assistant;
- Home Assistant's MQTT integration and discovery are enabled (default
discovery prefix `homeassistant`).
Only `lbo.ecouser.net` was queried in all boot captures. A compatibility setup
may also override `lbo.ecovacs.net`, but it is not required by this captured
N95 firmware. The XMPP domain `155.ecorobot.net` is a virtual/JID domain and
was never DNS-resolved by the robot.
Observed identity (use configuration/discovery rather than hard-coding where
possible):
| Field | Value |
|---|---|
| Product/platform | `wukong` |
| Device class | `155` |
| Serial / DID / XMPP authcid | `E2998877665544332211` |
| Resource | `atom` |
| Full bot JID | `E2998877665544332211@155.ecorobot.net/atom` |
| DHCP hostname | `deebot` |
| MAC in captures | `02:00:00:00:95:01` |
---
## 3. Network bootstrap
### 3.1 DHCP and DNS
The robot uses the first DNS server in DHCP option 6. At boot it sends two
near-identical A queries for:
```text
lbo.ecouser.net
```
The public response used a CNAME then A record, but a direct local A response
is sufficient. After obtaining an address the robot opens two parallel TCP
connections to port 8007.
Other boot traffic needing no bridge response:
- ARP probes/announcement for its DHCP address;
- IGMPv2 report to `226.1.1.1` (and after one reconnect, `224.0.0.1`).
### 3.2 Service discovery — HTTP port 8007
The robot sends one HTTP/1.0 `POST /lookup.do` per service, with no `Host`
header, and closes the connection. Accept that request. The two connections
are parallel; answer each on its own socket. Parse JSON semantically; the
observed requests are:
```json
{"todo":"FindBest","service":"EcoMsgNew"}
{"todo":"FindBest","service":"EcoUpdate"}
```
Respond `200 OK`, `Content-Type: application/json; charset=utf-8`, with compact
JSON, no spaces, and `port` as a JSON number:
```json
{"result":"ok","ip":"<bridge-ip>","port":5223}
{"result":"ok","ip":"<bridge-ip>","port":8005}
```
The response order follows the request, not a fixed connection order.
`EcoMsgNew` is the XMPP endpoint and `EcoUpdate` the firmware endpoint.
### 3.3 Firmware check — HTTP port 8005
The robot requests:
```http
GET /products/wukong/class/155/firmware/latest.json HTTP/1.0
Connection: Close
Accept: Application/json
Content-Type: application/json
```
Return HTTP 404 and close. The captured success-path response was
`Content-Type: text/plain; charset=utf-8` and body `Not Found` (9 bytes, no
newline). The robot opened XMPP about 180 ms later and did not retry. A JSON
body and a successful manifest were not observed. Do not serve one.
### 3.4 Endpoint caching
DNS, `/lookup.do`, and firmware lookup occur at power-on. After wifi loss or a
half-open XMPP failure, the robot reconnects directly to the cached
`EcoMsgNew` IP and port. It does not repeat DNS or HTTP discovery. Therefore:
- keep the bridge listener address stable;
- an address change requires a robot reboot to rediscover it;
- the HTTP services need not participate in ordinary XMPP reconnects.
---
## 4. XMPP transport and framing
Listen on TCP port 5223. Traffic is plaintext XML despite the conventional
port and despite the captured server advertising required STARTTLS. The robot
ignores STARTTLS and sends SASL PLAIN immediately. TLS is not required for
this model; advertising STARTTLS is optional fidelity.
XMPP is a continuous XML stream, not a complete XML document:
- the opening `<stream:stream>` remains unclosed while the session lives;
- a stanza can span TCP packets;
- several stanzas can share one TCP packet;
- TCP packet boundaries must never be treated as stanza boundaries;
- handle XML declaration, stream open, complete top-level stanzas and
`</stream:stream>` incrementally;
- tolerate namespace prefixes and attribute ordering differences;
- reject malformed XML safely, but log and ignore unknown valid stanzas rather
than terminating the session.
### 4.1 Complete handshake
```xml
C→S <?xml version='1.0'?><stream:stream
xmlns:stream='http://etherx.jabber.org/streams'
xmlns='jabber:client' to='155.ecorobot.net' version='1.0'>
S→C <stream:stream xmlns:stream="http://etherx.jabber.org/streams"
xmlns="jabber:client" version="1.0" id="{opaque-stream-id}"
from="155.ecorobot.net">
S→C <stream:features>
<auth xmlns="http://jabber.org/features/iq-auth"/>
<starttls xmlns="urn:ietf:params:xml:ns:xmpp-tls"><required/></starttls>
<mechanisms xmlns="urn:ietf:params:xml:ns:xmpp-sasl">
<mechanism>PLAIN</mechanism>
</mechanisms>
</stream:features>
C→S <auth xmlns='urn:ietf:params:xml:ns:xmpp-sasl'
mechanism='PLAIN'>{base64(NUL + serial + NUL + password)}</auth>
S→C <success xmlns="urn:ietf:params:xml:ns:xmpp-sasl"/>
C→S <?xml version='1.0'?><stream:stream ...
to='155.ecorobot.net' version='1.0'>
S→C <stream:stream ... id="{same-opaque-id}" from="155.ecorobot.net">
S→C <stream:features>
<bind xmlns="urn:ietf:params:xml:ns:xmpp-bind"/>
<session xmlns="urn:ietf:params:xml:ns:xmpp-session"/>
</stream:features>
C→S <iq type='set' id='{bot-id}'><bind
xmlns='urn:ietf:params:xml:ns:xmpp-bind'><resource>atom</resource></bind></iq>
S→C <iq type="result" id="{bot-id}"><bind
xmlns="urn:ietf:params:xml:ns:xmpp-bind"><jid>{serial}@155.ecorobot.net/atom</jid>
</bind></iq>
C→S <iq type='set' id='{next-bot-id}'><session
xmlns='urn:ietf:params:xml:ns:xmpp-session'/></iq>
S→C <iq type="result" id="{next-bot-id}"/>
C→S <presence><status>hello world</status></presence>
S→C <presence to="{full-bot-jid}"> dummy </presence>
```
Server rules:
- derive class/domain from initial stream `to=`;
- decode SASL enough to obtain authcid if desired, but accept all passwords;
- never log the raw or decoded SASL credential;
- use the requested resource (`atom`) in the returned JID;
- READY is reached after session result and `hello world` presence;
- mimic the literal dummy presence content, including the surrounding spaces;
- repeat the same stream id on the post-SASL stream open; a new TCP connection
gets a new id. The robot does not validate either;
- `iq-auth` is advertised but not used;
- robot iq ids are monotonic for a powered-on robot and continue across XMPP
reconnects; do not assume reset or small values;
- server-generated ids are opaque and only need to be unique among outstanding
requests.
### 4.2 Session replacement
There is one active connection per full bot JID. The captured server sent
`</stream:stream>` and FIN about 20 ms after the new session IQ and before
presence; the robot double-RSTs if that FIN arrives. Doing the replacement at
bind is early relative to that server and is still correct, because the new
connection is authoritative before any command. On a successful new bind:
1. atomically replace the JID→connection mapping;
2. make the new connection authoritative;
3. send `</stream:stream>` and close the old local socket if possible;
4. do not reject the new bind because an old socket still appears connected;
5. ignore later reads, closes or callbacks from the obsolete session generation.
The old path can be black-holed, so sending its stream close may never reach
the robot. Correctness must not depend on delivery.
---
## 5. Controller identity, report registration and pings
The bridge acts as a virtual XMPP controller. Use a stable syntactically valid
JID such as:
```text
n95bridge@ecouser.net/homeassistant
```
The exact localpart/resource are arbitrary. The bot learns where to send
responses and unsolicited reports from the `from=` of incoming controller
stanzas. This registration is per XMPP session.
Immediately after every READY, send:
```xml
<iq id="{sid}" to="{bot-jid}" from="{controller-jid}" type="get">
<ping xmlns="urn:xmpp:ping"/>
</iq>
```
The bot replies:
```xml
<iq type='result' from='{bot-jid}' to='{controller-jid}' id='{sid}'/>
```
A session that never sees a controller `from=` produces no `Sched2`,
`BatteryInfo`, `CleanReport`, `ChargeState`, or `error` (both reconnects in
capture 3). In the two boots the first push was `Sched2`, about 100 ms after
the `SetTime` result, not in the second between the announce ping and
`SetTime`. Send the ping and then `SetTime`. Re-announce after every re-bind.
The app's first ping was 12.9 s or 38.8 s after ready; that delay is the user
opening the app, not a robot timer. The app sometimes sent a second ping a
few hundred milliseconds later. One ping is enough.
### 5.1 Keepalive directions
- **Bot→server**, every ~120 s:
```xml
<iq from='{bot-jid}' to='155.ecorobot.net' id='{bot-id}' type='get'>
<ping xmlns='urn:xmpp:ping'/>
</iq>
```
Answer immediately:
```xml
<iq type="result" to="{bot-jid}" from="155.ecorobot.net" id="{bot-id}"/>
```
- **Bridge/controller→bot:** the observed app used ~90 s. The bridge should use
~60 s and require a matching result within 10–15 s. This provides faster HA
availability detection while retaining protocol semantics.
---
## 6. `com:ctl` envelope and correlation
### 6.1 Command
```xml
<iq id="{sid}" to="{bot-jid}" from="{controller-jid}" type="set">
<query xmlns="com:ctl">
<ctl td="{Command}" id="{cid}">...</ctl>
</query>
</iq>
```
- `sid`: XMPP stanza id chosen by bridge.
- `cid`: command correlation id; the app uses random zero-padded 8-digit
strings. Generate a unique value among outstanding commands.
- `Move` is the exception: its `<ctl>` has no `id`.
### 6.2 Two-part response
Most commands produce both:
1. stanza receipt ack, correlated by `iq/@id == sid`:
```xml
<iq to='{controller-jid}' type='result' id='{sid}'/>
```
2. command result in a new iq-set, correlated by `ctl/@id == cid`:
```xml
<iq to='{controller-jid}' type='set' id='{bot-id}'>
<query xmlns='com:ctl'>
<ctl id='{cid}' ret='ok'>...payload, and errno='' on queries...</ctl>
</query>
</iq>
```
`ret` was only `ok` in these captures. `errno=''` is present on `Get*` and
`SetCleanSpeed` results and omitted on `SetTime`, `Clean`, `Charge`,
`PlaySound`, `AddSched`, `ModSched`, and `DelSched`. A missing `errno` with
`ret='ok'` is success. `Move` produced only the stanza ack. The app reused
iq ids across later commands and reused a ctl id on an immediate status retry;
each request still had one result. Complete a cid once, and do not treat a
later request that reuses a finished id as a duplicate.
Do not confuse a command response's `ret/errno` with an unsolicited
`td='error'` event.
### 6.3 Unsolicited push
```xml
<iq to='{controller-jid}' type='set' id='{bot-id}'>
<query xmlns='com:ctl'><ctl td='{Report}'>...</ctl></query>
</iq>
```
Observed pushes have `to=` but no `from=`. Contrary to normal IQ-set practice,
the captured app sent no `iq result` ack for pushes; do not wait for or send
one.
Parser quirks:
- a bare `<battery power='NNN'/>` can appear directly under `<query>` without
`<ctl>`, as its own `<iq type='set'>` with a bot sequence id. It was seen
twice, 70–90 ms after a normal `GetBatteryInfo` result and carrying the same
power. Treat it as BatteryInfo and do not ack it;
- ignore whitespace text nodes;
- unknown children/attributes must not fail the session;
- publish/log an unparsed diagnostic for unknown valid messages (§13).
---
## 7. Half-open failure and automatic recovery
Observed complete behavior when router state was cleared:
1. last healthy bot ping was answered;
2. 120 s later, next bot ping was transmitted but black-holed;
3. the identical TCP segment (same sequence and XMPP id) was retransmitted.
With a short RTT the delays were +0.668, +2.342, +5.368, +11.426, +23.468,
+47.666 and +95.863 s. An earlier black-holed ping used a longer RTO
(+0.92, +3.00, +7.02, +15.04, +31.12, +63.29 s, capture ended during that
series). The backoff is adaptive. Do not hard-code it;
4. at +120.002 s the robot sent TCP FIN, without XMPP stream close;
5. it did not wait for FIN-ACK;
6. at FIN+4.999 s it opened a new TCP connection to cached IP:5223;
7. full stream/SASL/re-stream/bind/session/presence completed in 0.455 s
(0.42–0.48 s on the other reconnects);
8. there was no DNS, HTTP lookup, DHCP renewal, special resume token or
XEP-0198 stream resumption.
Required bridge behavior:
- mark MQTT availability offline on TCP close, XML stream close, or a bridge
ping missing its 10–15 s response deadline;
- accept the new TCP connection and full authentication unconditionally;
- replace stale session state atomically (§4.2);
- after READY rerun the complete initialization in §8;
- do not carry outstanding command correlations into the new session; fail
them locally as connection-lost;
- retained state may remain for display, but availability remains offline
until reinitialization succeeds.
The server cannot make firmware reconnect sooner once the network is
black-holed. Its own ping deadline can make Home Assistant show offline sooner.
---
## 8. READY initialization and availability
On every READY, in order:
1. controller announce ping (§5), wait for result;
2. send `SetTime` with current Unix seconds and local UTC offset;
3. issue `GetBatteryInfo`, `GetCleanState`, `GetChargeState`, `GetCleanSpeed`,
`GetSched`, and three `GetLifeSpan` requests;
4. process their result payloads through the same MQTT state/attribute mapping
used for pushes;
5. publish retained availability `online` after announce succeeds (the status
fan-out may finish asynchronously).
MQTT commands received while unavailable must not be silently queued: reject
or drop them and expose a diagnostic/last-error value. Commands are actions;
executing stale retained/queued commands after reconnect is unsafe.
Set MQTT Last Will and Testament to retained `offline` on the robot's
availability topic. Also publish retained `offline` on graceful bridge
shutdown.
---
## 9. MQTT namespace and delivery rules
Let:
```text
root = ecovacs/{did}
did = robot serial
```
For the observed robot, `root` is
`ecovacs/E2998877665544332211`.
| Topic | Direction | Payload | Retain |
|---|---|---|---|
| `{root}/command` | HA→bridge | HA vacuum command string | no |
| `{root}/set_fan_speed` | HA→bridge | `standard` or `strong` | no |
| `{root}/send_command` | HA/client→bridge | extension string/JSON | no |
| `{root}/state` | bridge→HA | HA vacuum state JSON | yes |
| `{root}/json_attributes` | bridge→HA | complete attributes JSON | yes |
| `{root}/availability` | bridge→HA | `online` / `offline` | yes |
| `{root}/raw` | bridge→diagnostics | normalized JSON containing raw XML and parse status | no |
| `{root}/command_result` | bridge→diagnostics | command correlation/result JSON | no |
QoS 0 is sufficient and matches native HA defaults. QoS 1 is permissible for
inbound command topics, but the bridge must deduplicate redelivery and never
retain command messages.
Always republish the complete `state` object and complete attributes object,
not partial JSON, so retained data remains coherent.
Recommended diagnostic envelopes:
```json
{"direction":"robot_to_bridge","kind":"unparsed","timestamp":"RFC3339",
"session":12,"xml":"<iq ...>...</iq>","reason":"unknown td"}
```
```json
{"sid":"1234","cid":"01234567","command":"Clean","phase":"result",
"ret":"ok","errno":"","timestamp":"RFC3339"}
```
Raw XML can contain identifiers; access-control the topic and never include
SASL auth data.
---
## 10. Home Assistant native vacuum contract
### 10.1 Discovery
Publish retained to the single-component discovery topic. The component is
the topic, so the payload has no `platform` key (`platform` belongs to
`homeassistant/device/...` payloads). Put the serial in the object id.
```text
homeassistant/vacuum/ecovacs_E2998877665544332211/config
```
Republish when the bridge reconnects to MQTT and when Home Assistant's birth
message `online` arrives on `homeassistant/status` (default).
```json
{
"name": "Deebot N95",
"unique_id": "ecovacs_E2998877665544332211",
"command_topic": "ecovacs/E2998877665544332211/command",
"set_fan_speed_topic": "ecovacs/E2998877665544332211/set_fan_speed",
"send_command_topic": "ecovacs/E2998877665544332211/send_command",
"state_topic": "ecovacs/E2998877665544332211/state",
"json_attributes_topic": "ecovacs/E2998877665544332211/json_attributes",
"availability_topic": "ecovacs/E2998877665544332211/availability",
"payload_available": "online",
"payload_not_available": "offline",
"fan_speed_list": ["standard", "strong"],
"supported_features": [
"start", "stop", "return_home", "status", "locate",
"clean_spot", "fan_speed", "send_command"
],
"device": {
"identifiers": ["E2998877665544332211"],
"manufacturer": "Ecovacs",
"model": "Deebot N95 (wukong/155)",
"serial_number": "E2998877665544332211",
"connections": [["mac", "02:00:00:00:95:01"]]
}
}
```
Do not advertise `pause` until `act='p'` is verified on this firmware. Do not
advertise segment cleaning; the N95 is non-mapping and no segments were
observed. Do not list `battery` in `supported_features`; current Home
Assistant rejects that value on MQTT vacuum. The state payload is only
`state` and `fan_speed`. Battery and consumables are attributes.
Optional, not required for conformance: also discover retained sensors on the
same device for `battery_level` (`device_class` `battery`, unit `%`) and the
three consumable percents, fed from `{root}/battery` and
`{root}/consumables/{side_brush,main_brush,filter}`. Those topics are in
addition to the canonical attribute object.
### 10.2 State payload
```json
{"state":"docked","fan_speed":"strong"}
```
Allowed HA states: `cleaning`, `docked`, `paused`, `idle`, `returning`,
`error`.
State derivation and precedence:
| Robot observation | HA state/action |
|---|---|
| `td=error`, `errno != 100` | `error`; store code |
| `td=error`, `errno=100` | clear active error; wait for/follow subsequent CleanReport/ChargeState |
| charge `SlotCharging` | `docked` |
| charge `going` | `returning` |
| charge `Idle` | set `charge_state` to `Idle` and re-derive. This ends `docked`. It does not start a clean |
| clean type `auto`, `border`, `spot`, `singleRoom` | `cleaning` |
| clean type `stop` | `docked` if last charge state is SlotCharging, otherwise `idle` |
| verified accepted `act=p` | `paused` |
| no state yet after initialization | `idle` |
When simultaneous information conflicts, precedence is:
```text
active error > docked > returning > paused > cleaning > idle
```
`errno=100` is an observed all-clear event, not a fault: in one capture it was
followed by resumed auto cleaning; in another by stop + SlotCharging.
### 10.3 Attributes
Example full retained object:
```json
{
"battery_level": 76,
"side_brush": 68,
"main_brush": 90,
"filter": 70,
"lifespan_total": {"side_brush":365,"main_brush":365,"filter":365},
"clean_type": "auto",
"charge_state": "Idle",
"last_error": null,
"last_command_error": null,
"schedules": [
{
"name":"17901970980846",
"on":true,
"time":"21:59",
"repeat":"0001000",
"flag":"p",
"action":{"td":"clean","type":"auto"}
}
]
}
```
Home Assistant's native vacuum state schema represents state and fan speed;
battery, consumables, schedules and diagnostics are attributes. A deployment
may additionally publish MQTT sensor discovery entities sourced from these
attributes, but they are optional and must not change the canonical topics or
units here: battery/consumables are integer percent; `total` is preserved as
reported because its unit was not established.
---
## 11. HA commands mapped to XMPP
Wrap every shown `<ctl>` in the command envelope from §6.1.
| MQTT input | XMPP `<ctl>` | Verification |
|---|---|---|
| `start` on `{root}/command` | `<ctl td="Clean" id="{cid}"><clean type="auto" speed="{fan}" act="s"/></ctl>` | observed |
| `stop` | `<ctl td="Clean" id="{cid}"><clean type="stop" speed="{fan}" act="h"/></ctl>` | observed |
| `return_to_base` | `<ctl td="Charge" id="{cid}"><charge type="go"/></ctl>` | observed |
| `clean_spot` | `<ctl td="Clean" id="{cid}"><clean type="spot" speed="{fan}" act="s"/></ctl>` | observed |
| `locate` | `<ctl td="PlaySound" sid="0" id="{cid}"/>` | observed |
| `pause` | `<ctl td="Clean" id="{cid}"><clean type="{current}" speed="{fan}" act="p"/></ctl>` | library-known, not captured; not advertised |
| `standard` / `strong` on set-fan topic | `<ctl td="SetCleanSpeed" id="{cid}" speed="{payload}"/>` | observed |
`{fan}` is last known `standard|strong`, defaulting to `standard` until queried.
`{current}` is the last active clean type.
Expected consequences:
- Clean commands: stanza ack, ctl result, then CleanReport.
- Charge `go`: stanza ack, ctl result, CleanReport stop, ChargeState going;
later SlotCharging when docked.
- SetCleanSpeed: stanza ack and ctl result `ret='ok' errno=''` with no speed
echo. Store `{fan}` on that result. During an active clean a CleanReport at
the new speed followed; while stopped it did not always follow.
- PlaySound: stanza ack + ctl result.
---
## 12. Extension commands and exact protocol vocabulary
HA's `vacuum.send_command` publishes either a string or flattened JSON to
`{root}/send_command`. Use JSON below.
### 12.1 Clean modes
```json
{"command":"clean","clean_type":"auto|border|spot|singleRoom"}
```
```xml
<ctl td="Clean" id="{cid}">
<clean type="{clean_type}" speed="{fan}" act="s"/>
</ctl>
```
Observed clean types, both as commands and as `CleanReport` values: `auto`,
`border`, `spot`, `singleRoom`, `stop`. `singleRoom` case is exact. `SpotArea`
exists in library code but was not observed and is not specified as supported
here. Schedule entries stored only `auto` (§12.7).
Resume (not captured; library-known):
```json
{"command":"resume"}
```
uses `act="r"`; expose only after validation.
### 12.2 Manual movement
```json
{"command":"move","action":"forward|SpinLeft|SpinRight|TurnAround|stop"}
```
```xml
<ctl td="Move"><move action="{action}"/></ctl>
```
Observed actions: `forward`, `SpinLeft`, `SpinRight`, `TurnAround`, `stop`.
`backward` is library-known but unobserved. Move has no ctl id and only an IQ
stanza ack. A new Move was sent without a `stop` first, including `forward`
followed by another `forward`. Do not insert a stop, and do not retain or
replay movement messages. Captured bursts lasted 0.24–3.1 s; a firmware
timeout beyond that was not observed.
### 12.3 Dock cancellation
```json
{"command":"cancel_return"}
```
```xml
<ctl td="Charge" id="{cid}"><charge type="stopGo"/></ctl>
```
Observed command charge values: `go`, `stopGo`.
Observed state values, including pushes: `Idle`, `going`, `SlotCharging`.
`go` produced `going` within about 50 ms and `SlotCharging` on arrival (21 s
later in one capture). `stopGo` produced `Idle`. Leaving the dock also pushed
`Idle`.
### 12.4 Time
```json
{"command":"set_time"}
```
```xml
<ctl td="SetTime" id="{cid}">
<time t="{unix-seconds}" tz="{signed-whole-hours}" tzm="{signed-minutes}"/>
</ctl>
```
Observed UTC+1 example used `tz="1" tzm="0"`. Compute both components from
local UTC offset; don't copy that example globally.
### 12.5 Status refresh
```json
{"command":"get_status"}
```
Fan out:
```xml
<ctl id="{cid}" td="GetBatteryInfo"/>
<ctl id="{cid}" td="GetCleanState"/>
<ctl id="{cid}" td="GetChargeState"/>
<ctl id="{cid}" td="GetCleanSpeed"/>
<ctl id="{cid}" td="GetSched"/>
```
Issue a unique cid for each request.
Expected result payloads:
```xml
<battery power='076'/>
<clean type='stop' speed='standard' st='h' t='' a=''/>
<charge type='Idle'/>
<ctl id='{cid}' ret='ok' errno='' speed='standard'/>
<s ...>...</s>
```
Map these exactly as equivalent pushes: response data is state data, not only
a command acknowledgement.
### 12.6 Consumable lifespan
```json
{"command":"get_lifespan"}
```
Send three requests:
```xml
<ctl id="{cid}" td="GetLifeSpan" type="SideBrush"/>
<ctl id="{cid}" td="GetLifeSpan" type="Brush"/>
<ctl id="{cid}" td="GetLifeSpan" type="DustCaseHeap"/>
```
Response:
```xml
<ctl id='{cid}' ret='ok' errno='' type='SideBrush' val='068' total='365'/>
```
Mapping: `SideBrush→side_brush`, `Brush→main_brush`,
`DustCaseHeap→filter`; parse `val` as integer percent and preserve `total` as
integer with unknown unit.
### 12.7 Schedule CRUD
Add:
```json
{"command":"add_sched","name":"17901970980846","on":"1",
"time":"21:30","repeat":"0111000","clean_type":"auto"}
```
```xml
<ctl td="AddSched" id="{cid}">
<sched name="{name}" on="{on}" time="{HH:MM}" repeat="{mask}">
<ctl td="Clean"><clean type="{clean_type}"/></ctl>
</sched>
</ctl>
```
Modify:
```json
{"command":"mod_sched","name":"17901970980846","on":"0",
"time":"21:30","repeat":"0111000","clean_type":"auto"}
```
```xml
<ctl td="ModSched" id="{cid}">
<ModSched name="{existing-name}">
<sched name="{name}" on="{on}" time="{HH:MM}" repeat="{mask}">
<ctl td="Clean"><clean type="{clean_type}"/></ctl>
</sched>
</ModSched>
</ctl>
```
Delete:
```json
{"command":"del_sched","name":"17901970980846"}
```
```xml
<ctl td="DelSched" id="{cid}"><DelSched name="{name}"/></ctl>
```
Get:
```json
{"command":"get_sched"}
```
```xml
<ctl td="GetSched" id="{cid}"/>
```
Schedule representation received in GetSched result or Sched2 push:
```xml
<s n='17901970980846' o='1' t='21:59' r='0001000' f='p'>
<ctl td='clean' type='auto'/>
</s>
```
Mapping:
| Wire | MQTT schedule field |
|---|---|
| `n` | `name` (opaque; preserve exactly) |
| `o='0'/'1'` | `on=false/true` |
| `t` | `time`, local `HH:MM` |
| `r` | `repeat`, 7 chars; index 0 Sunday through index 6 Saturday |
| `f` | `flag`; observed always `p`, preserve without interpretation |
| inner lowercase `ctl td='clean' type=X` | `action:{"td":"clean","type":X}` |
An empty `GetSched` result is `<ctl id ret='ok' errno=''/>` with no `<s>`
children. An empty `Sched2` push is `<ctl td='Sched2'/>`. Both mean an empty
schedule array. `Sched2` is pushed after the controller is known (in both
boots, about 100 ms after the first `SetTime`, not at XMPP ready), after every
mutation, and when a schedule fires. The `21:59` / `0001000` entry fired at
21:58:59 local: `Sched2`, then `CleanReport auto` 73 ms later. Each `<s>` has a
space before its inner `<ctl>` and before `</s>`; there is no text between
adjacent `<s>` elements. Only `type='auto'` was stored. `ModSched` edits
changed `on`, `time`, and `repeat`; the inner name always matched the wrapper
name, so a rename is unverified. `f` was `p` on every entry.
### 12.8 Raw diagnostic command
```json
{"command":"raw","xml":"<ctl .../>"}
```
This optional expert interface wraps valid ctl XML in §6.1. It must be access
controlled, size limited, XML parsed (not string-concatenated), prohibited
from injecting stream/auth elements, and disabled by default.
---
## 13. Robot reports and complete MQTT mapping
| XMPP message | Observed payload | MQTT analogue/action |
|---|---|---|
| stream open/features/auth/bind/session | §4 | internal session state; no direct MQTT; failures affect availability and diagnostics |
| `hello world` presence | §4 | READY transition; run §8; then availability online |
| dummy presence | server→bot | no MQTT output |
| controller ping/result | §5 | report registration/liveness; no state topic; timeout→offline |
| bot ping/result | §5 | server liveness; answer; no MQTT output |
| command stanza ack | iq result by `sid` | `{root}/command_result`, phase `ack`; internal correlation |
| ctl command result | `ret/errno`, by `cid` | `{root}/command_result`, phase `result`; update `last_command_error`; parse any payload below |
| `GetBatteryInfo` result | `<battery power='NNN'/>` | attributes `battery_level` integer |
| bare query battery quirk | `<query><battery power='NNN'/></query>` | same battery mapping; diagnostic may flag quirk but it is parsed |
| `GetCleanState` result | clean `type/speed/st/t/a` | derive HA state; set `clean_type`, `fan_speed`; preserve unknown nonempty fields diagnostically |
| `GetChargeState` result | charge type | derive HA state; set `charge_state` |
| `GetCleanSpeed` result | ctl `speed` | HA state JSON `fan_speed` |
| `GetLifeSpan` result | type/val/total | consumable attributes |
| `GetSched` result | zero or more `<s>` | complete schedules attribute |
| `CleanReport` push | clean `type/speed/st/rsn` | derive HA state, fan_speed, clean_type; preserve nonblank unknown st/rsn diagnostically |
| `ChargeState` push | `Idle/going/SlotCharging` | derive HA state + charge_state |
| `BatteryInfo` push | battery power | battery attribute |
| `Sched2` push | schedule list | replace complete schedules attribute |
| `error` push errno 103 | cliff/stair halt observed | HA `error`, `last_error="103"`; following reports may change motion state but error remains until clear 100 |
| `error` push errno 100 | all-clear observed | clear `last_error`; following reports determine HA state |
| TCP/XML session close | FIN/RST/EOF/stream close | retained availability offline |
| unknown valid stanza/td/field | any | `{root}/raw` unparsed diagnostic; do not terminate session |
All message forms observed across the four captures therefore have either:
- a Home Assistant state/attribute/availability mapping;
- a command/diagnostic MQTT mapping; or
- an explicitly documented internal transport role with no meaningful HA
state analogue.
There are **no observed application messages left silently unmapped**.
### 13.1 Explicitly not mapped to HA vacuum state
These are intentionally internal or diagnostic because HA's existing vacuum
schema has no equivalent:
- XMPP stream negotiation, SASL, bind/session, iq receipt acks and pings;
- raw `sid`/`cid`, `ret`, command errno (published on `command_result` and
attributes, not vacuum state);
- schedule flag `f='p'` (preserved as `flag`);
- unknown `CleanReport st/rsn` values (blank in captures; preserve/log if not);
- GetCleanState `t/a` (empty in captures; preserve/log if not);
- lifespan `total` unit (value retained; unit intentionally unspecified);
- IGMP and ARP traffic;
- OTA success manifest, which was never observed and is unnecessary.
---
## 14. Validation, error handling and security requirements
- Accept any provisioned robot's SASL password, but require valid SASL PLAIN
structure and a nonempty authcid to avoid parser ambiguity.
- Never publish/log auth stanzas, decoded passwords, broker credentials or
other secrets.
- XML-escape all dynamic identifiers and attribute values; never construct XML
from untrusted raw fragments except the disabled expert interface.
- Validate MQTT enum values, schedule masks (`^[01]{7}$`), time (`HH:MM`),
booleans and payload sizes before issuing commands.
- Complete each cid once. A later command may reuse a finished id; the captured
app also reused a ctl id before its result. MQTT redelivery of a command is
still deduplicated.
- Put finite deadlines on pending sid/cid correlations and fail them on session
replacement.
- Unknown well-formed XML is nonfatal and goes to `{root}/raw`; malformed XML
terminates only the affected session and marks availability offline.
- Do not retain command topics or raw diagnostics.
- Subscribe to MQTT commands before publishing discovery/online so HA can
control immediately.
- On MQTT reconnect, republish discovery, availability and current full
retained state/attributes as needed.
- State mutations should be serialized per robot to keep reports, responses
and session replacement ordered.
---
## 15. Implementation acceptance checklist
### Bootstrap and XMPP
- [ ] DNS A override for `lbo.ecouser.net` reaches a stable listener address.
- [ ] 8007 returns exact compact numeric-port FindBest responses.
- [ ] 8005 returns acceptable 404 for firmware path.
- [ ] 5223 implements streaming plaintext XMPP and complete handshake.
- [ ] SASL values are accepted but never logged.
- [ ] New same-JID bind atomically supersedes stale/half-open sessions.
- [ ] Bot domain pings are answered exactly with matching id.
- [ ] Controller JID is announced after every READY.
- [ ] Half-open connection becomes MQTT offline before robot recovery where
possible; recovered full handshake becomes online again.
### Commands and state
- [ ] All HA commands in §11 emit exact `com:ctl` forms.
- [ ] All extensions in §12 validate inputs and correlate results.
- [ ] Both ack (`sid`) and command result (`cid`) phases are handled.
- [ ] Response payloads and pushes share the mappings in §13.
- [ ] `errno=100` clears rather than creates an error.
- [ ] Bare battery iq stanzas are parsed and not acked. A ctl id reused before
its result still completes once.
- [ ] Schedule mask is Sunday-first and schedule reports replace the full list.
- [ ] Unknown messages are logged/published without disconnecting.
### MQTT and Home Assistant
- [ ] Discovery payload is retained under configured HA discovery prefix.
- [ ] State JSON always contains a valid HA vacuum state.
- [ ] Fan speed is `standard|strong` and discovery list matches.
- [ ] Availability and LWT are retained; commands are not retained.
- [ ] Offline commands are rejected/dropped, never replayed after reconnect.
- [ ] Battery, consumables, schedules and errors remain visible as attributes.
- [ ] `pause`, resume and unobserved actions are not advertised as verified.
Passing this checklist is sufficient to begin end-to-end testing with the
provisioned N95 and Home Assistant without dependence on Ecovacs cloud
services after the DNS redirect.
+685
View File
@@ -0,0 +1,685 @@
# Ecovacs legacy XMPP protocol — full specification (robot-facing side)
> **Privacy notice:** Device serials, MAC addresses, authentication values, controller identifiers, and other identifying values shown here have been consistently replaced with synthetic values.
Derived from four captures of a **wukong**-platform Ecovacs robot, device **class 155**:
| Capture | Window | Duration | Pkts | Notes |
|---|---|---|---|---|
| `packetcapture-ix1.12-20260923211202.pcap` | 21:12:21–21:14:24 | ~2 min | 326 | boot + basic commands |
| `packetcapture-ix1.12-20260923214817.pcap` | 21:48:43–21:59:15 | ~10.5 min | 841 | clean state; schedules, stair-protection, errors, dock charge |
| `packetcapture-ix1.12-20260923223825.pcap` | 22:38:44–22:46:01 | ~7.3 min | 100 | forced wifi reconnects, cached endpoint reconnect, black-holed ping |
| `packetcapture-ix1.12-20260923230041.pcap` | 23:01:02.724–23:05:08.209 | 245.5 s | 37 | half-open timeout through complete automatic recovery (35 IP frames + 2 ARP) |
**Device identity**
- Model platform: `wukong`, device **class `155`** (appears as XMPP domain `155.ecorobot.net` and in the firmware URL)
- Serial / DID / XMPP uid: `E2998877665544332211`
- MAC: `02:00:00:00:95:01`; DHCP hostname request: **`deebot`**
- SASL credential: `a4f19c2e7b603d58e1c947ab25d8063f` — constant across all captured handshakes, therefore a persistent device credential rather than a session token
- Robot IP: `172.20.20.12` via DHCP; gateway/DNS `172.20.0.1`
**Scope of this document:** everything needed to implement the *robot-facing* side
of a replacement server — i.e. the services a **provisioned** robot contacts from
power-on until it is ready to accept commands. The app-facing side (portal REST
API, user auth) is *not* required for the robot to connect and is covered only
where the robot's behavior depends on it.
---
## 1. Boot sequence — shared order (captures 1 and 2)
The two power-on captures follow the same order. They are not packet-identical:
capture 1 was offered a DHCP lease of **6761 s** and sent **one** IGMPv2 report;
capture 2 was offered **7200 s** and sent **two** reports about 0.21 s apart.
DNS retransmit gaps were 4.2 ms and 7.4 ms. Absolute times below are capture 1;
capture 2 matches to within a few tens of milliseconds.
```
t+0.00 DHCP Offer + ACK (hostname "deebot"; Discover/Request not in the capture)
t+0.01 ARP probe, t+0.49 second probe, t+1.08 announcement
t+2.00 ARP who-has gateway → gateway MAC 02:00:00:00:00:01
t+2.00 IGMPv2 report → 226.1.1.1 (×1 in capture 1, ×2 in capture 2)
t+2.00 DNS A? lbo.ecouser.net (2 queries, src port 4096, 4–7 ms apart, before the answer)
t+3.00 TCP → {lbo-ip}:8007 ×2 parallel POST /lookup.do (EcoMsgNew & EcoUpdate)
t+3.19 TCP → {update-ip}:8005 GET …/firmware/latest.json
t+3.53 TCP → {xmpp-ip}:5223 XMPP stream open
t+3.83 SASL PLAIN → success → re-stream → bind 'atom' → session → presence
t+4.03 READY (hello-world presence). No controller traffic until the app connects.
```
Boot→ready ≈ **4.0 s** (presence at 4.026 s and 3.984 s). A stale session for the
same JID is closed with `</stream:stream>` + FIN **after the new session IQ**,
about 20 ms after that IQ and before `hello world` (capture 1: old port 16128,
FIN at t=4.021, session result in the same millisecond; bind result was 24 ms
earlier). The robot answers that FIN with two RSTs. Closing at bind is early
relative to the captured server but is safe: the kick is done before presence.
**Dependencies:** only **DNS** and the **`lookup.do` host** are strictly needed
to redirect the robot — everything downstream is learned from `lookup.do`
responses. Firmware check may fail/404 without consequence.
---
## 2. DNS
- Resolver used: the **first** server from DHCP option 6 (here `172.20.0.1`;
DHCP offered `172.20.0.1, 192.168.0.5, 192.168.0.4`).
- Query: `A? lbo.ecouser.net` — sent twice, 4.2 ms apart in capture 1 and 7.4 ms apart in capture 2, both before the response arrives.
- Source port: `4096` on both queries in both boots.
- Response observed: `CNAME slb-oldweb-iot-eu.ww.ecouser.net` → `A 8.211.16.120`. That address is also the `EcoUpdate` host returned by `lookup.do`. The XMPP host is a different literal address (`47.87.130.1`).
- To redirect: resolve `lbo.ecouser.net` (and `lbo.ecovacs.net` for safety — see
bumper docs) to the replacement server's IP. A wildcard
`address=/ecouser.net/{ip}` + `/ecovacs.net/{ip}` + `/ecovacs.com/{ip}`
covers everything, including `{class}.ecorobot.net` should the robot ever
resolve it (it does **not** — it connects to the IP from `lookup.do`).
## 3. Service discovery — `POST /lookup.do` (plaintext HTTP, port 8007)
The robot opens **one TCP connection per service, in parallel** (EcoMsgNew's
SYN leads EcoUpdate by under 1 ms; either response may arrive first). HTTP/1.0,
no `Host` header, tiny JSON body (newline + tab indent, and a tab between
`:` and the value), one request per connection. A replacement listener must
accept HTTP/1.0 without `Host`. The captured server answered HTTP/1.1 with
`Connection: close` and FIN; the robot then RSTs. One of the two lookup
sockets was reset by the robot without a captured server FIN. Send the
response and close; tolerate an immediate RST.
```http
POST /lookup.do HTTP/1.0
Content-Length: 48
Accept: Application/json
Content-Type: application/json
{
"todo": "FindBest",
"service": "EcoMsgNew"
}
```
```http
HTTP/1.1 200 OK
Content-Type: application/json; charset=utf-8
Content-Length: 46
{"result":"ok","ip":"47.87.130.1","port":5223}
```
| `service` | Returns | Used for |
|---|---|---|
| `EcoMsgNew` | `{"result":"ok","ip":"<xmpp-ip>","port":5223}` | XMPP control channel |
| `EcoUpdate` | `{"result":"ok","ip":"<ota-ip>","port":8005}` | firmware check host |
**Critical formatting:** body must be compact JSON with **no spaces** and
`port` as a JSON **number** (not string). (Bumper comment: bot is "very picky"
about this.)
For a local server: answer `EcoMsgNew → {your-ip}, 5223`. `EcoUpdate` may point
anywhere — the robot tolerates a failed/absent update check, but pointing it at
the local server and answering §4 keeps the boot path fully on-LAN.
## 4. Firmware check — `GET` on the `EcoUpdate` host (port 8005)
```http
GET /products/{product}/class/{class}/firmware/latest.json HTTP/1.0
Connection: Close
Accept: Application/json
Content-Type: application/json
```
→ here `/products/wukong/class/155/firmware/latest.json` → **`404`**, accepted
gracefully; robot proceeds to XMPP about 180 ms later without retry. The
captured 404 was `HTTP/1.1`, `Content-Type: text/plain; charset=utf-8`,
`Content-Length: 9`, body exactly `Not Found` (no trailing newline). A JSON
body or a successful manifest was not observed; do not invent one. XMPP starts
only after this response.
## 5. XMPP service — TCP 5223, **plaintext**
Despite 5223 being the legacy SSL port, **no TLS is negotiated**: the server
advertises `<starttls><required/></starttls>` and the robot ignores it,
proceeding straight to SASL PLAIN. A replacement server does not need TLS at
all for this robot; advertising `starttls` is optional cosmetic fidelity.
### 5.1 Handshake — verbatim
```xml
C→S <?xml version='1.0'?><stream:stream xmlns:stream='http://etherx.jabber.org/streams'
xmlns='jabber:client' to='155.ecorobot.net' version='1.0'>
S→C <stream:stream xmlns:stream="http://etherx.jabber.org/streams"
xmlns="jabber:client" version="1.0" id="{32hex}" from="155.ecorobot.net">
S→C <stream:features>
<auth xmlns="http://jabber.org/features/iq-auth"/>
<starttls xmlns="urn:ietf:params:xml:ns:xmpp-tls"><required/></starttls>
<mechanisms xmlns="urn:ietf:params:xml:ns:xmpp-sasl"><mechanism>PLAIN</mechanism></mechanisms>
</stream:features>
C→S <auth xmlns='urn:ietf:params:xml:ns:xmpp-sasl' mechanism='PLAIN'>{base64}</auth>
S→C <success xmlns="urn:ietf:params:xml:ns:xmpp-sasl"/>
C→S <?xml version='1.0'?><stream:stream … to='155.ecorobot.net' version='1.0'> <!-- reopen -->
S→C <stream:stream … id="{same-32hex}" from="155.ecorobot.net">
S→C <stream:features><bind xmlns="urn:ietf:params:xml:ns:xmpp-bind"/>
<session xmlns="urn:ietf:params:xml:ns:xmpp-session"/></stream:features>
C→S <iq type='set' id='0'><bind xmlns='urn:ietf:params:xml:ns:xmpp-bind'>
<resource>atom</resource></bind></iq>
S→C <iq type="result" id="0"><bind xmlns="urn:ietf:params:xml:ns:xmpp-bind">
<jid>E2998877665544332211@155.ecorobot.net/atom</jid></bind></iq>
C→S <iq type='set' id='1'><session xmlns='urn:ietf:params:xml:ns:xmpp-session'/></iq>
S→C <iq type="result" id="1"/>
C→S <presence><status>hello world</status></presence>
S→C <presence to="E2998877665544332211@155.ecorobot.net/atom"> dummy </presence>
```
### 5.2 Details a server must get right
- **Stream `to=`** carries `{class}.ecorobot.net` — this is how the server
learns the device class. Parse it out of the (unclosed) stream tag.
- **SASL PLAIN** payload = `base64("\0" + serial + "\0" + token)`; here
`\0E2998877665544332211\0a4f19c2e7b603d58e1c947ab25d8063f`. Authcid = serial =
uid. The password is a persistent factory/account credential. A local server
accepts the presented password unconditionally.
- **JID:** `{serial}@{class}.ecorobot.net/atom`. Resource is always `atom`.
- **Stream `id`:** 32 hex digits. The post-SASL `<stream:stream>` repeats the
same id. The next TCP connection gets a new id. The robot does not check it.
- **iq `id`:** stanzas the robot originates (bind, session, pings, pushes,
ctl results) use a counter that is **monotonic for the whole boot, with no
gaps**, not per session. Result acks echo the controller's iq id and are not
part of that counter. Bind=0/session=1 only on the first connection after
power-on. Capture 1 ran 0..51, capture 2 ran 0..172, capture 3 continued
201..213 across two reconnects (bind `208`/`210`, session `209`/`211`), and
capture 4 continued `222`..`225`. The jump 213→222 is the eight uncaptured
120 s pings between those files, not a reset. Don't assume small numbers.
The app's own iq ids are decimal strings of 3–8 digits and **were reused**
(`6677` after 16.9 s, `7689` after 315 s), once per later command, each with
one ack. There was no duplicate ack of a single request.
- After `session` → READY. The robot emits `<presence><status>hello world
</status></presence>`; answer with a `<presence> dummy </presence>` addressed
to its full JID.
- `<auth>` (iq-auth) is advertised but never used by the robot.
- **Session uniqueness:** on a new session for the same JID, close the previous
connection (`</stream:stream>` + FIN). Observed ~20 ms after the robot's
session IQ and before presence. The robot double-RSTs if that FIN arrives.
If the old path is black-holed the FIN is never delivered; the new bind must
still be accepted.
- **Framing:** stanzas are a continuous non-document XML stream — implement a
streaming parser: the robot may send several stanzas in one TCP segment, and
a stanza may straddle segments. Treat `<?xml …?>` + `<stream:stream>` (never
closed by the client) and a lone `</stream:stream>` (session end) specially.
---
## 6. The `com:ctl` command layer
### 6.1 Addressing
- **Bot JID:** `{serial}@{class}.ecorobot.net/atom`
- **Controller JID:** `{uid}@ecouser.net/{resource}` — e.g. the real app used
`demouser01234567@ecouser.net/LABclient001`. The server relays stanzas
between these JIDs; the robot directs all its responses/reports to the
controller JID learned from incoming `from=` attributes. With one controller
it is unambiguous; with several, route reports to the controller that last
commanded the bot (or broadcast — observed data can't distinguish).
- **Controller announcement:** the app announces itself with
`<iq type='get' to='{bot-jid}' from='{ctl-jid}'><ping xmlns='urn:xmpp:ping'/></iq>`.
That is the first post-ready traffic, but the delay is when the user opened
the app, not a robot timer: **12.877 s** in capture 2 and **38.751 s** in
capture 1. The app then sent a second ping (0.13–0.27 s later) and `SetTime`
about 1 s after the first ping. Steady-state controller pings are **~90 s**
(89.96–92.21 s in capture 2). The bot answers `type='result'` to that JID.
**No report was emitted on a session that never received a controller
`from=`** (capture 3, both reconnects, including several minutes on the
second). The first report in both boots was an empty `Sched2`, ~100 ms after
the `SetTime` result and not in the gap between the announce pings and
`SetTime`. A bridge must send its own `from=` on every session (ping, then
`SetTime`, matching the app). The learned JID does not survive a re-bind.
### 6.2 Command (controller → bot)
```xml
<iq id="{sid}" to="{bot-jid}" from="{ctl-jid}" type="set">
<query xmlns="com:ctl"><ctl td="{Command}" id="{cid}">…payload…</ctl></query>
</iq>
```
- `sid` = stanza id (controller-chosen, 3–8 digits observed). `cid` = **ctl
correlation id**. Every captured cid is a zero-padded 8-digit decimal
(`01410553`, `00027119`). The robot echoes it verbatim. The app reused a cid
on an immediate status retry before the first response (`GetCleanState`
`46393039` and `57306986`); one ctl result then arrived. A bridge should keep
cids unique among outstanding commands and still accept one result for a cid
that was issued twice.
- Several complete iq stanzas were written in one TCP segment (two `Get*`s, and
once those two plus a ping). Segment boundaries are not stanza boundaries.
- `Move` commands are sent **without** `id` on `<ctl>` and produce only the
stanza ack (no ctl response). All others carry `id`. A new `Move` was sent
without a preceding `stop` (`TurnAround` then `SpinLeft`, `forward` then
`forward`). Observed move bursts lasted 0.24–3.1 s; no firmware auto-stop
was seen inside that window.
### 6.3 Response — asymmetric **two-stanza** pattern
1. Stanza-level ack: `<iq type='result' … id='{sid}'/>`
2. Payload response as a **new `<iq type='set'>`** bot → controller:
```xml
<iq to='{ctl-jid}' type='set' id='{bot-seq}'>
<query xmlns='com:ctl'><ctl id='{cid}' ret='ok' errno=''>…payload…</ctl></query>
</iq>
```
Correlation is by **`ctl/@id`** — not the iq id. `ret` was only `ok` in these
captures (`fail` was not seen). `errno=''` is present on `Get*` and
`SetCleanSpeed` results and **absent** on `SetTime`, `Clean`, `Charge`,
`PlaySound`, `AddSched`, `ModSched`, and `DelSched` (`<ctl id='…' ret='ok'/>`).
Treat a missing `errno` as no error. Do not require the attribute.
### 6.4 Reports (bot → controller, unsolicited)
```xml
<iq to='{ctl-jid}' type='set' id='{bot-seq}'>
<query xmlns='com:ctl'><ctl td='{Report}'>…</ctl></query></iq>
```
`td` ∈ `Sched2`, `CleanReport`, `ChargeState`, `BatteryInfo`, `error` —
detailed in §8/§9.
**Push stanzas carry `to=` but no `from=`** (don't require one when parsing),
and — contrary to normal XMPP iq semantics — **the controller never acks
them**: zero `type='result'` replies to push ids appear in ~10.5 min of
capture. The bot neither expects nor notices. A bridge must not wait for acks
on pushes and need not emit them.
### 6.5 Pings
- Controller→bot: `<iq type='get' to='{bot-jid}' from='{ctl-jid}'><ping
xmlns='urn:xmpp:ping'/></iq>` — the app's announce/keepalive, sent **every
~90 s** (also the first post-session stanza, §6.1); bot answers
`type='result'` to the controller JID.
- Bot→server: `<iq from='{bot-jid}' to='155.ecorobot.net' type='get'><ping
xmlns='urn:xmpp:ping'/></iq>` every **~120 s**; server must answer
`<iq type='result' from='155.ecorobot.net' to='{bot-jid}' id='{n}'/>`.
- The server itself never pings the bot (all bot-ward pings carry the
controller `from=`).
---
## 7. Command reference (all `td` values seen on the wire)
Request/response XML is verbatim. `{cid}` = ctl id echoed in response.
### 7.1 `SetTime` — clock sync (always the first command after connect)
```xml
<ctl td="SetTime" id="{cid}"><time t="1790194386" tz="1" tzm="0"/></ctl>
→ <ctl id='{cid}' ret='ok'/>
```
`t`=epoch seconds, `tz`=tz hours offset, `tzm`=tz minutes offset (UTC+1 here).
**A standalone server should emit this itself** once a bot session is ready.
### 7.2 `GetBatteryInfo`
```xml
<ctl id="{cid}" td="GetBatteryInfo"/>
→ <ctl id='{cid}' ret='ok' errno=''><battery power='076'/></ctl>
```
`power` = 0–100, zero-padded to 3 digits.
### 7.3 `GetCleanState`
```xml
<ctl id="{cid}" td="GetCleanState"/>
→ <ctl id='{cid}' ret='ok' errno=''><clean type='stop' speed='standard' st='h' t='' a=''/></ctl>
```
`type` = current/last clean mode. Observed snapshots: `type='stop' st='h'`
and `type='auto' st='s'`. `p`/`r` were not in a `GetCleanState` reply.
`t` and `a` were empty strings in every reply (not the single space used by
`CleanReport`'s `st`/`rsn`).
### 7.4 `GetChargeState`
```xml
<ctl id="{cid}" td="GetChargeState"/>
→ <ctl id='{cid}' ret='ok' errno=''><charge type='Idle'/></ctl>
```
`type` ∈ `Idle` (not docked/not charging), `going` (returning to dock),
`SlotCharging` (on dock, charging).
### 7.5 `GetCleanSpeed` / `SetCleanSpeed`
```xml
<ctl id="{cid}" td="GetCleanSpeed"/> → <ctl id='{cid}' ret='ok' errno='' speed='standard'/>
<ctl id="{cid}" td="SetCleanSpeed" speed="strong"/> → <ctl id='{cid}' ret='ok' errno=''/>
```
`speed` ∈ `standard`, `strong`. The ctl result is `ret='ok' errno=''` and does
**not** echo `speed`. During an active auto clean, `SetCleanSpeed` was followed
by a `CleanReport` at the new speed. While `type='stop'`, a speed change was
not consistently followed by a report. Apply the new speed when `ret='ok'`,
and also accept a later `CleanReport` or `GetCleanSpeed`.
### 7.6 `GetLifeSpan` — consumable life
```xml
<ctl id="{cid}" td="GetLifeSpan" type="SideBrush"/>
→ <ctl id='{cid}' ret='ok' errno='' type='SideBrush' val='068' total='365'/>
```
`type` ∈ `SideBrush`, `Brush`, `DustCaseHeap` (filter). `val` = % remaining
(zero-padded, `068` = 68 %). `total` was `365` for all three types. The unit
was not established; keep the integer and do not label it hours.
### 7.7 `Move` — manual driving
```xml
<ctl td="Move"><move action="forward"/></ctl>
```
`action` ∈ `forward`, `backward` (sucks-known, not exercised), `SpinLeft`,
`SpinRight`, `TurnAround`, `stop`. **No `id` on ctl → stanza-ack only.**
### 7.8 `Clean` — start/stop a clean
```xml
<ctl id="{cid}" td="Clean"><clean type="auto" speed="strong" act="s"/></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ CleanReport push)
```
`type` ∈ `auto`, `border` (edge), `spot`, `singleRoom` (camelCase on the wire —
sucks maps `singleroom`, a real vocab discrepancy), also `stop` for the
stop-command itself. `act` = `s` start / `h` halt; `p`,`r` (pause/resume) known
from sucks. `speed` ∈ `standard|strong`.
### 7.9 `Charge` — dock control
```xml
<ctl id="{cid}" td="Charge"><charge type="go"/></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ CleanReport stop + ChargeState 'going')
```
`type` = `go` (return to dock) / `stopGo` (cancel return). `go` is followed
within ~50 ms by `CleanReport stop` and `ChargeState going`, and by
`SlotCharging` when the robot is on the dock (21.4 s later in capture 2;
again, with errno 100 and `CleanReport stop`, at the start of capture 3).
`stopGo` is followed by `CleanReport stop` and `ChargeState Idle`.
### 7.10 `PlaySound` — find-me beep
```xml
<ctl id="{cid}" td="PlaySound" sid="0"/> → <ctl id='{cid}' ret='ok'/>
```
---
## 8. Schedule subsystem — fully captured
### 8.1 `AddSched`
```xml
<ctl id="{cid}" td="AddSched">
<sched name="17901966021514" on="1" time="19:30" repeat="0001000">
<ctl td="Clean"><clean type="auto"/></ctl>
</sched></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ Sched2 push)
```
### 8.2 `ModSched`
```xml
<ctl id="{cid}" td="ModSched">
<ModSched name="17901966021514">
<sched name="17901966021514" on="0" time="19:30" repeat="0001000">
<ctl td="Clean"><clean type="auto"/></ctl>
</sched></ModSched></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ Sched2 push)
```
Wrapper `<ModSched name="{existing-name}">` selects the entry to replace.
Captured edits changed `on`, `time`, and `repeat`. The inner `<sched name>`
was always equal to the wrapper name; a rename was not tested. The inner
clean type was always `auto`.
### 8.3 `DelSched`
```xml
<ctl id="{cid}" td="DelSched"><DelSched name="17901966423286"/></ctl>
→ <ctl id='{cid}' ret='ok'/> (+ Sched2 push)
```
### 8.4 `GetSched` and the `Sched2` report — `<s>` element format
```xml
<ctl id='{cid}' ret='ok' errno=''>
<s n='17901966302986' o='1' t='01:50' r='1101011' f='p'> <ctl td='clean' type='auto'/> </s>
<s n='17901966021514' o='1' t='21:30' r='0111000' f='p'> <ctl td='clean' type='auto'/> </s>
</ctl>
```
| attr | meaning |
|---|---|
| `n` | schedule name/id — app-generated unique string ≈ `epoch_seconds*10^4 + suffix`; opaque to the robot, echoed verbatim |
| `o` | on/enabled `0`/`1` |
| `t` | local time `HH:MM` |
| `r` | 7-char repeat bitmask **index 0 = Sunday … index 6 = Saturday**. Verified: `0001000` (index 3, Wednesday) fired on Wednesday 23 Sep 2026. Also observed: `0111000`, `1101011`, `1111111` |
| `f` | flag, constant `'p'` in every `<s>` (meaning unknown; echo as stored) |
- The inner action is `<ctl td='clean' type='auto'/>` — **lowercase `clean`,
`type` directly on `ctl`**. Only `auto` was stored. The command dialect uses
a `<clean type=…/>` child and `td="Clean"`. Each `<s>` contains a space
before the inner `<ctl>` and a space before `</s>`. Adjacent `<s>` elements
are concatenated with no text between them (`</s><s`).
- `Sched2` is pushed: **(a)** once a controller is known, immediately after the
first `SetTime` (empty `<ctl td='Sched2'/>` when the table is empty — this is
not at XMPP ready, which was 13–39 s earlier), **(b)** after every
Add/Mod/Del, **(c)** when a schedule fires. At 21:58:59 local the bot pushed
`Sched2` for the `21:59` / `0001000` entry and, 73 ms later, `CleanReport`
`auto`. That is under a second before the scheduled minute.
- `GetSched` returns the same `<s>` children inside the ctl response. An empty
table is a ctl result with `ret='ok' errno=''` and no `<s>` children, which
is a different element from an empty `Sched2` push.
- `DelSched` of the last entry → `Sched2` pushed with **no** `<s>` children
(`<ctl td='Sched2'/>`), observed twice (after the first batch and after the
21:59 entry was disabled and deleted).
---
## 9. Robot-initiated reports (pushes)
All are `<iq type='set' to='{ctl-jid}'><query xmlns='com:ctl'><ctl td=…/>`:
| `td` | Payload | When emitted |
|---|---|---|
| `Sched2` | `<s …/>` children or empty | after the controller is known (first one ~100 ms after SetTime), after sched mutations, when a schedule fires |
| `CleanReport` | `<clean type='{mode}' speed='{spd}' st=' ' rsn=' '/>` | after Clean and Charge, and on autonomous transitions. `st` and `rsn` were a single space in every report, including while running. `h`/`s` appear only in `GetCleanState`. Observed report types: `auto`, `border`, `spot`, `singleRoom`, `stop`, speeds `standard` and `strong`. `SetCleanSpeed` updated a following report during an active clean only |
| `ChargeState` | `<charge type='Idle'/'going'/'SlotCharging'/>` | `go` → `going` within ~50 ms of `CleanReport stop`, then `SlotCharging` on arrival (21.4 s later in one run). `stopGo` → `CleanReport stop` + `Idle`. `Idle` is also pushed on leaving the dock (35 s after `SlotCharging`, with `CleanReport stop`). `Idle` while off the dock is also the `GetChargeState` answer during a clean |
| `BatteryInfo` | `<battery power='NNN'/>` | about every **25.00 s** both while cleaning and while `SlotCharging`. Capture 2 fell from 77 to 68 over the clean, with one +1 step (70→71, which is the sample after docking). Capture 3 rose 67→68 on the dock. Slots were sometimes skipped (gaps of 50 s and 75 s) |
| `error` | `<ctl td='error' errno='NNN'/>` | **error events** — see §10 |
| *(ping)* | `<ping xmlns='urn:xmpp:ping'/>` to `{class}.ecorobot.net` | every ~120 s |
## 10. Error reporting — observed
`<ctl td='error' errno='N'/>` is a **push**, not a command response:
| errno | Context observed | Behavior |
|---|---|---|
| `103` | mid auto-clean; `CleanReport stop` 24 ms later | clean aborted — stair/cliff protection halt (per capture context). No command ctl-result carried errno 103 |
| `100` | 15.771 s after 103, then `CleanReport auto` 24 ms later (capture 2). Separately, at the start of capture 3: `CleanReport stop` 23 ms later and `SlotCharging` 56 ms later | **all-clear / error-cleared beacon**, not a new fault. The reports that follow carry the new motion state. The stop 4.4 s after the capture-2 beacon was a later `Clean` `act='h'`, not part of the beacon |
**`errno='100'` means "error cleared", not "error".** In capture 2 it preceded
a resumed `CleanReport auto`; in capture 3 it preceded `CleanReport stop` +
`SlotCharging`. Consumers should treat it as clearing a prior fault and derive
state from the `CleanReport`/`ChargeState` that follow, never as an error
itself.
Command-level errors also exist (sucks/bumper knowledge, not exercised here):
`ret='fail'` + `errno` on ctl responses (`3`,`5`,`8` per sucks charge handling;
`103` = permission denied on *command responses* — different from the `td=error`
push!). **Do not conflate:** `td='error'` pushes are device fault reports.
## 11. Session lifecycle & timing behavior
- **Boot→ready ~4 s**; XMPP session is long-lived.
- Bot→server ping every ~120 s (`to='{class}.ecorobot.net'`).
- Controller pings relayed on demand; robot always answers `result`.
- **One session per JID:** the new session kicks the old connection
(`</stream:stream>`+FIN) about 20 ms after the session IQ and before
presence — observed in captures 1 and 3. The robot double-RSTs when that
FIN arrives.
- Session end: `</stream:stream>` from either side; robot just drops TCP on
power-off (no graceful close observed at shutdown).
- Robot iq `id` sequence is strictly incrementing **per boot**, continuing
across reconnects (§5.2).
### 11.1 Reconnect behavior (capture 3 — forced wifi drops + router state clears)
- **Reconnects skip the entire bootstrap.** After a wifi drop the robot does
DHCP renew (leases 7200 s then 7183 s) + ARP probe + one IGMPv2 report, then
opens a fresh TCP connection **directly to the cached `EcoMsgNew` IP:5223**
— no DNS query, no `lookup.do`, no firmware check. The first reconnect
reported `226.1.1.1`; the second reported `224.0.0.1`. The `lookup.do`
result is cached for the life of the boot. The DNS/`8007`/`8005` services
are only needed at power-on. **If the bridge IP changes, the robot cannot
rediscover it without a reboot.**
- **Every reconnect is a full re-handshake**: stream → SASL PLAIN (same
factory token) → re-stream (same stream id as that connection's first open)
→ bind `atom` → session → `hello world`. SYN to dummy presence was 0.42 s,
0.48 s, and 0.45 s. No credential or endpoint renegotiation exists.
- **Stale-session kick on each new session** — `</stream:stream>`+FIN about
20 ms after the new session IQ, before presence. Not at the bind result.
- **No reports on a session with no announced controller** — conns B and C in
capture 3 produced *zero* pushes (the app was gone and never re-announced).
Confirms §6.1: the learned controller JID is per-session and must be
re-established after every reconnect.
- **Dead-path detection is pure TCP, and the RTO is adaptive.** Capture 4 is
the complete timeout. The last healthy bot ping (`id=222`) was acknowledged.
Exactly 120.000 s later the bot sent ping `id=223`. With the return path
black-holed, that identical 130-byte segment (same sequence, same XMPP id)
was retransmitted at `+0.668, +2.342, +5.368, +11.426, +23.468, +47.666,
+95.863 s`. No second XMPP stanza was generated. Capture 3 had already
started this for ping `id=213` and was still retransmitting when the file
ended, on a longer RTO: `+0.918, +2.999, +7.016, +15.043, +31.116, +63.289 s`.
Do not hard-code either series.
- **The robot abandons the half-open socket after 120.002 s** (capture 4,
measured from the original ping). It sent TCP FIN, did not wait for FIN-ACK,
and did not send `</stream:stream>`.
- **Fresh connect after 4.999 s:** new TCP connection to the same cached
endpoint. SYN to dummy presence was 0.455 s; bind/session ids continued as
`224`/`225`. No DNS, `lookup.do`, DHCP, firmware lookup, or XEP-0198 resume.
The first bot ping of a new session is not immediate (about 95 s after
presence on capture 3's second session); on a stable session the period is
120.000–120.002 s.
- The old server-side socket may remain half-open because neither its close nor
the robot's FIN can traverse the cleared state. The new bind must atomically
replace the JID→connection mapping and close/discard the old local socket;
never reject the new bind merely because that JID appears connected.
### 11.2 Required reconnect state machine
```text
ESTABLISHED
bot ping every 120 s
ping write/ack failure → kernel TCP retransmission
120 s without delivery → robot sends FIN, abandons socket
wait ~5 s
TCP connect cached EcoMsgNew IP:port
full XMPP authentication/bind/session (not XEP-0198 stream resumption)
READY
```
A server cannot shorten the robot firmware's client-side 120 s timeout once
packets are black-holed. It can improve Home Assistant accuracy independently
by detecting its own failed controller pings/TCP keepalive and publishing
`offline` before the robot reconnects.
**Bridge-side improvements over what the real server demonstrated:**
1. Detect zombie sessions before the robot's own ~125 s ping-timeout/reconnect
cycle: send controller pings every ~60 s and mark offline when one misses a
10–15 s result deadline. Optional TCP keepalive is an additional signal.
2. On every new READY: re-announce + `SetTime` + status fan-out (mandatory —
the per-session learned JID is gone).
3. Availability will flap offline→online across reconnects; the retained
state topics keep HA's entity populated throughout.
## 12. Value enumerations (complete observed set plus marked library values)
```
clean.type auto | border | spot | singleRoom | stop [SpotArea library-known]
clean.act s (start) | h (halt) [p,r library-known]
clean.st s (running, GetCleanState) | h (halted, GetCleanState) | ' ' (every CleanReport)
clean.speed / SetCleanSpeed.speed / GetCleanSpeed standard | strong
move.action forward | SpinLeft | SpinRight | TurnAround | stop [backward library-known]
charge.type go | stopGo (command)
charge state Idle | going | SlotCharging (reports/queries)
lifespan.type SideBrush | Brush | DustCaseHeap
ctl.ret ok | fail ; ctl.errno '' | <numeric>
s.o 0|1 ; s.r 7 chars, index 0 = Sunday … index 6 = Saturday ; s.f 'p'
```
## 13. Firmware quirks to tolerate
1. **STARTTLS ignored** despite `<required/>` — never wait for it.
2. Two identical DNS queries 4–7 ms apart, both before the answer; two parallel `lookup.do` connections. No `Host` header.
3. `<query>` may contain a **bare `<battery power='…'/>` with no `<ctl>`**.
Seen twice, both times a full `<iq type='set' id='{bot-seq}'>` (capture 1
id `46` power `076`; capture 2 id `60` power `077`), 70–90 ms after a normal
`GetBatteryInfo` result with the same power. The app sent no iq result for
those ids. Parse the battery and do not ack.
4. The app reused iq ids and ctl ids (see §5.2 and §6.2). Each request still
got one result. There was **no** duplicate ack of one stanza ~600 ms apart.
Complete a cid once; a later request may legally reuse it after the first
result, and the captured app sometimes reused a ctl id before the result.
5. `Move` ctl has no `id`; never expect a ctl response for it.
6. Sched `<s>` elements contain literal-space text nodes and the inner action
uses lowercase `td='clean'` with `type` on `ctl`.
7. `hello world` presence has no `type`; answer with ` dummy ` presence.
8. HTTP/1.0 requests, `Accept: Application/json` (capital A), tiny bodies;
responses must be space-free JSON with numeric `port`.
9. Stream `from=`/`id=` values are not validated by the robot (real server:
`from="{class}.ecorobot.net"`, random hex id — mimic for fidelity).
10. Stanza boundaries ≠ TCP segment boundaries in both directions.
## 14. Minimum server checklist (robot-facing)
| # | Service | Required behavior |
|---|---|---|
| 1 | DNS | `lbo.ecouser.net` (and `lbo.ecovacs.net`) → server IP |
| 2 | TCP 8007 | `POST /lookup.do` `FindBest`: `EcoMsgNew`→`{ip,5223}`, `EcoUpdate`→`{ip,8005}`; compact JSON |
| 3 | TCP 8005 | `GET /products/*/class/*/firmware/latest.json` → 404 (or canned manifest) |
| 4 | TCP 5223 | XMPP stream: features (+iq-auth, optional starttls advert, PLAIN), SASL accept-all, bind→`{serial}@{class}.ecorobot.net/atom`, session result, dummy presence |
| 5 | XMPP | Parse `to='{class}.ecorobot.net'` for devclass; keep uid/JID table |
| 6 | XMPP | `urn:xmpp:ping`: answer server-domain pings; relay controller pings |
| 7 | com:ctl | Route `iq/query/ctl` between controller JIDs and bot JIDs **verbatim** (no schema validation); both `type=set` and `result` |
| 8 | com:ctl | Optionally inject own commands from a virtual controller JID (app-free control) — the bot answers to `from=`; first `from=` seen per session registers the report-push destination (announce with a `urn:xmpp:ping`, repeat ~90 s); pushes need no ack |
| 9 | XMPP | Atomically replace the JID→connection mapping and close/discard the stale local socket on same-JID re-bind; always accept the new connection even if the old half-open socket cannot receive its close |
| 10 | — | Track last-reporting state (`CleanReport`/`ChargeState`/`BatteryInfo`/`Sched2`/`error`) for a status API |
Nothing else is required: no TLS, no HTTP 443, no MQTT, no app auth — the robot
is fully served by the above.
## 15. Bumper coverage map (updated after capture 2)
**Already covered:** `lookup.do`+`FindBest`/`EcoMsgNew` (confserver.py:422),
plaintext XMPP handshake, devclass extraction from `to=`, SASL-accept for bots,
`atom` bind → correct JID form, session/presence, ping handling both ways,
transparent `com:ctl` relay both directions (including `type='set'` responses
and all push reports — Sched2/error/CleanReport relay fine), errno=103
*command-response* repair flow (AddUser/SetAC/GetUserInfo).
**Gaps/risks:**
1. `EcoUpdate` lookup → hardcoded real Ecovacs `47.88.66.164:8005`; no local
8005 listener or `latest.json` route.
2. `lbo.ecouser.net` absent from DNS docs (wildcard covers it).
3. **`errno='103'` substring match in `_handle_result` matches `td='error'`
fault pushes.** Those pushes have neither an `error` nor an `admin`
attribute, so `adminuser` is never set and the handler raises. It should
ignore `td='error'` pushes and only run the AddUser path for a ctl command
response that actually carries `error` or `admin`.
4. Stream `from=` domain (`ecouser.net` vs real `{class}.ecorobot.net`) and
static stream id `"1"` — cosmetic.
5. No stale-session kick on same-JID rebind.
6. `_handle_ctl` crashes on `to`-less stanzas (all observed stanzas have `to`).
7. Bumper sends `GetDeviceInfo` post-presence — not in the captured server
behavior, and this N95 never sent or answered that command in these files.
The response schema is unknown. Do not depend on it.
8. sucks vocab: `singleroom` vs wire `singleRoom`; `SetTime` missing `tzm`;
sched commands (`AddSched`/`ModSched`/`DelSched`/`Sched2`) **absent from
sucks entirely** — documented here for the first time.
**Bottom line:** the protocol is now documented completely enough to implement
the robot-facing server from scratch — bootstrap, discovery, handshake,
command/response correlation, full `td` vocabulary including the schedule
subsystem, all push reports, error semantics, keepalive, and firmware quirks.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.