Absence is not a measurement
A silent KNX address may be waiting on bus timing, contain no value or be unreachable. A command-only relay has no reporting path at all. A cloud subscription can lapse while the physical device continues working. An offline inverter says nothing about the independent energy meter beside it.
The FY2026 state-feedback investigation treated these as different evidence conditions. The system must avoid both false confidence and false fault reports. One fleet-wide timeout cannot supply the correct meaning for every transport and device class.
Preserve the distinction between command and observation
On the KNX path, null or absent values had been admitted as readings. Filtering group addresses without a value prevents “nothing was returned” from becoming a measurement. On a command-only path, silence cannot confirm the requested outcome because no confirmation channel exists.
A timed-out control command now marks the affected station unavailable rather than leaving the last commanded value appearing confirmed. Here a station means the gateway and the devices reached behind it; an entity is an individual state or control point. Those scopes matter when deciding what a failure should affect.
Isolate independent sources
An observed failure coupled a Tuya energy meter’s output to the availability of a Modbus inverter. When the inverter went offline, valid meter readings disappeared as well. The two sources were decoupled so that each could continue publishing on its own account.
The principle is broader than that incident: missing data from one publisher should not erase valid evidence from another. Source identity, availability and measurement meaning must survive aggregation. Otherwise an apparently tidy combined state hides the very observations needed to diagnose a partial outage.
| Device or transport class | Meaning of missing feedback | Policy established in the investigation |
|---|---|---|
| KNX group address | An absent value is not an observation. | Filter absent readings; distinguish delay through class-specific handling. |
| Tuya subscription | The subscription may have lapsed while the device operates. | Surface timeouts and isolate accounts; recover the subscription. |
| Modbus inverter | The source is offline or missing data. | Keep independent meter publications available. |
| Command-only output | No response is structurally expected. | Do not infer confirmation; mark station unavailable on command timeout. |
A healthy log can conceal a blind subscription
Tuya subscription timeouts were found to disappear before becoming visible in the failure record. The absence of logged errors had therefore been mistaken for coverage of the failure mode. Surfacing those timeouts and separating subscriptions per account made the failure observable and limited its scope.
An operation timeout supports re-establishing a lapsed subscription. This is a subscription-lifecycle condition, not proof that the physical device has failed. The documented classification applies to the characterised device classes; the report does not supply independently derived thresholds for every BACnet or LifeSmart path.
Reload is a reconciliation operation
The edge lifecycle investigation compared three independently drifting states: what the cloud wants, what Home Assistant has registered and what the installation is doing. A reload could report success while a KNX cache retained values from the previous definition. Restarting transports was therefore insufficient.
The revised path compares desired and loaded state subsystem by subsystem. Newly required subsystems are set up, retired ones are torn down and unchanged ones remain active in incremental mode. A forced mode provides an explicit rebuild path. This limits disturbance to healthy subsystems instead of treating every configuration change as a whole-plugin restart.
Teardown owns the entities it leaves behind
Retiring a subsystem without removing its registered entities left persistent unavailable “ghosts” across switch, light, cover, climate and sensor platforms. These survived restarts. Registry cleanup therefore became part of teardown, so a later rebuild could start from a coherent entity set.
Other failures exposed ordering and identity assumptions: a synchronous listener blocked bootstrap, and a DHCP address change left the cloud holding an obsolete host address. The documented changes moved the listener to a background task, compared local and recorded addresses, and prevented BACnet binding to an address absent from the host.
Recovery needs evidence, not a success label
A returning device can have liveness re-probed in place rather than require a restart that disturbs registry state. Resynchronisation must also account for platform load order: a present device can look absent if registration occurs before its runtime platform is ready.
The evidence base consists of field disagreements, reproduction cases, recorded changes and targeted regressions. It supports the described policies; it does not establish a numerical fleet-wide reliability guarantee or prove convergence under every concurrent failure. Hardening overlapping start/stop sequences and making consistency continuously observable remain development directions. Home Assistant provides local automation execution on this path; Tecco’s work described here governs the surrounding model, feedback and lifecycle behaviour.
