Matter devices are unavailable

Even without API access, an app can still f.e. be turned into a residential proxy or botnet, since internet access is always allowed on Homey. And since only manifest changes require review, a malicious actor could just publish a malicious code/update to the store immediately after the first version is approved.

In my opinion, Athom should just review every update, not only manifest changes.

Hi. Thanks for the mention. I noticed this app yesterday and thought it may help. I will try to get some stuff moved back over into Homey and use this tool to see if it shows anything

The zigbee one helped me find a device offline that I thought was working.

Cheers

Yes, that’s absolutely true.

Maybe a @moderators can split these posts so we can discuss this further in a separate topic.

Over a week since I moved devices over to the HA ZBT-2

No failed flows or devices dropping off the mesh

Everything is rock solid and devices are much faster to respond

I’m currently just using devices with the official HA app, but ideally I’d like to reset Homey and get it to join the HA thread network

Does anyone have a solid workflow for doing that? I’m assuming I need to do a full factory reset and start from scratch and can’t use a backup

Has there been any improvement with v13.4 and the ‘fixes’ for device reporting intervals?

What I can say from my own experience is this:

When I first bought Homey, I had these problems almost every day and was close to smashing the device against the wall. Then an update arrived, and Homey ran perfectly for about three weeks without a single issue or reboot. My automatic reboot check was still active, as was my logging, so I am quite confident that it was genuinely stable during that period.

Then one of the 13.x updates arrived, and the problems started again on a daily basis. Once again, I reached the point where I wanted to throw Homey against the wall.

The latest update took a long time to arrive. After installing it, all the devices that had previously failed to come back—and that I had become tired of repairing again and again—suddenly reappeared. That made me cautiously optimistic that the problem had finally been solved.

But last night Homey rebooted again because one of the devices monitored by my reboot flow had become unavailable.

So yes, the update may improve the situation, but in my case it has not solved the underlying problem.

What really disappoints me is:

  • Homey worked perfectly with one of the earlier 13.x firmware versions.
  • A later update reintroduced the problem.
  • It then took a very long time to investigate something that, from a basic IT troubleshooting perspective, should have started with a simple question: what changed between the last stable firmware and the first problematic one?
  • Support continued to suggest that the cause was probably the charger, the cable, or something specific to my home environment.

Even assuming that Homey now becomes stable, how am I supposed to trust that a future update will not break it again?

I am also still missing transparency about the reporting-interval changes. I understand that the update reduces communication traffic, but I have not received an answer to some very basic questions:

  • Are messages queued and delivered later?
  • Are intermediate messages dropped?
  • Are only sufficiently large value changes reported?
  • Are identical consecutive reports suppressed?
  • Are event-based messages treated differently from regular state updates?

The technical fix is only part of the issue. The bigger problem is that I have lost confidence in the platform, in the update process, and in the defensive way support has handled this from the beginning.

I had another crack at this since matter updates have been pushed out. Got my Nuki lock in immediately, got my onvis thread plugs in immediately and got one of my tado thermostats in. All seem to be stable . Not sure the issue with Tado but so far I can say matter / thread has improved and hopefully that will continue to be the case

I have pretty much the same experience with all my IKEA Matter devices, even the problematic MYGGBETT have been stable for long time now.

Four people have answered that in the last day and all four say some version of yes. I cannot add my own answer, because I do not own a Homey anymore and cannot test anything, so everything below is reading public sources rather than measuring.

What I would like is for this moment to turn into something Athom can actually build on, and I think there are three things in it. More insight into what is happening rather than how it feels. Logging, so the state is recorded instead of remembered. And data going to Athom. The third one has a deadline, so I will start there.

None of this is about talking anyone off Homey. It is the opposite. It is about handing Athom the kind of shared picture a community can build when enough people post what they actually see, which is how these things tend to get cornered on any platform, including the one I ended up on.

The report almost nobody sends

Everyone sends a diagnostic report when something is broken. Almost nobody sends one while it works.

Doekse has said it twice in this thread, in #220 and #227, that the diagnostic report to support is the channel he can actually act on. Right now several people here have a working Matter setup after a long broken stretch. That is a control group from exactly the houses that were failing, on the same hardware, the same network and the same devices as the broken reports already on file. A report from a house that was always fine does not give Athom that. A report from a house that was broken last month and works today does.

So a genuine ask to @Welshsmarthome, @Jocke_Wallen, @Pascal_Nohl and anyone else currently in a good state. Send a diagnostic report now, while it works, and write in the ticket that it is a working state and not a fault, with a pointer to this thread. It takes a couple of minutes and it is the one sample that expires. The moment something drifts we are back to sending broken reports, and there will be no before to hold them against.

Worth saying plainly, because it decides who can help here. The diagnostic report is a menu item. It needs no scripting, no Web API and no logging setup. The most valuable thing anyone in this thread can do right now happens to be the easiest thing in it, and that means it is not limited to the few people here who script.

Why I am pushing data instead of just agreeing

Because more than one thing changed at the same time.

v13.4 landed. But people also re-paired devices in the same week, some devices carry their own firmware, and at least one setup here had devices sitting on two different Thread networks. Any of those could be carrying the improvement. I believe every one of the four reports, and “it got better” still cannot tell Athom which change to keep.

@Jocke_Wallen, this is where your setup can say more than most, and I mean that as an invitation rather than a challenge. In your own topic “Thread network question” from June you described two IKEA GRILLPLATS plugs that kept ending up on a Google Thread network called Google-2B8F instead of Homey’s, however often you removed and re-added them, and Peter_Kawa pointed at commissioning through my.homey.app so Homey’s own border router gets used. Did those plugs end up back on Homey’s Thread network in the end? If they did, your improvement has two changes behind it and the order they happened in would be worth knowing. If they did not, that is just as useful, because then it is the firmware on its own. Either answer helps, and thank you for posting in the first place.

Insight and logging

The useful report back on v13.4 is not whether it feels better. It is what actually happened over a set period. How many devices went unavailable in a week, which ones, Thread or Wi-Fi, and whether they still answered when you pressed the button during the quiet stretches. A device that stayed reachable the whole time is the fix working. A device that was dead and simply not flagged is not. Nothing in that needs a script, it needs a note on paper and a date.

Pascal’s newest post is already that shape, and it is the most useful thing posted here in a while. All his Matter over Thread devices came back with no repairing, his two Matter over Wi-Fi devices are still not right, and his Homey still rebooted last night on an unavailable device. Thread and Wi-Fi behaving differently is a lead. Two dedicated test devices, one on each transport, is the right instrument. Pascal, if you post those two timelines after a week or two, several of us can read them with you.

The logging half is the part I would still like from the platform rather than from us. Pascal only knows about that three week stable period because his own reboot check and his own logging were running. That should not be a thing users have to build. I made the wider point in #204 and will not repeat it, and the two Insights apps from #233 and #234 are worth a run as a stopgap.

One question to Athom, and one warning

v12.4.8 added “Adds functionality to mark a Matter device as unavailable when it is offline”. That is the mechanism behind the title of this thread. v13.4.0 raises the maximum subscription interval. So @Doekse, did the window Homey uses to decide a Matter device is offline move together with the subscription interval, or is it set independently of it?

I ask because the two outcomes look the same from outside for the first few weeks. Fewer devices genuinely dropping is a fix. The same devices dropping but Homey taking longer to flag them is not. Both read as “fewer red triangles since the update”. Pascal’s report answers part of it already, since his devices actually came back and needed no repairing, which points at a real fix. But the flag still fires, because his Homey still rebooted on an unavailable device.

The warning is small and worth getting ahead of. v13.4.0 does contain “Fixes a race condition that could cause a “device unavailable” error”, but that line sits under Bluetooth, not under Matter. Given the title of this thread it will get read as a Matter fix, and in the changelog it is not one. The Matter section of v13.4.0 is three lines, the two subscription interval ones and a robot vacuum mapping fix.

One last note for Pascal’s process point from #246, which I think is fair. The changelog is the only artifact those of us outside Athom can try to answer “what changed between the last stable firmware and the first problematic one” from, and the public feed at ota-api.homeypro.net/api/v1/changelog carries 286 entries with no release date on any of them. You cannot line a firmware up against the week your house started misbehaving. That is a small fix on Athom’s side and it would make every report in this thread more useful.

I would suggest disabling this Flow for a while and see what happens. A reboot might seem helpful but when your Thread Network tries to stabilize put you are continuously “pulling the plug” that might also make things worse.

As mentioned in our public changelog, the subscription interval has changed for Matter devices. This change was implemented on behavior we saw with @deejayreissue’s setup.

In layman’s terms this means that Homey is allowing devices more time between subscription reports. If a device fails multiple of these reports it will be marked as unavailable. The user can now also adjust this value within the device settings.

Looking at my Motionblinds for example the default minimum is 0 and the default maximum is 300 seconds. Homey can increase the maximum value for up to 20% to make sure devices don’t all check in at the same time. This data is all visible within the advanced settings of a device.

Yes, all devices are on ‘Homey Pro’ Thread network, working well and behaving as expected.
I check on Dev Tools/Matter every now and then just to see that everything is as configured.
After I “rebuilt” my Matter network is has been rock solid as far as I can see even though it isn’t that big (only 11 devices, all from IKEA).

(Link to Thread network question thread)

Hi Ulrik,

Hi Ulrik,

We exchange a lot privately, and I genuinely appreciate your help and knowledge.

However, I cannot really describe the current situation as a “working state.” My Homey still reboots because devices continue to become unavailable. The fact that most devices may be available at the exact moment when I generate the diagnostic report does not mean that the underlying problem has been resolved.

Hi. I gave it another go. Got my onvis plugs in without issue and my nuki lock. So far they are on, working and stable. So on my part i can say so far yes its an improvement and hopefully it will continue in that direction

These are all matter thread devices but the only I have issues with Tado valves. im contacting support for that as Im Not convinced this is a homey issue.

As I’m a total novice on Matter/Thread networking I just asked Google Gemini if it makes any difference id devices are on different Thread networks:

Does it matter if they are on different Thread networks?

Yes, it matters for the efficiency and reliability of your smart home. Having separate Thread networks splits your devices into isolated groups, causing several real-world disadvantages:

  • Weaker Mesh Network: Thread is a mesh technology; every mains-powered device (like a smart plug or light switch) acts as a repeater to extend signal range. If your devices are split into two networks, they cannot repeat signals for each other, resulting in smaller, weaker networks with shorter range.
  • No Battery-Device Redundancy: Battery-powered devices (like motion or door sensors) rely on nearby mains-powered Thread routers to pass their signals to a hub. If a sensor is on “Network A” but its closest smart plug is on “Network B”, the sensor cannot use that plug to talk to the hub.
  • Increased Wireless Congestion: Multiple Thread networks mean multiple master controllers transmitting signals simultaneously. This creates unnecessary traffic on the 2.4GHz radio frequency spectrum, potentially causing interference with your Wi-Fi.
  • Slight Speed Latency: If two devices on the same Thread network talk to each other, the signal stays purely within the fast, low-power mesh. If they are on different networks, the signal must travel from the first device to its Border Router, jump through your home Wi-Fi/Ethernet router, go to the second Border Router, and finally go down to the second device.

Coincidence or not: my Matter network became stable when all devices are on the “Homey Pro” segment.
So maybe (take this from a newbie) it’s woth checking and aligning devices to be on “Honey Pro” network?

Thank you for going and checking, and for looking it up afterwards. And drop the novice disclaimer, I had an AI read the whole changelog for me before I posted mine. It is just a better Google search when the subject is technical, it still needs checking afterwards.

And Pascal, you are right that I aimed that at the wrong person. I did not check it properly before posting, sorry for mentioning you.

I hope I can help put an end to all this endless discussion:

Homey Pro/mini/HSHS 13.3.0 is officially Matter 1.5 certified

How does that solve anything?

The problem lies more with the person sitting in front of the screen than with the ‘incompetent’ (sorry, Athom) developers. After all, certificates aren’t just handed out to the first comer.

From what I’m reading about the certification process, it mostly tests for protocol conformance. I’m sure Homey is very good at conforming to the Matter protocol. It just seems to have a very hard time maintaining stable device connections, which is hardly a PEBKAC issue.

That is certainly good news, but it does not end this discussion.

Matter 1.5 certification confirms that Homey’s submitted implementation complies with the applicable Matter 1.5 requirements and has passed the required certification tests. It shows that Homey has correctly implemented the certified parts of the standard and provides a defined level of interoperability.

However, it does not necessarily mean that every Matter 1.5 device type and every optional capability is already fully supported by Homey. Nor does it prove how the implementation behaves in a large and complex real-world installation over an extended period.

The certification test is obviously far more meaningful than a YouTube review in which someone unboxes a contact sensor, pairs it with a Homey one metre away, opens and closes it twice and concludes that it “works perfectly and is super fast” — probably with almost no other devices connected to that demonstration Homey. :wink:

Nevertheless, certification remains a defined conformance and interoperability test. It does not automatically validate:

  • long-term stability;
  • the number of devices that can be handled reliably;
  • radio range and performance in a real home;
  • recovery after communication failures;
  • the stability of a large Matter fabric or Thread network;
  • or the behaviour of subscriptions over days and weeks.

Therefore, Matter 1.5 certification is a positive and important achievement, but it does not answer the stability questions being discussed in this thread. In particular, it does not explain why devices repeatedly become unavailable or why rebooting Homey makes them return.

So no, the certification cannot simply stop this discussion.