Hi Abe,
Here is a simple evidence for your question. I have only one Thread Mesh and all TBR are in the same Network.
I can send you the copy without anonymised information per PM if you need them.
Which is severely impacted by Homey Pro’s Thread chip and antenna, which isn’t powerful enough in some scenarios.
I don’t understand why we are dancing around this. Well I do understand from a business perspective, but it’s becoming increasingly tiresome at this stage.
Sorry to keep going on about it, but with the same network stack and the exact same devices as before (no new firmware updates), but using a different Thread implementation to what I have before with the ZBT-2, I have no Thread issues whatsoever.
Need I remind you?
All the different symptoms are clearly laid out in my message. If you’re going to assign blame, please bring the receipts.
@Doekse The complete scan has found 20 Extender Devices and 44 End Devices, all reachable.
Which Thread Router is currently acting as the primary Thread Border Router? Do note, the Thread Tools app doesn’t always use clear or recognizable names, so it can take a bit of detective work to figure out which physical device it actually refers to.
Normally the one on top near the network sign and that is Homey.
You mean every TBR. The standard says 250.
That’s not the case here. On my network, the current Thread Leader is Nest Wifi Mesh #F4D5, while my OpenThread Border Router is listed at the top of the Thread Tools app.
What do you mean with The standard?
I agree. However, during more than a year with Apple Home, the Thread topology must also have changed many times. An Apple TV was occasionally disconnected, HomePods were unplugged for cleaning, and individual Border Routers joined and left the network.
As a project manager, my first question when a previously stable system becomes unstable is always:
What changed immediately before the problems started?
In this case, the most obvious difference is clear:
That does not by itself prove that Homey is the sole cause, but it makes Homey’s participation, Matter implementation and interaction with the other Border Routers an obvious area to investigate.
I can also go back to my earlier troubleshooting.
At one point, I disconnected every Apple Thread Border Router and restarted Homey. The unavailable devices returned. Because this worked three or four times, I initially concluded that the Apple Border Routers might be causing the problem.
However, the same procedure later stopped working.
There were situations in which the devices did not return to Homey even after several Homey restarts. Between restarts, I sometimes waited for more than an hour to give the Thread network enough time to rebuild and stabilise itself.
These situations occurred repeatedly over time.
I even left all Apple Thread Border Routers disconnected for more than a week, despite having to explain to my wife why the Apple TV and HomePods were unavailable. During that period, Homey still lost devices at least once a day, and sometimes more frequently.
This is why I no longer believe that the mere presence of Apple Thread Border Routers explains the problem.
They may influence the topology or expose a weakness, but the instability also occurred when Homey was the only available Thread Border Router. That suggests that the root cause lies deeper than simply having multiple Border Routers on the network.
Can you explain how I can find the leader ?
You were talking about the possibility that Homey was not the Thread Leader and that another Thread Border Router may have taken that role.
Does Homey log its own Thread role in the diagnostic report, including whether it was acting as Leader, Router or Child?
Does the report also allow you to see whether the Leader changed frequently, and whether such a change occurred shortly before the devices became unavailable?
If this information is only available for the moment when the report is generated, it does not help much with an intermittent problem. This is exactly why a long-term diagnostic mode would be useful: it could record role changes, Leader changes and device failures over time, allowing you to verify whether the unavailable states consistently occur after a change of Leader.
Without that historical correlation, the idea that another Leader or Border Router caused the problem remains only a possibility, not a conclusion supported by the report.
“Standard” was not the most precise term. I was referring to the official statement made by the Thread Group for smart homes: more than 250 devices in a low-power wireless mesh network.
I am looking at this from the perspective of an ordinary customer.
Many users buy one Homey, one Google Nest Hub or another single smart-home hub. They do not know how many Thread Border Routers are involved, how Router and End Device roles are assigned, or how several Border Routers may cooperate within a network.
When the official Thread Group tells consumers that Thread can connect more than 250 devices in a smart home, the natural and reasonable interpretation is that a Thread-based smart-home system should be able to handle a network of that size.
The statement does not say:
“More than 250 devices, provided that the user installs several Thread Border Routers, carefully distributes the load and understands the internal topology.”
It simply advertises more than 250 devices for the smart home.
Of course, this is not a strict legal requirement forcing every manufacturer to support exactly 250 Matter devices under every possible condition. Different products have different hardware, software and practical limitations.
However, it is a very strong guideline and creates a legitimate customer expectation.
The Thread architecture itself goes considerably further, with up to 32 Routers and an address space of up to 511 End Devices per Router. The theoretical maximum is therefore 16,384 nodes. I fully understand that this is an architectural ceiling and not a realistic smart-home target.
That is precisely why the Thread Group communicates the much more practical figure of 250+ devices for homes.
Therefore, when Athom suggests that Homey might already be reaching its limits in a network containing far fewer devices, it is not sufficient to refer vaguely to “limits.”
Athom should clearly state:
how many Matter-over-Thread devices Homey Pro is designed and tested to support reliably;
whether that figure assumes Homey is the only Thread Border Router;
whether additional Border Routers are required or merely optional;
and which specific resource becomes the limiting factor.
If Homey’s practical limit is substantially below the 250+ devices promoted for Thread-based smart homes, that limit should be clearly documented before purchase, not discovered by users only after their systems become unreliable.
Sorry, somewhat off-topic, but may be relevant if Homey hits a hardware or software limit.
Could you run Homey.thread.executeCommand({ command: 'leaderdata' }) in the Web API Playground, which you can find in the Homey Developer Tools? This will return the router ID of the current Thread Leader.
You can then run Homey.thread.executeCommand({ command: 'router table' }) to get an overview of all Thread Routers and their RLOC16 addresses. These can be cross-referenced with the information shown in the Thread Tools app to identify the leader device.
I don’t disagree with the Thread Group’s statement at all. Thread is indeed designed to scale to well and theoratically support many devices.
The important part, however, is that the Thread Group is describing the capabilities of a Thread mesh network, not guaranteeing that every possible deployment with 250 devices will perform identically. The same documentation also explains that Thread networks consist of End Devices, Routers and one or more Border Routers, and that scalability depends on having sufficient routing infrastructure and a healthy mesh.
In practice, the achievable size of a Thread network is still influenced by many of the same factors we’ve been discussing throughout this thread: the RF environment, network topology, the placement and number of Routers and Border Routers, the mix of mains-powered and battery-powered devices, reporting intervals, the amount of traffic being generated and the overall quality of the mesh. Two homes with exactly the same number of Thread devices can therefore behave very differently.
That’s also why we don’t publish a single maximum number of Matter-over-Thread devices for Homey. There simply isn’t one that would be technically meaningful across all installations. The limiting factor is rarely just the number of devices; it’s the overall characteristics of the network.
The leader is an Apple TBR. Homey has the role of a router.
Now the monkey comes from the sleeve we Dutchies would say. Which exact Apple device is the Leader? If this is a WiFi connected device, such as your Homepods, that could certainly be an interesting lead here.
The current Thread Leader is an Apple TV connected to the network through Ethernet.
I understand that “now the monkey comes out of the sleeve” means that the real explanation has finally become visible. But what exactly do you believe this finding explains?
It certainly explains why some Thread traffic reaches Homey through its eth0 interface. Homey is the Matter Controller, while the Apple TV is currently acting as the Thread Leader and as one of the Thread Border Routers. Traffic reaching Homey through another Border Router is therefore expected behaviour in a Thread network with multiple Border Routers.
However, this does not yet explain why Matter devices become unavailable.
The fact that the Apple TV is currently the Leader is only a snapshot. To establish a connection with the failures, we would need to know whether:
the Leader changed shortly before devices became unavailable;
Homey changed its Thread role;
the Thread partition changed or split;
routing through the Apple TV actually failed;
or the same failures also occurred while Homey was the Leader.
Without such a correlation, identifying the Apple TV as the current Leader explains the route taken by some traffic, but not the reliability problem itself.
I would like to return to the diagnostic logs.
Do Homey’s logs show which route was being used for the affected devices before they became unavailable? More specifically, is there evidence that their Matter traffic was reaching Homey through eth0 at the time the failures occurred?
Without that information, linking the problem to the Apple TV is still only speculation.
We should also remember that the same devices previously operated reliably in an environment consisting entirely of Apple Thread Border Routers. Therefore, the mere involvement of an Apple TBR cannot by itself explain the instability.
The “monkey” could just as easily be a problem on Homey’s side when handling Matter communication arriving through eth0, rather than through its own Thread radio. But that is also speculation.
To distinguish between these possibilities, the logs would need to show:
which Border Router or route was used by each affected device;
whether that route changed shortly before the failure;
whether communication through eth0 stopped or produced errors;
and whether the same devices failed while communicating through Homey’s own Thread interface.
Without such historical routing data, we are not identifying a root cause. We are only guessing based on the current topology.
And finally, all devices are currently available, at the Moment of test.
Okay, that’s good. In that case Wi-Fi does not appear to be the issue here. But the network itself still might be a variable.
What you could try is creating a HomeyScript that runs every minute or so, retrieves the Thread Leader ID, and stores it in a variable or log. This would allow you to track whether and when the leader changes, and potentially correlate those changes with the issues you are experiencing.
I do not think this is a reasonable request.
I am willing to provide diagnostic reports, describe the failures precisely and perform clearly defined troubleshooting steps. However, I am not willing to develop and maintain a custom monitoring solution to compensate for missing diagnostic functionality in Homey.
Homey’s Thread stack already knows the current Leader, its own role, the partition and other relevant network information. If Athom believes that Leader changes may be related to the failures, these events should be logged by Homey with timestamps and included in the diagnostic data.
Polling one value every minute with a user-created HomeyScript would also be technically inadequate. It could miss changes between polls, and I would still have to develop the script, organise persistent data storage and analyse the collected results myself.
I purchased Homey as a product. I am not employed by Athom to design and implement the diagnostic tools required to investigate its reliability problems.
If Athom provides a ready-to-use diagnostic tool, whether as a Homey app or a script that can simply be installed and started without requiring me to write code or organise the data storage myself, I am willing to run it and provide the resulting data.
The fact that a user-created script is being suggested instead rather confirms the point I have already made: the current diagnostic tools are insufficient for investigating intermittent Matter and Thread failures.
I don’t think that’s an entirely fair characterization.
The suggestion to periodically query the Thread Leader wasn’t intended to replace Homey’s diagnostics or ask you to develop a permanent monitoring solution. It was simply a practical way to validate a specific hypothesis: whether Thread Leader changes correlate with the moments devices become unavailable.
While Homey indeed knows the current Thread Leader, the standard diagnostics are intended to provide a broad snapshot of the system rather than continuously log every aspect of the Thread network over an extended period. Logging every topology change indefinitely would come with storage and performance implications, especially on an embedded device.
A simple HomeyScript is therefore not being proposed because Homey lacks this information internally, but because it’s a lightweight way of gathering additional data for this particular investigation without requiring custom firmware or engineering builds.
I also agree that polling once per minute wouldn’t capture every transient Leader change. However, it doesn’t necessarily have to. The goal isn’t to produce a complete history of the Thread network, but to determine whether there is any meaningful correlation between larger topology changes and the reported outages. Even an imperfect correlation can significantly narrow the scope of an investigation.
Every diagnostic tool has limits. The current diagnostics have already helped identify several useful observations, such as indications of multiple Thread networks, communication over eth0, channel access failures and incomplete Thread transactions. They haven’t yet identified a definitive root cause, but that doesn’t necessarily mean they’re insufficient; it simply means the available data doesn’t yet point conclusively to one explanation.
If you want me to perform this test, please provide the ready-to-run script and exact instructions for logging, storing and exporting the data. I will run the tool and return the results.
B.T.W. The HUE Matter bridge is also connected over eth0