We play No Man’s Sky co-op from two machines in the same flat, my son on Windows and me on a Mac, and I set up every session myself. On the first one Windows asked whether nms.exe should be allowed on the network, and because he plays under a standard account the dialogue wanted an administrator password. I typed mine in. We played for about ten minutes, and then his session disappeared from my game and would not come back until we both restarted.
It happened again on the next session, and again after that. Everything those two evenings turned up was real, I fixed all of it, and the session kept dropping.
The thing in the way sat two layers of address translation above my flat, at my internet provider, and no amount of competence on my side of the wall could reach it. Each of those fixes was genuine work on a genuine defect, which is exactly what made them so easy to count as progress. The only progress is movement in the symptom.
Defaults that assume an administrator is at the keyboard
The password prompt was the first thing to look at, and it turned out to be hiding two separate defects.
Windows Firewall’s “do you want to allow” dialogue writes two rules when you answer it, and both defaults are poor. It scopes them to the Private profile alone, and it sets edge traversal to DeferToApp. The first matters because the Ethernet interface passes through Public for a few seconds on every boot and every network re-identification, and inbound to the game is dropped in that window. The second matters because edge traversal is precisely the unsolicited inbound that peer-to-peer hole punching depends on.
Reading the firewall’s own event log made the evening legible. At 19:46 the Xbox app re-registered the game’s package after an update. At 19:55 the firewall service created the rules with the action set to block, which is what it holds while the dialogue waits for an answer. Twelve seconds later my administrator password flipped them to allow. Two hours after that the game asked separately for edge traversal, and the dialogue came back a second time.
[past-max]The same machine was already carrying the evidence of how this ends. Two permanent block rules sat there for viber.exe under another family member’s account, left behind when somebody dismissed the same dialogue months earlier. Inbound Viber had been silently dead ever since and nobody had connected the two.
- scoped to the Private profile only, so a boot that passes through Public drops inbound
- edge traversal left at DeferToApp, so NAT hole punching is not permitted
- held as an explicit block rule for as long as the dialogue is unanswered
- becomes a permanent block if anyone clicks cancel, with no further prompt
- needs an administrator physically present the next time it appears
- scoped to every profile, so the transient Public window changes nothing
- edge traversal set to allow, which is what peer connections require
- written once, in advance, with the firewall policy exported first
- survives dismissal because there is nothing left to dismiss
- the standard account launches the game and nobody is asked for anything
The replacement is two rules written once, New-NetFirewallRule -Program <the exe> -Profile Any -EdgeTraversalPolicy Allow, one for TCP and one for UDP, with the firewall policy exported first so there is a way back. Binding to the package identity instead would be tidier and does not work: Game Pass titles are full-trust packaged applications with no AppContainer SID, so a package-scoped rule matches none of their traffic. The executable’s path is the right anchor because the Xbox app pins its install root.
The prompt never came back. We dropped again the next day.
What a packet count settles that a theory cannot
The physical layer went first, because a previous incident cost me most of a Saturday for want of checking it. The Windows machine’s link was up at a gigabit with zero errors and zero discards, and there had been no link-down event in fifteen days. My Mac sat on Wi-Fi two rooms away; a hundred and fifty pings between them lost nothing.
[crossref]Then I captured packets, which I should have done on the first evening. pktmon ships with Windows, takes a filter on a single address, and needed twelve seconds to end the guessing. I ran it from my Mac over SSH while the two of us were standing next to each other in the same save, which is the only reason the sample is worth anything.
Two dialogues, not one repeated
The channel is localised, so matching on the message text finds nothing
and reading the XML finds everything. Decoded, Action 2 is
block and 3 is allow, and the process that made the change is the tell:
svchost.exe is the firewall service acting on its own,
dllhost.exe is a person answering the dialogue. Two
dllhost entries, nearly two hours apart, are two separate
questions rather than one impatient repeat.
Asking the same machine for every inbound block rule takes one line and returns things nobody remembers agreeing to. Two of them had been cutting inbound Viber for months under an account that never opens a terminal.
Both defects were real and both were fixed. By the next evening the machine held zero inbound block rules and the prompt never came back. The session dropped anyway.
One flat, one switch, and nothing between them
Two machines on the same switch, in the same flat, with both of us standing in the same place in the same game. Over twelve seconds the only things they said to each other were a device-discovery broadcast and Syncthing announcing itself on the local network. The game exchanged nothing at all.
That single count did more than every theory I had built in two evenings. It ruled out the firewall, the wireless link, the switch and the cable in one measurement, because none of those can be the reason two peers never address each other in the first place. Whatever the game was doing, it was doing it somewhere out on the internet.
I had also convinced myself an hour earlier that the two machines were talking, because the Windows neighbour table kept my Mac’s entry in the reachable state across three consecutive samples. That was my own SSH session holding the entry warm. The lesson is cheap and I keep relearning it: when you are the one connected to the machine, you are part of its traffic.
The address my router believes in is not mine
The router reported its WAN address as 100.66.81.157. That
range, 100.64.0.0/10, is reserved for exactly one purpose:
providers who have run out of public addresses and put their subscribers
behind a second layer of translation. The address the internet sees for me
was a different one entirely, and a traceroute showed a hop between the two
that declines to identify itself.
The three connection tests say the rest. Across the LAN, fine. Out to my own router’s WAN address and back in, fine, which means my router performs hairpin translation correctly. Out to the public address and back, nothing. The provider’s translator will not turn my own packets around and hand them back to me.
That is the whole failure. The game’s matchmaking sees two players arriving from one public address and gives each of us the other’s public endpoint, because from the outside that is what we look like. Both of us then try to reach a door that only opens from a side neither of us is on.
Which is why the session goes out to Hello Games’ own relay instead, and why we get to play at all. That relay session hangs off a UDP mapping inside the carrier’s translator, and when the translator rebuilds its mappings the session goes with them and does not come back on its own. Ten minutes is how long one of them lasted.
The boundary of what you can fix is a line on a tariff
My router had done its part. Its UPnP daemon held live mappings for the game, labelled Microsoft Multiplayer, forwarding UDP 3074 and 47601 to the Windows machine. Every one of them was decoration. A port opened on your own router is opened on the wrong router when there is a carrier translator above it, and nothing on my side of that translator can change what it does.
I went looking for a public address on the fibre line. The provider sells one, on its own site, as an add-on. My account cannot buy it, and a subscriber on the same plan explained why in one line on a forum: the combo tariffs have no dedicated IP. Bundling television, internet and mobile onto one promotional price removed the option, nothing in the account says so, and getting it back means dissolving the contract and signing a new one.
The backup line was translated too, in an ordinary private range. That provider sells the public address for about a euro a month, and it appeared minutes after I bought it. The router’s WAN address and the address the internet sees became the same string, which is the entire technical content of the fix.
I could not have engineered my way to that. It was a checkbox on somebody else’s billing page.
The checkbox cost me a perimeter. Behind the carrier’s translator, inbound traffic from the internet could not reach my LAN at all, and I had been leaning on that deliberately: it is the one piece of protection you cannot misconfigure, because you never configured it. A routable address ends it. Nothing is forwarded through today, but the assumptions in my own security notes were written on the far side of that line, and going back through them is a job now rather than a formality.
A capability arrives configured by whoever delivered it
Unlike the fibre provider, the other one also hands out IPv6. My router had been sitting on Native with prefix delegation enabled for months with nothing to delegate: no global address on any client, no default route, ping6 answering no route to host.
The moment a prefix arrived, the router started advertising itself to every device as the IPv6 resolver. Its own resolver forwards to the provider. macOS follows the first nameserver it is given without arguing, so every internal hostname in the house began resolving through the provider to a public wildcard record, and the Pi-hole that filters advertising and telemetry for every device fell out of the path entirely. Nothing alerted. It surfaced because a phone could not open Home Assistant.
- ● 01clienttakes the first nameserver it is handed
- ● 02routeradvertised itself the hour a prefix landed
- – 03Pi-holenever asked
- ● 04providerwhere the router forwards instead
- ✕ 05wildcardevery internal name, one public answer
The fix I reached for was to pin the advertisement to the Pi-hole’s own IPv6 addresses. I flagged the prefix’s instability in the same line as the recommendation, and recommended it anyway. The prefix moved twice inside a few hours, so the addresses I had written down were dead before the evening was. A risk you name and do not design around is just a risk you have written down.
[max]IPv6 is off for now. Stock firmware on this router has no setting that stops it advertising itself, which is a documented bug rather than a thing I missed, and a community firmware fork does expose the control. Flashing a fork onto the one box the whole household routes through, for a capability nothing currently uses, is a bad trade. The only reason IPv6 was ever switched on was a smart bulb that refused to pair, and a Zigbee stick made that problem disappear months ago.
The public address and the uptime are on different wires
There is a gateway appliance in my backlog, wanted for a firewall at the edge and failover between the two lines. Its value moved this week, for a third reason that only appeared once the public address landed on the wrong wire.
The public address now lives on the backup line. That line is twisted pair, and the switch it depends on runs off a modest battery in the next stairwell, which in a city with scheduled blackouts makes it the first thing to go dark. The fibre is what stays up. So the connection I need for uptime and the connection I need for a reachable address are two different wires, and neither is sufficient alone.
That is a routing problem, and consumer firmware solves the wrong half of it. It offers failover, which sends everything down one line when the other dies. What this needs is a policy: the gaming machine leaves through the line with the public address, the rest of the house leaves through the fibre, and each keeps the property it was chosen for. Owning the gateway is what makes that expressible.
The router still resolves through the provider rather than through the Pi-hole, so anything that falls back to it loses filtering. I made that check state-tracked rather than a standing alert, on the reasoning that no client uses the router as a resolver. Then I found the Windows machine doing exactly that over IPv6, which means the premise I built the check on was already false when I wrote it.
While we play, the whole flat is flipped onto the backup line at the switch in the cabinet, so both machines sit behind the public address. We played for two hours and the session held.