Backup Internet Connection: How to Test a Real Failover
Unplug the modem. Watch the laptop's status icon stutter, blink, and come back a few seconds later on the second line. That is the test almost everyone runs, and it establishes the one thing nobody actually doubted: that the router will eventually notice.
The thing you bought the backup for is different. It is the meeting at 10 a.m. where the person on the other end never learns that anything happened. And whether that works turns on two entirely separate events that the words "automatic failover" collapse into one. First, the router has to decide the primary line is dead and move traffic. Second, the conversation already in flight has to survive being moved. A product page will describe the first and let you assume the second.
They can fail independently, and they usually do. Set what the router vendors publish about their own failover behaviour against what the protocol specifications say about sessions crossing an address change, and the gap between the two is where a backup line quietly becomes a thing that keeps email working while your calls still drop. Below is what those documents say, the arithmetic buried in the shipped defaults, and a test that takes forty minutes and one cooperative colleague.
The router documentation uses the word "killing"
Start with the vendor's own description of the mechanism, because it is more honest than any review.
OPNsense has a per-gateway option called Failover States. The gateway documentation, read 29 September 2026, explains it in one sentence: "If this gateway goes down, force clients to reconnect over a different online gateway by killing states associated with this gateway."
Killing. That is the firewall's term for discarding the connection-tracking entries that make your existing sessions work. The option exists because the alternative is worse — without it, clients keep trying to use a path that is gone — but read plainly, it says that a successful failover deliberately terminates the connections that were open when the line died. The design intent is that applications reconnect. Some applications reconnect invisibly. A call does not.
pfSense exposes the same behaviour from the other end. Its gateway groups page, which the site says was last updated on 19 December 2025, offers Keep Failover States with two choices, and the description of the second one is the sentence to sit with: "States created by policy routing rules using this gateway group are killed when a higher-priority gateway returns to an online state."
Note the direction. That is not about the outage. That is about the recovery. Choose the wrong setting and a call that survived the failover gets cut when the fibre comes back, which produces the single most confusing symptom in this whole area: the meeting drops at the moment the internet is repaired. The other option, "Keep states on gateway recovery," leaves those connections on the backup until they end naturally. Both are defensible. Only one of them matches what you want during a meeting, and nothing about the outage itself tells you which one you have configured.
Then there is the setting that decides whether any of this happens at all. OPNsense's documentation states that "by default the system only chooses a (new) default gateway on startup or when an interface is connected or disconnected," and that if you want the default gateway to change when the current one stops being reachable, "you can enable 'Gateway switching' in System->Settings->General." In other words: monitoring can be configured, alerts can be firing, the second WAN can be perfectly healthy, and traffic still will not move, because the switch that moves it is a separate checkbox that defaults to off. A modem that is unplugged at the wall does trigger a link-down event and will fail over. An upstream outage that leaves your Ethernet link up and your gateway unreachable is the common case, and it is exactly the case that default configuration does not handle.
The two settings are also coupled, which is easy to miss. The same page's description of Failover States — the state-killing option above — adds that it "requires 'default gateway switching' to be enabled." So the option whose whole purpose is to push your sessions onto the surviving line is itself inert until you turn on the checkbox that nobody mentions. Enabling one without the other gets you a router that has an opinion about your dead gateway and does nothing with it.
If you own a dual-WAN router, this is the first thing to check, before any measurement. A backup that has never engaged is indistinguishable from a backup that engages perfectly, right up until the day you need it.
Twelve seconds of arithmetic sitting in the shipped defaults
Detection is not instant, and the shipped numbers tell you roughly how slow it is before you measure anything.
pfSense's gateway configuration documentation, last updated on 29 August 2025 and read on 29 September 2026, lists the monitoring defaults for the dpinger daemon:
| Setting | Default | What it governs |
|---|---|---|
| Probe Interval | 500 ms | How often a ping goes to the monitor address — "the default is to ping twice per second" |
| Loss Interval | 2000 ms | How long before a probe counts as lost |
| Time Period | 60000 ms | The window over which results are averaged |
| Alert Interval | 1000 ms | How often the daemon checks for an alert condition |
| Latency Thresholds | From 200 ms, To 500 ms | Warning at the lower, gateway down above the upper |
| Packet Loss Thresholds | From 10 %, To 20 % | Warning at the lower, gateway down above the upper |
Now multiply. Two probes a second across a sixty-second averaging window is 120 probes. A line that has gone completely dark stops answering all of them, so the averaged loss figure climbs by roughly 1.7 percentage points per second. To exceed the 20 percent threshold that marks the gateway down, about 24 probes have to be missing — call it twelve seconds of total outage. The documentation also notes the constraint that keeps this honest: "The Time Period must be greater than twice the sum of the Probe Interval and Loss Interval, otherwise there may not be at least one completed probe."
Twelve seconds is a long time in a conversation. It is long enough for everyone to say "you're frozen," long enough for a person to hang up and rejoin, and long enough that a media session will have given up on its own.
There is a lever. pfSense's gateway group settings include a Trigger Level, and the default, Member Down, is described as marking the gateway down "only when it is completely down, past one or both of the higher thresholds" — which the same page admits "may miss more subtle issues with the circuit that can make it unusable long before the gateway reaches that level." Choosing Packet Loss instead marks it down when loss crosses the lower alert threshold, 10 percent by default. Twelve lost probes instead of twenty-four. On the same arithmetic that is about six seconds rather than twelve.
Treat all of that as an estimate of the order of magnitude, not a measurement. Averaging implementations differ, your box may not be running defaults, and the only figure that governs your Tuesday morning is the one your own stopwatch produces. But it does tell you what regime you are in. Sub-second failover is not something a loss-averaging ICMP monitor with these defaults produces, and no amount of tightening the probe interval alone will get you there while the averaging window stays at a minute.
For contrast, the protocol built specifically to detect forwarding failures fast is RFC 5880, Bidirectional Forwarding Detection, published in June 2010. Its introduction is a direct complaint about exactly this problem: networks use "relatively slow 'Hello' mechanisms" and "the time to detect failures ('Detection Times') available in the existing protocols are no better than a second, which is far too long for some applications." BFD's detection time is not averaged at all — it is, per section 6.8.4, the detect multiplier times the negotiated transmit interval, where the multiplier is "roughly speaking, due to jitter, the number of packets that have to be missed in a row to declare the session to be down." Consecutive misses, not a percentage of a minute. That is the difference between a design that expects to fail over during a phone call and a design that expects to fail over during a workday, and a consumer dual-WAN box running default gateway monitoring is the second one.
Your public address changes, and that is the part the call notices
Suppose detection is fast and the router moves traffic in two seconds. The call can still end, for a reason that has nothing to do with the router.
Traffic leaving by the second WAN carries a different source address. And the identity of a connection includes that address. RFC 9293, the 2022 consolidated TCP specification, puts it in the glossary: a connection is "a logical communication path identified by a pair of sockets." Change one socket and it is not the same path any more — it is a new connection that the far end has never heard of, and the old one sits there until something times it out. No router setting repairs that, because there is nothing wrong with the router.
Real-time media has a tighter clock on top. Modern conferencing runs over WebRTC, and WebRTC requires continuous proof that the far end still wants your packets. RFC 7675, STUN Usage for Consent Freshness, establishes what it calls "a 30-second expiry time on consent," chosen "to balance the need to minimize the time taken to respond to a loss of consent with the desire to reduce the occurrence of spurious failures." The operative rule is unambiguous:
Consent expires after 30 seconds. That is, if a valid STUN binding response has not been received from the remote peer's transport address in 30 seconds, the endpoint MUST cease transmission on that 5-tuple.
To keep consent alive, the document says implementations "SHOULD set a default interval of 5 seconds," randomised between 0.8 and 1.2 times that, and "MUST NOT set the period between checks to less than 4 seconds." Those checks belong to a 5-tuple. Your failover changed the tuple. So the thirty-second timer is now running on a path that can no longer be validated, and at the end of it the endpoint is required to stop sending.
Media flows are allowed to move, but only through a specific renegotiation. RFC 8445, the 2018 ICE specification, reserves changing where a stream is sent to an ICE restart, and section 9 says an ICE restart "causes all previous states of the data streams, excluding the roles of the agents, to be flushed," and that "to restart ICE, an agent MUST change both the password and the username fragment for the data stream(s) being restarted." That is a signalling exchange between the two clients and the service, not something your router can do on their behalf. Whether your conferencing app performs one promptly, and how long it takes, is a property of that app — which is why two people on the same home network can lose a call and rejoin at visibly different speeds.
Three things genuinely are built to survive an address change, and knowing which of them you are using explains most of the variation people see:
- QUIC. RFC 9000 opens section 9 with the design goal stated outright: "The use of a connection ID allows connections to survive changes to endpoint addresses (IP address and port), such as those caused by an endpoint migrating to a new network." It is not free — an endpoint "MUST perform path validation ... if it detects any change to a peer's address, unless it has previously validated that address" — and a server can forbid migration outright by sending the
disable_active_migrationtransport parameter. But a QUIC-based download or stream can cross a failover intact where a TCP one cannot. - Multipath TCP. RFC 8684 specifies TCP that runs over several paths at once, and is explicitly designed so that "MPTCP's connection will stay alive at the data level, in order to permit break-before-make handover between subflows." It requires support at both ends, which in practice means a small number of services.
- An IPsec VPN running MOBIKE. This one matters for anyone whose work day lives inside a corporate tunnel. RFC 4555 exists because of a limitation it states plainly about base IKEv2: "Currently, it is not possible to change these addresses after the IKE_SA has been created." MOBIKE is the extension that lifts that restriction, letting "a mobile Virtual Private Network (VPN) client keep the connection with the VPN gateway active while moving from one address to another," and letting "a multihomed host move the traffic to a different interface if, for instance, the one currently being used stops working." Without it, your tunnel drops at failover and takes every application inside it down with it, including the softphone.
This is also the honest explanation of why bonding products cost what they cost. Peplink's SpeedFusion page, read 29 September 2026, describes its Hot Failover as working differently from "traditional failover methods that wait for a link to drop before reacting," and claims it "ensures that active sessions remain online without interruption." That is the vendor's own marketing claim and not a measurement, and I have not tested one. But the architecture behind such claims is not mysterious: traffic is wrapped in a tunnel to a remote endpoint, so the address your applications see belongs to the tunnel and never changes when the underlying link does. The session survives because, from the application's point of view, nothing moved. Any product that promises calls surviving failover is doing some version of that, and any product that does not is promising you the first event only.
Two lines, or one line with two bills
None of the above matters if both connections die together, and this is where money gets spent on redundancy that is not redundant.
Netgate's multi-WAN documentation, read 29 September 2026, states it without hedging: "Two connections of the same type cannot be relied upon to provide redundancy in most cases. An ISP outage or cable cut will commonly take down all connections of the same type." It goes further on the geography, warning that even two different technologies from one large provider, while they "usually traverse significantly different networks until reaching core parts of the network," commonly "utilize the same cable path," which "still leaves a site vulnerable to extended outages from cable cuts." The same page opens with the practical version of the point, written by people who have watched it happen: it is "highly desirable to obtain connectivity choices for a multi-WAN deployment which utilize disparate cabling paths."
So the question is not whose logo is on each bill. It is what the two services have in common:
| Shared element | The failure it fails to survive | How to check in ten minutes |
|---|---|---|
| Same cable, pole or conduit to the street | A cut, a car into a pole, a contractor's backhoe | Look at where each service physically enters the building; a coax and a phone pair on the same messenger wire share one accident |
| Same cell tower | A sector outage, a tower on generator, congestion during the local event | Fixed wireless and a phone hotspot on the same carrier are usually the same radio; put the backup on a different carrier |
| Same upstream transit | A routing problem you cannot see or fix | Run a traceroute on each line and compare the first four or five hops; if they converge immediately, so will the outage |
| Same power strip | The most common home outage there is | Both modems and the router on one small UPS, not on the same unprotected strip |
| Same account | A billing suspension or an accidental disconnect order | Two providers means two independent ways to stay connected while you argue with one |
The power line is not a throwaway. A backup connection that goes dark in the same second as the primary because a breaker tripped is a design flaw, not bad luck, and it costs less to fix than any of the others.
What a cellular backup is actually rated to carry
If the backup is cellular — a phone hotspot, a dedicated hotspot device, or a fixed wireless gateway — two published numbers decide whether it can hold a meeting, and neither is the headline speed.
The first is what the carrier says the service typically delivers, upstream. T-Mobile's expected speeds and performance metrics page, read 29 September 2026, breaks its 5G figures out by usage rather than by plan name, and the row that covers hotspots is separate from the row that covers phones. For "Temp Fixed Wireless & Hotspot/Tethering" it gives download speeds "typically between 118 – 402 Mbps," upload speeds "typically between 6 – 33 Mbps," and latency "typically between 15 – 27 ms." The page explains what those bounds mean: "These ranges are projections based on roughly the 25th and 75th percentiles of network tests." So the 6 is not a floor — roughly a quarter of the tests behind it came in lower — and the latency band is genuinely good, better than a lot of the DSL it might be backing up.
The same page also lists what moves you around inside those ranges, and one item on the list is worth reading twice. Among the factors affecting your experience it names "uses that affect your network prioritization, such as whether you are using Smartphone Mobile HotSpot (tethering) or if you are a Heavy Data User." Tethering is itself a prioritisation factor. Your backup traffic is not treated the same as the traffic from the same phone's own screen, by policy, all the time — not only when the tower is busy.
The second number is the one that actually ends careers on video, and it lives in a bulleted list on a support page rather than anywhere near a speed claim. Verizon's mobile hotspot FAQ, read 29 September 2026, states that after a plan's premium hotspot allowance is spent you keep working at a reduced rate, and spells the rates out: on the Simplicity plan, 10 GB and then "speeds reduced to up to 1 Mbps for the rest of your billing month"; on Unlimited Ultimate, up to 200 GB and then "unlimited data at 4 Mbps"; on Unlimited Plus, 30 GB and then "unlimited data at 3 Mbps when on 5G Ultra Wideband and 600 Kbps when on 5G/4G LTE."
Six hundred kilobits per second. Set that against what a meeting asks of the upstream and the backup has already failed before any router gets involved — that arithmetic, using the conferencing vendors' own per-endpoint figures, is in the piece on how much upload a working day really needs. A throttled hotspot is not a slow backup. For a call, it is no backup.
So find the equivalent sentence for your own plan today, while nothing is broken, and write two numbers on the same line: your allowance, and the speed after it. Then compare that allowance against a realistic outage. An eight-hour workday of video on a line that has to carry everything is not measured in hundreds of megabytes. If your allowance is 10 GB, one bad week can spend it, and the second half of the outage happens at the throttled rate.
Two adjacent problems are worth naming here, both covered elsewhere on this site. A fixed wireless gateway may be refused at your address even where coverage exists, for reasons of sector capacity rather than signal, and the refusal has a documentary value of its own — that is the qualification gate. And if satellite is the candidate for the second path, the recurring cost and the physical siting requirements are laid out in what satellite actually costs every month.
The procedure: a stopwatch, a continuous ping, and somebody willing to stay on the line
Everything above is documentation. This is the part that produces a number about your house.
Do it on a weekday, in a window when an outage costs nothing, and tell the person helping you what is about to happen. Before the test, establish a baseline for the backup on its own, using the same method you would use for any other line — the server-choice and timing rules that make a result reproducible are in the repeatable speed test method — because a failover that works perfectly onto a line that cannot carry the load has told you nothing.
Then run this:
- Start a timestamped continuous ping to a stable address, logging to a file, one packet per second. On Windows,
ping -t 1.1.1.1and on macOS or Linuxping -D 1.1.1.1will do; what matters is that every reply and every gap carries a time. Leave it running for the whole test. - Start a second ping, bound to the backup line if you can, so you can see that the second path was healthy the entire time. If the backup is a phone hotspot, a ping from the phone itself is close enough.
- Join a real call with your accomplice. Camera on, screen share running, in the application you actually use for work. A browser tab playing video is not a substitute; it has no upstream and no consent timer.
- Break the primary line the way it actually breaks. Not at the router's power switch. Disconnect the coax or fibre at the modem, or disable the WAN interface in the router, so the failure looks like an upstream loss rather than a device reboot. Note the wall-clock second you did it. That is t0.
- Record four times. t1, when the router's status page or log shows the gateway marked down. t2, when the ping log starts getting replies again. t3, when audio is actually usable again in both directions — ask your accomplice, because your own client will show its own recovery before theirs does. And t4, if it happens: the moment the call ended and had to be rejoined.
- Then test the recovery, which is the step everybody skips. Reconnect the primary and start the clock again. This is where "kill states on gateway recovery" shows itself, and where a call that survived the outage sometimes dies.
- Do the whole thing three times, because one run tells you nothing about variance and two runs tell you very little.
Record it like this, so the file is worth something in three months:
| Run | t0 cut | t1 gateway down | t2 packets flow | t3 audio back | t4 call dropped? | Failback gap | Notes |
|---|---|---|---|---|---|---|---|
| 1 | |||||||
| 2 | |||||||
| 3 |
One caution on interpretation. Pings coming back at t2 is the event most people stop at, and it is the weakest of the four. ICMP replies mean packets are moving; they say nothing about whether your sessions moved with them. The pair that decides whether the backup did its job is t3 and t4, and only the other person on the call can give you t3 honestly.
Reading the four timestamps, and which gap is yours to fix
Each gap points at a different thing, and the whole value of the exercise is that they point in different directions.
A long t0 to t1 is detection, and it is configuration. This is the twelve seconds of averaging from the section above. It lives in your gateway monitoring thresholds, the trigger level on your gateway group, and — if traffic never moved at all — in whether gateway switching is switched on. It is the most fixable gap on the list and the one that needs no purchase.
A long t1 to t2 is the backup coming up cold: a cellular modem that had no session established, a DHCP lease being acquired, a hotspot that had gone to sleep. The fix is to keep the backup warm rather than idle, by having something small and continuous use it even while the primary is healthy — a monitoring ping out that interface is enough. A backup that has to boot is a backup that arrives after the meeting.
A long t2 to t3, or a t4 at all, is the session layer, and no router setting in the world will shorten it. This is the 30-second consent expiry, the ICE restart your application either performs quickly or does not, the TCP connections identified by a pair of sockets that no longer exists, the VPN tunnel that did not have MOBIKE. Your options here are different in kind: a tunnel or bonding product that hides the address change, a conferencing client that renegotiates fast, or a corporate VPN configuration that supports mobility. If this is your big gap, buying a faster backup line changes nothing.
A drop at failback is a settings choice, and it is a one-line change. Decide whether you want states killed on recovery or kept, and set it deliberately rather than inheriting it.
And if the call quality on the backup is bad rather than absent, the problem is not failover at all — it is what the second line does under load, which is measured with the jitter and loss numbers rather than the throughput ones. Those thresholds, and how to log them continuously, are in the piece on what actually breaks video calls; and if the wireless leg in your own house is the variable, cut that possibility out first with the twenty-minute wired diagnosis before blaming either line.
One more thing worth saying plainly, because it is cheap and it works. Not every household needs automatic failover. A hotspot on the desk, already powered, already paired, with the laptop's Wi-Fi able to switch to it in about fifteen seconds, will get you back into a meeting faster than a badly configured dual-WAN router will — and it moves the decision to a human who can see that the call matters. Manual failover is not a failure of engineering. It is a legitimate design, and for one person in one room it is often the right one.
Book forty minutes this week and produce the three rows. Put the stopwatch times in a file next to your speed log, and add two things that take five minutes each: the post-allowance speed printed on your own cellular plan's support page, and a traceroute from each line saved side by side so you can see whether the two paths are really two. Then diary one recurring entry — a quarterly repeat of the same test — because carriers re-home cells, firmware updates reset trigger levels, and the only failover you can trust is one you have timed since the last thing changed.
Frequently asked questions
Will my video call survive an internet failover?
Only if something is deliberately preserving the session, because your public IP address changes when the traffic moves to the second line. A TCP connection is, in RFC 9293's own definition, a logical communication path identified by a pair of sockets, and a socket includes the address — change it and the connection is a different connection. For WebRTC media there is a hard clock on top of that: RFC 7675 sets a 30-second expiry on consent to send, and says that if a valid STUN binding response has not arrived from the peer's transport address within that window the endpoint MUST cease transmission on that 5-tuple. Some things do survive an address change by design, including QUIC via its connection ID and IPsec VPNs running the MOBIKE extension. Plain TCP sessions and an unrenegotiated media flow do not.
How long does a dual WAN router take to detect that the primary line is down?
Far longer than the marketing implies, and the arithmetic is in the defaults. pfSense ships gateway monitoring at a 500 ms probe interval with results averaged over a 60,000 ms window, and the packet loss threshold that marks a gateway down is 20 percent. Two probes a second over sixty seconds is 120 probes, so a total blackout has to accumulate about 24 lost probes, or roughly twelve seconds, before the average crosses the line. Switching the gateway group's trigger level to the lower 10 percent alert threshold roughly halves it. Those are documented defaults rather than a measurement of your own box, which is why the only figure that describes your own line is the one a stopwatch produces.
Is a phone hotspot good enough as a backup for working from home?
It depends on a number most people have never looked up: the speed your plan falls back to after the hotspot allowance is used. Verizon's mobile hotspot support page, read 29 September 2026, states that on Unlimited Plus you get 30 GB of premium hotspot data and after that unlimited data at 3 Mbps on 5G Ultra Wideband and 600 Kbps on 5G or 4G LTE. A 600 Kbps ceiling will not carry the meeting you bought the backup for. Find the equivalent line for your own plan and write the post-allowance figure next to your monthly usage before you rely on it.
Can I use two connections from the same provider as primary and backup?
Netgate's multi-WAN documentation is blunt about it: two connections of the same type cannot be relied upon to provide redundancy in most cases, because an ISP outage or cable cut will commonly take down all connections of the same type. The same page notes that even two different technologies from one large provider often share a cable path, which leaves the site exposed to a cut. The test to apply is not whose logo is on the bill but whether the two services share a cable, a pole, a tower, a power strip or an upstream network.