People • Technology • Possibilities+91 74199 74199[email protected]
Home / Insights / SD-WAN Failover
Network Resilience

Multi-ISP SD-WAN Failover: Design and Testing

How to design two internet links that fail independently, how FortiGate SD-WAN health checks decide a link is down, and how to test failover so you know the real recovery time.

Dual ISP SD-WAN failover
SD-WAN & Internet Resilience

As of 11 October 2026. Configuration below is illustrative FortiOS 7.6 syntax and has not yet been validated in our lab. Failover times depend on your timers, links and providers, so measure them rather than relying on any published figure.

Two internet links are not automatically resilient

Many offices pay for a second ISP and discover during the first outage that both circuits ran through the same duct, the same exchange or the same upstream carrier, or that traffic did not move because nothing was checking whether the internet beyond the modem was reachable. SD-WAN solves the steering problem; design and testing solve the rest.

Design the physical diversity first

  • Different last-mile media where possible, for example fibre plus a licensed wireless or 5G link, rather than two fibres in the same trench.
  • Different building entry points and exchanges. Ask each provider for the route and the exchange they terminate on; get it in writing.
  • Different upstream carriers. Two resellers of the same backbone fail together.
  • Separate power paths for both modems or ONTs and the firewall, on UPS.

How FortiGate SD-WAN decides a link is down

On a FortiGate, each WAN interface becomes an SD-WAN member, members are grouped into zones, a performance SLA (health check) probes a target over each member, and SD-WAN rules choose which member carries which traffic. Fortinet documents three health-check timers under config system sdwan › config health-check:

SettingMeaningDefault
intervalMilliseconds between probes (20–3600)500
failtimeConsecutive failed probes before the member is marked dead5
recoverytimeConsecutive successful probes before the member is marked alive again5

From those, Fortinet gives the detection times:

alive → dead  =  failtime     × interval / 1000   seconds
dead  → alive =  recoverytime × interval / 1000   seconds

Defaults:  5 × 500 / 1000 = 2.5 s to declare a link dead
           5 × 500 / 1000 = 2.5 s to declare it alive again

Fast detection is not always better. A short recoverytime on an unstable link makes traffic flap back and forth; a long one keeps traffic off a repaired link for minutes. Fortinet’s own worked example (interval 2,000 ms, failtime 2, recoverytime 60) takes 4 seconds to declare a link dead and 120 seconds to bring it back. Separately from dead/alive, SLA thresholds for latency, jitter and packet loss mark a link out of SLA; packet loss is measured over the last 100 probes, so recovery from heavy loss is gradual. Fortinet notes that from FortiOS 7.6.4 a member recovering from dead goes straight to in-SLA once its metrics are within thresholds.

Illustrative configuration

Illustrative example (FortiOS 7.6 syntax, not lab-validated). Interface names, the probe target and thresholds are placeholders. Probe a target beyond each provider’s network, not the provider’s own gateway, or a failure upstream of the gateway will go unnoticed.

config system sdwan
    set status enable
    config zone
        edit "internet"
        next
    end
    config members
        edit 1
            set interface "wan1"
            set zone "internet"
            set gateway <ISP-A-gateway>
        next
        edit 2
            set interface "wan2"
            set zone "internet"
            set gateway <ISP-B-gateway>
        next
    end
    config health-check
        edit "internet-sla"
            set server "<probe-target-1>" "<probe-target-2>"
            set protocol ping
            set interval 500
            set failtime 5
            set recoverytime 10
            set members 1 2
            config sla
                edit 1
                    set latency-threshold 150
                    set jitter-threshold 30
                    set packetloss-threshold 5
                next
            end
        next
    end
    config service
        edit 1
            set name "business-apps"
            set mode sla
            set dst "all"
            config sla
                edit "internet-sla"
                    set id 1
                next
            end
            set priority-members 1 2
        next
    end
end

Before change: export the current configuration. Rollback: restore the backup or disable the SD-WAN rule, which returns traffic to the default route behaviour.

Test it like an outage

Run each test in a maintenance window with a continuous, timestamped ping and a long-running application session (for example a video call or a large download) open from a test laptop.

TestHowRecord
Hard failureUnplug the primary WAN cableSeconds until ping recovers on the secondary; whether the video call survived
Upstream failureBlock the probe targets on the primary path onlyWhether the SLA marks the link dead even though the cable is up
DegradationAdd latency or loss with a link emulator, or test during a known bad periodTime to move traffic out of SLA; any flapping
RecoveryRestore the primary linkSeconds until traffic returns; whether it flaps
Inbound servicesReach published services (VPN, web) from outside during failoverWhether DNS, NAT and certificates work on the secondary IP

Repeat each test at least three times and keep the results. Those numbers, not the vendor defaults, are what you can tell the business. Inbound services are where failover most often disappoints: a site-to-site VPN or published server bound to the primary IP will not follow the outbound traffic unless it has been designed to.

Sources

Primary sources used for this article (checked 11 October 2026):

Frequently Asked Questions

About the Author

Lalit Bhardwaj — Founder & Technology Strategist, XOOPIE. Lalit leads XOOPIE with a hands-on technology strategy and infrastructure engineering approach, focused on understanding how an organization actually operates and translating that reality into an appropriate, resilient technical design.

Read Lalit Bhardwaj's full profile →