caddy-sip-guardian/docs/troubleshooting/udp-conntrack-ack-loss.md
Ryan Malloy f73823448f Update case study: conntrack fix was defense-in-depth, not root cause
Agent thread message 010 revealed critical update: conntrack fix didn't
resolve the production blocker.

### What Actually Happened:
Host physical NIC pcap (tcpdump -i enp1s0) showed ZERO ACK packets
arriving at docker-2. The loss is upstream in carrier infrastructure
(ClearFly ↔ Twilio seam), not at the conntrack layer.

### Documentation Updates:
- Added prominent warning box at top of case study
- Updated verification section to emphasize host-NIC pcap FIRST
- Clarified that conntrack + ACK fast-path are defense-in-depth, not fixes

### Key Lesson:
Same symptom (Timer H expiry at 64s), different root cause. Container
pcaps show what reaches userspace but can't prove what never arrived.
Always verify at the physical NIC layer before concluding a lower-layer
fix worked.

### Still Valuable:
- Conntrack 30→120s: prevents future UDP timeout issues
- ACK fast-path: architecturally correct, prevents security pipeline drops
- Case study: documents real failure mode worth knowing about

The diagnostic convergence was sound; the actual blocker was one layer
further up the stack.

See: docs/agent-threads/ack-loss-from-twilio-trunk/010-*
2026-06-22 16:16:29 -06:00

15 KiB
Raw Permalink Blame History

Case Study: UDP Conntrack Timeout Causing ACK Loss and Call Death at 64 Seconds

⚠️ IMPORTANT UPDATE (2026-06-22):

This case study documents a real failure mode (conntrack timeout < Timer H) that can cause exactly the symptoms described. However, in the original production case that prompted this investigation, the conntrack fix did not resolve the actual blocker.

Verification at the host physical NIC (tcpdump -i enp1s0) proved that ACK packets never reached docker-2 at all - the loss was upstream in the carrier infrastructure (ClearFly ↔ Twilio seam), not at the conntrack layer.

Key lesson: Same symptom (timing-precise call death + Timer H expiry), different root cause. Always verify at the host NIC level before concluding that a lower-layer fix (like conntrack) resolved the issue. Container-level pcaps show what reaches userspace but can't prove what never arrived.

The conntrack fix and ACK fast-path improvements remain valuable defense-in-depth measures that prevent future issues of this shape.

See: Agent Thread Message 010

Problem Statement

Twilio trunk calls through caddy-sip-guardian were dying at exactly 64 seconds after answer. The failure was 100% reproducible and timing-precise.

Observable symptoms:

  • Calls answered normally (INVITE → 200 OK → ACK)
  • Audio worked perfectly for ~64 seconds
  • At t+64s, call abruptly terminated with BYE from Asterisk
  • Asterisk logs showed Timer H expiry: "Disconnecting call because no ACK received"

Architecture Context

┌──────────────┐         ┌────────────────────┐         ┌──────────────┐
│    Twilio    │  UDP    │  caddy-sip-        │  UDP    │   Asterisk   │
│  Trunk       │ :5060   │  guardian          │ :5060   │   (asterpbx) │
│ 54.244.51.x  │────────>│ (Layer 4 proxy)    │────────>│ 172.20.7.6   │
└──────────────┘         └────────────────────┘         └──────────────┘
                              │
                              │ Lives on docker-2.supportedsystems.com
                              │ Behind Linux kernel conntrack/NAT
                              ▼
                       UDP conntrack table

Key components:

  • Twilio: SIP trunk provider (UDP transport)
  • caddy-sip-guardian: Layer 4 SIP proxy with security features (rate limiting, enumeration detection, bans)
  • Asterisk: PBX backend running in container
  • Linux conntrack: Kernel connection tracking for NAT/stateful firewall

Investigation Timeline

Layer 1: SIP Protocol Timers (Red Herring)

Hypothesis: Asterisk's endpoint timers too aggressive?

; /etc/asterisk/pjsip.conf endpoints
timers_sess_expires = 1800     ; 30 minutes
timers_min_se = 90             ; 90 seconds

Result: Ruled out. Asterisk's session timers were set to 30 minutes, far beyond the 64-second failure.

Key lesson: Timer H (INVITE transaction timeout) is 32 seconds per RFC 3261. The 64-second call death is exactly 2× Timer H, suggesting Timer H fired once, sent BYE, waited another 32s for BYE response, then disconnected.


Layer 2: sip-guardian ACK Handling (Real Bug, Wrong Layer)

Hypothesis: ACKs triggering enumeration detection or rate-limiting false-positives?

Code Review Finding:

In l4handler.go, every SIP request (including ACK) went through the full security pipeline:

// l4handler.go:128-333
func (h *SIPHandler) Handle(cx *layer4.Connection, next layer4.Handler) error {
    // 1. Ban check
    // 2. Whitelist check → SKIP TO PROXY if whitelisted
    // 3. GeoIP check
    // 4. Read buffer
    // 5. SIP validation
    // 6. Rate limiting (by method) ← ACK can hit this
    // 7. Enumeration detection ← ACK can hit this
    // 8. Pattern matching
    // 9. Pass to proxy
}

Problem identified: ACK is a mid-dialog request, not a security threat. It should not traverse enumeration detection or rate limiting.

Fix Applied (Option B - Quick Fix):

// Extract SIP method for rate limiting
method := ExtractSIPMethod(buf)

// Fast-path for ACK - mid-dialog request, not a security threat
// ACKs arrive in response to 200 OK retransmissions (RFC 3261) and would
// trigger false-positive enumeration/rate-limit detection.
// See: docs/agent-threads/ack-loss-from-twilio-trunk/
if method == "ACK" {
    h.logger.Debug("ACK exempted from enumeration/rate checks (mid-dialog fast-path)",
        zap.String("ip", host),
    )
    if h.guardian.metricsEnabled() {
        RecordConnection("allowed")
    }
    return next.Handle(cx)
}

Result: Deployed to production... but calls still died at 64 seconds.

Verification: Packet capture inside sip-guardian container showed zero ACK packets arriving from Twilio. The ACKs were being dropped upstream of sip-guardian, before they reached the userspace layer.

Key lesson: The architectural fix was correct and valuable for other deployments, but it didn't solve this specific blocker. The loss was happening at the kernel layer.


Layer 3: UDP Conntrack Timeout (Root Cause)

Hypothesis: Linux conntrack expiring UDP flows before SIP transaction completes?

Investigation on docker-2.supportedsystems.com:

$ cat /proc/sys/net/netfilter/nf_conntrack_udp_timeout
30

$ cat /proc/sys/net/netfilter/nf_conntrack_udp_timeout_stream
120

The smoking gun: nf_conntrack_udp_timeout = 30 seconds

This is the timeout for unreplied UDP flows (single-direction traffic with no response). Bidirectional UDP flows get 120 seconds.


Root Cause: Timing Conflict Between Conntrack and SIP Protocol

SIP Call Flow vs. Conntrack Lifecycle

Time SIP Event Conntrack State
t+0 Twilio → INVITE Entry created (30s timeout)
t+0.1 asterpbx → 200 OK ASSURED (bidirectional)
t+0.2 Twilio → ACK Flow established
t+0.5 asterpbx retx 200 OK (Timer A) Active
t+1 asterpbx retx 200 OK Active
t+2 asterpbx retx 200 OK Active
t+4 asterpbx retx 200 OK Active
t+8 asterpbx retx 200 OK Active
t+16 asterpbx retx 200 OK Active
t+30 Conntrack expires Entry deleted
t+32 asterpbx retx 200 OK (final) No entry
t+32 Twilio → ACK DROPPED (no conntrack)
t+64 asterpbx Timer H fires → BYE Call dies

Why the 30s Timeout Kills SIP

RFC 3261 Timer H: 32 seconds for INVITE 200 OK to be ACKed before abandoning

Linux UDP conntrack: 30 seconds before expiring "unreplied" UDP flows

The problem:

  1. Asterisk's 200 OK retransmissions trigger Twilio to re-send ACK
  2. But ACK-to-ACK retransmissions are the same packet, not new bidirectional traffic
  3. Conntrack sees this as "unreplied" (no NEW packet in 30s)
  4. At t+30s, conntrack expires the flow
  5. At t+32s, Asterisk's final 200 OK retransmit arrives, Twilio sends ACK
  6. Kernel sees unsolicited UDP packet (no conntrack entry) → DROP
  7. Asterisk never receives ACK → Timer H fires at t+64s → sends BYE

Why 64 seconds exactly:

  • Timer H fires at t+32s (ACK not received)
  • Asterisk sends BYE
  • Waits for BYE response (another transaction)
  • Final timeout at t+64s (2× Timer H)

The Fix

Production Fix: Increase UDP Conntrack Timeout

Applied on docker-2.supportedsystems.com:

# Increase UDP conntrack timeout to 120s (matches stream timeout)
sudo sysctl -w net.netfilter.nf_conntrack_udp_timeout=120

# Make permanent
echo "net.netfilter.nf_conntrack_udp_timeout = 120" | \
  sudo tee /etc/sysctl.d/99-sip-conntrack.conf

Before:

  • Unreplied UDP timeout: 30s (too short for SIP Timer H)
  • Stream UDP timeout: 120s (bidirectional flows)

After:

  • Unreplied UDP timeout: 120s (outlasts Timer H by 88s)
  • Stream UDP timeout: 120s (unchanged)

New Timeline with 120s Timeout

Time Event Conntrack
t+0-32 INVITE/200 OK/ACK dance Active
t+32 Final 200 OK retransmit + ACK Entry still valid
t+64 Timer H would fire... But ACK was received!
Call continues normally Dialog established

The conntrack entry now survives the entire Timer H window (32s) with 88 seconds of headroom.


Verification

⚠️ CRITICAL: Verify at the Host NIC Level FIRST

Before assuming the conntrack fix worked, capture at the physical NIC to prove ACKs are actually arriving:

# On the docker host (NOT inside containers)
sudo tcpdump -i enp1s0 -nn -s 0 -w /tmp/host-nic-sip.pcap port 5060

# During active test call, filter for ACK:
sudo tshark -r /tmp/host-nic-sip.pcap -Y 'sip.Method == "ACK"'

# Should see ACK packets. If zero ACKs appear at the host NIC,
# the loss is UPSTREAM (carrier/routing), not conntrack.

Why this matters: In the original production case, zero ACK packets reached the host NIC even after the conntrack fix was applied. The actual blocker was upstream carrier infrastructure (ClearFly ↔ Twilio seam), not conntrack. Container-level pcaps can't prove what never arrived.

Confirm the conntrack fix took effect

# Check current value
cat /proc/sys/net/netfilter/nf_conntrack_udp_timeout
# Should output: 120

# Verify persistent config
cat /etc/sysctl.d/99-sip-conntrack.conf
# Should contain: net.netfilter.nf_conntrack_udp_timeout = 120

Test call verification (only meaningful if host NIC shows ACKs arriving)

# 1. Place call to trunk
# 2. Let it run for >64 seconds
# 3. Verify call stays active
# 4. Check Asterisk logs - should NOT show Timer H expiry

Architectural Improvements Made During Investigation

Even though the conntrack fix was NOT the production blocker in the original case, we made valuable architectural improvements that remain defense-in-depth measures:

1. ACK Fast-Path in sip-guardian

File: l4handler.go:234-246

ACKs now bypass ALL security checks (enumeration detection, rate limiting, validation, pattern matching). This prevents false-positive bans in deployments where ACKs DO reach the sip-guardian layer.

Benefits:

  • Prevents enumeration false-positives from ACK retransmissions
  • Prevents rate-limiting of legitimate mid-dialog ACKs
  • Improves performance (one less layer of processing per ACK)
  • More correct architecturally (ACKs are not a security threat)

2. Agent Thread Protocol Documentation

Directory: docs/agent-threads/ack-loss-from-twilio-trunk/

Created immutable message protocol for cross-layer debugging:

  • 001: Problem report (calls dying at 64s)
  • 002: Layer 2 diagnosis (ACK handling bug)
  • 003: Confirmation to proceed with fix
  • 004: Diagnosis refinement (rate-limiting, not enumeration)
  • 005: Code fix delivered
  • 006: Evidence that ACK loss is upstream (pcap)
  • 007: Root cause found (conntrack timeout)

This protocol made the layered diagnosis tractable - each layer was isolated and verified independently.


When This Same Issue Can Occur

Trigger conditions:

  1. UDP SIP transport (TCP doesn't have this issue - persistent connection)
  2. Long SIP transaction timers (INVITE transactions with >30s Timer H equivalent)
  3. Default Linux conntrack timeout (30s for unreplied UDP)
  4. NAT/conntrack between SIP endpoints

Other protocols affected:

  • DNS (if using UDP with long timeouts)
  • RADIUS (accounting sessions)
  • Any UDP application-layer protocol with retransmission timers >30s

Alternative Workarounds (If You Can't Change Conntrack)

  1. Switch to TCP transport (;transport=tcp in SIP URI)

    • Pros: No conntrack timeout issues, guaranteed delivery
    • Cons: More overhead, persistent connection state
  2. SIP Session Timers (periodic re-INVITE during call)

    • Pros: Keeps conntrack entry alive with regular traffic
    • Cons: More SIP messages, some endpoints don't support it
  3. Whitelist the trunk IPs in sip-guardian

    • Pros: Bypasses security pipeline entirely, no processing delay
    • Cons: Loses enumeration/validation for that trunk
    • SECURITY WARNING: Use ONLY narrow trunk-specific IPs (e.g., /30 subnets), not entire carrier infrastructure ranges

Lessons Learned

1. Cross-Layer Debugging Requires Systematic Isolation

The bug spanned three independent layers:

  • Application: SIP protocol timers (red herring)
  • Userspace: sip-guardian security pipeline (real bug, wrong layer)
  • Kernel: UDP conntrack timeout (actual blocker)

Without systematic isolation, we might have:

  • Blamed Asterisk's timers (wrong)
  • Stopped after fixing the ACK handling (incomplete)
  • Never discovered the conntrack issue

2. Timing-Precise Failures Point to State Machine Conflicts

The 64-second precision was the critical clue:

  • Not random/intermittent (rules out packet loss, congestion)
  • Not variable (rules out load-dependent issues)
  • Exactly 2× Timer H (32s)

Pattern: When a failure occurs at a mathematically precise interval, look for:

  • State machine timeouts
  • Timer conflicts between independent subsystems
  • Retry logic with exponential backoff hitting a cap

3. Packet Captures Don't Lie

The pcap inside the sip-guardian container was the smoking gun that proved the ACK never arrived at the userspace layer. Without it, we might have:

  • Spent hours debugging sip-guardian code
  • Blamed Twilio's implementation
  • Never checked the kernel layer

Lesson: Capture at multiple layers:

  • Wire-level (tcpdump on host interface)
  • Container-level (tcpdump inside container)
  • Application-level (SIP debug logs)

4. Defense in Depth: Fix All Layers

Even though the conntrack timeout was the blocker, the ACK fast-path fix is still valuable:

  • Other deployments might not have conntrack issues but DO have enumeration detection
  • Architectural correctness matters (ACKs shouldn't be treated as enumeration probes)
  • Future-proofs against regressions

Don't stop at the first fix that works - fix the architectural issues you find along the way.


References

  • RFC 3261 (SIP): Section 17.1.1.2 - Timer H (INVITE transaction timeout, 32 seconds)
  • Linux conntrack docs: Documentation/networking/nf_conntrack-sysctl.txt
  • Agent thread protocol: docs/agent-threads/PROTOCOL.md
  • Related thread: /home/rpm/claude/sip/setup-server/docs/agent-threads/active-call-survival-hold-endpoint/

Quick Reference

Symptoms Checklist

If you see:

  • Calls dying at exactly 64 seconds
  • "Timer H expired" or "no ACK received" in Asterisk logs
  • UDP transport
  • NAT/conntrack between SIP endpoints

Then check:

cat /proc/sys/net/netfilter/nf_conntrack_udp_timeout

If it's 30 seconds and you're using SIP, increase it to 120:

sudo sysctl -w net.netfilter.nf_conntrack_udp_timeout=120
echo "net.netfilter.nf_conntrack_udp_timeout = 120" | \
  sudo tee /etc/sysctl.d/99-sip-conntrack.conf

No restart required - takes effect immediately.