Welcome to Tuning Linux TCP Stack for High-Throughput Low-Latency Workloads. Out of the box, the Linux kernel TCP/IP stack is optimized for general-purpose networkingβbalancing reliability, fairness, and modest memory usage. However, for applications demanding millions of requests per second, such as ad-exchanges or high-frequency trading platforms, the default `sysctl` parameters are a major bottleneck.
1. Expanding Buffer Sizes (BDP)
The Bandwidth-Delay Product (BDP) dictates how much unacknowledged data can be in transit on a network link. The default Linux TCP read/write buffers are too small for high-bandwidth, high-latency connections. You must increase `net.core.rmem_max` and `net.core.wmem_max` to at least 16MB (or higher for 10Gbps+ links), and adjust the `net.ipv4.tcp_rmem` and `net.ipv4.tcp_wmem` auto-tuning arrays to allow the kernel to scale up buffer sizes dynamically for demanding connections.
2. Choosing the Right Congestion Control Algorithm
Historically, Linux used CUBIC as the default TCP congestion control algorithm, which aggressively cuts the congestion window when it detects packet loss. For modern high-speed networks, Google's BBR (Bottleneck Bandwidth and Round-trip propagation time) is vastly superior. BBR measures the actual delivery rate and round-trip time instead of relying on packet loss, allowing it to sustain high throughput even on networks with minor packet loss. Enable it via `net.ipv4.tcp_congestion_control=bbr`.
3. Optimizing the Time Wait State
In high-throughput environments where short-lived connections are common (like REST API gateways), the system can quickly exhaust all available ephemeral ports, leading to "Cannot assign requested address" errors. Tuning `net.ipv4.ip_local_port_range` to the maximum (1024 to 65535) and enabling `net.ipv4.tcp_tw_reuse` allows the kernel to safely reuse sockets in the TIME_WAIT state for new outbound connections.
4. Offloading to Hardware
Software interrupt handling destroys CPU performance. Ensure that hardware offloading features like TCP Segmentation Offload (TSO), Generic Receive Offload (GRO), and Receive Side Scaling (RSS) are enabled on the NIC via `ethtool`. This pushes the heavy lifting of packet segmentation and reassembly down to the NIC hardware, freeing up CPU cycles for application logic.
Conclusion
Modifying kernel parameters carries risk, but for systems pushing the limits of gigabit networking, it is unavoidable. By calculating the correct buffer sizes, leveraging BBR, and properly managing socket states, system administrators can squeeze line-rate performance out of standard Linux distributions.