Latency and bandwidth — study guide
The concept's fragments, read in order.
Two words people use as one
People call a network connection "fast" and mean one of two unrelated things. One is how long a single message takes to make the trip; the other is how much data the link can move each second. Treating both as a single quantity called speed is where most misunderstanding about networks starts.
The first is latency — a time, measured in milliseconds. The second is bandwidth — a rate, measured in bits per second. They answer different questions: latency asks how soon the first byte shows up, bandwidth asks how fast the bytes pile up once they do. Nothing forces the two to move together.
Keeping them apart is not pedantry. A slow-feeling application and a slow-downloading one have different diseases, and the cures do not overlap — buying a wider pipe does nothing for a link whose problem is distance. Which of the two measurements matters is the whole question this concept turns on.
How long one trip takes
Latency is the time it takes a packet to get from here to there. In practice the number people quote is round-trip time (RTT): send a packet to a destination and measure how long until its reply comes back. It is a duration, reported in milliseconds (ms), and it says nothing at all about how much data you are sending.
The trip takes as long as it takes like timing a single round trip on a road: it takes as long as the journey takes, no matter how many travelers set out after the first, no matter how much follows behind it. Whether you send a single byte or begin an enormous transfer, the first packet still has to make the journey, and adding more data behind it does not make the front of the line arrive any sooner.
Round-trip time is exactly what the ping command reports, one line per reply in milliseconds. Latency to a machine on the same local network is small; latency to one on another continent is much larger, because the distance is real and the trip is physical. It is the delay you pay before any answer can come back at all.
How much fits per second
Bandwidth is the maximum amount of data a link can carry per second. Where latency is a time, bandwidth is a rate, measured in bits per second and usually quoted in megabits or gigabits per second (Mbps, Gbps). It describes the capacity of the pipe, not the speed of any one drop moving through it.
Widening a road lets more cars pass each minute like the number of lanes on a highway: more lanes let more cars pass each minute, but they do not change how fast any one car drives, but it does not make any single car drive faster. Bandwidth works the same way: more of it means more bits delivered each second, while the time for the first bit to arrive is a separate matter entirely.
Bandwidth is a ceiling — the most the link can do. A high-bandwidth link can move a large file quickly once the data is flowing, but it makes no promise about how long you wait for that flow to begin. Capacity and delay are simply different dimensions of the same connection.
A fat pipe can still be slow
Latency and bandwidth are independent measurements: knowing one tells you nothing about the other. A link can be wide and far, or narrow and near, in any combination. Picture the pipe's width as its bandwidth and its length as its latency like a water pipe whose width sets how much flows per second and whose length sets how long the water takes to arrive, two separate things — widening a pipe moves more water per second but does not make the water's journey any shorter.
The reason is that the two are set by different things. Bandwidth is a property of the link's data rate, which has nothing to do with distance; latency is dominated by how far the signal must travel. A satellite link can carry plenty of data per second and still make every packet wait through a long trip up to orbit and back down.
So a fat pipe — one with bandwidth to spare — can still feel sluggish if it is long, and a modest link right next to you can feel instant. Upgrading the dimension you are not short on buys nothing.
What the round trip is spent on
Where does the time in a round trip actually go? Latency is not one thing but a sum of several delays, and knowing which one dominates tells you whether anything can be done about it.
Propagation delay is the time the signal spends in transit, set by distance and the speed it travels through the medium. Transmission delay is the time to push a packet's bits onto the link, set by the packet's size and the link's data rate. Queuing delay is the time a packet waits in line behind others before it can be sent. A smaller processing delay covers the handling of each packet's header. Together they add up to the round-trip time.
Propagation is the one physics puts a floor under. Light in optical fiber travels at about two-thirds of the speed of light in a vacuum — roughly 200,000 kilometers per second, or about five microseconds per kilometer. No equipment upgrade beats that floor; the only way to cut propagation delay is to shorten the distance.
Which delay dominates depends on the link. Across a long distance, propagation rules and distance is the only lever; under heavy traffic, queuing swells as packets back up; transmission delay stands out only when a large packet is serialized onto a slow link. Diagnosing latency starts with asking which of these the time is being spent on.
How much data is in flight
At any instant, a link is holding some amount of data that has been sent but not yet acknowledged — bits in flight. How many? That number is the bandwidth-delay product: the link's bandwidth multiplied by its round-trip time.
The formula has the right shape. Bits per second times seconds gives bits, so bandwidth times round-trip time is a count of bits — the most a sender can put on the wire before the first acknowledgement could possibly come back. A far-away link with high bandwidth has a large product, because a lot of data fits inside the long trip.
To keep such a link busy, the sender has to keep at least a bandwidth-delay product of data outstanding at once. TCP's in-flight window is exactly this limit; if it is smaller than the product, the sender runs out of permitted data and stalls waiting for acknowledgements, leaving capacity idle. A fat, far pipe left underfilled delivers far less than it could.
What you get versus what the pipe allows
Bandwidth is what the link allows; throughput is what you actually get. Throughput is the data rate a transfer really achieves, and it sits at or below the rated bandwidth — never above it.
The gap comes from everything the rated figure ignores. Protocol headers take up room that is not your data; congestion makes senders slow down to share the link; and lost packets have to be sent again, costing both time and capacity. Each one shaves the delivered rate down from the theoretical ceiling.
So bandwidth is a promise about the maximum, not the delivered amount. Quoting a link's bandwidth tells you the best it could ever do; measuring throughput tells you what it did. The first is a specification, the second is a result.
Picking the right thing to optimize
Once latency and bandwidth are separate in your head, a practical rule falls out: work out which one your workload actually needs, because optimizing the other buys nothing.
An interactive workload lives and dies by latency. A game, a video call, or a chain of small requests in an agent pipeline spends its time waiting for one round trip after another, so the delays add up and a wider pipe does not help. A bulk transfer is the opposite: a large file, a backup, or a data sync needs bandwidth, wants the pipe as wide as it can get, and barely notices the first packet's delay.
And some latency you cannot buy your way out of with a bigger link. If your users are far from your server, the round trip is long because the distance is long; the fix is to move the server closer — a region near the users — not to widen the connection. Distance is the one cost a fatter pipe never pays down.