Constraint, Load, and Tolerance
Every system has a load it can carry and a point past which it stops bending and starts breaking.
Every system has a load it can carry and a point past which it stops bending and starts breaking.
A company I have in mind doubled its sales team over a year, with a clear and reasonable expectation. Twice the salespeople should produce something close to twice the results. The hiring went well, the new people were good, and the output did not double. It rose by perhaps a third, and then the per-person numbers, which had looked healthy before, began to sag across the whole team, the veterans included. Nobody had gotten worse at their job. The thing that had changed was that the system feeding those salespeople with qualified opportunities had a capacity, and the old team had been sitting comfortably under it, and the new team pushed the whole operation up against that limit and then past it. Beyond the limit, the reps were competing for a supply of opportunities that had not grown, and the experience of the whole team degraded at once, because they were all drawing from the same strained source.
This is one of the most common and least understood patterns in go-to-market, and it is an instance of something every engineering discipline takes as foundational and this one has barely begun to think about. Every system has a load it is being asked to carry, a capacity it is able to carry, and a relationship between the two that governs how it behaves. Below capacity it behaves one way. Near capacity it behaves another, usually worse than people expect. Past capacity it stops slowing down in proportion and starts to break, and the breaking is often sudden and disproportionate to the small extra push that caused it. A field that cannot reason about load and capacity is a field that will keep being surprised by failures that were, in a precise sense, predictable.
The three quantities
Let me put down the three quantities the whole essay turns on, because naming them is most of the work, and the rest follows once they are clear.
The first is load, which I will write as . The load on a go-to-market system is the demand it is being asked to handle, the volume of opportunities flowing to the sales team, the number of new users arriving at onboarding, the quantity of interest the system has to process and convert. Load is measured in the demand quantities from which the rest of this is built, a flow of demand per unit of time.
The second is capacity, which I will write as . Capacity is the load the system is actually able to carry while still performing as designed, the most it can take before its outputs start to degrade. A sales team has a capacity to work opportunities well, set by how many people it has and how much time each good opportunity requires. An onboarding system has a capacity to bring users to their first real success, set by whatever the bottleneck resource happens to be. Capacity is real and finite and usually unmeasured, which is the root of the trouble, because a limit you have not measured is a limit you will discover by crossing it.
It is worth sitting with how strange it is that capacity goes unmeasured, because in most engineered systems it is the first thing anyone establishes. No one builds an elevator without knowing the weight it can carry, and the limit is posted on the wall where everyone can see it. Go-to-market builds the equivalent of elevators with no posted limit all the time, sales teams whose real capacity to work opportunities well has never been measured, onboarding flows whose capacity to bring users to a first success is unknown, support systems sized by feel. The capacity is no less real for being unposted. The only difference is that the people loading the system cannot see the limit they are approaching, so they approach it blind and learn where it was only when the system gives way.
The third quantity is the one that does the real work, and it is the ratio of the first two. I will call it utilization and write it as
the fraction of capacity that the current load is using. When is well below one, the system has slack, it is carrying less than it can handle, and it tends to perform comfortably. When approaches one, the system is running near its limit. When exceeds one, the load is greater than the capacity, and the system cannot keep up, and something has to give. The entire behavior of a loaded system can be read off, to a first approximation, from where sits, which is why it is the quantity worth watching and almost no one in go-to-market watches it.
What happens near the limit
The reason utilization matters so much, and the reason the sales story played out the way it did, is that systems do not degrade gracefully as they approach their capacity. They degrade with accelerating severity, and there is a precise and well-studied account of why.
The account comes from the theory of queues, the mathematics of what happens when work arrives at a resource that can only handle so much at a time, which is exactly the situation of opportunities arriving at a sales team or users arriving at an onboarding flow. Queueing theory was built for telephone exchanges a century ago and has described congestion of every kind ever since, and its central result is one of the most useful and least known facts about loaded systems. Under a standard set of assumptions, the average time a unit of work waits before it is handled grows in proportion to
The shape of that expression is the whole point. The waiting time does not rise steadily as utilization climbs. It rises gently while is small and then explodes as approaches one, because the denominator is heading toward zero. Put some numbers through it and the behavior becomes vivid. At a utilization of one half, the factor is two. At nine tenths, it is ten. At ninety-nine hundredths, it is a hundred. The last sliver of utilization, the move from ninety percent to ninety-nine percent, does more damage to the waiting time than the entire journey from zero to ninety did. A system can look fine at eighty percent of capacity and be in crisis at ninety-eight, and the difference between those two states is a small change in load that pushed it into the region where the curve turns vertical.
This is what happened to the sales team. Adding people raised the load on the opportunity-generating system, utilization climbed toward one, and the time each opportunity spent waiting to be worked, and the quality of attention it received, fell off a cliff that had been invisible while there was slack in the system. The veterans got hurt along with the newcomers because the cliff is a property of the whole system at high utilization, and it punishes everyone drawing on the strained resource. Nobody was failing. The system had been pushed into the steep part of the curve, where small increases in load produce large increases in delay and large drops in quality.
The leadership read it, at first, as a performance problem. Perhaps the new hires were weaker than they had looked, perhaps the managers needed to push harder, perhaps the market had softened. They reached for the usual remedies, more coaching, higher activity targets, more pressure, and none of it helped, because none of it touched the actual cause. The cause was not in any person. It lived in the relationship between how much the team could absorb and how much it was being asked to absorb, and no remedy aimed at the people could fix a problem that lived in the structure.
It helps to have an intuition for why the curve turns vertical, because the formula can feel like a trick otherwise. When a system has plenty of slack, a burst of arrivals is absorbed quickly, because the idle capacity catches up before the next burst lands. As utilization rises, the idle capacity shrinks, so each burst takes longer to clear, and the backlog from one burst is often still being worked off when the next one arrives. Near full utilization there is almost no spare capacity to catch up with, so backlogs build on backlogs and the waiting time runs away. The blowup is the basic logic of a system with no room to recover sitting in front of a load that keeps coming. Variability is the accomplice. If work arrived in a perfectly smooth trickle, a system could run near full utilization safely, and the trouble is that real demand arrives in bursts, and bursts into a nearly full system are what produce the runaway.
There is a companion result worth knowing, because it ties this directly to the quantities a go-to-market system can already count. It is a remarkably general law of queues, holding under almost no assumptions at all, that the average number of items being handled inside a system equals the rate at which they arrive multiplied by the average time each one spends inside. The arrival rate is the demand flow, the same flow a go-to-market system already tracks as opportunities or sign-ups per week, and the average time inside is something it can observe. Together they pin down how much work is sitting in the system at once, which is the quantity that determines whether anyone is keeping up. The law is exact and dimensionally clean, and it means a team that knows its inflow and its handling time already holds, with no new instrumentation, most of what it needs to see how loaded it actually is. The information is usually right there, uncollected, because no one was looking for it.
I want to be careful about the model rather than oversell it. The clean expression above comes from specific assumptions, work arriving at random in a particular statistical sense, a single server, a simple discipline for choosing what to handle next, and real go-to-market systems violate these assumptions in various ways. The exact curve for a real system will differ from the simplest formula. What does not differ, what is robust across all the realistic variations, is the qualitative fact that delay and degradation blow up as utilization approaches capacity, that the blowup is nonlinear, and that the dangerous region is the last stretch before reaches one. You can argue about the precise shape. You cannot, if you take the mathematics seriously, expect a system run at ninety-eight percent of capacity to feel like a system run at seventy, and that single piece of understanding would prevent a large fraction of the self-inflicted go-to-market failures I have watched.
Tolerance and the rated band
Engineers handle this reality with two ideas that go-to-market has no working equivalent of, and both are worth importing by name.
The first is the idea of a rated band, the range of load over which the system is designed to perform to specification. A component is characterized by a range rather than by a single number, the conditions under which it will do what it promises, and operating inside that range is normal while operating outside it is, by definition, operating the system in a regime it was not built for. A go-to-market system has a rated band too, whether or not anyone has identified it, a range of load within which it creates and captures demand well, below which it wastes capacity it is paying for and above which it degrades. The band exists. The difference between an engineered system and an improvised one is whether anyone knows where the edges are.
The second idea is tolerance in the deeper sense, the acknowledgment that load is never known exactly and varies in ways you cannot fully predict, so a system has to be built to perform across a range of conditions rather than at a single expected point. Demand arrives unevenly, in bursts and lulls, and a system designed only for the average load will spend much of its time either starved or overwhelmed, because the average is a number the actual load rarely equals. Designing for tolerance means designing for the variation, for the peaks and not only the means, which is a different and more demanding standard than designing for the number on the forecast.
A concrete picture helps. Suppose demand averages a hundred units a week but actually arrives as sixty in a quiet week and a hundred and sixty in a busy one. A system sized for the average of a hundred runs at comfortable utilization in the quiet weeks and well past its capacity in the busy ones, so it spends part of its life idle and part of its life in crisis, and its average performance, the thing the forecast quietly promised, is a level it almost never actually delivers. The average load turns out to be the one figure the system is least often carrying. Designing for the average is therefore close to designing for a condition that rarely occurs, and the variation around the average, the very thing a single forecast number throws away, is what decides whether the system holds.
These two ideas reframe what it means to scale a go-to-market system. The instinct the field has, captured in the phrase about doing more of what works, treats scaling as turning a dial, as though a system that performs at one level of load will perform the same way at triple the load. The mathematics of utilization says this is close to never true, that every system has a band, that pushing load up toward capacity moves the system into the steep part of the curve, and that real scaling means raising capacity deliberately and keeping utilization in the healthy range, rather than simply pouring on more load and expecting the old behavior to hold.
The margin nobody leaves
There is a third engineering habit that follows directly, and its absence in go-to-market is almost an emblem of the whole problem. Engineers do not design a system to carry exactly the load they expect, because the expected load is uncertain and the cost of being wrong is failure, so they build in a margin, a deliberate gap between the capacity they install and the peak load they anticipate. The ratio of the two is the safety factor,
the amount by which rated capacity exceeds the worst load the system is expected to face. A bridge built with a safety factor of two is built to carry twice the heaviest load anyone expects to put on it, and the doubling buys the one thing that matters, which is staying up when the load turns out to be heavier than expected or the structure weaker than hoped, both of which happen.
Go-to-market, as a rule, runs with a safety factor close to one, which is to say with no margin at all. It staffs the sales team to handle exactly the forecast volume, sizes the onboarding system for the expected number of users, plans the whole apparatus around the load it hopes to see, and then it meets a good month, a load above the forecast, and discovers that a system running at a safety factor of one has nowhere to go when the load rises, because every increase in load is an increase in utilization toward the vertical part of the curve. The cruel irony is that success is the trigger. The better the demand-generation works, the higher the load, and a system with no margin converts its own success into congestion, so that a great quarter for the top of the funnel becomes a terrible quarter for everything downstream, and no one connects the two because no one was thinking in terms of load and capacity and margin at all.
The pattern is easy to recognize once you have the language for it. A company gets a burst of attention, a launch that lands, a piece of content that travels, a moment of demand it did not entirely plan for, and the very thing everyone wanted becomes the thing that breaks the operation, because the surge drives utilization past one across every system that has to handle it. Sign-ups arrive faster than onboarding can absorb them, so the new users have a poor first experience and drift away. Inquiries arrive faster than sales can work them, so good opportunities sit untouched and go cold. The company got exactly the demand it had been trying to create and could not hold it, and the post-mortem, where there is one, tends to blame execution rather than the absent margin that guaranteed the collapse the instant the load spiked.
Capacity is something you design
It would be easy to take all of this as counsel to keep utilization low and leave it there, but that misreads the point, because capacity is not a fixed fact of nature. It is a property of how the system is built, and it can be raised. This matters because the goal is rarely to run a small system at a comfortable utilization forever, the goal is to grow, and growth means raising capacity ahead of load so that utilization stays in the healthy band even as the load climbs. The engineering question is therefore not only how loaded the system is now, but how its capacity can be increased, at what cost, and how quickly relative to the growth in load.
This reframes a great deal. When a sales system approaches its capacity, the instinct is to push the existing team harder, which only drives utilization toward the cliff. The engineered response is to raise capacity, by adding people sometimes, but more fundamentally by removing whatever the binding constraint actually is, which is frequently not headcount at all. If the constraint is the supply of qualified opportunities, more salespeople do not raise throughput at all, and they make the experience worse, because they raise the load against a capacity that is set somewhere else entirely. Finding the true constraint, the single resource that actually caps the system’s throughput, is the central skill, because raising capacity anywhere other than the binding constraint spends money and moves nothing. A discipline would locate the constraint before it scaled. The field, more often, scales the part that is easy to scale and is puzzled when the system as a whole refuses to respond.
A quick worked case shows the discipline. Suppose a sales organization can source two hundred qualified opportunities a week, each rep can work twenty-five of them well, and there are six reps. The reps can handle a hundred and fifty, the sourcing supplies two hundred, so the binding constraint is the reps, and a seventh and eighth rep will turn the surplus opportunities into real throughput. Now change one number and suppose there are ten reps able to handle two hundred and fifty, while sourcing still supplies two hundred. The binding constraint has moved to sourcing, and hiring an eleventh rep adds capacity to the part of the system that already has slack, which raises cost and changes nothing, because the two hundred opportunities were the ceiling all along. Same company, same instinct to add reps, opposite correct answer, and the only way to tell the two situations apart is to have located the binding constraint before deciding what to scale.
The hypothesis
Let me state the claim of this essay as a hypothesis, sharp enough to be tested and to be wrong.
The hypothesis is that the performance of a go-to-market system degrades as a definite and computable function of its utilization , rising slowly while there is slack and then sharply as approaches one, in the manner the theory of queues describes, so that the failures which the field experiences as sudden and mysterious are in fact the predictable consequence of running a system into the steep region of a known curve. The second part of the claim is empirical and specific. Most go-to-market systems are run with a safety factor near one, with little or no deliberate margin between capacity and expected peak load, which is why a surprising share of go-to-market crises are triggered by success rather than failure, by a load that rose past a capacity nobody had measured.
This is testable, and the test is concrete. Identify the capacity of a stage of a go-to-market system, measure the load over time, compute utilization, and watch what happens to delay and quality as utilization moves. The hypothesis predicts a specific shape, healthy performance through the middle range and rapidly worsening performance as utilization approaches one, and it predicts that the systems suffering the worst congestion will be the ones with the least margin running closest to their limit. If you ran that analysis and found performance falling smoothly and linearly with load, with no acceleration near capacity, the queueing account would be wrong for go-to-market systems, and that would be worth knowing. I do not expect that result, because the sales team that doubled and gained a third is the ordinary signature of a system that crossed into the steep part of the curve with no margin to absorb the load, rather than an unusual story.
What it changes
The point of all this is not to add three more metrics to a dashboard. It is to change the questions the field knows how to ask. A discipline that thought in terms of load and capacity would ask more than whether a channel is working. It would ask what load the system can carry before it degrades, where the rated band is, how much margin exists between the expected peak and the capacity, and how close to one the utilization is being allowed to climb. It would treat capacity as a thing to be measured and designed rather than discovered by collapse, and it would treat margin as a cost worth paying rather than a slack to be eliminated in the name of efficiency. Above all, it would stop being surprised when a system run at the edge of its capacity fails under a load that was, by the plain arithmetic of utilization, always going to break it.
There is a thread that pulls forward from here. A system with a capacity and a load and a utilization is a system with a state, a condition at each moment that determines how it will behave, and a state is only useful if you can see it, if the system tells you where it is on the curve before it goes over the edge. That is the question of sensing and feedback, of building a system that reports its own utilization in time to act on it, and it turns out that most go-to-market systems are flying with almost no instruments at all, carrying loads they cannot see toward limits they have not measured, which is a strange way to operate anything you hoped to call engineered.
References
- On queueing theory and the behavior of loaded systems, including the growth of delay as utilization approaches capacity, see the foundational work originating with A. K. Erlang and standard treatments of queueing in operations research; Little’s law (J. D. C. Little, 1961), relating the average number of items in a system, the arrival rate, and the average time each spends in the system, is the clean dimensional anchor connecting these quantities.
- On safety factors, rated conditions, and design margins, see standard references in structural and mechanical engineering design.