Skip to main content

The in-depth guide · Last updated August 9, 2026

What is pump by exception?The mechanics, the ceiling, and the ladder past it.

Pump by exception is an operating model that replaces fixed-route well checks with event-driven visits: SCADA watches every well continuously and raises an exception only when a reading leaves its normal window, so the field stops driving to wells that are running fine.

It replaced the calendar with the signal, and it earned its place as the dominant field operating model of the late 2010s. At scale, most operators hit the same wall: the alarms pile up, nothing on the screen is ranked in dollars, and the queue in year five is no smarter than the queue in year one.

This guide covers why the model exists, the three layers it runs on, what it got right, the five ways it stalls, and the four rungs between a threshold queue and a ranked daily plan. It is written for the person who has to decide what to do next, not for the person who has to be sold.

Section 1 · Why the model exists

The fixed-route era allocated attention by geography and calendar, not by need.

For decades a pumper visited every well on a set schedule whether anything was wrong or not. The model had four expensive flaws, and pump by exception was built to fix all four at once.

Wasteful

By the industry’s own accounting of the fixed-route era, up to 60% of scheduled visits found nothing wrong. Trucks rolled, fuel burned, and hours passed with no productive outcome. (Published figure, WorkSync Research Volume IV.)

Exposed

More windshield time means more exposure to road hazards, and rural lease roads in poor conditions are where that exposure concentrates.

Slow to respond

A well could go down Monday morning and wait until its Wednesday slot to be discovered, because the calendar decided when someone looked.

Does not scale

When operators consolidated through acquisition, well counts doubled and field organizations did not. Fixed routes became physically impossible before they became unpopular.

Exception-based surveillance inverted the default: visit nothing unless the data says otherwise. The published operational record on that inversion is strong. The canonical industry analysis found lease-operator value-added time moving from roughly 25% to roughly 60% of the day under exception-based surveillance, with production downtime reduced by roughly a third through earlier intervention on degradation signals already present in the data (Alvarez & Marsal, 2015).

That is a real result, and nothing in this guide takes it back. The question this guide answers is what happens after the inversion works, when the queue it created grows past the attention available to read it. More on exception-based surveillance.

Section 2 · How it works

Three layers, and every later evolution reuses all three.

A working pump-by-exception deployment already has the architecture that everything after it stands on. Knowing which layer a problem lives in is what keeps the next step from being sold internally as a rip-and-replace.

Layer 01

Sensing

Pressure transducers, flow meters, tank-level sensors, and RTUs stream to a historian at second-to-minute cadence over cellular, radio, or satellite. This layer answers what the well is doing right now.

Layer 02

Rules

Configurable thresholds define a normal operating window for each parameter. A reading outside its window generates an exception. More sophisticated engines support compound rules, for example high casing pressure together with low production rate.

Layer 03

Delivery

Exceptions are pushed to the field by mobile app, text, or email. A lease operator reviews the alert, decides whether to drive out, and logs the outcome. Dispatch is manual in basic deployments and proximity-based in more advanced ones.

Nothing in the rest of this guide removes a layer. Every move described below is additive: it changes what flows through the layers and how the field consumes the result. That matters for budgeting and for politics. See what connects into the layers you already run.

Section 3 · Partial credit, honestly given

What pump by exception got right, and why that matters for what comes next.

It replaced the calendar with the signal. Attention stopped being allocated by geography and started being allocated by evidence. Wherever alarm discipline held, that change was worth real money and real exposure hours.

It built the three-layer architecture everything else stands on. Sensing, rules, delivery. An operator running exceptions today already owns the hardest and slowest part of the build.

It taught the organization to trust data over habit. This is the least-discussed achievement and probably the most valuable. A company that has run exceptions for several years already fought and won the hard argument: that a signal in the historian can outrank a habit in the truck. The next argument is the same one applied a layer up, that a dollar estimate in the queue can outrank a gut ranking of the queue. Field organizations that lived through the first transition are the easiest population in the industry to bring along to the second.

Section 4 · The ceiling

Five failure modes you can see from the truck.

These are well documented in the alarm-management literature. They are described here the way they present in the field, because the field presentation is what any next step has to fix. Each one comes with its operational tell, the thing you can look for in your own operation this week.

01Alarm overload

Exception volume scales with well count, sensor count, and rule count. Human attention does not. A queue of 30 to 50 active exceptions per lease operator per shift is routine in mature deployments (ANSI/ISA-18.2 context, as cited in WorkSync Research Volume IV). Past roughly a dozen items, a queue stops being a work list and becomes background noise with occasional spikes.

The tell
Alarms acknowledged in batches without being read, and foremen keeping a private mental list of the wells that actually matter because the official queue stopped carrying that information.

02Every exception looks equal

A low-pressure alarm on a 5 BOPD stripper renders identically to a pump failure on a 200 BOPD producer: same color, same row height, same notification sound. The engine has no concept of working interest, commodity price, lifting cost, or deferment risk. The alarm does not know what a barrel is worth, so the queue cannot know what an hour of field time is worth.

The tell
The morning argument about where to send the workover rig gets settled by whoever argues loudest, because no number on the screen can settle it.

03Thresholds drift away from declining wells

Every threshold was correct on the day it was set. Wells decline; thresholds do not. A static low-rate alarm set against month-one production is, by month eighteen, either firing constantly and being ignored or widened so far it can no longer detect a real problem. Re-tuning limits by hand across thousands of wells is a job nobody is staffed for, so in practice it happens rarely.

The tell
Wells whose alarm limits have been edited more than twice, and wells where the limit now sits below the well’s own economic limit.

04No economic context, so no defensible deferral

When everything is urgent, deferral is guesswork, and guesswork cannot be defended in the variance review. Skipping a low-value alarm to chase a high-value one is usually the right call, made with no record of why it was right.

The tell
Inconsistency between crews, and variance-report archaeology: discovering two billing cycles later that a high-value well sat compromised while crews worked nuisance items.

05The system never learns

The most expensive failure mode is the quietest. The field closes out hundreds of work orders a month, and each one contains the answer to the question the system should be asking: was this exception worth the visit? In a classical deployment that answer goes nowhere.

The tell
False-alarm rates that do not fall, severity estimates that do not improve, and a queue in year five no smarter than the queue in year one.

The pattern. None of the five is fixed by more SCADA, more rules, or more dashboards. Each one is a missing layer rather than a missing signal: a priority layer, a value layer, a planning layer, and a learning layer. That is what the rest of this guide describes.

Section 5 · The maturity ladder

Four rungs between a threshold queue and a ranked day.

The most useful question is not which platform to buy. It is which rung you are standing on, because that determines what the next move actually is. Read the right column and find the one that describes your 6 AM.

RungNameWhat you see at 6 AM
0Pump by exceptionAn unranked queue of threshold breaches.
1Prioritized exceptionsA shorter queue: rationalized alarms, per-well bands that follow decline, persistence rules that suppress transients.
2Scored exceptionsThe same queue with a dollar figure and a confidence band on every row.
3Ranked daily planA route, not a queue: the day’s work ordered by value under constraints, with the reasons attached.
4Pump by priorityThe closed loop: outcomes feed back, scoring improves, the plan re-ranks mid-shift, and measurement work competes with production work on the same ruler.

Published in WorkSync Research, Volume IV

Each rung pays for itself

Queue rationalization alone recovers attention. Scoring alone changes the workover-rig argument. The ranked plan alone moves drive time. No rung’s business case depends on reaching the next one, which is why this is a ladder and not a platform migration.

The rungs are ordered by trust, not technology

Scoring an unrationalized queue produces precise rankings of garbage. Ranking an unscored queue produces efficient routes to the wrong wells. The order is not a sequencing preference, it is the only order that works.

Rung 4 is a different kind of step

Rungs 1 to 3 are operating-practice changes an organization can drive with its existing systems plus a scoring and planning layer. Rung 4 adds machinery: closed-loop learning, multi-source fusion, and value-of-information sensing.

If you would rather answer six questions than read a table, the Field Ops Blind Spot Test places your operation on the adjacent maturity frame in about two minutes.

Section 6 · Rung 2, in arithmetic

What a dollar figure on the row actually changes.

The obvious way to rank a queue is production times price: a 100 BOPD well at $70/bbl is $7,000 a day, so work that one first. That back-of-the-envelope version ignores the terms that decide whether fixing a well moves your own cash flow. Working interest decides how much of the barrel is yours. Lifting cost decides how much of the barrel is margin. Trend decides how much is still at risk tomorrow. And self-resolution probability decides whether the trip was needed at all.

Two wells go down the same morning. The arithmetic below is illustrative, at an assumed $70/bbl realization and before royalty and taxes, and it is written out so you can run it against your own two wells.

TermWell AWell B
Gross production150 BOPD40 BOPD
Working interest12%100%
Lifting cost$28/bbl$8/bbl
Gross revenue at risk$10,500/day$2,800/day
Net margin per barrel$42$62
Net cash flow at risk$756/day$2,480/day

A gross-volume queue sends the truck to Well A, which looks almost four times larger. A scored queue sends it to Well B, which is worth more than three times as much to the company that owns it. Same two alarms, same morning, opposite decision.

Add the fifth term and the gap widens. If Well A’s exception is an intermittent trip with a meaningful chance of clearing on its own, and Well B’s is a rod pump failure that will not recover without a visit, then the expected value of the Well B trip rises again while the expected value of the Well A trip falls. That is the whole of rung 2: the row carries a number a controller would defend, so the deferral becomes defensible too.

Ask the question the other way and most operations answer it honestly: does your downtime reporting distinguish wells by revenue at risk, or does it all roll up as BOE? An hour down is not the same hour everywhere. More on pricing deferred production.

Section 7 · Rung 3, and whether anyone uses it

A ranked plan only survives contact with the field if the field can argue with it.

Every element on rungs 1 to 3 has been deployed somewhere and quietly abandoned. The abandonment is rarely technical. Experienced field team members are the experts on what is actually happening at the well, and a plan that treats their judgment as noise gets routed around within a month. Six mechanics decide whether it holds.

A flag can be disputed

A field team member who does not believe a flag can dispute it. The dispute lifts the item to the engineer or the foreman and holds it there until it is resolved, instead of bombarding someone with a flag nobody trusts.

A flag can be deferred, with a reason

A well already scheduled for workover can be deferred with a reason and stops generating noise until the reason expires. Deferral becomes a recorded decision rather than an undocumented skip.

Nothing is silently suppressed

A list people stop checking is how real problems get masked, so items are disputed, deferred, or resolved, and never quietly dropped. When the control room spots something the system missed, they add the flag and the system forces the reconciliation.

The plan shows its reasoning

A ranked plan that arrives as an oracle gets ignored by week three. Each line carries what tripped, why it ranks where it ranks, and what a good outcome looks like. People accept being re-ordered by a system that can explain itself.

The override is a feature, and it is data

The field knows things the model does not: the lease road that washed out, the bearing somebody heard yesterday, the tank that gauges differently than it reads. Overriding is one tap with a reason code, and the override is recorded. A well the field consistently pulls forward is carrying information the scoring has not captured.

The plan re-ranks mid-shift, within limits

The 6 AM plan is a forecast, and by 10 AM part of it is wrong. Re-ranking fires on a defined trigger set and pushes only to affected crews, with a minimum improvement bar so small re-orderings are not worth the cognitive cost, and a rule that in-progress work is never interrupted by anything short of a safety gate.

Taken together these describe a system that does the assembling, the sorting, and the remembering, and leaves the judgment where it belongs, with the person standing at the well. The override stream that comes back is the single richest feedback signal the company owns, and classical deployments throw it away.

Section 8 · Where this goes next

The industry has run rung 0 for a decade. WorkSync productized rungs 2 through 4.

The largest operators built working versions of the upper rungs internally, by hand, with large teams. Everyone below them was left without a productized path, which is the gap WorkSync was built to close: every well ranked daily by dollars and risk, the plan in the truck by 6 AM, running on top of the SCADA, historian, and alarm discipline already in place. Nothing gets migrated.

At the deployed reference, a top 25 private producer running 5,000+ wells across three basins, the measured deployment figures are 15% more free cash flow on the same crew, 35% fewer miles driven for the same coverage, and a TRIR move from 1.8 to 0.3. Those are that operator’s figures, measured in live deployment rather than modeled. Whether they transfer at that magnitude is what a pilot on one of your own fields is for.

WorkSync Research · Volume IV

The full playbook, in 15 pages

The ladder above is summarized from the operator-level playbook for this exact evolution: alarm rationalization, economic scoring, the ranked daily plan, and the operating-model changes (the morning meeting, the foreman’s role, the KPI swap) that make it stick, with a 90-day path.

Read “Taking Pump by Exception to the Next Level” →

Pump by exception, common questions

What is pump by exception?

Pump by exception is an operating model that replaces fixed-route well checks with event-driven visits. SCADA monitors well parameters continuously and raises an exception only when a reading falls outside its normal window, so the field stops driving to wells that are running fine. It is the oilfield’s implementation of management by exception, the older management philosophy that human attention should go only to deviations that matter.

How does pump by exception actually work?

Three layers. Sensing: transducers, flow meters, tank-level sensors, and RTUs stream to a historian at second-to-minute cadence. Rules: configurable thresholds define a normal operating window per parameter, and a breach generates an exception, with compound rules in more capable engines. Delivery: the exception is pushed to the field by mobile app, text, or email, and a lease operator decides whether to drive out and logs the outcome. Every later evolution reuses all three layers rather than replacing them.

What SCADA systems work with pump by exception?

Most modern SCADA platforms support exception-based workflows. The requirement is the ability to set configurable alarm thresholds and deliver alerts to mobile devices. If your SCADA can do that, it can run pump by exception, and the layers above it can read from what you already have without a migration.

How much driving does pump by exception actually remove?

Reductions vary widely with well density, failure rates, and how carefully alarm thresholds are calibrated, and published deployment summaries differ, so no single number describes detection on its own. What is worth knowing is that the gain erodes over time unless a prioritization layer is added on top, because alarm fatigue puts the wasted trips back. Where the full loop is running, the measured figure at the deployed reference is 35% fewer miles driven for the same coverage, a deployment figure measured across 5,000+ wells at a top 25 private producer rather than a projection.

How is pump by exception different from pump by priority?

Pump by exception flags a problem when a sensor threshold is breached. Pump by priority goes further and ranks every flagged issue by its estimated economic impact, so the highest-value work is addressed first rather than every alarm being treated as equal. Detection tells you something happened. A ranked plan tells you what it is worth and who goes first.

What does AI change about pump by exception?

It replaces the static threshold with a model of what each well should be doing. Data quality checks validate incoming SCADA readings against expected patterns so sensor drift and communication errors stop generating false exceptions. Anomaly detection learns each well’s operating signature, so deviation is measured against the well rather than against a number somebody set at commissioning, which is what stops threshold drift. Failure-pattern models move some work from reactive to scheduled. Economic scoring then puts a dollar figure on what is left. The point is not the technique, it is that all four fix layers the classical architecture never had.

What kinds of pump by exception software exist?

The market splits into recognizable patterns: dedicated first-generation pump-by-exception platforms with configurable rules engines; mobile-first field capture apps with growing exception capabilities; production accounting suites with exception workflows layered on; equipment-level optimizers for a specific artificial-lift class; and SCADA vendor monitoring stacks with basic alerting in the cloud layer. Each has real strengths and a shared ceiling: they stop at the alarm. The layers that price the alarm, sequence it into a drivable day, and learn from what the crew found sit above all of them and read from whichever one you already run.

Why does pump by exception stall at scale?

Five failure modes repeat. Alarm overload: exception volume scales with well count and rule count while human attention does not. Every exception looks equal: the engine has no concept of working interest, price, or lifting cost. Threshold drift: limits set on a new well go stale as the well declines and nobody is staffed to re-tune thousands of them. No defensible deferral: skipping an item is the right call with no record of why. And the system never learns from what the field found. None of the five is fixed by more SCADA, more rules, or more dashboards. Each is a missing layer, not a missing signal.

What comes after pump by exception?

A four-rung ladder. Rung 1 rationalizes the queue with per-well bands that follow decline. Rung 2 puts a dollar figure and a confidence band on every row. Rung 3 turns the scored queue into a constraint-sequenced route rather than a list. Rung 4 closes the loop, so outcomes retrain the scoring and the plan re-ranks mid-shift. Each rung pays for itself and no rung’s business case depends on reaching the next one.

Do we have to replace our exception system to move up the ladder?

No. The moves are additive. The SCADA stays, the historian stays, the alarm discipline stays, and the three layers keep doing their jobs. What changes is what flows through them and how the field consumes it. Teams that present the evolution internally as a rip-and-replace create resistance they did not need to create.

How do you set exception thresholds without creating alarm fatigue?

Three rules do most of the work. Set the band per well rather than per fleet, because a limit that is right for a new well is wrong for the same well two years later. Let the band follow the decline forecast, so the definition of normal moves as the well moves and nobody has to re-tune thousands of limits by hand. And require persistence before an item becomes an exception, so a transient that clears on its own never reaches a queue. The fourth rule is the one operators skip: validate the reading first. Calibration drift, stuck values, communication gaps, and physically impossible spikes should be caught and tagged before they are allowed to earn a truck roll.

Does pump by exception satisfy required well inspections?

Regulatory and lease-required visits keep their own cadence and are planned as their own class of work. Exception-driven visits answer the question of where discretionary field attention goes; they do not stand in for a visit a regulator or a lease agreement requires on a fixed interval. In a working deployment both live in one plan: required visits carry their deadline as an absolute constraint, and the value ranking orders everything around them, so a compliance round and a production exception on the same lease are sequenced into the same drive rather than into two trips.

How many wells can one field team member cover under pump by exception?

The binding constraint is exception volume and attention, not well count. A queue of 30 to 50 active exceptions per lease operator per shift is routine in mature deployments (ANSI/ISA-18.2 context, as cited in WorkSync Research Volume IV), and past roughly a dozen items a queue stops functioning as a work list. That is why span of control expands with the ranking layer rather than with the detection layer: once the queue is rationalized, scored, and sequenced, the same person spends the day on the items that carry the value instead of triaging a list. Any specific well-per-person figure depends on well density, drive times, failure rates, and instrumentation coverage, so it is worth measuring on your own field rather than adopting from someone else.

Are your best people working on your most valuable work today?

Keep the exceptions. Add the price and the route.

Bring one field’s data. We will show you the ranked plan your crew would have run today, what each stop was worth, and which rung it puts you on.

Continue the cluster: Production Surveillance · Lease Operator Routes · WellOPS, the field operations platform