Skip to main content
← White papers
WorkSync · Research · Volume IV · June 2026

Taking Pump by Exception to the Next Level

A practical playbook for evolving exception-based operations into prioritized, economically ranked daily plans.

Pump by exception was the right idea with a ceiling: queues that outrun attention, a 5 BOPD stripper alarming like a 200 BOPD producer, thresholds drifting off declining wells. This paper is the operator-level playbook for the next level. Three moves, in order: prioritize the exceptions, score them in dollars, and turn the scored queue into a ranked daily plan, plus the operating-model changes that make the moves stick. Light on math by design: the formal treatment is Volume I, Pump by Priority. This is the field manual.

3
moves: prioritize, score, rank
90
days, gated implementation path
+15%
free cash flow, same crew
35%
fewer miles driven
1.8 → 0.3
TRIR at the reference deployment
5,000+
wells, three basins

Abstract

Pump by exception was the right idea. It replaced the fixed route with the signal, cut routine site visits, and moved lease operators away from windshield time and toward productive work. It also has a ceiling, and the operators with the cleanest exception-based deployments are the first to feel it: alarm queues that outrun operator attention, a 5 BOPD stripper alarming with the same urgency as a 200 BOPD producer, thresholds that drift away from declining wells, and a system that never learns from what the field found.

This paper is the practical, operator-level playbook for taking pump by exception to its next level without discarding the SCADA investment, the rules engine, or the field workflows that already work. We present three moves, in the order operating teams should make them: (1) prioritized exceptions, the alarm-rationalization and per-well-band discipline that cuts queue volume before anything else changes; (2) economic scoring of alarms, the dollar substrate that lets a foreman distinguish a $200-per-day nuisance from a $15,000-per-day production loss; and (3) the ranked daily plan, which converts the scored queue into a constraint-respecting route that is in the truck cab by 6 AM.

Around the three moves we describe the operating-model changes that make them stick: how the morning meeting changes, how the foreman’s job changes, which KPIs to retire and which to adopt, and a 90-day implementation path gated by measurable checkpoints. The destination of this evolution is the closed-loop operating model we call Pump by Priority, formalized mathematically in Volume I of this series; the present paper is deliberately light on mathematics and heavy on operating practice. Reference outcomes measured at a 5,000-plus-well, three-basin deployment of the full model, against pre-deployment baselines: 15% free-cash-flow uplift on the same crew, 35% fewer miles driven, and TRIR from 1.8 to 0.3.

Contents of the 15-page playbook

  1. 1Introduction: The Queue Nobody Can Triage
  2. 2What Pump by Exception Got Right
  3. 3The Ceiling: Five Failure Modes You Can See from the Truck
  4. 4The Playbook at a Glance: Three Moves and a Ladder
  5. 5Move One: Prioritized Exceptions
  6. 6Move Two: Economic Scoring of Alarms
  7. 7Move Three: The Ranked Daily Plan
  8. 8The Operating-Model Change That Makes It Stick
  9. 9A 90-Day Implementation Path
  10. 10Measuring the Move
  11. 11The Bridge to Pump by Priority
  12. 12Conclusion
  13. ·References

Every section is published in full further down this page.

Preview · First 3 pages15 pages total
Taking Pump by Exception to the Next Level: white paper coverP.1
Taking Pump by Exception to the Next Level: abstract and introductionP.2
Download

Get the formatted PDF

The paper is published in full on this page. This download is the designed, printable document. Get all 15 pages: the three moves (prioritized exceptions, economic scoring of alarms, the ranked daily plan), the operating-model changes that make them stick, the KPI swap, and the gated 90-day implementation path. We will follow up only if it makes sense.

15 pages · PDF

Read the full paper

Get the formatted PDF ↑

The complete 15-page paper, published in full below. Prefer the designed, printable document? Request the PDF above.

Keywords: pump by exception, exception-based surveillance, alarm rationalization, economic scoring, ranked daily plan, pump by priority, field operations, upstream oil and gas, operating model.

1Introduction: The Queue Nobody Can Triage

It is 5:40 on a Tuesday morning. A lease operator responsible for 150 wells opens the tablet and finds the overnight queue: 30 to 50 active exceptions, which is a normal Tuesday (ANSI/ISA, 2016). High casing pressure on a well that always runs high casing pressure. A tank-level alarm that may be a stuck float. A communication failure on an RTU that has failed communication four times this month. A low-rate flag on a well that, although nobody in the queue can see it, is the highest-cash-flow asset in the operator's patch and has been drifting off forecast for three days.

The operator does what experienced operators do. He triages by proximity, by recency, by which well burned him last quarter, and by instinct. Some mornings the instinct is right. Across a full year of mornings, multiplied across every operator in the company, it measurably is not: the dollar spread between the most and least valuable exception in a typical queue exceeds two orders of magnitude, and no human carrying 150 wells can hold that ranking in his head while driving a lease road.

This is not a failure of the operator, and it is not a failure of pump by exception. It is the predictable ceiling of an operating model that was designed to answer one question (which wells changed?) being asked to answer a different one (what is the most valuable thing my crew can do today?).

1.1What this paper is, and is not

This paper is an implementation playbook. It is written for the operations manager, field foreman, production superintendent, and VP of Operations who already run exception-based surveillance, who are broadly happy with the decision to adopt it, and who can feel its ceiling. It describes, at the operating level, how to evolve a working pump-by-exception deployment into a prioritized, economically scored, ranked-plan operation, what changes in the field and in the morning meeting, what to measure, and in what order to make the moves.

It is deliberately not a mathematical treatment. The formal successor operating model, Pump by Priority, is specified rigorously in Volume I of this series (WorkSync Research Team, 2026a): the Bayesian detection layer, the risk-adjusted value estimation, the constraint-aware routing solver, and the closed-loop learning machinery, with theorems and benchmarks. Readers who want the proofs should read Volume I. Readers who want to know what their foremen should do differently in March than they did in February should read this paper first. The two documents are companions: Volume I is the why and the what; this volume is the how.

1.2The thesis in three sentences

Pump by exception got the architecture right and stopped one layer short: it detects change but does not value it. The next level is reached by three sequential moves, each of which pays for itself before the next begins: prioritize the exceptions, score them in dollars, and convert the scored queue into a ranked daily plan that respects real-world constraints. The moves are 20% technology and 80% operating model, which is why this playbook spends most of its pages on the operating model.

2What Pump by Exception Got Right

Any honest playbook for evolving a model should begin by crediting it. Pump by exception earned its place as the dominant field operating model of the late 2010s for good reasons, and the evolution described in this paper preserves every one of them.

2.1It replaced the calendar with the signal

The fixed-route model that preceded exception-based operations had a simple, expensive flaw: attention was allocated by geography and calendar rather than by need. By the industry's own accounting of the fixed-route era, up to 60% of scheduled visits found nothing wrong. Trucks rolled, fuel burned, exposure hours accumulated on rural lease roads, and the well that failed on Monday waited until its Wednesday slot to be discovered. Exception-based surveillance inverted the default: visit nothing unless the data says otherwise.

The published operational record on that inversion is strong. The canonical industry analysis found lease-operator value-added time moving from roughly 25% to roughly 60% of the day under exception-based surveillance, with production downtime reduced by roughly a third through earlier intervention on degradation signals already present in the data (Alvarez & Marsal, 2015). The mechanism behind both figures is the same: attention stops being allocated to wells the data says are fine, and starts being allocated where a signal says something changed.

2.2It built the three-layer architecture everything else stands on

A working pump-by-exception deployment already has the three layers that every subsequent evolution reuses:

  1. Sensing. SCADA, pressure transducers, flow meters, tank-level sensors, and RTUs streaming to a historian at second-to-minute cadence.
  2. Rules. Configurable thresholds defining a normal operating window per parameter, with compound rules in more sophisticated engines.
  3. Delivery. Mobile push of exceptions to the field, with manual or proximity-based dispatch.

Nothing in this playbook removes a layer. The moves in Sections 5 through 7 are additive: they change what flows through the layers and how the field consumes it. This matters for budgeting and for politics. The evolution described here is not a rip-and-replace, and the teams that present it internally as one will create resistance they did not need to create.

2.3It taught the organization to trust data over habit

The least-discussed achievement of pump by exception is cultural. An organization that has run exceptions for several years has already fought and won the hard argument: that a signal in the historian can outrank a habit in the truck. Operators who lived through that transition are the easiest population in the industry to bring along to the next level, because the next level is the same argument applied one layer up: a dollar estimate in the queue can outrank a gut ranking of the queue.

3The Ceiling: Five Failure Modes You Can See from the Truck

The failure modes of mature exception-based operations are well documented (Alvarez & Marsal, 2015; ANSI/ISA, 2016; EEMUA, 2013), and Volume I treats them formally. Here we describe them the way they present in the field, because the field presentation is what the playbook has to fix.

3.1Alarm overload

Exception volume scales with well count, sensor count, and rule count; operator attention does not. A queue of 30 to 50 active exceptions per operator per shift is routine in mature deployments. Past roughly a dozen items, the queue stops functioning as a work list and starts functioning as background noise with occasional spikes. The operational tell: operators who acknowledge alarms in batches without reading them, and foremen who maintain a private mental list of "wells that actually matter" because the official queue stopped carrying that information.

3.2All exceptions look equal

A low-pressure alarm on a 5 BOPD stripper renders identically to a pump failure on a 200 BOPD producer: same color, same row height, same notification sound. The system has no concept of working interest, commodity price, lifting cost, or deferment risk. The alarm does not know what a barrel is worth, so the queue cannot know what an hour of crew time is worth. The operational tell: the morning argument about where to send the workover rig is settled by whoever argues loudest, because no number on the screen can settle it.

3.3Thresholds drift away from declining wells

Every threshold was correct on the day it was set. Wells decline; thresholds do not. A static low-rate alarm set against month-one production is, by month eighteen, either firing constantly (and being ignored) or has been manually widened so far that it can no longer detect a real problem. Re-tuning thresholds by hand across thousands of wells is a job nobody is staffed for, so in practice it happens rarely. The operational tell: wells with alarm limits that have been edited more than twice, and wells where the limit is now below the well's economic limit.

3.4No economic context, no defensible deferral

When everything is urgent, deferral is guesswork, and guesswork cannot be defended in the variance review. An operator who skips a low-value alarm to chase a high-value one is making the right call with no record of why. The organization experiences this as inconsistency between crews and as variance-report archaeology: discovering two billing cycles later that a high-value well sat compromised while crews worked nuisance items.

3.5The system never learns

The most expensive failure mode is the quietest. The field closes out hundreds of work orders a month. Each one contains the answer to the question the system should be asking: was this exception worth the visit? In a classical deployment that answer goes nowhere. False-alarm rates do not fall, severity estimates do not improve, and the queue in year five is no smarter than the queue in year one.

Remark. None of the five failure modes is fixed by more SCADA, more rules, or more dashboards. Each is a missing layer, not a missing signal: a priority layer, a value layer, a planning layer, and a learning layer. That is the architecture of the next level.

4The Playbook at a Glance: Three Moves and a Ladder

4.1The maturity ladder

We find it useful to put four rungs on the ladder between classical pump by exception and the full closed-loop model:

Table 1: The maturity ladder from pump by exception to Pump by Priority.
RungNameWhat the operator sees at 6 AM
0Pump by exceptionAn unranked queue of threshold breaches.
1Prioritized exceptionsA shorter queue: rationalized alarms, per-well bands that follow decline, persistence rules suppressing transients.
2Scored exceptionsThe same queue with a dollar figure and a confidence band on every row.
3Ranked daily planA route, not a queue: the day's work ordered by value under constraints, with reasons attached.
4Pump by PriorityThe closed loop: outcomes feed back, scoring improves, the plan re-ranks mid-shift, measurement tasks compete with production tasks. Volume I's subject.

Three properties of the ladder matter more than its labels. First, each rung pays for itself: queue rationalization alone recovers operator attention; scoring alone changes the workover-rig argument; the ranked plan alone moves drive time. No rung's business case depends on reaching the next. Second, the rungs are ordered by trust, not by technology. Scoring an unrationalized queue produces precise rankings of garbage; ranking an unscored queue produces routes to the wrong wells. Third, rung 4 is a different kind of step: rungs 1 to 3 are operating-practice changes an organization can drive with its existing systems plus a scoring and planning layer; rung 4 adds machinery (closed-loop learning, multi-source fusion, value-of-information sensing) that justifies the formal treatment in Volume I.

4.2The three moves

The remainder of the paper takes the moves in order (Sections 5 to 7), then the operating-model changes that hold them in place (Section 8), the 90-day path (Section 9), measurement (Section 10), and the bridge to the full model (Section 11).

Move 1
Prioritized exceptions
Cut the noise.
Weeks 1 to 4.
Move 2
Economic scoring of alarms
Put dollars on the queue.
Weeks 3 to 8.
Move 3
The ranked daily plan
Turn the queue into a route.
Weeks 7 to 13.
↩ field close-out feedback tightens all three
Figure 1: The three moves, in order, with overlapping implementation windows. The gold return path is the habit (Section 8) that prepares an organization for the full closed loop.

5Move One: Prioritized Exceptions

Move one is queue hygiene, and it is deliberately unglamorous. Its goal is a queue small enough and honest enough to be worth scoring. Experience with alarm management in the process industries (ANSI/ISA, 2016; EEMUA, 2013) provides the discipline; the upstream-specific work is adapting that discipline to thousands of small, declining, heterogeneous assets rather than one large plant.

5.1Rationalize the alarm base

Run the classical rationalization questions over every configured alarm, per well: Does a defined operator response exist? Is the consequence of inaction stated? Is the limit value justified by the well's current behavior rather than its behavior at completion? Alarms failing the first two questions are candidates for demotion to logged-only events. In mature upstream deployments it is common for a substantial fraction of configured alarms to fail rationalization, most often because they were copied from a template well and never revisited.

Practical guidance for the rationalization pass:

  • Work by failure mode, not by well. Rationalizing "high casing pressure" across the whole field in one sitting produces consistent logic and takes days. Rationalizing well-by-well produces drift and takes months.
  • Put a foreman and a production engineer in the same room. The foreman knows which alarms the field ignores; the engineer knows which signals predict real failures. The intersection is the keep list.
  • Record the response with the alarm. An alarm whose defined response is "call the foreman to ask what to do" is not yet rationalized.

5.2Let the band follow the decline

The single highest-leverage technical change in move one is replacing static thresholds with per-well expected-behavior bands that update as the well declines. The principle: the question is never "is rate below X?" but "is rate below what this well should be doing this week?" A band fitted to the well's own decline trend and noise envelope, re-fit on a regular cadence, fires when the well departs from its own expectation. Wells on decline stop generating creeping false alarms; wells that step change get caught the day it happens, not when they finally cross a stale limit. Volume I specifies the full Bayesian machinery (decline-curve priors, heteroscedastic confidence bands); for the purposes of this playbook, what matters is the operating property: nobody has to re-tune thresholds by hand again, which is the reason threshold drift was never fixed under the manual regime.

5.3Suppress transients with persistence rules

Single-reading spikes (communication hiccups, gauger activity on location, momentary slugging) should not page a human. A persistence requirement, where deviation must hold across a defined window before an exception is raised, removes a large class of nuisance alarms at the cost of minutes of detection latency, a trade that is almost always correct for non-safety signals. Safety-class alarms are explicitly excluded: they keep their immediate paths, unmodified. That exclusion is not a footnote; it is what makes the rest of the rationalization defensible to the HSE organization.

5.4What to measure in move one

Three numbers, sampled weekly, tell the team whether move one is working: average active exceptions per operator per shift (should fall steadily toward a queue a human can actually read); percentage of exceptions that resulted in field action when visited (should rise as nuisance volume falls); and count of alarm-limit manual edits (should fall toward zero as bands take over). These are also the trust-building numbers for the field: operators who watch the queue get shorter and truer will follow the team to move two. Operators handed a dollar-ranked version of a queue they already distrust will not.

6Move Two: Economic Scoring of Alarms

Move two attaches a dollar figure to every item in the rationalized queue. This is the substrate change: the moment the queue stops being a list of engineering deviations and becomes a list of business decisions.

6.1The one formula this paper needs

This playbook needs exactly one formula, and the field version of it fits on an index card. The value at risk of deferring exception i for Δ days is approximately

Vi(Δ)  ≈  qloss × P × WI × NRI factordaily revenue exposure × Δ  +  pesc(Δ) × Cescescalation risk  −  Cvisit,
(1)

where qloss is the production affected (full rate for a down well, partial for a compromised one), P is price, WI and NRI capture the operator's economic interest, pesc(Δ) is the probability the condition escalates into a costlier failure if left for Δ days, Cesc is the cost of that escalation, and Cvisit is the cost of the visit itself. Volume I develops the rigorous form (a discounted expected-lost-cash-flow integral with risk adjustment and Monte Carlo evaluation over the posterior); equation (1) is its operating approximation, and the difference rarely changes the top five rows of a queue.

6.2A worked illustration

The following two-row comparison is an illustration with invented inputs, constructed to show the mechanics; it is not customer data.

Table 2: Illustration only: the same queue, before and after scoring. Inputs are invented for the example.
ExceptionAffected rateEscalation riskScore (eq. 1)Unscored rank
Low rate, flagship producer drifting 8% off its bandpartial, largemoderatehigh (thousands/day)#7 of 12
Low pressure, stripper well, recurringfull, smalllowlow (tens/day)#2 of 12

Under timestamp ordering the stripper well gets visited first because it alarmed at 4 AM and renders at the top. Under scoring, the flagship producer's drift, two orders of magnitude more valuable, is the first stop on someone's route. Nothing about the detection changed. What changed is that the queue now answers the question the foreman was actually asking.

6.3Severity tiers the field can hold

Raw dollar figures rank well but communicate poorly at 6 AM. Most teams bucket scores into a small number of tiers with attached service-level expectations. A pattern that has worked in practice:

Table 3: A four-tier severity scheme. Tier boundaries are operator-calibrated; the structure is the point.
TierMeaningResponse expectation
T1Major economic exposure or fast escalation riskFirst stops of today's plan
T2Material exposure, slow escalationToday, after T1
T3Real but small; batches well geographicallyThis week, clustered with nearby work
T4Logged; below visit economicsNo visit; watch for promotion

T4 deserves emphasis because it is the tier that did not exist under classical pump by exception: the explicit, recorded, defensible decision not to roll a truck, made by arithmetic rather than by fatigue. When equation (1) evaluates negative (the visit costs more than the exposure), the system should say so out loud. A large share of the 35% reduction in miles driven at the reference deployment (Section 10) is T4 working as designed.

6.4Safety and compliance are gates, not scores

One rule must be stated unambiguously, in the scoring standard and in training: safety-critical and regulatory items are never ranked by dollars. They sit above the economic queue as hard gates: a leak indication, an H2S alarm, a compliance task inside its regulatory window. These dispatch first regardless of any score, and the scoring layer is forbidden from trading them against production value. The same posture appears as hard constraints in the routing formulation of Volume I and Volume II (WorkSync Research Team, 2026a,b); at the playbook level it is simpler to state: dollars rank the discretionary; gates own the non-discretionary. Teams that blur this distinction lose the HSE organization and deserve to.

6.5Where the inputs come from

Every input to equation (1) already exists somewhere in the operator's systems: rates and trends in the historian and production database, prices in the marketing or accounting system, WI/NRI in land or accounting, lifting and visit costs in finance's LOE model, escalation probabilities initially from engineering judgment by failure mode (later, at rung 4, learned from outcomes). The integration is read-only. No system of record is replaced, and the accounting team's numbers are used rather than re-invented, which is also the political path of least resistance.

7Move Three: The Ranked Daily Plan

A scored queue is still a queue. Move three converts it into the artifact that actually changes a shift: a per-crew, ordered plan, in the cab by 6 AM, that respects the constraints a dispatcher has to respect and shows its reasoning.

7.1From queue to route

The conversion is a constrained assignment problem: which crew takes which tasks, in what order, such that expected value captured is maximized subject to hard constraints. The constraint classes that matter in the field:

  • Qualification. The operator without the right certification is never assigned the task that requires it.
  • Regulatory windows. Inspections and compliance tasks with deadlines enter the plan as fixed obligations, not as suggestions.
  • Geography and shift length. Drive time is a cost on the objective; the shift end is a boundary, with the realities of schedule patterns (5/2, 7/7, 14/14) respected.
  • Equipment and synchronization. Tasks requiring the hot-oiler, the workover rig, or two people on location are scheduled when the dependency is available, not before.

Volume II treats this problem formally (a multi-class vehicle-routing formulation and the solver machinery to re-plan in seconds (WorkSync Research Team, 2026b)); operationally, what the team needs to know is that the technology to do this conversion in minutes, and to re-do it mid-shift when the world changes, exists and is field-proven. The playbook question is not whether the route can be computed. It is whether the organization will run on it, which is Section 8's subject.

7.2The plan must show its reasoning

A ranked plan that arrives as an oracle gets ignored by week three. Each line of the plan should carry, in plain language: what tripped (the exception and its band departure), why it ranks here (the dollar figure and tier), and what good looks like (the defined response from rationalization). Operators accept being re-ordered by a system that can explain itself, the same psychology that made exception-based surveillance adoptable in the first place.

7.3The override is a feature

The field knows things the model does not: the lease road that washed out, the operator who heard the bearing yesterday, the tank that gauges differently than it reads. The plan must be overridable, and the override must be one tap, with an optional reason code. Then, critically, the override is recorded as data. A well that operators consistently pull forward is carrying information the scoring has not captured; a recommendation consistently deferred without consequence is over-valued. At rungs 1 to 3 a human reviews override patterns monthly; at rung 4 the closed loop consumes them automatically. Either way, the override stream is the single richest feedback signal the organization owns, and classical deployments throw it away.

7.4Mid-shift re-ranking

The 6 AM plan is a forecast, and by 10 AM it is partly wrong: a new T1 fired, a job ran long, a part is not on the truck. The plan should re-rank on a defined trigger set (new high-tier exception, task completion variance beyond a limit, crew availability change), pushing a revised remainder-of-day to affected crews only. Two disciplines keep this from becoming churn: a minimum improvement bar before a revision is pushed (small re-orderings are not worth the cognitive cost), and a rule that in-progress tasks are never interrupted by anything short of a safety gate.

8The Operating-Model Change That Makes It Stick

Every element in Sections 5 to 7 has been deployed somewhere and quietly abandoned, because the technology was installed and the operating model was not changed around it. This section is the part of the playbook most implementations skip.

8.1The morning meeting inverts

Under pump by exception, the morning meeting assembles the plan: the foreman reads the queue, debates, assigns. Under the ranked plan, the plan already exists at 6 AM, and the meeting's job inverts to exception handling on the plan: review the T1s, hear the overrides and their reasons, confirm the gates (safety, regulatory) are staffed, and release. Teams report the meeting shortening materially, but the more important change is qualitative: the conversation moves from "what should we do?" to "is the plan wrong anywhere, and why?", which is precisely the feedback the system needs to improve.

8.2The foreman's job moves up a level

The foreman under classical PBE is a human router: consuming alarms, applying tribal knowledge, dispatching. The moves do not eliminate that judgment; they relocate it.

Table 4: The foreman's day, before and after the three moves.
Before: human routerAfter: plan editor and teacher
Reads the raw queue at 5:30, triages by memoryReviews a ranked plan at 6:00, edits the places where the field knows better
Settles dispatch arguments by senioritySettles them by the score, overrides with a reason when the score is wrong
Re-tunes alarm limits when complaints accumulateReviews band exceptions monthly; the bands maintain themselves
Carries the "wells that matter" list in his headAudits the system's ranking against that list, and files the differences as feedback

The framing that lands with field leadership: the system is taking the clerical part of the job (reading, sorting, remembering) and leaving the judgment part (knowing when the model is wrong). Given the demographics of the lease-operator workforce, a meaningful share of whom are within sight of retirement (U.S. Bureau of Labor Statistics, 2025), encoding the triage layer in a system rather than in retiring heads is also a continuity decision, and experienced foremen, asked to spend their last working years teaching the system rather than racing the queue, tend to hear that correctly when it is said plainly.

8.3Retire the old KPIs deliberately

Operating models follow measurement. If the field is still measured on visits per day and route completion, the ranked plan will lose, because the ranked plan deliberately reduces visits. The KPI swap should be explicit, announced, and dated:

Table 5: The KPI swap. Retired metrics reward motion; adopted metrics reward value.
RetireAdopt
Visits per operator-dayValue-weighted plan completion (the dollars on the plan that got captured)
Average alarm count (as a workload boast)Actionable-exception rate (visits that found real work)
Route completion percentageDeferment hours on top-tier wells
Time-to-acknowledge (alone)Time-to-restore on T1/T2, and windshield hours per BOE

8.4Close-out codes: the habit that feeds the loop

The cheapest high-value change in this entire paper: a structured close-out on every visited exception, three taps in the field app. Found-as-flagged / found-different / found-nothing, plus actual production impact where known. This stream is what turns scoring from configured arithmetic into a calibrated instrument: found-nothing rates by alarm type identify the rationalization backlog; found-different rates identify scoring blind spots. At rung 4 this stream becomes the training signal for closed-loop learning (WorkSync Research Team, 2026a). Organizations that build the close-out habit during move one arrive at the full model with a year of labeled outcomes already banked; it is the single best preparation for the closed loop, and it costs the field seconds per job.

9A 90-Day Implementation Path

The playbook compresses into a 90-day path with measurable gates. Two design rules: pilot on one route team, not the whole field (champions volunteer; skeptics watch the scoreboard); and never advance through a gate on the calendar alone, advance on the gate's numbers.

Table 6: The 90-day path. Gates advance on evidence, not on schedule.
WindowWorkGate to advance
Weeks 1 to 2Baseline: queue volumes, actionable rates, deferment on top wells, current KPIs. Read-only integration to historian, production DB, accounting.Baseline published and agreed; nobody disputes the starting numbers later.
Weeks 3 to 4Move one on the pilot route: rationalization by failure mode, bands on, persistence rules on. Close-out codes begin.Queue per operator-shift down materially; actionable rate rising; field reports the queue "reads true."
Weeks 5 to 8Move two: scoring live on the rationalized queue, tiers defined, safety/regulatory gates wired and tested. Scores visible alongside the familiar queue first (shadow mode), then ranked view becomes default.Foreman uses the score unprompted in the morning meeting; T4 no-visit decisions being made and recorded.
Weeks 9 to 13Move three: ranked plan to the pilot crew's cabs at 6 AM, overrides with reason codes, mid-shift re-rank on triggers. KPI swap announced for the pilot team.Plan adherence with explained overrides; windshield hours falling; zero gate misses; pilot team would not give it back.

Two failure patterns account for most stalled implementations, and both are avoidable by fiat. Skipping move one: scoring a noisy queue teaches the field that the dollar figures are precise nonsense, and trust, once spent there, is expensive to re-earn. Running shadow mode forever: a scored view that never becomes the default is a dashboard, and dashboards do not change Tuesday. The deployment record across exception-based programs is blunt on this point: forecasts that do not change tomorrow's work order do not change the numbers. Set the date on which the ranked view becomes the default, and keep it.

After day 90, expansion is route team by route team, each inheriting calibrated bands and scoring from the pilot, each faster than the last. The full closed loop (rung 4) becomes the natural next conversation once two or three route teams are running ranked plans and the close-out stream has accumulated.

10Measuring the Move

10.1The scoreboard

Four families of metrics, all derivable from systems the operator already runs, cover the playbook end to end: attention (exceptions per operator-shift, actionable-exception rate), economics (deferment hours on top-tier wells, value-weighted plan completion, LOE per BOE trend, free cash flow on the same crew), exposure (windshield hours, miles driven per crew-shift, TRIR trend), and system health (override rate with reasons, found-nothing rate, band re-fit currency). The baseline from weeks 1 to 2 is what makes the scoreboard honest; operators who skip the baseline end up arguing about counterfactuals in the quarterly review.

10.2The reference deployment

For the destination numbers, the full Pump by Priority model, of which the three moves are the implementation on-ramp, has a published reference deployment at a top 25 private producer operating 5,000-plus wells across three basins (Western Anadarko, Permian, and Wyoming), measured against pre-deployment baselines (WorkSync Research Team, 2026a,b):

  • 15% free-cash-flow uplift on the same crew: more value captured per crew-day, not more crew-days.
  • 35% fewer miles driven: the compounding of T4 no-visit decisions, geographic batching of T3 work, and visits that stopped finding nothing.
  • TRIR from 1.8 to 0.3: fewer truck hours and qualification gates enforced at dispatch; the safety case and the economic case are the same case.

Outcomes are specific to that deployment and its calibration; a new basin, equipment mix, or operating philosophy requires a calibration period before results can be compared. What transfers is the structure: every mechanism behind those numbers is one of the three moves running at full maturity, plus the closed loop tightening them over time.

11The Bridge to Pump by Priority

An organization that has run the three moves for two or three quarters has, without ceremony, built most of the prerequisites for the closed-loop model: a rationalized signal base, a dollar substrate the field trusts, a constraint-respecting plan the field runs, and an accumulating stream of labeled outcomes. What remains is the machinery that Volume I formalizes, and it is worth naming plainly what each piece adds:

  • Closed-loop learning. Scoring weights update from outcomes automatically, with drift detection, rather than via monthly human review. The system in year two is measurably better calibrated than in year one.
  • Multi-source fusion. SCADA, manual gauges, run tickets, and allocation data are reconciled into one estimate with a disagreement test, so the plan stops inheriting single-source errors, and discrepancies (theft, leaks, miscalibration) surface as ranked work.
  • Value-of-information sensing. The system starts asking for measurements (gauge this tank today, the uncertainty is decision-relevant) on the same dollar substrate as production work.
  • Confidence-aware handoff. Low-confidence recommendations announce themselves as such and route to a human, which is the honest version of automation.

One more reason not to stop at rung 3, stated without adjectives: the learning loop compounds. Two operators can buy identical software; the one whose operating model feeds outcomes back accumulates calibration the other never gets, and the gap between them widens every quarter the loop runs. In a consolidating basin, that compounding is the difference between the acquirer's operating model and the acquired's. The window in which moving early is a differentiator rather than a catch-up is, on the current evidence, measured in a small number of years (BCG, 2025; McKinsey, 2024).

12Conclusion

Pump by exception was the right idea, executed one layer short of its potential. The ceiling its best operators feel is not a reason to abandon the model; it is the model asking for its next layer: priority on the queue, dollars on the priorities, a plan built from the dollars, and an operating model that feeds what the field learns back into the system.

The playbook is deliberately modest in its demands. No SCADA replacement. No data-cleanup year. No bet on a distant platform program. Three moves, each self-funding, each measurable inside a quarter, run on one pilot route by a foreman who can override anything the system says, provided he tells it why. The mathematics that waits at the end of the path is in Volume I for the readers who want it. The foremen do not need the theorems. They need a queue that reads true, a score that settles arguments, and a plan in the cab at 6 AM, and that, concretely, is what taking pump by exception to the next level means.

References

Alvarez & Marsal. The Advantages of Exception-Based Surveillance. Industry whitepaper, 2015.

ANSI/ISA. ANSI/ISA-18.2-2016: Management of Alarm Systems for the Process Industries. International Society of Automation, 2016.

EEMUA. Publication 191: Alarm Systems, a Guide to Design, Management and Procurement. Engineering Equipment and Materials Users Association, 2013.

P. F. Drucker. The Practice of Management. Harper & Brothers, 1954.

J. J. Arps. Analysis of decline curves. Transactions of the AIME, 160(1):228–247, 1945.

U.S. Bureau of Labor Statistics. Labor force statistics: age distribution of oil and gas extraction occupations. 2025.

Boston Consulting Group. The industrial AI value pool in energy operations. Industry analysis, 2025.

McKinsey & Company. Agentic operations: the next operating-model shift in heavy industry. Industry analysis, 2024.

WorkSync Research Team. Pump by Priority: closed-loop AI-driven operational execution in upstream oil and gas. WorkSync Industry Whitepaper, Volume I, 2026. https://www.work-sync.ai/white-papers/pump-by-priority.

WorkSync Research Team. Optimal route planning in oil and gas: multi-class field work and constraint-aware crew dispatch. WorkSync Industry Whitepaper, Volume II, 2026. https://www.work-sync.ai/white-papers/optimal-route-planning.

Where the playbook leads

The destination numbers, measured at the full Pump by Priority reference deployment: a top-25 private producer across three basins (Western Anadarko, Permian, and Wyoming), 5,000+ wells, same crews, against pre-deployment baselines.

MetricBaselineOutcome
Free cash flow, same crewbaseline+15%
Miles driven, deployment figurebaseline-35%
TRIR (recordables / 200k hrs)1.80.3

Outcomes are specific to that deployment and calibration. New basins, equipment classes, or operating philosophies require a calibration period before results match these benchmarks.

Keep your exception system. Add the layers it is missing.

The playbook is the how. WellOPS is the system that runs it: the scored, constraint-aware, ranked plan in every truck cab by 6 AM, on top of the SCADA and rules you already own.