Skip to main content

Management by Exception · Operating Model

Management by Exception in Oil & Gas Operations. Alarm only on what matters. Rank by what pays.

Management by exception is the operating model behind every efficient upstream operation: limited human attention goes only where the data says it pays, and asset performance improves on the wells you already own. The discipline traces from Taylor and Drucker through manufacturing and IT ops into oilfield SCADA. Its upstream form, pump by exception, is where most operators start, and it earned real gains wherever alarm discipline held. After a decade-plus of hard-won industry lessons, WorkSync productized the next step as pump by priority. Whether you are a large operator stalled mid-rollout with the field pushing back, or a smaller operator with no clear place to start, the model is now something you implement in weeks, not years.

Definition

What is management by exception in oil and gas?

Management by exception (MBE) is an operating model where leaders and field crews focus attention only on items that have deviated meaningfully from expected performance. Routine, in-spec activity is trusted to continue on its own. The limited human attention available goes to the small set that has moved out of bounds. Operators reach for it when the goal is to improve asset performance on existing wells: less downtime, less deferred production, lower unit operating cost.

The inversion matters. Manage-by-walk-around scales linearly with headcount: if you double your wells you double your pumpers. Management by exception scales with software: an exception-based system that continuously monitors 5,000 wells is the same shape as one that monitors 500, only the input data set grows.

In upstream oil and gas this manifests as two specific applications: pump by exception (alarm and visit only deviating wells) and pump by priority (rank the deviations by economic impact, sequence them into a constraint-aware route, learn from every outcome). Together they replace fixed Monday-Wednesday-Friday route loops with a daily plan that only visits wells where attention is worth the windshield time.

The evolution

MBE to pump by exception to pump by priority.

One idea, three generations. Each generation keeps what the previous one got right and fixes where it stopped.

The management philosophy

Management by exception

Human attention goes only to meaningful deviations; routine, in-spec activity runs on its own. Older than the oilfield: Taylor wrote it down in 1903, and every efficient operation practices some version of it.

The oilfield implementation

Pump by exception

SCADA brings the philosophy to the wellhead: alarm and visit only deviating wells. Real progress past fixed routes, and the place most operators start. Its limit is that it stops at detection: the exception arrives with no dollar figure, no route, and no memory. The in-depth guide walks the three layers and the ceiling.

The next step: priced and routed

Pump by priority

Keeps the exception discipline and adds what detection lacked: economic scoring on every deviation, constraint-aware routing into a drivable day, and closed-loop learning from every outcome. The full side-by-side is on Pump by Exception vs Pump by Priority.

The lineage

From Taylor to the oilfield.

Exception-based management is older than the SCADA that runs your wells. Knowing the lineage matters because it tells you what the discipline already knows about its own failure modes.

1903
Taylor

Frederick W. Taylor's Shop Management originates the idea: reports condensed to the exceptions, management attention only on deviations.

1954
Drucker

Peter Drucker carries management by exception into the modern canon in "The Practice of Management".

1970s
TPS + SPC

Toyota Production System and statistical process control embed exception-based management into manufacturing lines.

1990s
IT Ops

IT operations adopt exception-based alarm management (Nagios, Tivoli, OpenView).

2000s
Pump by Exception

SCADA-driven exception alarms enter upstream oil and gas. The "pump by exception" workflow is born.

2020s
Pump by Priority

Economic scoring + ML anomaly detection layer ranking onto pump-by-exception. The next era.

Now
Closed-Loop

Outcomes feed back into scoring every shift, and Willie builds the ranked plan your pumpers run.

Where WorkSync fits

The end-to-end management by exception platform with full closed-loop connectivity between your core corporate functions and your field expertise.

We connect your team and deliver cash-flow-optimized, risk-adjusted work plans before the work day starts. Everyone closes part of the loop. Corporate-workflow agents stop at the office. Lift optimizers stop at the pump. WorkSync closes the whole loop: corporate systems to the pumper's ranked plan and back, every shift.

WellOPS is the world’s most advanced Pump by Exception platform, because it goes beyond flagging exceptions to pricing and routing them. Pump by priority is the next step from pump by exception, and the two are widely treated as the same thing. If you arrived here on the oilfield term rather than the management one, the software decision itself is laid out on the pump by exception software page.

In practice, that closed loop shows up as four outcomes on the same crew:

Outcome 01

Tackle high-value deferred production

The wells with the most production down and the fastest payback surface at the top of the plan, not wherever the route loop happens to reach them.

Outcome 02

Escalate issues efficiently

Exceptions route to the right person with dollar context attached, instead of dying in a group text or a Monday morning meeting.

Outcome 03

Manage liquid inventory

Tank levels and haul timing are worked into the same ranked plan, so loads move before tanks top out and wells shut in.

Outcome 04

Know when a well stops paying its way

The economic-limit call, made on data: stop putting money into a well that will not pay it back, and redeploy the crew where it will.

The failure modes

Where exception-based management goes wrong, and how WorkSync fixes it.

Failure mode 01

Alarm fatigue

The failure

Most SCADA-driven exception systems generate hundreds of alarms a shift. Without ranking, crews triage by recency or loudness; the highest-value alarm gets buried.

The WorkSync fix

Continuous economic scoring: every flagged exception carries a dollar-impact estimate and a tier (P1 / P2 / HIGH / MED). The morning ranked plan is sorted by $, not by timestamp.

Failure mode 02

Stale "normal"

The failure

Wells decline. Equipment ages. Operating bands shift. Fixed alarm thresholds, set during commissioning and never updated, produce false negatives (real issues hidden behind a thresholds that has crept out from under the actual performance) and false positives (alarms that mean nothing).

The WorkSync fix

Continuously updated Arps decline forecasts per well, confidence bands per signal, and ML anomaly detection that learns each asset's normal individually rather than against a fleet-wide rule.

Failure mode 03

Missing economic ranking

The failure

Exception management without economic ranking is just an alarm list. Two wells deviating at the same time may have $12,500/day vs $90/day of revenue at risk. The 10x spread is invisible without scoring.

The WorkSync fix

Cash-flow-weighted task ranking. Every potential field task scored by dollar-impact. The top 220 stops across 10 crews surface as the daily ranked plan.

Failure mode 04

No closed-loop learning

The failure

Exception lists that don't learn from outcomes get progressively worse. When a flagged exception doesn't pan out, the next similar pattern keeps getting flagged at the same priority.

The WorkSync fix

Reinforcement learning closes the loop. Every completed task feeds outcomes back into the scoring models. Each week the plan ranks the right work more accurately.

WorkSync Research · Volume IV · June 2026

The playbook for fixing all four failure modes, in 15 pages

Taking Pump by Exception to the Next Level is the operator-level implementation guide: prioritized exceptions, economic scoring of alarms, the ranked daily plan, the KPI swap, and a 90-day path gated on evidence rather than the calendar.

Read the white paper →

Continue the cluster

Exception management is one piece of the upstream optimization discipline.

Frequently asked

What VPs of Ops ask about exception-based management.

What is management by exception in oil and gas?

Management by exception is an operating model where attention goes only to the assets that have moved meaningfully away from expected performance. Routine, in-spec activity is left to run. In oil and gas it is adopted to improve performance on wells an operator already owns: less deferred production, fewer avoidable failures, lower unit operating cost from the same crew. The idea originates with Frederick W. Taylor's Shop Management (1903), was carried into the modern management canon by Peter Drucker, and reached upstream through SCADA-driven exception alarms in the 2000s.

How is management by exception different from pump by exception and pump by priority?

They are three rungs of the same ladder. Management by exception is the operating model: act only where the data says action is warranted. Pump by exception is the industry's first-generation implementation of it: SCADA thresholds flag wells that deviate, and crews visit what is flagged. Pump by priority is the productized next step: the flagged exceptions are scored in dollars, sequenced into a constraint-aware route for the crews and qualifications actually available, and every field outcome is fed back so next week's ranking is better. An operator can be doing pump by exception and still not be managing by exception, because an unranked alarm list does not allocate the day.

What do we need in place before we can run it?

Three things, and most operators have only the first. (1) A definition of normal per asset: a continuously updated decline forecast, an expected pressure band, an expected runtime envelope. Fixed thresholds set at commissioning are not a definition of normal, because wells decline underneath them. (2) A scoring layer that converts each deviation into dollars at risk, so a separator nuisance alarm and a high-rate well drifting off forecast stop arriving with equal urgency. (3) A delivery layer that puts the ranked result in the field team's hands before the shift starts. WorkSync reads the SCADA or historian, production accounting, EAM or CMMS, and GIS systems already in place, read-only, and the ranked plan is live in 4 weeks.

Does this work if only part of our field has SCADA?

Yes, with a limit worth stating plainly. Exception detection is only as good as the signal, so a well with no telemetry cannot generate a real-time exception. Those wells are still ranked, using what does exist: production accounting volumes, run tickets, gauge and haul history, work-order and failure history, and the decline forecast. They stay on a time-based cadence until better signal exists, and they are visibly labelled that way rather than quietly treated as healthy. Partial coverage is the normal starting condition, not a disqualifier, and the ranking improves as instrumentation grows.

What changes for a field team member day to day?

The work does not get heavier, the order changes. Instead of a standing route driven Monday, Wednesday, and Friday, the crew opens a plan that was rebuilt overnight: the wells whose signal moved are at the top with the dollars at risk attached, the wells whose signal did not move are not on the list. Escalation carries the same economic context, so an exception routes to the right person with a number beside it. In the reference deployment at a top 25 private producer, the same crew drove 35% fewer miles (deployment figure) and the operation measured a 15% free cash flow uplift on the same crew (deployment figure).

What are the advantages and disadvantages of management by exception?

The advantage is leverage. Attention is finite and wells are not, so pointing the attention at the small set that has moved buys capacity that headcount would otherwise have to buy. It also makes the operation legible: what got worked, and why, becomes a record instead of a habit. The limits are real and worth naming before you build on it. The model is only as good as the definition of normal behind it, so a stale baseline hides problems as effectively as it raises false ones. It says nothing about magnitude on its own, so a deviation worth $90 a day and a deviation worth $12,500 a day arrive with equal urgency until something prices them (illustrative pair). It can starve the routine work that keeps deviations rare, which is why preventive and regulatory work has to compete in the same ranking rather than sit outside it. And it needs signal: an asset with no telemetry generates no exception, so partial coverage has to be visible rather than read as health. Each of those limits is a missing layer above the detection, which is what the ranking, scoring, and feedback layers supply.

How do you measure whether management by exception is working?

Pick the operating metric first and baseline it before anything is connected. Four measures separate a working program from a busy one. Deferred production, because the point of ranking is that the expensive deviations get worked first. Operating cost per well on the same crew, because the model claims capacity rather than headcount. Miles or hours driven per unit of economic coverage, because a plan that removes low-value stops shows up here before it shows up anywhere else. And the share of exceptions that turned out to be worth the visit, which is the only measure that tells you whether the queue itself is improving. At the deployed reference, a top 25 private producer running 5,000+ wells, the measured deployment figures are a 15% free cash flow uplift on the same crew and 35% fewer miles driven for the same coverage.

What usually goes wrong when operators try this?

Four failure modes account for most of it. Alarm fatigue: hundreds of exceptions a shift with no ranking, so crews triage by recency and the highest-value item gets buried. Stale normal: baselines that do not move as wells decline, producing both false alarms and missed real ones. Missing economic ranking: exceptions worked in the order they arrived rather than the order they pay, which is the difference between a busy field and a productive one. No closed-loop learning: when a flagged exception turns out to be nothing, the next identical pattern is flagged at exactly the same priority. All four are ranking and feedback problems rather than sensor problems, which is why buying more instrumentation rarely fixes them on its own.

See exception-based ranking on your data.

Talk to our team about how management by exception can be run on your wells. The 4-week stand-up is credited toward your first license. Move the metric you anchored on, or you owe no license fee.