Management by Exception · Operating Model
Management by Exception in Oil & Gas Operations. Alarm only on what matters. Rank by what pays.
Management by exception is the operating model behind every efficient upstream operation: limited human attention goes only where the data says it pays, and asset performance improves on the wells you already own. The discipline traces from Taylor and Drucker through manufacturing and IT ops into oilfield SCADA. Its upstream form, pump by exception, is where most operators start, and it earned real gains wherever alarm discipline held. After a decade-plus of hard-won industry lessons, WorkSync productized the next step as pump by priority. Whether you are a large operator stalled mid-rollout with the field pushing back, or a smaller operator with no clear place to start, the model is now something you implement in weeks, not years.
Definition
What is management by exception in oil and gas?
Management by exception (MBE) is an operating model where leaders and field crews focus attention only on items that have deviated meaningfully from expected performance. Routine, in-spec activity is trusted to continue on its own. The limited human attention available goes to the small set that has moved out of bounds. Operators reach for it when the goal is to improve asset performance on existing wells: less downtime, less deferred production, lower unit operating cost.
The inversion matters. Manage-by-walk-around scales linearly with headcount: if you double your wells you double your pumpers. Management by exception scales with software: an exception-based system that continuously monitors 5,000 wells is the same shape as one that monitors 500, only the input data set grows.
In upstream oil and gas this manifests as two specific applications: pump by exception (alarm and visit only deviating wells) and pump by priority (rank the deviations by economic impact, sequence them into a constraint-aware route, learn from every outcome). Together they replace fixed Monday-Wednesday-Friday route loops with a daily plan that only visits wells where attention is worth the windshield time.
The evolution
MBE to pump by exception to pump by priority.
One idea, three generations. Each generation keeps what the previous one got right and fixes where it stopped.
Management by exception
Human attention goes only to meaningful deviations; routine, in-spec activity runs on its own. Older than the oilfield: Taylor wrote it down in 1903, and every efficient operation practices some version of it.
Pump by exception
SCADA brings the philosophy to the wellhead: alarm and visit only deviating wells. Real progress past fixed routes, and the place most operators start. Its limit is that it stops at detection: the exception arrives with no dollar figure, no route, and no memory. The in-depth guide walks the three layers and the ceiling.
Pump by priority
Keeps the exception discipline and adds what detection lacked: economic scoring on every deviation, constraint-aware routing into a drivable day, and closed-loop learning from every outcome. The full side-by-side is on Pump by Exception vs Pump by Priority.
The lineage
From Taylor to the oilfield.
Exception-based management is older than the SCADA that runs your wells. Knowing the lineage matters because it tells you what the discipline already knows about its own failure modes.
Frederick W. Taylor's Shop Management originates the idea: reports condensed to the exceptions, management attention only on deviations.
Peter Drucker carries management by exception into the modern canon in "The Practice of Management".
Toyota Production System and statistical process control embed exception-based management into manufacturing lines.
IT operations adopt exception-based alarm management (Nagios, Tivoli, OpenView).
SCADA-driven exception alarms enter upstream oil and gas. The "pump by exception" workflow is born.
Economic scoring + ML anomaly detection layer ranking onto pump-by-exception. The next era.
Outcomes feed back into scoring every shift, and Willie builds the ranked plan your pumpers run.
Where WorkSync fits
The end-to-end management by exception platform with full closed-loop connectivity between your core corporate functions and your field expertise.
We connect your team and deliver cash-flow-optimized, risk-adjusted work plans before the work day starts. Everyone closes part of the loop. Corporate-workflow agents stop at the office. Lift optimizers stop at the pump. WorkSync closes the whole loop: corporate systems to the pumper's ranked plan and back, every shift.
WellOPS is the world’s most advanced Pump by Exception platform, because it goes beyond flagging exceptions to pricing and routing them. Pump by priority is the next step from pump by exception, and the two are widely treated as the same thing. If you arrived here on the oilfield term rather than the management one, the software decision itself is laid out on the pump by exception software page.
In practice, that closed loop shows up as four outcomes on the same crew:
Tackle high-value deferred production
The wells with the most production down and the fastest payback surface at the top of the plan, not wherever the route loop happens to reach them.
Escalate issues efficiently
Exceptions route to the right person with dollar context attached, instead of dying in a group text or a Monday morning meeting.
Manage liquid inventory
Tank levels and haul timing are worked into the same ranked plan, so loads move before tanks top out and wells shut in.
Know when a well stops paying its way
The economic-limit call, made on data: stop putting money into a well that will not pay it back, and redeploy the crew where it will.
The failure modes
Where exception-based management goes wrong, and how WorkSync fixes it.
Failure mode 01
Alarm fatigue
The failure
Most SCADA-driven exception systems generate hundreds of alarms a shift. Without ranking, crews triage by recency or loudness; the highest-value alarm gets buried.
The WorkSync fix
Continuous economic scoring: every flagged exception carries a dollar-impact estimate and a tier (P1 / P2 / HIGH / MED). The morning ranked plan is sorted by $, not by timestamp.
Failure mode 02
Stale "normal"
The failure
Wells decline. Equipment ages. Operating bands shift. Fixed alarm thresholds, set during commissioning and never updated, produce false negatives (real issues hidden behind a thresholds that has crept out from under the actual performance) and false positives (alarms that mean nothing).
The WorkSync fix
Continuously updated Arps decline forecasts per well, confidence bands per signal, and ML anomaly detection that learns each asset's normal individually rather than against a fleet-wide rule.
Failure mode 03
Missing economic ranking
The failure
Exception management without economic ranking is just an alarm list. Two wells deviating at the same time may have $12,500/day vs $90/day of revenue at risk. The 10x spread is invisible without scoring.
The WorkSync fix
Cash-flow-weighted task ranking. Every potential field task scored by dollar-impact. The top 220 stops across 10 crews surface as the daily ranked plan.
Failure mode 04
No closed-loop learning
The failure
Exception lists that don't learn from outcomes get progressively worse. When a flagged exception doesn't pan out, the next similar pattern keeps getting flagged at the same priority.
The WorkSync fix
Reinforcement learning closes the loop. Every completed task feeds outcomes back into the scoring models. Each week the plan ranks the right work more accurately.
WorkSync Research · Volume IV · June 2026
The playbook for fixing all four failure modes, in 15 pages
Taking Pump by Exception to the Next Level is the operator-level implementation guide: prioritized exceptions, economic scoring of alarms, the ranked daily plan, the KPI swap, and a 90-day path gated on evidence rather than the calendar.
Read the white paper →Continue the cluster
Exception management is one piece of the upstream optimization discipline.
Go deeper
The management by exception guides
- How to Implement Management by Exception in Oil and Gas→
The four hard requirements and the path that works
- Pump by Exception vs Pump by Priority→
One lineage, three generations, defined
- Start Without a Data Team→
One module, one field, no data lake required
- Why Tank Gauging Is Dangerous→
The safety case for removing the routine visit
Companion pillars
The cluster
- Upstream Optimization→
The umbrella discipline
- What Is Pump by Exception?→
The in-depth guide: three layers, five failure modes, four rungs
- Pump by Exception Software→
What the software does with the queue you already have
- Pump by Priority→
Tasks ranked by $ impact
- Four Eras of Field Ops→
Where exception fits in the bigger arc
- 90-day Path Era 2 → 4→
How to upgrade your operating model
- LOE Reduction pillar→
Where exception management pays the bill
How the loop closes
Capability deep-dives
Modules that run it
Where to start
- The Productized Rollout→
The stand-up playbook: sprints, gates, crews live in weeks
- Work Engine→
The ranked plan in every truck cab
- Route Optimizer→
35% fewer miles driven
- Oilfield Route Planning→
The category guide: routing, dispatch, and how to evaluate it
- Field Data Capture→
Mobile, voice-enabled, offline
- Field Work Management→
Permits, MOC, lone worker
- DataHub→
Integration backbone, included in your WellOPS or FlowSync subscription
- Field Ops Blind Spot Test (no cost)→
Six questions, find your operating blind spots
Frequently asked
What VPs of Ops ask about exception-based management.
What is management by exception in oil and gas?
Management by exception is an operating model where attention goes only to the assets that have moved meaningfully away from expected performance. Routine, in-spec activity is left to run. In oil and gas it is adopted to improve performance on wells an operator already owns: less deferred production, fewer avoidable failures, lower unit operating cost from the same crew. The idea originates with Frederick W. Taylor's Shop Management (1903), was carried into the modern management canon by Peter Drucker, and reached upstream through SCADA-driven exception alarms in the 2000s.
How is management by exception different from pump by exception and pump by priority?
They are three rungs of the same ladder. Management by exception is the operating model: act only where the data says action is warranted. Pump by exception is the industry's first-generation implementation of it: SCADA thresholds flag wells that deviate, and crews visit what is flagged. Pump by priority is the productized next step: the flagged exceptions are scored in dollars, sequenced into a constraint-aware route for the crews and qualifications actually available, and every field outcome is fed back so next week's ranking is better. An operator can be doing pump by exception and still not be managing by exception, because an unranked alarm list does not allocate the day.
What do we need in place before we can run it?
Three things, and most operators have only the first. (1) A definition of normal per asset: a continuously updated decline forecast, an expected pressure band, an expected runtime envelope. Fixed thresholds set at commissioning are not a definition of normal, because wells decline underneath them. (2) A scoring layer that converts each deviation into dollars at risk, so a separator nuisance alarm and a high-rate well drifting off forecast stop arriving with equal urgency. (3) A delivery layer that puts the ranked result in the field team's hands before the shift starts. WorkSync reads the SCADA or historian, production accounting, EAM or CMMS, and GIS systems already in place, read-only, and the ranked plan is live in 4 weeks.
Does this work if only part of our field has SCADA?
Yes, with a limit worth stating plainly. Exception detection is only as good as the signal, so a well with no telemetry cannot generate a real-time exception. Those wells are still ranked, using what does exist: production accounting volumes, run tickets, gauge and haul history, work-order and failure history, and the decline forecast. They stay on a time-based cadence until better signal exists, and they are visibly labelled that way rather than quietly treated as healthy. Partial coverage is the normal starting condition, not a disqualifier, and the ranking improves as instrumentation grows.
What changes for a field team member day to day?
The work does not get heavier, the order changes. Instead of a standing route driven Monday, Wednesday, and Friday, the crew opens a plan that was rebuilt overnight: the wells whose signal moved are at the top with the dollars at risk attached, the wells whose signal did not move are not on the list. Escalation carries the same economic context, so an exception routes to the right person with a number beside it. In the reference deployment at a top 25 private producer, the same crew drove 35% fewer miles (deployment figure) and the operation measured a 15% free cash flow uplift on the same crew (deployment figure).
What are the advantages and disadvantages of management by exception?
The advantage is leverage. Attention is finite and wells are not, so pointing the attention at the small set that has moved buys capacity that headcount would otherwise have to buy. It also makes the operation legible: what got worked, and why, becomes a record instead of a habit. The limits are real and worth naming before you build on it. The model is only as good as the definition of normal behind it, so a stale baseline hides problems as effectively as it raises false ones. It says nothing about magnitude on its own, so a deviation worth $90 a day and a deviation worth $12,500 a day arrive with equal urgency until something prices them (illustrative pair). It can starve the routine work that keeps deviations rare, which is why preventive and regulatory work has to compete in the same ranking rather than sit outside it. And it needs signal: an asset with no telemetry generates no exception, so partial coverage has to be visible rather than read as health. Each of those limits is a missing layer above the detection, which is what the ranking, scoring, and feedback layers supply.
How do you measure whether management by exception is working?
Pick the operating metric first and baseline it before anything is connected. Four measures separate a working program from a busy one. Deferred production, because the point of ranking is that the expensive deviations get worked first. Operating cost per well on the same crew, because the model claims capacity rather than headcount. Miles or hours driven per unit of economic coverage, because a plan that removes low-value stops shows up here before it shows up anywhere else. And the share of exceptions that turned out to be worth the visit, which is the only measure that tells you whether the queue itself is improving. At the deployed reference, a top 25 private producer running 5,000+ wells, the measured deployment figures are a 15% free cash flow uplift on the same crew and 35% fewer miles driven for the same coverage.
What usually goes wrong when operators try this?
Four failure modes account for most of it. Alarm fatigue: hundreds of exceptions a shift with no ranking, so crews triage by recency and the highest-value item gets buried. Stale normal: baselines that do not move as wells decline, producing both false alarms and missed real ones. Missing economic ranking: exceptions worked in the order they arrived rather than the order they pay, which is the difference between a busy field and a productive one. No closed-loop learning: when a flagged exception turns out to be nothing, the next identical pattern is flagged at exactly the same priority. All four are ranking and feedback problems rather than sensor problems, which is why buying more instrumentation rarely fixes them on its own.
See exception-based ranking on your data.
Talk to our team about how management by exception can be run on your wells. The 4-week stand-up is credited toward your first license. Move the metric you anchored on, or you owe no license fee.