MTBF and MTTR for oil & gas maintenance teams.

Four numbers carry most of the weight in a reliability review. This page says exactly how RelyentX computes each one and which denominator it used, because that detail decides whether the number means what you think it means.

The four numbers, plainly

Mean time between failures (MTBF) is the average time from one failure to the next, across a population of equipment. Mean time to repair (MTTR) is the average time a failure takes to put right. Mean time to failure (MTTF) is the average time equipment spends running between repairs, which is MTBF minus the repair time. Availability is the share of the period the equipment was not under repair.

They answer different questions. MTBF tells you how often you are interrupted, MTTR how expensive each interruption is to recover from, and availability what the two together cost you in the end. A fleet with a long MTBF and a terrible MTTR is a fleet that fails rarely and then stays down for weeks.

Where the numbers come from in RelyentX

A failure is a corrective job card. Nothing else counts: preventive and modification work is excluded from the failure count, so planned work does not depress your MTBF.

Repair time comes off the card itself. Each closed corrective card carries a per-card MTTR, which is its close date minus its start date in whole days. Read that literally: it is how long the card was open, so it includes waiting for a part, a permit or a crew, not only the hours someone spent on the tools. A repair that took an afternoon but waited nine days for a seal counts as ten days.

From there the aggregates are three divisions. MTTR is downtime over failures. MTBF is calendar time over failures. MTTF is uptime over failures, where uptime is calendar time minus downtime, floored at zero. Availability is uptime over calendar time, as a percentage.

All of it is computed when you ask, from the job cards and the equipment master, and none of it is stored. There is no nightly rollup to go stale and no reconciliation step to forget: change a close date and the number changes with it.

Our denominator, in full

Calendar time here is the equipment count multiplied by the number of days in the window. Ten pumps over ninety days is nine hundred equipment-days, whether those pumps ran continuously, ran on rotation, or sat on a rack the whole quarter.

That is a calendar basis, not an operating-hours basis. RelyentX does not currently take a working-hours or runtime-meter feed, so it cannot tell a pump that ran 24/7 from one that was a cold standby. We say so here rather than let you infer an operating-hours figure from a calendar one.

What follows from that is worth knowing before you compare vendors. On a calendar basis, adding spare equipment to a class raises its MTBF and its availability without anything mechanical improving, because the denominator grew and the failure count did not. Compare like with like: the same equipment type, the same window, and a population that has not changed size underneath you.

One more thing the window does. Unless you set both ends of a date range, the window is derived from the earliest job card in view, so changing a filter changes the denominator and therefore the figure. That is correct behaviour and it surprises people: two screens showing different MTBFs for the same equipment are usually two different windows, not a bug.

The honest use of a calendar-based number is as a trend and a comparison between types, not as an absolute to put in a contract. It is the number your job cards actually support. An operating-hours figure would need a data source that is not there, and inventing one would make the number look better and mean less.

Availability is derived, not measured

The availability percentage here is uptime over calendar time, and uptime is whatever was not spent on a closed corrective card. It is not a reading from a control system, a historian or a sensor, and it does not know about a process upset that left equipment idle without a repair.

This means availability inherits every property of the inputs. A corrective card left open past the repair does not accrue downtime until it closes, so a backlog of unclosed cards flatters the number. That is a property of deriving availability from work orders rather than reading it off an instrument, and it is worth knowing before the figure is quoted at anyone.

PM compliance, and why it is labelled a proxy

Preventive-maintenance compliance is reported as the share of preventive cards in the window that were closed on or before their expected close date. An open card counts against it, as does one closed late.

That is a stand-in, and it is marked as one in the product rather than presented as true schedule adherence. Genuine plan-versus-completed compliance needs modelled PM plans and schedules, which is part of the committed model and is not yet on the shipped surface. Until it is, the proxy answers a narrower question honestly instead of a broader one badly.

What these numbers will not do

They will not predict a failure. RelyentX measures reliability from the events your work already produced; there is no forecasting model and no anomaly detection in the product, and the site does not claim any.

They will also not tell you why. That is what the work history is for, and it is the question the copilot is actually good at: ask why an asset keeps failing and the answer is retrieved from that asset’s own closed cards, certificates and status history, with those records shown next to it so you can check them.

Common questions

Is MTBF calculated on operating hours or calendar days?

Calendar days. Calendar time is the equipment count multiplied by the days in the window, and RelyentX does not take a runtime-meter or working-hours feed today, so it cannot distinguish running time from idle time. The figures are days, and they are comparable across equipment types and over time rather than quotable as absolute operating-hours reliability.

Does preventive work count as a failure?

No. Only corrective job cards count as failures. Preventive and modification cards are counted separately, so planned maintenance does not damage your MTBF.

Why did our availability improve when we added equipment?

Because availability is uptime over calendar time, and calendar time is the equipment count multiplied by the window. Adding equipment to a type grows the denominator immediately while the failure count takes time to follow. Compare windows where the population is stable, or compare types against each other rather than a type against its own past.

Are the numbers stored or recomputed?

Recomputed on every read, from the job cards and the equipment master. Nothing is persisted, so correcting a close date corrects every figure that depended on it, with no rebuild step.

See it against your own equipment data.

A walkthrough on your asset classes, your certificate disciplines, your job cards, not a canned demo.