The blank field is not a discipline problem
Every plant has had the conversation. Management asks why the failure codes are blank. Maintenance says they are busy. Management sends a reminder, compliance improves for three weeks, and then the field is blank again. The cycle repeats about annually and is universally read as a culture problem, which is why it never gets fixed.
It is not a culture problem. It is an exchange with nothing on one side. Filling in that field costs a technician three minutes at the end of a long job, on a terminal in an office they have to walk to, and returns them precisely nothing. The data goes into a report that a reliability engineer may read next quarter. Nobody who fills in the form ever sees anything come back. Every system built that way decays, in every industry, regardless of how many reminders are sent.
Which means the fix is not enforcement. It is changing the exchange so that making the record gives the person making it something they want, within the same shift.
You are probably here because
- A predictive maintenance proposal assumed a failure history you do not have
- Your work order history is mostly free text and blank codes
- The same machine keeps failing and nobody can prove it from records
- You are deciding whether to buy sensors or fix the data first
The exchange section is the root cause. The section on what is possible with thin history is the one that saves a year of waiting.
Give something back in the same shift
The change that works is to put information the technician wants on the same screen where the record is made, so opening the job is useful before anything is typed.
The single highest-value item is the last three things that were done to this asset, with dates, what was found, and photographs. A technician walking up to a machine they have not touched in eight months wants exactly that, and today they get it by finding somebody who remembers. Put it on the screen and the screen becomes worth opening. Then add the parts that fit this asset with current stock and location, the manual page for this component, the torque and clearance specifications, and any open note left by the previous shift.
Once the screen is genuinely useful, the record becomes a by-product of using it rather than a tax on finishing. That is the whole mechanism, and it is more effective than any policy. A technician who has just been saved twenty minutes of hunting will spend forty seconds telling the system what they found, because they can see where it went.
Four fields at the machine beat twenty-two at a desk
The second change is arithmetic. A twenty-two field form takes several minutes and will be abandoned or filled with defaults. Four structured fields on a phone, at the machine, take under a minute and will be completed.
Structure only what you will actually query. In our experience that is: which asset, what was observed, what was found, and what was done. Everything else — the story, the caveats, what the operator said, the thing that will be needed next time — belongs in free text, and free text is now genuinely easy to capture because dictation works well on a phone in a noisy plant, far better than typing with gloves on.
Two details make the difference between adoption and abandonment. It has to work with no signal, because the basement and the far end of the plant have none, and a form that fails to submit is a form nobody uses twice. And photographs must be one tap, attached to the asset rather than to a message thread, because a photograph of the failed part is often more informative than any paragraph and takes a fraction of the time.

What makes a maintenance record actually get filled in — our judgment
Our judgment of what drives completion rates, based on building these interfaces. Not a survey. The bottom row is the design most systems ship with by default.
The asset register is the foundation, and it is usually wrong
None of this matters if the record cannot be attached to a specific thing. A failure logged against “Line 3” is a note. A failure logged against a specific gearbox, with a stable identifier that survives being moved to another line, is a data point in an actuarial record. The difference decides whether you can ever answer a reliability question.
Most asset registers have four problems at once. Assets that were scrapped years ago are still listed. The same physical machine appears twice under two spellings. The hierarchy stops at the machine and never reaches the component that actually fails. And the identifier encodes the location, so it becomes wrong the moment the machine moves.
Cleaning that up is unglamorous work with a large payoff and it is mostly walking the floor with a tablet. Two design rules make it durable. Identifiers are opaque and permanent — they mean nothing, they never change, and the location is an attribute rather than part of the name. And the hierarchy goes at least one level below the machine, because “pump P-114” failing tells you far less than “the mechanical seal on pump P-114,” and seals and bearings and couplings have completely different failure behavior.
For the taxonomy itself, there is a good free starting point. The international standard on collecting reliability and maintenance data for equipment — written for the petroleum and petrochemical industries and widely borrowed outside them — publishes equipment hierarchies, failure modes and failure mechanisms that are already thought through. Copying a structure that works and trimming it to your equipment is a much better use of a week than designing a taxonomy from scratch, and it has the side benefit of being comparable to something outside your own plant.
Failure coding that people will actually use
Codes fail for the same reason forms fail: too many options, written in a vocabulary the user does not think in.
The structure that survives contact is three short questions rather than one long list. What was observed — the symptom that brought you here, which the operator can also report. What was found — the actual condition, which only the technician knows. What was done — the action, which is what drives parts and labor. Each one gets a list of roughly ten to fifteen options for the equipment class, in the words people already use at that plant.
Then watch the share coded “Other,” the same way you would watch it in a downtime system. A rising Other rate is not laziness; it is the list failing to describe what is happening. Review it quarterly with two technicians and a planner in the room, and change the list. A taxonomy that never changes is a taxonomy that stopped matching the plant some time ago.
What is possible before you have a failure history
Here is the part that saves a year. A common piece of advice is to fix the data and come back in eighteen months. That is unnecessary, because several genuinely useful things do not need failure history at all.
| Approach | What it needs | What it actually tells you |
|---|---|---|
| Run-hour based intervals | A run signal per asset, nothing else | Whether a calendar PM is early or overdue for how much the asset actually ran |
| Threshold condition monitoring | Vibration, temperature, pressure, motor current | That a measured condition has left its normal band |
| Baseline anomaly detection | A long period of known-healthy operation | That behavior is different from usual — not that it is failing |
| Population reliability analysis | Many identical assets, replacement dates | Whether the fleet shows an age relationship worth acting on |
| Supervised failure prediction | Many labeled failures per mode, per asset class | Genuine lead time — and it is the one thing you cannot do yet |
Run hours first. Converting calendar-based preventive maintenance to usage-based is often the largest single improvement available, and it needs only a run signal. A machine that ran two hundred hours last month and one that ran fifteen do not need the same service interval, and treating them identically means over-servicing one and under-servicing the other. Over-servicing is not harmless either: taking a working machine apart introduces failures that would not otherwise have happened.
Anomaly detection is worth understanding precisely, because it is frequently oversold. Trained on healthy operation, it flags behavior that is different from usual. Different is not the same as failing. A new product, a new operator, a colder month and a genuine developing fault all look like anomalies. It is a useful attention-directing tool if a person reviews the alerts and it earns its keep by pointing at things worth looking at — and it becomes noise the moment it is treated as a failure prediction and alerts nobody triages.
Supervised prediction is the one that needs the history you do not have. To learn what precedes a specific failure mode, a model needs many examples of that mode, not two. Most plants have two, on the assets that matter most, because the important machines fail rarely by design. This is not a modeling problem to be engineered around; it is a data volume problem, and the honest answer is that supervised prediction on those assets is a project for later, after the recording has been fixed. Fleets of identical small assets are the exception: fifty similar pumps generate enough events to learn from in a reasonable period.
The interval that decides whether monitoring is worth anything
One concept from reliability practice does more work than any algorithm. Between the moment a failure first becomes detectable and the moment the equipment actually fails, there is an interval — sometimes weeks, sometimes minutes. Detection is only useful if you check more often than that interval, and the value of any monitoring depends entirely on it.
Some failures develop over months and are detectable early by vibration or oil analysis, and monthly or continuous monitoring buys real lead time. Others go from fine to broken in a shift, and no inspection frequency you can afford will catch them. For those, monitoring is not the answer; redundancy, a spare on the shelf, or a design change is.
So before instrumenting an asset, ask what the dominant failure mode is and roughly how fast it develops. If nobody knows, the maintenance crew usually has a good opinion, and a rough answer changes the decision more than a precise model would. This is also why the reliability-centred maintenance literature is worth reading before buying sensors: the airline studies behind it found that most failure modes examined showed no strong relationship to operating age, which is why scheduled overhauls so often fail to help and why condition-based approaches took over. The point is not that time-based maintenance is wrong. It is that it only helps for failure modes that actually wear out.
What the sensor packages are, honestly
The market is full of wireless sensor packages sold as artificial intelligence for maintenance. Most of them are a good sensor, a gateway, a threshold, and a dashboard, sometimes with anomaly detection on top.
That is not a criticism. A good vibration sensor with correctly set alarm bands is a genuinely useful instrument, and the published vibration evaluation standards give defensible starting bands by machine class. It is worth buying for what it is. It is not worth buying under the impression that it will predict a specific failure with a specific lead time on your particular equipment, because that claim requires evidence from your plant that does not exist yet.
Two questions sort the offerings quickly. Can we export our own raw measurements, in an open format, including the full history. And what exactly does the alert mean — a fixed threshold, a learned baseline, or a model trained on failures from other customers' equipment, which may or may not resemble yours. Honest vendors answer both immediately.
The first ninety days
- Clean the asset register — retire dead assets, merge duplicates, add one level below the machine
- Give assets permanent opaque identifiers with location as an attribute, and label them physically
- Put entry on a phone at the machine, working offline, with one-tap photos
- Add the last-three-repairs screen before asking anyone to fill in anything new
- Cut the form to four structured fields and let dictation carry the rest
- Instrument run hours on the twenty assets that cause the most downtime
- Start a proper failure log on the constraint, even if it is one page per event
- Report the “Other” rate monthly and revise the code lists when it climbs
Notice that none of that is a model, and most of it is not software. It is also the work that decides whether anything modelled two years from now will be worth trusting.
The mistakes we see
- Buying a predictive platform before the asset register is clean — the platform will faithfully aggregate failures across three duplicate records of one machine
- Mandating fields instead of redesigning the exchange — produces defaults, not data
- A taxonomy designed by an engineer who does not do the work — precise, complete, and abandoned in a month
- Treating anomaly alerts as failure predictions — the alerts become noise and then get muted
- Instrumenting assets whose failures develop faster than the inspection interval — a sensor that cannot arrive in time
- Discarding raw condition data after alarming on it — that history is what makes later analysis possible
- Blaming the technicians — they are responding rationally to a system that gives them nothing
When you do not need us
If your maintenance system is well configured and the actual problem is spare parts — long lead times, stock-outs on critical items, no criticality ranking — that is an inventory and procurement problem, and it is usually a bigger lever than anything predictive. Fix that first and you will get more uptime for less money.
Likewise, if your CMMS vendor already offers mobile entry and you have not turned it on, turn it on before commissioning anything custom. A great many plants are paying for a mobile module that was never deployed because the rollout needed two weeks of attention nobody had. That is a fortnight of internal work, not a project.
Bottom line
Maintenance records are blank because filling them in costs the technician time and returns nothing. Change that exchange — put the last three repairs, the parts, the manual page and the specifications on the same screen — and completion follows without a policy. Cut the structured fields to four, put them on a phone at the machine, and let dictation and photographs carry the rest. Clean the asset register first, because a record that cannot be attached to one specific component is not evidence of anything. Then do the reliability work that does not need history: usage-based intervals, condition thresholds, and population analysis where you have fleets. Supervised failure prediction on critical assets is a later project, and saying so plainly is more useful than promising it now.
Frequently asked questions
More than most plants have, and the honest unit is failures per mode rather than years of records. Learning what precedes a specific failure mode needs many examples of that mode; a handful across a decade is not enough, and no amount of modeling technique compensates. Fleets of identical small assets are the practical exception, because fifty similar units generate events at a useful rate. For critical single assets, condition monitoring and physics-based thresholds are the realistic path.
Partly, and it is worth an afternoon to find out. Language models are reasonably good at sorting “replaced bearing, noisy” into a structured code, so a retrospective pass over a few years of text can produce a rough failure history. Two warnings. Extraction quality has to be measured against a hand-coded sample before anyone relies on it, and the text often omits the very thing you want — if nobody wrote down which component failed, no amount of processing recovers it.
On the right assets, yes. The right assets are rotating equipment that is critical, expensive to fail, and fails in ways that develop slowly enough to be detected in time. That last condition is the one people skip. Start with a handful on the worst offenders, use the published evaluation bands for the machine class as a starting point rather than an answer, and insist on exporting your own raw data so the history is yours.
Usually not, and replacing it will not fix the completion problem — a new system with the same twenty-two field form produces the same blank fields. The two things worth checking are whether it has a usable mobile entry path and whether you can get your own data out through an interface. If both are true, the improvements described here can be built alongside it. If you cannot export your own history, that is a genuine reason to move, and it will still be a reason next year.
Putting the last three repairs on this asset, with photographs, in front of the technician before they start. It costs one screen, it changes the entry from a tax into a trade, and it improves the work itself on the first day — which is the only kind of change that survives a busy quarter. Cleaning the asset register is a close second and is the prerequisite for everything after.
