He had more to contribute than the system let him
Watching a safety advisor work through an incident convinced me reports are not set up to support how we learn. While it was a serendipitous 20-minute encounter, it is one of those experiences that shaped my perspective on safety work ever since.
During a discovery project, I was waiting for the safety advisor to take me to a job site. He came up to me apologising that he had to close out an incident report first. He assured me it wouldn't take long, maybe 20 minutes. I would miss some field time, but as fieldwork comes with more data than you can hope to analyse anyway, I tend not to get too upset about such changes. To make the most of our time, I asked him if I could observe him working through the incident. The safety advisor visibly relaxed, realising he wouldn't be wasting my time, and happily agreed. We went to his office, where he opened the reporting system.
The safety advisor needed to work through a reported near-miss involving an isolation switch handle falling to the ground off an electricity pole. The switch handle had broken off while an operator had opened it with a switch stick, but had not hit anyone. A 10 cm metal bar falling from the conductor height and landing on someone could hurt, but is unlikely to really injure someone. In addition, the switch operator had been wearing a hard hat, further reducing the chance of harm.
In the system, the safety advisor was running through the different text boxes, dropdown boxes, and checkboxes that had already been completed. Occasionally he fixed a broken sentence, or adjusted a rating.
There is a bit of a question of what a safety advisor can contribute in these situations. They have less knowledge of the event than the original operator involved, and are dealing with second-hand knowledge. While not adding anything new, the safety advisor was aligning language and ratings to make it easier for the wider organisation to digest. Not particularly exciting, but not meaningless work. Then, as the safety advisor kept going, I was surprised.
As the safety advisor got to the main incident description, he paused. I asked what he was thinking of, and he responded that there might have been a few things going on. This was a new type of switch, he said, and offered three ideas on what might have happened: 1) maybe this new type of switch was no good, 2) maybe the crew did not know how to operate these switches, 3) maybe it was just a single bad switch of an otherwise good type of switch.
After thinking for a moment, the safety advisor decided to give the operator involved a call. He told the operator he was looking through the incident and asked him quickly to run through the incident.
Here the safety advisor caught my attention more. The safety advisor connecting the events to something else he knew, he was making parts meet. It looked like he would be moving the understanding of the event forward. I might be able to observe the development of new insights, and what follow-up this creates. In addition, this part of generating hypotheses is not described in investigation processes or reporting systems. This was the safety advisor's initiative, based on what he understood to be meaningful.
After the operator had gone through the incident, the safety advisor added that he had changed the severity rating, and was about to close out the incident. Hanging up the phone, the safety advisor was still looking uneasy. He had not made progress validating any of his hunches. Using the intranet, he searched someone else's phone number, but then closed the page without making a call. I asked what was going on, and the safety advisor hesitantly responded that this engineer was involved with these new switches, and added that the engineer would not be able to tell him more based on just this incident. After clicking back and forth between some browser tabs, the safety advisor said it was not really worth spending more time on this.
Naturally, this was a less exciting turn. I wasn't going to see him reach new insights. As an outsider, I couldn't judge whether it was reasonable or not to halt the inquiry, but the reasoning made sense. There will be incidents or risks that are too small to keep chasing, or questions that cannot be answered by looking deeper into a single event. The distinction between whether this was a bad type of switch, or this was a single dud, cannot be answered from a single event. It would require ongoing monitoring of a lot of the switches over a longer period.
Then the final question before closing out the report was root cause. This was a dropdown box with 5 options, of which the safety advisor had to select one. I saw both 'equipment failure' and 'operator error' on this list, and wondered which he was going to pick. To my surprise, the safety advisor selected 'Failure in Risk Management' and closed out the incident. This had no relation to anything he had told me so far. Trying to hide my confusion, I calmly asked why he had chosen that option. While I was convinced my voice had not given anything away, the safety advisor immediately spoke to the discrepancy with his earlier hypotheses. He explained that he couldn't prove either the equipment or training part, and then added that if there was an incident, it must include a problem with risk management.
This felt crazy to me. The safety advisor's explanation suggested he deliberately picked the least informative cause. The option he selected, as he understood it, only said there was an accident. The classification was driven by being defensible, and being unfalsifiable, rather than by trying to capture anything about the incident. Imagine the board report saying 30% of accidents are due to failure in risk management. Does this tell them anything about the incidents that have been happening? Or is there any decision that information could be helpful for?
While the choice for 'Failure in Risk Management' was odd to say the least, can we really put it on the safety advisor here? First, he cared. He wasn't switched off. He had been going out of his way to help me with the discovery work. And he tried with this incident. He tried to understand what happened based on what he considered reasonable. If he wanted to be done quickly, he could have skipped all the pauses and thinking. He wanted more, but did not have any good choices.
It is hard not to trace this back to the incident reporting design. The most obvious part here is, of course, the dropdown box. Suggesting an incident has a single root cause is ridiculous, which has been well covered by others. Even if it was possible to select multiple predetermined options, it would have been ridiculous. It basically meant that the assumption is there is nothing qualitatively unique about the 'roots' of any incident. That there was nothing you didn't already know that mattered before you had the incident. In that sense, you could have already made the investment in safety without having the incident, let alone examine the incident. But the specificity of causation is just the beginning.
The really painful part of the situation is what the reporting system did not offer. There was nothing to support or capture the safety advisor's efforts to learn. The safety advisor had thought of hypotheses, tried to collect data on them, but then had nowhere to put that. The questions, thinking, considerations, were lost with the closing of the incident. There was no opportunity to acknowledge uncertainty; everything needed to be presented as if it was resolved, taking away any reason for further anxiety. If the safety advisor wanted to continue exploring these hypotheses, he would have to do it in spite of the reporting system, rather than the reporting system supporting this learning. The reporting system was now saying the incident was resolved and the root cause had been found. Nothing to see here any more, why do you keep looking?
Now someone might say that this is the price to pay for aggregation. The incident records offer a larger-scale view, and to do that, they need to be made to look neat and stripped of context. That way reporting sets can go across incidents to draw new conclusions we otherwise couldn't reach. Putting the futility of the use of "Failure in Risk Management" category aside, there is some truth to this. No doubt there are genuine use-cases where counting predetermined boxes helps aggregate data in a meaningful way.
But this misses what is already there. The actions of the safety advisor contain an alternative. That of hypotheses and testing them. Charles Sanders Peirce called this abduction, which is the basis of most forms of inquiry, from science to police investigations. Rather than asking for 'root cause' with limited dropdown boxes, the reporting system could collect questions and hypotheses, against which each (new) incident could be tested to see whether it helps answer or verify them. This is what curious people already do, not just this safety advisor. Through both observation and stories, I've learned about many investigators who keep notes with questions and hypotheses about incidents, often before they even start an investigation. A reporting system could help connect these questions to people who have more data on this. When you start with a question of hypotheses, both data on incidents and everyday work can be equally valid. Rather than forcing curious people to conclude incidents are due to a "Failure in Risk Management," one incident at a time, the reporting system could hold the cognitive work, that is, the hypotheses and half-followed leads, and connect them to other people, who can offer the relevant data and experiences they have to move them forward.