Featured Analysis

How PFD Monitor automated a national safety thematic review in around 5 minutes

In a peer-reviewed study, PFD Monitor reviewed 4,730 coroners' reports in minutes, identified two-thirds more relevant reports than an earlier national review, and agreed with clinicians 97% of the time.

Portrait of Sam Osian
By Sam Osian Founder of PFD Monitor
Published 11 August 2026 7 min read researchautomationmental health

Five minutes and 29 seconds.

That was how long PFD Monitor took to review 4,730 Prevention of Future Deaths reports, find those concerning children who died by suicide, identify the coroners' concerns and organise the results into a table for analysis.

The work was published in BMJ Mental Health. It tested whether a national safety review, previously carried out manually by the Office for National Statistics, could be completed using PFD Monitor.

The results went beyond speed. Within the period covered by the ONS review, PFD Toolkit identified 62 relevant reports. The ONS had identified 37. The Toolkit produced a set two-thirds larger, and when its decisions were independently checked by three psychiatrists, the AI and clinicians agreed 97% of the time.

The study shows how a large and difficult public archive can become usable safety intelligence. Reviews that once required researchers to find, open and code documents one by one can now be run across the entire collection in minutes, with expert checking built into the research design.

Why these reviews are difficult

Prevention of Future Deaths reports contain coroners' concerns about risks that could lead to further deaths. Read together, they can reveal recurring problems across services, places and organisations.

They were never published as a ready-made research dataset. The archive mixes ordinary text documents with scanned files. Descriptions and metadata vary, and around 70% of the reports examined in the study had no category tag at all.

Researchers have therefore had to narrow the archive before they can begin a thematic review. That can mean relying on website categories or search terms, then manually reading and recording information from every candidate report. It is slow work, and relevant reports can sit outside the categories a researcher expects to find them in.

Human review can only assess the reports that make it into the candidate set. When categories are missing or assigned inconsistently, a report may be excluded before a researcher has the chance to read it. Manual coding then introduces a second challenge: applying the same definitions consistently across hundreds or thousands of documents.

The ONS review provided a useful test. Its researchers had studied PFD reports about children who died by suicide between January 2015 and November 2023. They identified 37 reports and analysed the concerns raised by coroners across 23 different topics.

PFD Monitor was given the same broad task, but applied it to the full archive available for the study.

What PFD Monitor did

The Toolkit worked through 4,730 reports published between July 2013 and November 2023. It read the contents of each report rather than depending on the categories attached to it.

For every report, it assessed whether the case met the study's definition. For the reports that did, it recorded who received the coroner's report and whether each of the 23 concern topics used in the ONS review was present. It handled both scanned and machine-readable documents and produced a structured table rather than a collection of AI-written summaries.

The whole process — reviewing the archive, identifying the relevant reports, coding the concerns and producing the table — took 5 minutes 29 seconds on a consumer laptop.

Thousands of reports can therefore be examined in the time it would ordinarily take a researcher to read only a handful. The same eligibility rule and the same 23 concern topics are applied to every report, including those without a useful website category.

The results were checked by clinicians

Speed alone would not make the approach useful for safety research. The study therefore asked three psychiatrists to review a sample of 146 reports without seeing the Toolkit's decisions.

The sample included every report the Toolkit had identified as relevant, along with 73 other reports selected from the archive. Where the clinicians initially disagreed, they discussed the report and reached a shared decision.

The Toolkit and the clinicians agreed on the final classification in 97% of cases.

The study also recorded no failures when reports were opened and processed. Together, those findings show that the Toolkit can do more than make a quick first pass. It can produce a reliable candidate dataset that experts can check, challenge and use for further analysis.

PFD Monitor found 25 more reports in the same period

PFD Monitor identified 73 reports concerning children who died by suicide. Sixty-two fell within the period covered by the ONS review. Another 11 were published before the relevant categories were introduced on the Judiciary website.

Against the ONS total of 37, finding 62 reports in the same period is a large difference: 25 additional reports, or a set around two-thirds larger.

The report-level ONS data were not published, so the two lists could not be compared one by one. It is therefore not possible to say that every additional report was an ONS error. The size of the gap nevertheless exposes a weakness in the manual route. The ONS researchers built their candidate set using the categories on the Judiciary website. PFD Monitor read the content of every report instead.

Around 70% of the study archive had no category tag. A careful manual review cannot recover a relevant report that was never placed in front of the researcher. By screening the complete archive, PFD Monitor reduces that dependence on earlier human decisions about how a report should be labelled.

The resulting table could already be used to examine the concerns coroners raised. Common topics included problems with processes, training, risk assessment, communication between services, referral delays and disconnected care.

These findings describe the reports, rather than the prevalence of a problem across the wider population. They give researchers and policy teams a much stronger starting point: a broader set of source documents, coded in a consistent way, with every result open to inspection.

From a one-off review to a repeatable capability

National thematic reviews are usually treated as substantial one-off projects. The PFD Monitor study points towards a different model.

Once a review question and its categories have been defined, the same method can be applied whenever new reports are published. Researchers can update an existing review, compare time periods or test a new question without rebuilding the document collection and coding process from the beginning.

For policy teams, that could support:

  • faster national reviews when a safety concern emerges;
  • regular updates instead of findings that become dated between projects;
  • consistent comparison across services, organisations or regions;
  • earlier identification of recurring concerns; and
  • a traceable body of source material for expert scrutiny.

The paper suggests that the same approach could be adapted for deaths in custody, medication safety, care-home incidents and maternal mortality. Other questions could be defined around age, care setting, the organisations involved or the type of concern raised.

Researchers would still decide the question, define the categories, check the outputs and interpret the evidence. PFD Monitor removes much of the repetitive document work that makes broad or regularly repeated reviews difficult to commission.

What the study does and does not show

The results come from one defined research task using one version of the AI. Other questions will need their own testing, particularly where reports do not contain all the information needed to make a decision.

The Toolkit analyses what is written in published documents. It cannot on its own establish why a death occurred, decide whether a problem is common in the wider population, or verify that an organisation carried out a promised action. Those questions require other evidence and human judgment.

What the paper demonstrates is both narrower and more practical: PFD Monitor can turn thousands of unstructured coroners' reports into a usable thematic dataset in around five minutes, and its decisions can stand up to independent clinical checking.

That opens the way to more frequent, more ambitious and more responsive use of PFD reports in public policy and safety research.

PFD Monitor's source code and research materials are openly available.


References

  1. Osian S, Dutta A, Bhandari S, Buchan IE and Joyce DW. Automating thematic review of prevention of future deaths reports: concordance study of a child-suicide analysis using large language models. BMJ Mental Health. 2026;29:e302212.
  2. Sharland E, Wallace E, Revie L, et al. A thematic analysis of Prevention of Future Death reports for children who died by suicide in England and Wales: January 2015 to November 2023. British Journal of Psychiatry.
  3. PFD Monitor source code and research materials.