Package documentation

📆 Changelog

Welcome to the project changelog. All notable changes to this project will be documented below.

0.4.1 - 2026-03-02

  • Area matching now uses exact matching with fuzzy fallback by default, reducing cases where small spelling or wording variations were previously classified as "Other".
  • Updated coroner area canonicalisation to align more closely with the current official list, with legacy area names now recoded to current canonical areas.
  • Added Scraper.rescrape_fields() so individual columns such as area can be refreshed across an existing dataset without re-scraping the full archive.
  • Improved recipient cleaning so boilerplate recipients such as Chief Coroner are removed, role-led recipient strings are reduced to organisations or departments, and common formatting variants such as NHS Trust/NHS Foundation Trust and selected department/highway names are normalised more consistently.

0.4.0 - 2025-11-30

  • discover_themes() now defaults to using untrimmed report text, with optional truncation controlled by trim_approach="truncate" and new max_tokens/max_words parameters. Switch to trim_approach="summarise" to re-enable LLM summarisation, which now uses summarise_intensity settings.
  • Users have reported that running extract_features() on a list of themes from discover_themes() seems to assign reports too liberally. We have made some changes to make the model a bit more conservative, only assigning a report with a theme if there is sufficient evidence that it is well represented in the underlying report. We'll continue monitoring this behaviour to make sure we've got the balance right.
  • Added a truncation-based alternative to LLM summarisation and new report-length diagnostics consolidated under count() for word or token charts/tables with threshold guidance.
  • count() now also supports a stats mode producing markdown summary statistics across word or token lengths.
  • count() statistics now render as a readable table without console prints, and the chart option is renamed to as_="hist" for clarity.

0.3.7 - 2025-09-02

  • In August 2025, the judiciary.uk website made some subtle changes that broke PFD Toolkit's scraper, meaning that we were unable to collect newly published reports. This issue has now been resolved, and all previously missed reports have now been added.

0.3.6 - 2025-08-03

  • Improve reliability and performance of the Scraper and Cleaner modules.
  • The Cleaner module now standardises each report 'area' to one of 77 official jurisdictions (e.g. "Liverpool and the Wirral"), so minor variations and typos are automatically corrected for consistent regional filtering.
  • load_reports() now refreshes the dataset by default. Pass refresh=False to use a previously cached copy instead of downloading again.

0.3.5 - 2025-07-07

  • Fixed issue where PFD Toolkit refused to run in Google Colab

0.3.4 - 2025-07-07

  • Deprecated user_query in Screener in favour of search_query. user_query will be removed in a future release.
  • Dropping spans in extract_features() no longer removes spans added during screening.
  • Downgraded pandas from 2.3.0 to 2.2.2
  • Fixed text cleaning bug that expanded dates and removed paragraph spacing.
  • Added tests covering span removal behaviour.

0.3.3 - 2025-06-25

  • Improved package installation time
  • Changed default LLM model from GPT-4.1-mini to GPT-4.1

0.3.2 - 2025-06-23

  • You no longer need to manually update the pfd_toolkit package to get access to freshly published reports. Instead, run load_reports(refresh=True).
  • Improve robustness of Scraping module in handling missing data between different scraping strategies.
  • Fixed typos and improve documentation.

0.3.1 - 2025-06-19

  • Improved reliability of weekly dataset top-ups.

0.3.0 - 2025-06-18

First public release! ✨

View this page in the package repository