📆 Changelog
Welcome to the project changelog. All notable changes to this project will be documented below.
0.4.1 - 2026-03-02
- Area matching now uses exact matching with fuzzy fallback by default, reducing cases where small spelling or wording variations were previously classified as
"Other". - Updated coroner area canonicalisation to align more closely with the current official list, with legacy area names now recoded to current canonical areas.
- Added
Scraper.rescrape_fields()so individual columns such asareacan be refreshed across an existing dataset without re-scraping the full archive. - Improved recipient cleaning so boilerplate recipients such as
Chief Coronerare removed, role-led recipient strings are reduced to organisations or departments, and common formatting variants such asNHS Trust/NHS Foundation Trustand selected department/highway names are normalised more consistently.
0.4.0 - 2025-11-30
discover_themes()now defaults to using untrimmed report text, with optional truncation controlled bytrim_approach="truncate"and newmax_tokens/max_wordsparameters. Switch totrim_approach="summarise"to re-enable LLM summarisation, which now usessummarise_intensitysettings.- Users have reported that running
extract_features()on a list of themes fromdiscover_themes()seems to assign reports too liberally. We have made some changes to make the model a bit more conservative, only assigning a report with a theme if there is sufficient evidence that it is well represented in the underlying report. We'll continue monitoring this behaviour to make sure we've got the balance right. - Added a truncation-based alternative to LLM summarisation and new report-length diagnostics consolidated under
count()for word or token charts/tables with threshold guidance. count()now also supports astatsmode producing markdown summary statistics across word or token lengths.count()statistics now render as a readable table without console prints, and the chart option is renamed toas_="hist"for clarity.
0.3.7 - 2025-09-02
- In August 2025, the judiciary.uk website made some subtle changes that broke PFD Toolkit's scraper, meaning that we were unable to collect newly published reports. This issue has now been resolved, and all previously missed reports have now been added.
0.3.6 - 2025-08-03
- Improve reliability and performance of the Scraper and Cleaner modules.
- The Cleaner module now standardises each report 'area' to one of 77 official jurisdictions (e.g. "Liverpool and the Wirral"), so minor variations and typos are automatically corrected for consistent regional filtering.
load_reports()now refreshes the dataset by default. Passrefresh=Falseto use a previously cached copy instead of downloading again.
0.3.5 - 2025-07-07
- Fixed issue where PFD Toolkit refused to run in Google Colab
0.3.4 - 2025-07-07
- Deprecated
user_queryinScreenerin favour ofsearch_query.user_querywill be removed in a future release. - Dropping spans in
extract_features()no longer removes spans added during screening. - Downgraded pandas from 2.3.0 to 2.2.2
- Fixed text cleaning bug that expanded dates and removed paragraph spacing.
- Added tests covering span removal behaviour.
0.3.3 - 2025-06-25
- Improved package installation time
- Changed default LLM model from GPT-4.1-mini to GPT-4.1
0.3.2 - 2025-06-23
- You no longer need to manually update the
pfd_toolkitpackage to get access to freshly published reports. Instead, runload_reports(refresh=True). - Improve robustness of Scraping module in handling missing data between different scraping strategies.
- Fixed typos and improve documentation.
0.3.1 - 2025-06-19
- Improved reliability of weekly dataset top-ups.
0.3.0 - 2025-06-18
First public release! ✨