Research

Scientific background, objectives, methodology, and related publications for the TRExt project.

Project Summary

TRExt makes the sensitive information held in clinical free-text available for federated research within Trusted Research Environments.

Clinical free-text such as patient letters, discharge summaries and other written clinical documents contain rich detail on symptoms, severity and treatment that is rarely captured in structured records, and current federated analytics pipelines cannot easily reach any of it. TRExt closes that gap by bringing together state-of-the-art natural language processing tools and data standardisation tools into a single pipeline, converting unstructured text into standardised, research-ready safe data. TRExt runs entirely inside the secure environment holding the data. Documents are never transmitted and no free-text leaves the environment. The pipeline is deployable next to the federated analytics infrastructure developed by the DARE UK TREvolution programme (called Five Safes TES), enabling this data to be analysed within standard data models with only approved safe research outputs leaving the secure environment.

Public Involvement & Engagement (PIE) activities are embedded throughout all stages of the project. See the PIE Activities page for details.

Research Themes

WP1

Clinical Natural Language Processing (NLP)

We are deploying CogStack, a system that reads through clinical documents and pulls out the meaningful details automatically. Its MedCAT component identifies the things being described, such as conditions, medications and symptoms, and matches each one to SNOMED CT, an internationally agreed set of clinical terms. Its RelCAT component then works out how those details connect, for example that a drug was given for a particular condition. The tools have been validated by clinicians across a wide range of specialties.

WP2

Data Mapping & Standardisation

We are deploying three tools to prepare this information for analysis. Carrot and Lettuce translate it into OMOP, a standard format widely used in health research, with Lettuce using AI to suggest translations and a trained person approving them. Where data needs to be used in clinical trials or submitted to regulators, the Unison platform can convert it into CDISC, the standard used for that purpose. Either format is ready for Five Safes TES, the technology that lets the data be analysed alongside data at other institutions without those records ever leaving the places that hold them.

WP3

Project Management & PIE

We are running the project through regular management meetings, chaired by a member of the public, along with weekly team check-ins and oversight from a Scientific Advisory Board. Alongside this, a dedicated specialist at SAIL Databank leads our public involvement programme, running online workshops with members of the public. These workshops shape how the technical work is designed, and what people tell us about their trust, concerns and recommendations feeds directly into the decisions we make.

Related Publications

Key publications from the TRExt team underpinning the project's methodology. Publications directly arising from TRExt will be added as they become available.

2024

TRE-FX: Enabling Federated Analytics Across Trusted Research Environments, Giles et al.

doi:10.5281/zenodo.10055353 ↗

Lettuce: LLM-Assisted OMOP Concept Mapping, Mitchell-White et al.

doi:10.48550/arXiv.2410.09076 ↗

Carrot: Software Tool for OMOP Data Transformation, Cox et al.

doi:10.2196/60917 ↗
2022

CogStack: Experiences of Deploying Integrated Information Retrieval and Extraction Services in Clinical Settings, Noor et al.

doi:10.2196/38122 ↗

MedCAT: Medical Concept Annotation Tool, Kraljevic et al.

doi:10.1016/j.artmed.2021.102083 ↗
2015

CRIS: Maudsley Biomedical Research Centre Clinical Record Interactive Search, Perera et al.

doi:10.1136/bmjopen-2015-008721 ↗