﻿<?xml version="1.0" encoding="UTF-8" ?>
<rdf:RDF xmlns:admin="http://webns.net/mvcb/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:prism="http://purl.org/rss/1.0/modules/prism/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:syn="http://purl.org/rss/1.0/modules/syndication/">
<channel rdf:about="http://medrxiv.org">
<admin:errorReportsTo rdf:resource="mailto:medrxiv@cshlpress.edu"/>
<title>medrxiv Subject Collection: Health Informatics</title>
<link>http://medrxiv.org</link>
<description>
This feed contains articles for medRxiv Subject Collection "Health Informatics"
</description>

<items>
<rdf:Seq>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.16.26363186v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.22.26363642v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.16.26363276v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.21.26363564v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.21.26363538v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.21.26363536v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.16.26363030v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.16.26363127v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.20.26363523v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.18.26363448v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.19.26363476v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.15.26362921v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.19.26363431v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.19.26363460v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.18.26363428v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.15.26363093v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.17.26363303v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.15.26362961v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.16.26363187v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.08.26361592v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.11.26362711v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.15.26363148v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.14.26363047v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.14.26363036v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.14.26363024v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.14.26363006v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.10.26362715v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.13.26362573v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.12.26362907v1?rss=1"/>
<rdf:li rdf:resource="https://www.medrxiv.org/content/10.64898/2026.09.12.26360876v1?rss=1"/>
</rdf:Seq>
</items>
<prism:eIssn/>
<prism:publicationName>medrxiv</prism:publicationName>
<prism:issn/>

<image rdf:resource=""/>
</channel>
<image rdf:about="">
<title>medrxiv</title>
<url>https://www.medrxiv.org/sites/default/files/medrxiv_internal_logo.png</url>
<link>http://medrxiv.org</link>
</image>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.16.26363186v1?rss=1">
<title>
<![CDATA[
A national-scale digital atlas of recorded care patterns across more than 1000 diagnostic categories using real-world health data: a retrospective observational study 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.16.26363186v1?rss=1
</link>
<description><![CDATA[
Background Routinely collected health data can reveal recorded care around diagnosis, but disease-specific curation does not scale across thousands of conditions. We scaled a previously developed automated framework that reduces heterogeneous clinical events into standardised, diagnosis-centred summaries of real-world care patterns. Methods In this retrospective observational study, we used EST-Health-30, a pseudonymised 30% random sample of Estonian residents with health-care records from 2012 to 2024. We constructed cohorts for ICD-10 three-character diagnostic categories using the first eligible observed diagnosis after a 3-year diagnosis-free observable lookback. For each category, we summarised recorded clinical events from a 90-day pre-index window, a 30-day post-index window, and a 365-day post-index window. An enrichment-based workflow compared earlier self-comparator periods and matched population controls, then filtered and aggregated concepts across clinical domains. Diagnosis-concept relationships were assessed by rubric-guided language-model classification, and atlas plausibility and utility by ten experts. Findings Of 1645 observed diagnostic categories, 1080 met inclusion criteria, yielding 3240 diagnosis-window summaries across 509,856 individuals. Each summary characterised diagnosis-associated events by prevalence, enrichment, timing, co-occurrence, and variation between patient groups after reducing candidate concepts by 97-98%. Among 296,303 retained diagnosis-concept pairs, the language model classified 24% as directly related, 55% as indirectly related, and 21% as noisy. Experts rated the atlas highly for generating research questions (median 6.5 of 7 [IQR 6-7]), providing information difficult to obtain conventionally (6 [5-7]), and supporting observational study design (5.5 [5-6]). Interpretation The atlas makes multidimensional diagnosis-centred patterns in longitudinal health records inspectable at population scale in order to support clinical orientation, hypothesis generation, and observational study planning. Outputs represent recorded care rather than individual patient trajectories or causal effects. Funding Estonian Research Council; European Union; Estonian Ministry of Education and Research; Innovative Medicines Initiative 2 Joint Undertaking; EFPIA.
]]></description>
<dc:creator><![CDATA[ Haug, M., Mooses, K., Oja, M., Heinsar, S., Lotman, E.-M., Pold, M., Pauklin, P., Loo, L., Ilves, P., Ruus, T., Grents, G., Uuskula, A., Arrak, M., Reisberg, S., Vilo, J., Kolde, R. ]]></dc:creator>
<dc:date>2026-09-23</dc:date>
<dc:identifier>doi:10.64898/2026.09.16.26363186</dc:identifier>
<dc:title><![CDATA[A national-scale digital atlas of recorded care patterns across more than 1000 diagnostic categories using real-world health data: a retrospective observational study]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-23</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.22.26363642v1?rss=1">
<title>
<![CDATA[
Mind the gap: emergent clinical risk at the interface of two individually safe AI systems in a multilingual ambient scribe 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.22.26363642v1?rss=1
</link>
<description><![CDATA[
Abstract Objectives To determine whether single-layer evaluation characterises clinical risk in the final note of a multilingual ambient AI scribe, and where serious errors arise. Design Two-arm evaluation on one common clinical-risk scale, using a frozen, reference-aligned synthetic corpus. Setting Scripted consultations spanning 77 languages and five complexity levels, from simple general practice to expert multidisciplinary-team handover. Main outcome measures Per-session incidence of at least one serious (HIGH or CRITICAL) note-layer error in intrinsic (reference script to note, n=385) and end-to-end (automatic speech recognition (ASR) transcript to note, n=2,302) generation; three LLM raters classified discrepancies using a four-mode taxonomy. Results Intrinsic note generation had a 3.4% serious-error rate, with no detectable gradient across complexity (1% to 6%; Cochran-Armitage p=0.42) or language resource (high 2%, medium 5%, low 4%). End-to-end serious errors rose steeply with complexity (L1 1% to L4/L5 24%) and language scarcity (high 7%, medium 10%, low 15%; OR 1.54 per tier, 95% CI 1.15-2.07; p=0.004, GEE clustered on language). Of serious in-note errors, 87% were ASR-derived and 7% note-originated; amplification exceeded correction threefold (fate entropy 1.19 bits). Repeated generation showed moderate reproducibility (Fleiss kappa 0.42); word error rate explained 24% of cross-language variance versus ~1% for transcription risk density. Conclusions Low serious-error rates in individual layers did not preclude higher end-to-end note risk. Evaluation should therefore include the clinician-facing note, because component-level metrics cannot fully capture risk created or transformed at the transcription-to-generation interface.
]]></description>
<dc:creator><![CDATA[ Bergman, H., Liu, V., Austin, B., Sanghera, R. ]]></dc:creator>
<dc:date>2026-09-23</dc:date>
<dc:identifier>doi:10.64898/2026.09.22.26363642</dc:identifier>
<dc:title><![CDATA[Mind the gap: emergent clinical risk at the interface of two individually safe AI systems in a multilingual ambient scribe]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-23</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.16.26363276v1?rss=1">
<title>
<![CDATA[
Evaluation of an AI-Powered Patient Education Application in Bariatric Surgery: A Prospective Mixed-Methods Feasibility Study 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.16.26363276v1?rss=1
</link>
<description><![CDATA[
Background Patient education is critical to safe and effective postoperative bariatric surgery care. Large language models (LLMs) may improve access to personalized patient education but evidence regarding direct patient use in clinical settings remains limited. We evaluated BariBot, an LLM-powered educational application for patients undergoing bariatric surgery. Methods We conducted a prospective, single-arm mixed-methods feasibility study at a tertiary academic medical center from May 2024 to May 2025. Adult patients undergoing bariatric surgery used BariBot independently on postoperative day 1. The application provided conversational postoperative education and adapted language complexity to participants' educational attainment. Outcomes included feasibility, usability, acceptability, perceived trustworthiness, and anticipated future use. Usability was assessed using the System Usability Scale (SUS). Immediately after BariBot use, participants completed a semi-structured qualitative interview assessing acceptability, perceived trustworthiness, and anticipated future use. Interview transcripts underwent thematic analysis. Results Sixteen participants completed the BariBot use session, SUS assessment, and post-intervention interview. No technical failures prevented application use, and no participant requested early discontinuation. The median SUS score was 91.3 (IQR, 81.3-98.1), consistent with excellent perceived usability. Qualitative analysis showed that participants valued the application's immediacy, structured responses, conversational context, ability to address follow-up questions, judgement-free nature, and future availability between clinical encounters. Trust was strengthened by alignment with prior clinical guidance and perceived institutional association, although participants emphasized the need for transparency, clinician oversight, and use as an adjunct rather than a replacement of clinician guidance. Conclusions BariBot was feasible to implement and was associated with high perceived usability and acceptability postoperatively after bariatric surgery. Larger, longitudinal, and comparative studies are needed to evaluate safety during unsupervised use and determine whether LLM-based educational applications improve patient understanding, engagement, and clinical outcomes. Keywords: Artificial intelligence, large language models, chatgpt, bariatric surgery, patient education, health literacy.
]]></description>
<dc:creator><![CDATA[ Samaan, J. S., Nguyen, S. T., Hamamah, S., Tran, A., Ceasar, R. C., Wolfe, N., Yu, T., Chavez, D., Rajeev, N., Advani, R., Watson, R., Abu Dayyeh, B., Sandhu, K., Wong, H. J., Tatonetti, N. P., Samakar, K. ]]></dc:creator>
<dc:date>2026-09-23</dc:date>
<dc:identifier>doi:10.64898/2026.09.16.26363276</dc:identifier>
<dc:title><![CDATA[Evaluation of an AI-Powered Patient Education Application in Bariatric Surgery: A Prospective Mixed-Methods Feasibility Study]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-23</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.21.26363564v1?rss=1">
<title>
<![CDATA[
Systematic Review and External Validation of Clinical Prediction Models for Adverse Pregnancy Outcomes Using Routinely Collected Pre-Conception and Early Pregnancy Data 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.21.26363564v1?rss=1
</link>
<description><![CDATA[
Objective Although numerous clinical prediction models (CPMs) have been developed to identify pregnancies at risk of serious adverse outcomes, most lack robust external validation, which restricts their reliability in clinical practice. This study aims to systematically review and externally validate existing clinical prediction models for Gestational Diabetes Mellitus (GDM), Pre-eclampsia (PE), Stillbirth, and Small-for-Gestational-Age (SGA) focusing on models for use in pre-conception or early pregnancy. The models were evaluated using a large, representative primary care cohort to assess their performance within a UK population. Method and Analysis A two-stage systematic review identified CPMs for GDM, PE, Stillbirth, and SGA. Eligible models used routinely collected maternal characteristics, fully reported the model equation, and included predictors commonly measured during the preconception or early pregnancy period. Validation was performed using the UK Clinical Practice Research Datalink (CPRD Aurum), including over 1.58 million pregnancies for women aged 14 - 49 between 2000 to 2020. Model performance was assessed through discrimination (C-statistic), calibration-in-the-large (CITL), calibration slope and calibration plot, with pooled estimates across 20 imputations. Results Of 30 studies included, 48 models were identified, comprising 47 binary outcome models and one continuous outcome model, and 23 models used UK cohorts. Models were developed in cohorts comprising 101 to 113,415 participants (median N=5,013). Whilst 29 models have been externally validated, only three have been validated in cohorts including UK population satisfying the requisite sample size. Across all outcomes, most models demonstrated limited generalisability in the UK population. The C-statistic for GDM ranged from 0.29 to 0.77, for PE from 0.42 to 0.71, whilst models for stillbirth and SGA showed weaker discrimination from 0.54 to 0.56 and 0.58 to 0.62, respectively. Models across all outcomes exhibited substantial miscalibration, with calibration-in-the-large ranging from -5.66 to 1.35, and calibration slope from -0.65 to 17.61. Conclusion Existing CPMs for pre-conception and early prediction of adverse pregnancy outcomes showed poor discrimination and frequent miscalibration in a large UK cohort. Whilst some models for SGA had excellent performance, most existing models for GDM, PE, and stillbirth demonstrated suboptimal generalisability, likely a result of differences in population characteristics and outcome prevalence. Therefore, most CPMs evaluated are not suitable for direct implementation in UK clinical practice and require further calibration, updating, even re-development using large, representative data before achieving clinical utility.
]]></description>
<dc:creator><![CDATA[ Zhang, Y., Martin, G. P., Darren, A., van Staa, T., Palin, V. ]]></dc:creator>
<dc:date>2026-09-22</dc:date>
<dc:identifier>doi:10.64898/2026.09.21.26363564</dc:identifier>
<dc:title><![CDATA[Systematic Review and External Validation of Clinical Prediction Models for Adverse Pregnancy Outcomes Using Routinely Collected Pre-Conception and Early Pregnancy Data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-22</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.21.26363538v1?rss=1">
<title>
<![CDATA[
An auditable evidence compiler for large language model-assisted systematic reviews 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.21.26363538v1?rss=1
</link>
<description><![CDATA[
Background: Large language models (LLMs) can support systematic reviews, but accurate individual outputs do not establish whether the final synthesis preserves the clinical question, accounts for statistical dependence and incorporates corrections. Objective: To develop and evaluate a framework linking LLM-assisted evidence processing to a versioned, auditable release of synthesis outputs. Methods: We used LLM agents to interpret sources and extract data. Deterministic code enforced statistical rules; investigators resolved material ambiguities and authorized release. We specified ten release properties covering evidence identities, statistical contributions and propagation of corrections. We retrospectively evaluated six integrity domains and historical failure events in one registered prognostic review, without an external comparator or held-out domain. Results: Fifty distinct root-cause events were documented, including 12 that had changed a pooled result before correction. Forty-six were resolved, and four remained disclosed limitations. The corpus comprised 454 reports, 445 studies, 441 cohort entities and 421 dependence clusters. Forty-one of 49 registered analyses were fitted, and eight retained explicit non-fitted states. All 39 source records across five principal analysis families reached a terminal source state. Two implementations within the project agreed across 1,217 numerical comparisons. All 94 file comparisons between release and publication packages were byte-identical. Two reviewers confirmed 39 principal records after seeing the same recommendations. Conclusions: This case provides a framework for inspecting synthesized evidence together with its provenance, statistical meaning and correction history. Comparative validity, generalizability and benefit in patient-centred care require independent evaluation.
]]></description>
<dc:creator><![CDATA[ Yin, C., Jing, Z., Zhang, Z. ]]></dc:creator>
<dc:date>2026-09-22</dc:date>
<dc:identifier>doi:10.64898/2026.09.21.26363538</dc:identifier>
<dc:title><![CDATA[An auditable evidence compiler for large language model-assisted systematic reviews]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-22</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.21.26363536v1?rss=1">
<title>
<![CDATA[
Scalable Causal-Interpretable Machine Learning for Cancer Prescreening Using Electronic Health Records 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.21.26363536v1?rss=1
</link>
<description><![CDATA[
Electronic health records provide a source of real-world information for disease risk prediction, but many machine learning models learn implicit associations that are difficult to inspect or use for clinical reasoning. Causal Bayesian networks (CBNs) offer a form of causal-interpretable machine learning by representing conditional dependencies, putative directional relations, and probabilistic evidence propagation in a directed acyclic graph. However, conventional CBN structure learning becomes unstable and computationally expensive when applied to large-scale EHR data. To address this limitation, we propose UPEBNL, a scalable framework for causal-interpretable cancer prescreening based on parallel CBN learning. UPEBNL integrates adaptive data slicing, quality-aware structure aggregation, and global DAG construction to learn stable and interpretable dependency structures from large-scale observational data. We evaluated UPEBNL in high-dimensional and multi-million-sample simulations and applied it to EHR-based prescreening for esophageal and colorectal cancer. In simulations, UPEBNL improved structural recovery accuracy by nearly 40% and achieved up to a 221.28-fold speedup over conventional CBN learning strategies. For cancer risk prediction, the learned CBNs provided interpretable evidence paths and achieved validation AUCs of 0.8171 for esophageal cancer and 0.784 for colorectal cancer. Calibration and decision curve analyses further supported the reliability and clinical utility of the models. These findings suggest that scalable CBN learning can support interpretable cancer prescreening from large-scale EHR data.
]]></description>
<dc:creator><![CDATA[ Zhang, S., Xue, F. ]]></dc:creator>
<dc:date>2026-09-22</dc:date>
<dc:identifier>doi:10.64898/2026.09.21.26363536</dc:identifier>
<dc:title><![CDATA[Scalable Causal-Interpretable Machine Learning for Cancer Prescreening Using Electronic Health Records]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-22</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.16.26363030v1?rss=1">
<title>
<![CDATA[
Uncoded Clinical Features from Multilingual Electronic Health Records in Catalonia: Development and Validation Study 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.16.26363030v1?rss=1
</link>
<description><![CDATA[
Background: Unstructured free-text narratives in electronic health records (EHRs) contain critical clinical information that is not captured by structured standard codes. In Catalonia, primary care notes follow a semi-structured format known as MEAP (in Catalan). However, extracting structured variables using cloud-hosted commercial large language models (LLMs) raises substantial data privacy concerns and incurs high operational costs. Objective: To develop and validate a privacy-preserving local hybrid pipeline combining a lightweight small language model (SLM) with deterministic regular expressions for extracting six uncoded urinary tract infection (UTI) clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from primary care EHR narratives. Methods: We deployed the open-weight model (Phi4-mini, 3.8B parameters) within the institutional firewall using Ollama. Prompts were iteratively refined in collaboration with clinicians, and regular expressions were applied post-inference. Validation was conducted across two independent arms: (1) a double-blind clinician gold standard review of 60 real patient records (consensus-resolved cases), and (2) an adversarial synthetic dataset of 720 notes enriched with linguistic noise, generated using GPT-4.1 and Grok-4.1. Point estimates and 95% confidence intervals (CIs) were computed for key performance metrics. Results: The pipeline processed 15,498 MEAP narratives from a matched cohort of 2,962 primary care patients and identified 4,663 clinical feature occurrences. Patients progressing to acute pyelonephritis (cases) presented a higher features burden than non-progressing controls (71.3% vs. 55.8%; standardized mean difference [SMD] = 0.326), particularly for fever (33.0% vs. 9.0%; SMD = 0.616) and lumbar pain (29.0% vs. 9.8%; SMD = 0.502). In real-world clinician validation, overall performance yielded 93.8% accuracy (95% CI 90.6%-96.1%), 83.6% sensitivity (95% CI 73.0%-91.2%), 96.6% specificity (95% CI 93.6%-98.4%), 87.1% positive predictive value (PPV; 95% CI 77.0%-93.9%), and 95.5% negative predictive value (NPV; 95% CI 92.3%-97.7%). In synthetic stress-testing, the pipeline demonstrated 87.2% accuracy (95% CI 84.6%-89.6%), 73.9% sensitivity (95% CI 68.9%-78.5%), near-perfect specificity (99.5%, 95% CI 98.1%-99.9%), and 99.2% PPV (95% CI 97.2%-99.9%). Conclusions: A privacy-preserving local hybrid framework combining a compact open-weight SLM with regular expression rules achieves high specificity and precision for extracting six uncoded clinical features from multilingual primary care narratives. Running entirely within institutional servers, this approach keeps patient data secure, meets privacy standards, and avoids Application Programming Interface (API) costs, making it a practical tool for EHR research, though with lower sensitivity for narratively complex descriptions.
]]></description>
<dc:creator><![CDATA[ Ouchi, D., Riudavets, S., Franquet, A., Fernandez-Garcia, S., Llor, C., Moragas Moreno, A., Giner-Soriano, M., Torres, F., Morros, R. ]]></dc:creator>
<dc:date>2026-09-22</dc:date>
<dc:identifier>doi:10.64898/2026.09.16.26363030</dc:identifier>
<dc:title><![CDATA[Uncoded Clinical Features from Multilingual Electronic Health Records in Catalonia: Development and Validation Study]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-22</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.16.26363127v1?rss=1">
<title>
<![CDATA[
Evaluation of linkage between health, education and social care administrative data for 21 million children in England 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.16.26363127v1?rss=1
</link>
<description><![CDATA[
Background Population-level administrative data cohorts hold huge potential for generating evidence to improve child health. Understanding the quality of these data is critical to enabling us to fully realise this potential. We aimed to evaluate linkage between children's health, education and social care administrative records (ECHILD) to identify possible sources of bias. Methods We created four cohorts to understand the drivers of inclusion/exclusion within ECHILD. For each cohort, we calculated the linkage rate by counting the total number of eligible individuals who appeared in the ECHILD linkage spine and any of the corresponding hospital/school component datasets. We then assessed whether linkage rates varied according to sociodemographic characteristics (year of birth, sex, ethnicity, region). Results ECHILD currently captures 34.2 million patients and 25.3 million pupils born between September 1984 and March 2023, aged 0-38 years. Over 90% of children with hospital birth records linked to a school record (rising to 93% for those born from 2014 onwards); 88% of children attending school in 2001/02 linked to a patient record (rising to 96% for those born in 2019/20), resulting in a total of 21 million linked individuals. Substantial variation existed between sociodemographic groups, with those living in London, residing in more deprived areas, or with recorded ethnicity other than White being less likely to be included. Conclusions Population-level datasets such as ECHILD often under-represent specific groups. Irrespective of the mechanisms by which individuals are excluded (e.g. opt outs, linkage errors or not interacting with services), those in minoritised ethnic groups and those living in more deprived areas are most affected. Linkage evaluations are critical for understanding who is and who is not included in analyses of these data so that researchers can account for potential sources of bias within analyses.
]]></description>
<dc:creator><![CDATA[ Nguyen, V. G., Lam, J., Stone, T., Blackburn, R., Gilbert, R., Ruiz Nishiki, M., Harron, K. ]]></dc:creator>
<dc:date>2026-09-22</dc:date>
<dc:identifier>doi:10.64898/2026.09.16.26363127</dc:identifier>
<dc:title><![CDATA[Evaluation of linkage between health, education and social care administrative data for 21 million children in England]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-22</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.20.26363523v1?rss=1">
<title>
<![CDATA[
Human-centred co-design of a dual-purpose heart failure dashboard 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.20.26363523v1?rss=1
</link>
<description><![CDATA[
Introduction: Heart failure care requires coordination across hospital and community settings, yet information is often fragmented across electronic medical records and clinical systems. Clinical dashboards can bring together and display information to support care delivery and service management; however, existing dashboards have generally focused on specific measures, interventions and monitoring pathways. This study aimed to co-design, iteratively develop and user-test an integrated heart failure dashboard that links patient-level clinical decision-making with service-level management. Methods: A human-centred design approach comprising needs identification, collaborative ideation, iterative prototype development and end-user testing was applied across two complementary operational and clinical dashboard streams. Thirty-six clinicians, health service managers, data and implementation scientists, and consumers from metropolitan and regional services participated. Results: For operational decision-making, participants prioritised real-time visibility of patients across heart failure services, patient trajectories, service-performance information, and identification of variation in guideline-directed care and outcomes. Clinical priorities included rapid synthesis of longitudinal information, optimisation of guideline-directed medical therapy, continuity across care settings, and clinical workload prioritisation. These requirements informed the dashboard prototypes. Many prioritised information elements were incompletely represented in structured data and distributed across disconnected systems and structured and unstructured clinical data sources. Conclusion: Human-centred co-design identified complementary patient- and service-level information needs and translated them into linked dashboard prototypes. The proposed dashboard brings together current clinical status and longitudinal heart failure care history at the patient-level alongside service-level patterns. Implementation is required to evaluate whether these linked views improve care processes and patient outcomes.
]]></description>
<dc:creator><![CDATA[ Blake, V. K., Craig, S. E., Carvalho, L., Thomson, M., Yu, J., Jorm, L., Lovell, N., Ooi, S.-Y. ]]></dc:creator>
<dc:date>2026-09-22</dc:date>
<dc:identifier>doi:10.64898/2026.09.20.26363523</dc:identifier>
<dc:title><![CDATA[Human-centred co-design of a dual-purpose heart failure dashboard]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-22</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.18.26363448v1?rss=1">
<title>
<![CDATA[
The Online Health Safety Gap: Consensus Alignment Does Not Imply Safety in Peer-to-Peer Health Narratives. 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.18.26363448v1?rss=1
</link>
<description><![CDATA[
Online health information systems evaluate content by its alignment with medical consensus, treating that alignment as a reliable signal of safety. In peer-to-peer health discourse, that assumption fails: advice that is factually accurate is not necessarily safe to act on. This paper identifies and empirically validates the Online Health Safety Gap, a structural misspecification in which medically accurate content can carry substantial health risk. Across 704 health narratives (699 classified), two scores computed from non-overlapping feature sets, Narrative Truth Distance for epistemic divergence and Narrative Risk Score for health risk potential, share only 4.9% of their variance (r=0.222, p<0.001). Consequently, 39.6% of narratives fall in the two off-diagonal quadrants that single-axis systems mishandle by construction: aligned-but-risky content (25.2%) invisible to fact-checking, and divergent-but-safe content (14.4%) incorrectly suppressed. The same independence constrains what existing benchmarks can measure. On an expert-labeled misinformation benchmark of 437 posts containing 127 misinformation instances, the highest discrimination attained is Youdens J of 0.349, by a supervised classifier trained on those labels directly, which neither frozen biomedical embeddings nor a prompted large language model improves upon. Labels encoding factual accuracy therefore carry little signal about behavioral risk, and benchmarks of this construction cannot evaluate risk-aware assessment. To operationalize the gap, we introduce the Classification Quadrant, which maps epistemic divergence and health risk potential onto four governance categories with distinct intervention implications, and argue for a shift from fact-centric evaluation to risk-aware assessment.
]]></description>
<dc:creator><![CDATA[ Clark, O., Joshi, K., Ahmed, Z. ]]></dc:creator>
<dc:date>2026-09-21</dc:date>
<dc:identifier>doi:10.64898/2026.09.18.26363448</dc:identifier>
<dc:title><![CDATA[The Online Health Safety Gap: Consensus Alignment Does Not Imply Safety in Peer-to-Peer Health Narratives.]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.19.26363476v1?rss=1">
<title>
<![CDATA[
Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.19.26363476v1?rss=1
</link>
<description><![CDATA[
The trustworthy use of large language models (LLMs) is a growing challenge in safeguarding sensitive patient data from leakage. We introduce and evaluate a privacy-preserving knowledge distillation framework for LLM-based clinical modeling, using multimorbidity scoring as a healthcare task. Although LLMs can encode rich clinical knowledge and improve upon traditional rule-based comorbidity scoring, their direct evaluation on large-scale biobank data remains constrained by patient privacy. In our framework, multimorbidity reasoning is distilled from state-of-the-art LLM teacher models into compact student models (CoLLMs) using synthetic cohorts that preserve UK Biobank distributions without exposing real patient data. This approach achieves high-fidelity knowledge transfer (Spearman rho = 0.75-0.89). Independent LLM-as-a-Judge evaluation confirms the clinical significance of the distilled knowledge and reveals substantial variability among teacher models. When applied to real UK Biobank data, CoLLM-derived multimorbidity scores improve survival prediction (C-index up to 0.91) and exhibit higher SNP heritability (h^2 approximately 0.05). Our work establishes a trustworthy, privacy-compliant pathway for large-scale healthcare applications of LLMs.
]]></description>
<dc:creator><![CDATA[ Awasthi, R., Yang, Y., Li, M., Zhu, X. ]]></dc:creator>
<dc:date>2026-09-21</dc:date>
<dc:identifier>doi:10.64898/2026.09.19.26363476</dc:identifier>
<dc:title><![CDATA[Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.15.26362921v1?rss=1">
<title>
<![CDATA[
MES: A Multi-Agent Evidence Synthesis System for Medical Decision-Making 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.15.26362921v1?rss=1
</link>
<description><![CDATA[
Medical evidence synthesis increasingly requires published studies, real-world clinical data and structured biomedical knowledge, yet most automated systems remain centered on literature retrieval and summarization. Here we present MES, a multi-agent framework for source-grounded medical evidence synthesis. MES coordinates six specialized agents to decompose clinical questions, retrieve literature and trial evidence, generate question-specific RWE by constructing and analyzing real-world cohorts, query biomedical knowledge graphs, integrate quantitative findings and screen generated claims against their cited evidence. The framework uses evidence-based medicine taxonomies to clarify underspecified questions, adapts evidence use when sources are absent or discordant and reports unresolved gaps rather than forcing consensus. MES also provides claim-level provenance tracking and cross-agent consistency checking, allowing final reports to distinguish trial evidence, real-world associations and knowledge-graph support. We evaluated MES in two clinical use cases and across 144 clinical queries spanning six evidence-based medicine categories. MES generated structured reports that preserved source traceability, identified cross-source disagreement and communicated uncertainty across diverse clinical question types. MES provides an auditable framework for organizing, testing and contextualizing heterogeneous clinical evidence while keeping the evidentiary basis of each conclusion explicit.
]]></description>
<dc:creator><![CDATA[ Li, H., Pan, W., Rajendran, S., Zang, C., Gabeskiria, J., Liu, V., Feng, V., Penmetsa, M., Lin, L., Jayakar, A. B., Boucher, A., Schenck, E. J., Yang, H. S., Wang, F. ]]></dc:creator>
<dc:date>2026-09-21</dc:date>
<dc:identifier>doi:10.64898/2026.09.15.26362921</dc:identifier>
<dc:title><![CDATA[MES: A Multi-Agent Evidence Synthesis System for Medical Decision-Making]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.19.26363431v1?rss=1">
<title>
<![CDATA[
Evaluating agentic simulation for local public health estimation 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.19.26363431v1?rss=1
</link>
<description><![CDATA[
Large language model (LLM)-based generative agents can reproduce aspects of individual human behavior, but whether they can be scaled to geographically grounded populations that reproduce real-world health behaviors remains unclear. Here, we introduce LLMPopSim, a generative population simulation framework that integrates U.S. Census and Centers for Disease Control and Prevention data to construct synthetic individuals and uses an LLM to simulate individual health behaviors whose aggregate outcomes can be evaluated at the community level. We developed the framework using historical Hawaii data and evaluated temporal and geographic generalizability using held-out 2022 cohorts from Hawaii and New York State, with colorectal cancer screening and mammography as proof-of-concept behaviors. Across the four state-outcome evaluations, mean absolute error ranged from 3.5 to 15.0 percentage points and correlations between simulated and observed ZCTA-level prevalence ranged from 0.26 to 0.69. Performance differed across dimensions of population fidelity: colorectal cancer screening predictions preserved geographic ranking more strongly but systematically overestimated prevalence and compressed geographic variation, whereas mammography achieved lower absolute error but weaker geographic correlation and inconsistent preservation of between-community variability. Prediction error was greatest in communities with lower observed screening prevalence and varied across community characteristics without a uniform socioeconomic gradient. These findings demonstrate that individually represented LLM-based synthetic agents can aggregate into population-level patterns that retain measurable features of real-world health behavior across temporal and geographic transfer, while identifying calibration, distributional fidelity and subgroup performance as key challenges for generative population simulation. LLMPopSim provides an empirical foundation for developing synthetic populations that may ultimately enable simulation of heterogeneous population responses to public health interventions.
]]></description>
<dc:creator><![CDATA[ Nakatsuka, M. A., Steele, R. J., Han, X., Oermann, E. K. ]]></dc:creator>
<dc:date>2026-09-21</dc:date>
<dc:identifier>doi:10.64898/2026.09.19.26363431</dc:identifier>
<dc:title><![CDATA[Evaluating agentic simulation for local public health estimation]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.19.26363460v1?rss=1">
<title>
<![CDATA[
A world model simulates the latent dynamics of human health 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.19.26363460v1?rss=1
</link>
<description><![CDATA[
Human health is a single underlying state that no measurement observes directly: diagnoses, blood tests, molecular profiles and images each capture one facet at separate times. Inferring health from such evidence requires a representation that integrates every modality and is carried forward and revised as observations arrive, which is the defining task of a world model. Here we introduce HealthFlux, a pan-modal world model that learns the latent dynamics of health from 5,647 features across eleven data domains, spanning clinical records, blood tests, genetics, proteomics, metabolomics and MRI, in 502,166 UK Biobank participants. Its hybrid state-space architecture combines ODE-based evolution between observations with continuous-time recurrent updates when new measurements arrive. In held-out participants, HealthFlux predicts 195 diseases and death over five years with a mean AUROC of 0.816, compared with 0.715 for the previous state-of-the-art model. These results remain true when validated in three independent cohorts, and HealthFlux also outperforms specialized clinical risk scores for disease and mortality. Simulated forward without further observations, the state continues to predict disease accurately up to a decade after the last measurement. HealthFlux predicts diseases excluded entirely from training, with a mean AUROC of 0.769, evidence that it has learned health itself rather than the diseases it was trained on. Each modality contributes information the others lack, and integrating them identifies individuals at risk whom single-modality models miss. HealthFlux thus makes health itself the object of prediction: one continuously updated state, informed by any measurement, from which the risk of any disease can be read years before diagnosis.
]]></description>
<dc:creator><![CDATA[ Wang, Z., Zhang, O., Wang, J., Jiang, Y., Ma, Z. C., Guo, M., Al Dajani, S., Eames, A., Moldakozhayev, A., Poganik, J. R., Moqri, M., Zhai, R., Isakova, A., Wu, Y., Yin, Z., Cong, L., Zou, J., Snyder, M. P., Ruohola-Baker, H., Gladyshev, V. N., Wyss-Coray, T., Ying, K. ]]></dc:creator>
<dc:date>2026-09-21</dc:date>
<dc:identifier>doi:10.64898/2026.09.19.26363460</dc:identifier>
<dc:title><![CDATA[A world model simulates the latent dynamics of human health]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.18.26363428v1?rss=1">
<title>
<![CDATA[
Identifying Family Relationships from Electronic Health Records: A Machine Learning Approach 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.18.26363428v1?rss=1
</link>
<description><![CDATA[
Objective: To develop and evaluate machine learning methods for automated identification of family relationships from electronic health records (EHRs). We employed Random Forest classifiers to identify five relationship types: Mother-Child, Father-Child, Sibling-Sibling, Twin-Twin, and Partner-Partner, using a four-stage iterative refinement process to improve precision for patient-centered outcomes research (PCOR) and healthcare applications. Materials and Methods: We used two large-scale Indiana datasets: the Indiana Network for Patient Care (INPC), comprising approximately 15 million unique individuals across approximately 45 million medical records, and the Indiana Natality dataset (birth certificates, 1970-2024), which served as the gold standard. Positive cases were derived by linking verified relationships from Natality records to INPC using a shared Global Identifier. Negative cases were drawn from a three-tier blocking strategy combined with sliding window restriction and similarity scoring, which together reduced the comparison space from over 1014 potential pairs to approximately 121 million candidates. Five Random Forest classifiers underwent four- stage iterative refinement: initial training, feature removal for multicollinearity, feature engineering to address systematic errors, and training-data refinement restricted to relationship pairs with at least one shared contact feature. Results: The models achieved precision of 0.92 to 1.00, recall of 0.97 to 1.00, and F1 scores of 0.94 to 1.00 across all relationship types (Mother-Child 0.97, Father-Child 0.98, Sibling-Sibling 0.98, Twin-Twin 1.00, Partner-Partner 0.94). Between 15% and 78% of Natality-verified relationships lacked any shared contact information in INPC and were structurally undetectable by record linkage; restricting training and evaluation to linkable pairs improved precision, and training-data refinement improved F1-scores by a further 0.02-0.04 for the Mother-Child, Father-Child, and Sibling models. High-confidence predictions (probability [&ge;] 0.9) captured 77-99% of true positives. Age difference was the primary predictor for parent- child and twin relationships, while phone number similarity was most important for siblings. Discussion: The framework identifies family relationships from standard demographic fields available in most EHR systems. The iterative refinement approach shows how systematic error analysis can guide 1 methodological improvements. Training-data refinement addresses a constraint in EHR-based family linkage: not all verified biological relationships have overlapping demographic footprints in healthcare data. Unlike rule-based approaches, the framework provides probability scores that enable configurable deployment thresholds. Conclusion: The Random Forest models demonstrate high performance suitable for large-scale research deployment and are ready for testing in clinical applications rather than immediate widespread clinical use. The methodology can be adapted to other EHR systems, supporting family-centered study design in patient- centered outcomes research. Keywords: family linkage, electronic health records, machine learning, Random Forest, record linkage, patient-centered outcomes research, health informatics
]]></description>
<dc:creator><![CDATA[ Pundir, A., Hill, A., Kahn, M. G., Kwan, B. M., Lindberg, D. M., Grannis, S. J., Schleyer, T. K., Schilling, L. M., Kautz, S. V., Ong, T. C. ]]></dc:creator>
<dc:date>2026-09-21</dc:date>
<dc:identifier>doi:10.64898/2026.09.18.26363428</dc:identifier>
<dc:title><![CDATA[Identifying Family Relationships from Electronic Health Records: A Machine Learning Approach]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.15.26363093v1?rss=1">
<title>
<![CDATA[
Federated learning in a regulator-audited secure processing environment: a multi-hospital deployment study 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.15.26363093v1?rss=1
</link>
<description><![CDATA[
Background. Federated learning enables collaborative model development without transferring patient level data between institutions. However, translation into routine healthcare practice remains limited because deployment requires more than distributed model training. Federated infrastructures must operate within secure processing environments, satisfy governance and security requirements, support auditability, and integrate with existing regulatory workflows for secondary use of health data. Although federated learning has been demonstrated in multiple clinical applications, evidence for its deployment, assessment, and operation within regulated health-data infrastructures remains limited. Methods. We implemented a decentralised swarm learning system for federated learning and integrated it into Acamedic, a certified secure processing environment (SPE) operated by Helsinki University Hospital (HUS). Introduction of the swarm learning coordination layer constituted a material modification to the previously certified HUS SPE and triggered a differential regulatory security assessment under Finland's Act on the Secondary Use of Health and Social Data and associated Findata requirements. The assessment evaluated controls relating to data isolation and locality, identity and access management, logging and monitoring, network security, and environment protection, while establishing a governance model that separates certification of infrastructure-level controls from study specific assessment of data, models, parameter exchanges, and outputs. The swarm learning architecture was developed and deployed across the university hospitals of Helsinki, Turku, and Tampere, with HUS serving as the regulator-assessed implementation within a certified SPE and partner sites operating under local institutional governance frameworks. As an operational exemplar, we trained federated DeepSurv survival models for acute myeloid leukaemia (AML) using harmonised longitudinal laboratory data. Findings. The differential security assessment identified one high-severity, two medium-severity, and two low-severity findings. All high- and medium-severity findings were remediated and verified, after which an independent certification report was issued for the assessed extension. The resulting federated-learning capability was authorised for secondary use of health data and operationalised as a reusable extension to the HUS secure processing environment. Swarm learning was successfully executed across three university hospitals without transferring patient-level data, achieving 100% parameter merge success. Training operations were traceable through infrastructure logs, container logs, and a distributed ledger coordination layer recording participant registration, parameter exchanges, and model-training lineage. On independent test sets from all three hospitals, the swarm-trained AML model demonstrated stronger risk stratification than locally trained reference models, with consistently stronger risk separation, higher log-hazard ratios (swarm vs local: 4.32 vs 1.79 at HUS, 7.20 vs 3.19 at TAYS, and 5.90 vs 2.91 at TYKS), and improved discrimination metrics. The resulting federated-learning capability was authorised for secondary use of health data, operationalised as a reusable extension to the HUS secure processing environment, and made available for future permit approved analyses. Interpretation. This study demonstrates that federated learning can be integrated into a certified secure processing environment, subjected to regulatory assessment, and operated across independent hospitals without centralising patient level data. The principal contribution is a reusable governance and infrastructure model that separates assessment of system-level controls from study specific evaluation of data, models, parameter exchanges, and outputs. By enabling federated learning to function as an assessed infrastructure capability rather than a project-specific exception, the approach provides a practical pathway for operationalising decentralised federated learning within regulated health-data environments. The findings are directly relevant to emerging European frameworks for secondary use of health data, including secure processing environments envisioned under the European Health Data Space, whose implementation builds on many of the same governance principles evaluated in this study.
]]></description>
<dc:creator><![CDATA[ Nieminen, V., Schultze, H., Kukkurainen, S., Rantala, H. J., Hakkarainen, L., Gustafsson, P.-E., Hakala, T., Turunen, E., Johansson, M., Laitinen, T., Tsallos, D., Porkka, K., Virkki, A., Schultze, J., Fey, E. ]]></dc:creator>
<dc:date>2026-09-21</dc:date>
<dc:identifier>doi:10.64898/2026.09.15.26363093</dc:identifier>
<dc:title><![CDATA[Federated learning in a regulator-audited secure processing environment: a multi-hospital deployment study]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.17.26363303v1?rss=1">
<title>
<![CDATA[
LLM-enabled Natural History Study Analysis to Support Rare Disease Research 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.17.26363303v1?rss=1
</link>
<description><![CDATA[
BackgroundRare diseases affect an estimated 300 million people worldwide, yet the research needed to guide diagnosis and treatment is often fragmented across multiple unstructured literature sources. Natural history studies (NHS) are a key source of this evidence, but manually extracting structured information from NHS publications can be tedious and does not scale.

MethodsWe have developed a proof-of-concept for an information extraction pipeline testing three open-source large-language models (LLMs) - Athena-v3-AWQ, Googles Gemma3-27B, and Metas Llama-3.1-70B-Instruct, to extract key NHS characteristics from PubMed abstracts curated from a Chan Zuckerberg Initiative disease research state model corpus (302 gold-standard and 8,338 full-corpus abstracts), and compared the models on efficiency, extraction completeness, and expert-rated accuracy.

ResultsAll three models processed abstracts with success rates exceeding 99%. However, Gemma achieved the best overall performance, with the highest expert-rated accuracy (68.0% of outputs rated "good" vs. 36.0% for Llama and 10.0% for Athena) and the fastest runtime on the full corpus ([~]16 minutes for 3,547 abstracts), despite Llama scoring higher on the automated Token F1 metric (0.874 vs. 0.723), highlighting a divergence between automated and human evaluation. Athenas lower performance was largely attributable to verbatim copying rather than synthesis of extracted content.

ConclusionsThese findings illustrate how locally deployed open-source LLMs can extract structured NHS characteristics at scale, thus supporting their use to accelerate evidence synthesis in rare disease research.
]]></description>
<dc:creator><![CDATA[ Li, K., Sid, E., Zhu, Q. ]]></dc:creator>
<dc:date>2026-09-20</dc:date>
<dc:identifier>doi:10.64898/2026.09.17.26363303</dc:identifier>
<dc:title><![CDATA[LLM-enabled Natural History Study Analysis to Support Rare Disease Research]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-20</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.15.26362961v1?rss=1">
<title>
<![CDATA[
Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.15.26362961v1?rss=1
</link>
<description><![CDATA[
BackgroundMuch of the symptom burden in colorectal cancer (CRC) patients is documented in unstructured discharge-note narrative, and manual extraction is not scalable. Whether large language models (LLMs) outperform rule-based and named entity recognition (NER) methods has not been rigorously benchmarked.

ObjectiveTo benchmark rule-based, NER, and zero-shot LLM methods for extracting 46 cancer-related symptoms from CRC discharge notes against an adjudicated ground truth.

MethodsWe analyzed 2,704 discharge notes from CRC patients in MIMIC-IV. A 46-symptom target list was built from the Memorial Symptom Assessment Scale and the EORTC QLQ-CR29. Four approaches -- dictionary-based rule matching, pretrained clinical NER, and zero-shot Claude Haiku and Gemini 3.5 Flash -- plus two hybrid variants (LLM output with post-hoc rule-based negation filtering) were evaluated against a 200-note gold standard adjudicated by two raters (pooled kappa=0.71, macro kappa=0.49), using Macro/Micro F1, precision, and recall.

ResultsGemini 3.5 Flash performed best (Macro F1=0.70, Micro F1=0.86, Macro Precision=0.74), followed by Claude Haiku (Macro F1=0.63, Macro Recall=0.71); both substantially outperformed rule-based (Macro F1=0.44) and NER (Macro F1=0.38) methods. Post-hoc negation filtering paradoxically degraded LLM performance (Gemini+Hybrid Macro F1=0.58; Claude+Hybrid Macro F1=0.54) by overriding correct predictions through rigid, fixed-window matching.

ConclusionsZero-shot LLMs substantially outperform rule-based and NER approaches for CRC symptom extraction; post-hoc negation correction should not be applied to LLM outputs without syntactic scope validation. Implications for Practice: Zero-shot LLM extraction offers a scalable, accurate alternative to manual chart review and traditional NLP pipelines for oncology symptom surveillance, without institution-specific rule development or model training.
]]></description>
<dc:creator><![CDATA[ Lee, Y., Dinov, I., Hu, X., Jiang, Y. ]]></dc:creator>
<dc:date>2026-09-20</dc:date>
<dc:identifier>doi:10.64898/2026.09.15.26362961</dc:identifier>
<dc:title><![CDATA[Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-20</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.16.26363187v1?rss=1">
<title>
<![CDATA[
A FAIR layer for the INHERENT haemoglobinopathy patient registry 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.16.26363187v1?rss=1
</link>
<description><![CDATA[
Haemoglobinopathy registries support research and outcome monitoring. Still, reuse is limited due to heterogeneous structures, registry-specific coding, and incomplete semantic representation. We designed and implemented a FAIRification workflow for the INHERENT haemoglobinopathy platform, an international genotype-phenotype registry, as part of the HemaFAIR project. The workflow was extended to the Cyprus Haemoglobinopathy Patient Registry to demonstrate its applicability across a second registry. Source data and metadata were transformed through two independent but complementary harmonisation branches executed in parallel: one producing an OMOP CDM representation, and the other generating a CARE-SM representation. The workflow generated graph-based semantic resources, predefined query services, public aggregate dashboards, and application programming interface (API) endpoints. Registry metadata were published through the European Rare Disease Registry Infrastructure and a FAIR Data Point. These outputs support findability, interoperable analysis and controlled reuse while preserving existing governance over patient-level data.
]]></description>
<dc:creator><![CDATA[ Tamana, S., Yiangou, C., Orphanou, K., Xenophontos, M., Papasavva, P. L., Bernabe, C., Roos, M., Wijnbergen, D., Kersloot, M. G., Cornet, R., Minaidou, A., Stephanou, C., Chatzimatthaiou, S., Landi, A., Giannuzzi, V., Bonifazi, F., Lederer, C. W., Kountouris, P. ]]></dc:creator>
<dc:date>2026-09-17</dc:date>
<dc:identifier>doi:10.64898/2026.09.16.26363187</dc:identifier>
<dc:title><![CDATA[A FAIR layer for the INHERENT haemoglobinopathy patient registry]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-17</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.08.26361592v1?rss=1">
<title>
<![CDATA[
Data Auditing and Quality Assurance in a Federated Learning Consortium; Getting the Best of Both Worlds from Cross-Institutional and In-House Data Quality Inspection 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.08.26361592v1?rss=1
</link>
<description><![CDATA[
IntroductionRare and heterogeneous disease research increasingly relies on privacy-enhancing technologies such as federated learning (FL) to enable cross-institutional collaboration across fragmented datasets. However, data quality assurance in FL can be limited, as individual-level data may not be directly accessible. While schema-dependent inspection offers a partial solution, it requires standardised schemas or resource-intensive frameworks, hindering scalability and collaboration. To address these challenges, we implemented two complementary dashboards and evaluated their interplay.

MethodologyWe adapted a recognised data quality control framework for re-use of electronic health record (EHR) data using breast cancer records from Centre Leon Berard into two dashboards: (1) an in-house dashboard for intra-clinic quality assessment, and (2) a federated dashboard for inter-clinic quality assessment. Artificial inconsistencies were introduced into distributed datasets mirroring the inhouse source to evaluate detection capabilities.

ResultsThe in-house dashboard provided granularity and reliability, pinpointing individual-level inconsistencies, while the federated dashboard enabled cross-institutional pattern detection - trade-offs inherent to their designs. The federated system revealed ecosystem-wide trends inaccessible to single-institution tools, whereas the in-house dashboard provided local validation and thoroughness.

DiscussionOur findings confirm a complementary relationship: FL dashboards provide scalable, collaborative oversight but may require additional quality checks, while in-house tools ensure thoroughness at the cost of scalability. A combined model seemingly offers the optimal balance, accommodating both institution-specific needs and collaborative research requirements in evolving, multi-institutional ecosystems.
]]></description>
<dc:creator><![CDATA[ Hogenboom, J., Perez, N., Filori, Q., Sans, A., Lobo Gomes, A., Dekker, A., van der Graaf, W., Husson, O., Crochet, H., Wee, L., Gouthamchand, V. ]]></dc:creator>
<dc:date>2026-09-17</dc:date>
<dc:identifier>doi:10.64898/2026.09.08.26361592</dc:identifier>
<dc:title><![CDATA[Data Auditing and Quality Assurance in a Federated Learning Consortium; Getting the Best of Both Worlds from Cross-Institutional and In-House Data Quality Inspection]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-17</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.11.26362711v1?rss=1">
<title>
<![CDATA[
Perceived value and stakeholder experience of a dedicated allied health professions informatics role in a specialist cancer centre: a mixed-methods service evaluation survey 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.11.26362711v1?rss=1
</link>
<description><![CDATA[
BackgroundDedicated allied health professions (AHP) informatics posts remain rare and largely unevaluated: one NHS trust in ten reports a formal AHP informatics lead.

ObjectivesTo evaluate stakeholder-perceived value and experience of a dedicated AHP Information Officer (AHPIO) role supporting an organisation-wide electronic health record (EHR), and to derive a working framework of the roles practice.

MethodsCross-sectional online survey within a registered service evaluation at a specialist cancer centre in England. A purposive census of 40 stakeholders (clinical and operational AHP staff, digital champions and EHR programme colleagues) rated six agreement items with four free-text questions. Analysis led with distributions, medians and Wilson intervals; inference was exploratory; free-text coding is fully audit-trailed.

ResultsTwenty-seven stakeholders responded (68%). Five items sat at a ceiling (medians 4 to 5; agreement 95% to 100%); all 25 answering the overall item agreed the role adds value, 80% strongly. Data and reporting access was the outlier (median 3.0; 52% neutral or below; agreement 48%, 95% CI 28% to 68%), ranking below every other item (Friedman chi-squared 36.3, df 5, 20 complete cases; Holm-adjusted post-hoc comparisons, rank-biserial -0.73 to -1.00). Nine themes located value in brokerage and routing (18/26), concrete outcome accounts (16/26) and a clinician who understands AHP work (12/26).

ConclusionsStakeholders located the roles value in brokerage, clinical-digital translation, representation and outcome response, resting on a named trusted contact; self-service data is the part-built pillar and development priority. The framework and method offer a replicable template for evaluating AHP informatics roles.
]]></description>
<dc:creator><![CDATA[ Bayquen, D. D. A. ]]></dc:creator>
<dc:date>2026-09-16</dc:date>
<dc:identifier>doi:10.64898/2026.09.11.26362711</dc:identifier>
<dc:title><![CDATA[Perceived value and stakeholder experience of a dedicated allied health professions informatics role in a specialist cancer centre: a mixed-methods service evaluation survey]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-16</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.15.26363148v1?rss=1">
<title>
<![CDATA[
Large Language Model-derived Symptom Clusters and Patient Outcomes in Colorectal Cancer from MIMIC-IV Clinical Notes 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.15.26363148v1?rss=1
</link>
<description><![CDATA[
BackgroundPrior research on symptom clusters (SCs) in colorectal cancer (CRC) has relied primarily on patient-reported outcome surveys, which capture symptom experience at discrete assessment points rather than the continuous documentation generated during routine care, leaving open whether SCs derived from electronic health record (EHR) text carry the same clinical meaning and predictive value. To construct and validate patient-level symptom co-occurrence networks from large language model (LLM)-extracted symptom data in CRC patients, and to test whether resulting SCs predict clinical outcomes.

MethodsUsing a zero-shot LLM extraction pipeline previously benchmarked against a manually annotated ground truth (Macro F1=0.70 for the best-performing model), we extracted 46 symptoms from 2,728 discharge notes of 1,507 CRC patients in MIMIC-IV. Patient-level symptom co-occurrence networks were constructed independently from Gemini 3.5 Flash and Claude Haiku extractions using phi correlation ([&ge;]0.10) and Louvain community detection, with sensitivity analyses across correlation thresholds, random seeds, note-aggregation strategy, and bootstrap resampling. Per-cluster symptom burden scores were tested as predictors of in-hospital mortality, 30-day readmission, and 1-year mortality using logistic regression adjusted for age, sex, and (in sensitivity models) metastatic disease.

ResultsBoth LLMs networks converged on three clinically coherent SCs -- Systemic, Gastrointestinal, and CRC Disease-Specific -- across sensitivity analyses (cross-model Adjusted Rand Index=0.727; 100-seed Louvain ARI=0.985; bootstrap ARI=0.733). The Systemic Symptom Cluster was the most consistent predictor of in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41), while the CRC Disease-Specific Cluster specifically and independently predicted 30-day readmission (OR=1.20); both associations were robust to adjustment for metastatic disease.

ConclusionLLM-extracted symptom data recover clinically coherent, reproducible SCs from unstructured discharge notes that carry independent prognostic value for mortality and readmission, supporting the clinical validity of automated, EHR-derived symptom profiling in CRC.

HighlightsO_LILLMs identified Systemic, CRC Disease-Specific, and Gastrointestinal clusters.
C_LIO_LIThree symptom clusters were reproducible across LLMs and sensitivity analyses.
C_LIO_LISystemic symptom burden predicted in-hospital and 1-year mortality.
C_LI
]]></description>
<dc:creator><![CDATA[ Lee, Y., Dinov, I., Hu, X., Jiang, Y. ]]></dc:creator>
<dc:date>2026-09-16</dc:date>
<dc:identifier>doi:10.64898/2026.09.15.26363148</dc:identifier>
<dc:title><![CDATA[Large Language Model-derived Symptom Clusters and Patient Outcomes in Colorectal Cancer from MIMIC-IV Clinical Notes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-16</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.14.26363047v1?rss=1">
<title>
<![CDATA[
Language models reflect clinical evidence but fail to adapt it to patients 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.14.26363047v1?rss=1
</link>
<description><![CDATA[
BackgroundClinical language models must use evidence to produce the number and action required by a particular case. Whether numerical knowledge reliably becomes a correct clinical response is unclear.

MethodsNUMBERS evaluated 16 model configurations on 1,300 questions linked to public clinical evidence. Linked experiments tested prevalence updating, patient-specific estimates and clinical actions. The direct-action experiment compared cutoff recall plus action selection with a supplied complete rule across 50 rules, five models and 7,500 calls.

ResultsAmong diagnostic estimates outside the source-result tolerance, 77.1% remained within the evidences central 80% predictive range. Models updated prevalence-dependent quantities correctly in 78.5% of comparisons with inputs and a calculation request, versus 33.7% with clinical wording. Supplying inputs and requesting calculation raised patient-level near-target answers from 29.7% to 78.8% across 2,340 pairs. Models selected an incorrect action despite stating a cutoff that implied the correct action in 413 of 3,750 recall-arm calls (11.01%; 95% confidence interval, 8.93 to 13.17). Supplying the complete rule raised action accuracy from 85.63% to 99.41%, an improvement of 13.79 percentage points (95% confidence interval, 11.65 to 15.95).

ConclusionsModels often produced evidence-consistent numbers but failed to adapt them to a case or act consistently with their own stated cutoff. Explicit inputs, calculations and complete rules substantially improved performance in controlled prompts.

FundingNational Academy of Medicine, Agreement No. 2026A008797
]]></description>
<dc:creator><![CDATA[ He, S., Joseph, J. W., Safari, P., Goff, A., Slusarz, P. J., Mohamed, A., Lord, S., Goldstein, J. N., Raja, A. S. ]]></dc:creator>
<dc:date>2026-09-15</dc:date>
<dc:identifier>doi:10.64898/2026.09.14.26363047</dc:identifier>
<dc:title><![CDATA[Language models reflect clinical evidence but fail to adapt it to patients]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-15</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.14.26363036v1?rss=1">
<title>
<![CDATA[
Interpretable Trajectory-Based Feature Extraction from Longitudinal Electronic Health Records 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.14.26363036v1?rss=1
</link>
<description><![CDATA[
Longitudinal electronic health records (EHRs) contain rich information about the evolution of a patients clinical state, but temporal measurements and events are often represented using summary statistics or complex learned representations that can be difficult to interpret clinically. We propose an interpretable trajectory-based feature representation that transforms heterogeneous longitudinal EHR data into compact, human-readable predictive features. Numerical temporal variables are represented by trajectory states such as INCREASING, DECREASING, and STABLE, while categorical clinical events are represented by interpretable states describing repeated occurrence, stability, or change. These trajectory features are combined with static patient and encounter characteristics and evaluated using L1-regularized logistic regression, which provides sparse feature selection and direct identification of positive and negative predictive features.

The approach is evaluated using MIMIC-IV in two clinically distinct prediction tasks: in-hospital mortality following ICU admission and hospital readmission within 30 days after discharge. For ICU mortality, multiple observation windows ranging from 1 to 48 hours are investigated to examine how the availability of longitudinal information affects prediction and the relative contribution of trajectory features. The highest AUROC of 0.8751 is obtained with a 32-hour observation window. For 30-day readmission, trajectories from the final 72 hours of the index hospitalization are evaluated using both three-state and five-state categorical representations. The two representations produce essentially identical discriminative performance (AUROC 0.6875 and 0.6876, respectively), indicating that increasing trajectory granularity provides little predictive benefit in this setting.

The results demonstrate that heterogeneous longitudinal EHR observations can be transformed into compact and directly interpretable temporal features while retaining useful predictive information. The proposed framework therefore provides a simple and model-compatible approach for identifying clinically interpretable patterns of patient evolution and can be extended with more detailed trajectory definitions or applied with other predictive models.
]]></description>
<dc:creator><![CDATA[ Bulgakov, V., Turchin, A. ]]></dc:creator>
<dc:date>2026-09-15</dc:date>
<dc:identifier>doi:10.64898/2026.09.14.26363036</dc:identifier>
<dc:title><![CDATA[Interpretable Trajectory-Based Feature Extraction from Longitudinal Electronic Health Records]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-15</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.14.26363024v1?rss=1">
<title>
<![CDATA[
Synthetic Echocardiograms from Diffusion Models in Rare Cardiovascular Disease 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.14.26363024v1?rss=1
</link>
<description><![CDATA[
The limited availability of imaging data for uncommon cardiovascular phenotypes constrains the development of robust imaging models. We evaluated whether class-conditional diffusion models can generate synthetic transthoracic echocardiograms that improve downstream cardiac imaging tasks. The primary application was cardiac amyloidosis detection in a Duke University cohort using a two-step classifier, with external analyses using EchoNet-Dynamic for image-fidelity assessment and EchoNet-LVH for wall-thickness phenotype classification. The Duke cohort was partitioned at the patient-encounter level into 70% training, 15% validation, and 15% test sets; generators were trained only on the training partition, augmentation levels were selected using validation AUROC, and final evaluation used held-out real test data. In the Duke all-view two-step analysis, adding synthetic images increased AUROC from 0.883 to 0.924, with an AUROC difference of 0.041 (95% CI, 0.013-0.069); in EchoNet-LVH, AUROC increased from 0.832 to 0.864, with an AUROC difference of 0.032 (95% CI, 0.021-0.044). Expert review found that synthetic images were sometimes difficult to identify as synthetic, but rated them lower for diagnostic adequacy. These findings suggest that diffusion-generated echocardiograms may provide a practical approach to augmenting limited training data for selected cardiac imaging tasks and motivate further evaluation across clinical settings.
]]></description>
<dc:creator><![CDATA[ Lyu, P., Henao, R., Kwee, L. C., Peng, F. Z., Vemulapalli, S., Shah, S. H., Khouri, M. G., Zhang, A. ]]></dc:creator>
<dc:date>2026-09-15</dc:date>
<dc:identifier>doi:10.64898/2026.09.14.26363024</dc:identifier>
<dc:title><![CDATA[Synthetic Echocardiograms from Diffusion Models in Rare Cardiovascular Disease]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-15</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.14.26363006v1?rss=1">
<title>
<![CDATA[
Spatially Context-Aware Transformers Facilitate Modeling-Based Anomaly Detection of Subtle Lesions in Brain MRI Images 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.14.26363006v1?rss=1
</link>
<description><![CDATA[
The detection of small and subtle lesions in high-resolution 3D volumes is a highly relevant, yet far from solved task in biomedical imaging. We here address a specific task in detecting certain types of epileptogenic lesions through our novel semi-supervised spatially context aware transformer (SpyCAT) approach to anomaly detection. SpyCAT is modeling-based in the sense that it builds on specific assumptions that constitute what is normal and what constitues relevant deviations from normality. We explicitly use these assumptions to justify the inductive bias of our anomaly detection approach. The resulting SpyCAT system is patch-based and uses a transformer architecture to process discrete tokens obtained from a vector quantizing variational autoencoder, which produces counterfactual patches through full 3D convolutions of each patch. We evaluate our approach on the grounds of point-annotations of two subtypes of epileptogenic lesions, using validation measures that build on the Metrics Reloaded framework, showing that SpyCAT can reliably identify and localize the lesion types under consideration, and outperforms state-of-the-art reference methods. Our code is available at https://github.com/johannesSX/SpyCAT.
]]></description>
<dc:creator><![CDATA[ Schwarz, J., Will, L., Wellmer, J., Mosig, A. ]]></dc:creator>
<dc:date>2026-09-15</dc:date>
<dc:identifier>doi:10.64898/2026.09.14.26363006</dc:identifier>
<dc:title><![CDATA[Spatially Context-Aware Transformers Facilitate Modeling-Based Anomaly Detection of Subtle Lesions in Brain MRI Images]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-15</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.10.26362715v1?rss=1">
<title>
<![CDATA[
Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.10.26362715v1?rss=1
</link>
<description><![CDATA[
ObjectiveAmbient AI scribes evaluated using frequency-based metrics such as word error rate (WER), which do not represent clinical consequence. We tested whether variation in these metrics tracks consequential transcription errors.

MethodsWe constructed a multilingual corpus from five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe. Six frequency metrics were computed. Three independent large language model raters from external providers assessed clinically meaningful error patterns in context using a Severity x Likelihood framework informed by UK digital clinical-safety-risk-management principles.

ResultsAcross 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH. None of six frequency metrics showed a statistically detectable association with serious clinical risk across languages; correlations were small (absolute Spearman {rho}<0.16). A Severity x Likelihood sum remained strongly correlated with WER ({rho}=0.80), showing that the aggregate remained dominated by benign errors. At complexity level 3, low-resource languages had worse WER than high-resource languages ({beta}=+0.078, 95% CI +0.045 to +0.111; p<0.0001), without a detectable difference in CRITICAL/HIGH risk (OR 1.21, 95% CI 0.43 to 3.43; p=0.72).

Consultation complexity was the principal predictor of serious risk (OR 3.06 per level, p<0.0001).

ConclusionAcross this controlled multilingual corpus, aggregate transcription-frequency metrics did not reliably track the sparse severe tail of clinically consequential errors. WER remains appropriate for transcription quality, but these data do not support its use alone as a proxy for clinical safety.

Context-aware assessment of error consequence provides complementary information that frequency measures can dilute.

HighlightsO_LIWER did not track serious clinical-risks rate across 77 languages
C_LIO_LIFive related frequency metrics showed the same dissociation
C_LIO_LISeverity-weighting remained dominated by common low-risk errors
C_LIO_LIQuality tracked language resource; serious risk tracked complexity
C_LIO_LIWe provide a reusable context-aware clinical-risk instrument
C_LI
]]></description>
<dc:creator><![CDATA[ Bergman, H., Liu, V., Austin, B., Sanghera, R. ]]></dc:creator>
<dc:date>2026-09-14</dc:date>
<dc:identifier>doi:10.64898/2026.09.10.26362715</dc:identifier>
<dc:title><![CDATA[Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-14</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.13.26362573v1?rss=1">
<title>
<![CDATA[
The information in diagnostic tests 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.13.26362573v1?rss=1
</link>
<description><![CDATA[
BackgroundDiagnostic accuracy measures describe test performance, but expected learning also depends on the probability of disease before testing. We explain diagnostic information as expected uncertainty reduction and construct an empirical reference for its interpretation.

MethodsWe reanalysed structured study-level diagnostic accuracy data from Cochrane reviews, public repositories and PubMed Central review tables. A common continuity-corrected bivariate random-effects model pooled sensitivity and specificity for defined diagnostic profiles. We calculated mutual information in bits and as the percentage of starting uncertainty resolved at stated disease probabilities. Profiles contributed equally to the reference distribution. Sensitivity analyses examined alternative estimators, equal-review weighting and whole-review resampling.

ResultsThe reference included 273 pooled profiles from 210 reviews, comprising 4104 study-result appearances. Median uncertainty resolved was 23.3%, 29.2% and 31.4% at starting probabilities of 5%, 20% and 50%, respectively. At 20%, the median represented 0.211 of 0.722 bits, with a profile-bootstrap 95% confidence interval of 27.6% to 32.7% for the percentage resolved. Equal-review weighting gave 30.0%; whole-review resampling gave an interval of 27.5% to 33.5%. Two components of the head impulse, nystagmus and test of skew examination with nearly equal Youden indices resolved 18.7% and 28.1% at a 5% starting probability. Carcinoembryonic-antigen thresholds illustrated probability-dependent information ordering; the Ottawa ankle rule distinguished expected learning from uncertainty change after one result.

ConclusionsDiagnostic information makes expected learning explicit alongside conventional accuracy measures. Reporting bits and percentage of starting uncertainty resolved, with the starting probability and clinical context, supports interpretation and comparison of diagnostic evidence.
]]></description>
<dc:creator><![CDATA[ He, S., Locke, B. W., Joseph, J. W., Liebovitz, D. M., Rohlfsen, C., Goff, A., Safari, P., Kabrhel, C., Mohamed, A., Raja, A. S., Goldstein, J. N. ]]></dc:creator>
<dc:date>2026-09-14</dc:date>
<dc:identifier>doi:10.64898/2026.09.13.26362573</dc:identifier>
<dc:title><![CDATA[The information in diagnostic tests]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-14</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.12.26362907v1?rss=1">
<title>
<![CDATA[
App-Guided Hepatitis Self-Testing Versus Provider-Administered Rapid Testing Among Out-of-School Young Adults in Ibadan, Nigeria 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.12.26362907v1?rss=1
</link>
<description><![CDATA[
BackgroundHepatitis B virus (HBV) seroprevalence reaches 10-12% among out-of-school youths in urban Southwest Nigeria, yet fewer than three in ten in this demographic have ever undergone HBV testing. E-health self-testing applications offer an accessible alternative delivery modality for hepatitis screening, but direct evidence comparing their acceptance against traditional laboratory-based delivery in informal commercial settings is lacking.

ObjectiveTo determine whether e-health app-guided hepatitis self-testing achieves equivalent disease detection and superior acceptance outcomes compared with traditional laboratory-based screening among out-of-school youths in a high-footfall commercial site in Ibadan, Nigeria.

MethodsA cross-sectional parallel-arm pilot study recruited 120 out-of-school youths at commercial premises on Iwo Road, Ibadan (test group: n=60, e-health app-guided hepatitis B and C self-testing; control group: n=60, traditional laboratory-based testing using identical rapid diagnostic kits). Pre- and post-test structured questionnaires assessed technology readiness, behavioural intention, perceived ease of use, and adoption barriers. The Technology Acceptance Model (TAM) and Unified Theory of Acceptance and Use of Technology (UTAUT) provided the analytical framework. Group comparisons used chi-square and Fishers exact tests (=0.05).

ResultsHBV seroprevalence was 10.0% in the test group and 11.7% in the control group ({chi}{superscript 2}=0.087, p=0.769), confirming statistically equivalent detection across modalities. Hepatitis C virus was undetected in both groups. The test group demonstrated greater willingness to use e-health testing over traditional laboratory services (81.6% vs 70.0%), higher comfort with mobile health technology (75.0% vs 71.7%), and equivalent recommendation intent (93.3% vs 95.0%). Prior e-health service use was significantly higher in the test group (28.3% vs 13.3%; {chi}{superscript 2}=4.09, p=0.043). Anticipated ease of use (75.0% "very comfortable") substantially exceeded experienced ease post-use (41.7% "very or extremely easy"). Privacy concern was the leading adoption barrier in both groups (30.0% vs 36.7%), outranking cost (25.0% vs 15.0%).

ConclusionsE-health self-testing achieves non-inferior hepatitis B detection compared with traditional laboratory delivery while generating comparable acceptance outcomes. The gap between anticipated and experienced ease of use identifies usability as the primary design target. Privacy-centred tool design, not cost reduction alone, is the critical condition for scaling hepatitis e-health self-testing among out-of-school youths in urban Nigeria.
]]></description>
<dc:creator><![CDATA[ Adeluwoye, A. O., Rufai, Y. M., Olayinka, A. R., Adeluwoye, N. N., Alabetutu, A. ]]></dc:creator>
<dc:date>2026-09-14</dc:date>
<dc:identifier>doi:10.64898/2026.09.12.26362907</dc:identifier>
<dc:title><![CDATA[App-Guided Hepatitis Self-Testing Versus Provider-Administered Rapid Testing Among Out-of-School Young Adults in Ibadan, Nigeria]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-14</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.medrxiv.org/content/10.64898/2026.09.12.26360876v1?rss=1">
<title>
<![CDATA[
AI-Based Synthetic Data in Biomedicine: A Decade of Growth and a Persistent Translation Gap 
]]>
</title>
<link>
https://www.medrxiv.org/content/10.64898/2026.09.12.26360876v1?rss=1
</link>
<description><![CDATA[
AI-generated synthetic data are increasingly used to address data scarcity, privacy constraints and experimental limitations in biomedicine, but how far these methods have translated into practice remains unclear. We conducted a systematic mapping and bibliometric analysis of 4,143 publications spanning 2015-2025, combining expert annotation with LLM-assisted classification across data modality, medical domain, paper type, deployment status and research stance. Publication volume grew continuously; 77.8% of papers were strongly supportive while critical work remained below 1%. Medical imaging dominated the corpus, consistent with well-characterized transformation-group invariances supporting data augmentation and generative modeling. Highly cited primary research concentrated disproportionately in molecular and pharmaceutical applications, where SE(3)-equivariant architectures and structure-prediction models accelerated generative approaches. Only 27 publications reported operational use; omics and tabular clinical data, lacking well-characterized invariance structures, remained underrepresented. These findings reveal a gap between methodological growth and deployment, motivating investment in evaluation standards, deployment reporting and encoding domain-relevant invariances.
]]></description>
<dc:creator><![CDATA[ Asgari, N., Perez, I. F., Epelde, G., Zhang, L., Horesh, L., Saab, C., Muszkat, M., Rosen-Zvi, M. ]]></dc:creator>
<dc:date>2026-09-14</dc:date>
<dc:identifier>doi:10.64898/2026.09.12.26360876</dc:identifier>
<dc:title><![CDATA[AI-Based Synthetic Data in Biomedicine: A Decade of Growth and a Persistent Translation Gap]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-14</prism:publicationDate>
<prism:section></prism:section>
</item>
</rdf:RDF>
