Life sciences · Preprint
arXiv · September 3, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint describes a methodological study comparing large language models to traditional machine learning for CKD screening using zero-shot and few-shot learning. LLMs achieved competitive performance in low-data settings but showed model-dependent and unstable results as complexity increased. The work is exploratory and does not establish clinical utility or real-world screening effectiveness.
Retrospective algorithmic evaluation with multiple model comparisons. Intervention: Large language models (LLMs) with clinically selected tabular features and structured prompt templates, evaluated in zero-shot and few-shot in-context learning settings. Compared with: Traditional machine learning, deep learning, tabular foundation models (TFM), and existing CKD screening tools.
LLMs achieved competitive performance using only a small number of examples, often matching or outperforming traditional approaches in low-data settings LLM performance remains model-dependent and less stable as input complexity increases ML, DL, and TFM models show more consistent improvement with larger training data
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an uncontrolled, methodological evaluation of LLM performance on a screening task using retrospective data, without clinical validation, prospective deployment, or comparison to a gold standard in a real-world screening cohort.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning (ML) and deep learning (DL) approaches require labeled data and model training, limiting their use in real-world screening settings. This study evaluates the effectiveness of large language models (LLMs) for CKD screening under zero-shot and few-shot in-context learning settings and compares them with traditional ML and DL methods. We propose a framework that uses clinically selected tabular features and structured prompt templates to enable LLM-based inference without task-specific training. LLM performance is evaluated across multiple prompt styles, feature configurations, and data settings, and compared with standard ML, DL, and tabular foundation model (TFM) baselines, and existing CKD screening tools. The results show that LLMs can achieve competitive performance using only a small number of examples, often matching or outperforming traditional approaches in low-data settings. However, their performance remains model-dependent and less stable as input complexity increases. In contrast, ML, DL, and TFM models show more consistent improvement with larger training data. Overall, the findings highlight a trade-off between data efficiency and stability, suggesting that LLMs may serve as a flexible complementary approach for CKD screening when labeled data are limited.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.